Learning focus monitoring method and system for learning tablet

CN122796579APending Publication Date: 2026-09-22JIANGXI BUTIAN INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610777262.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-01
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0005]本发明提供用于学习平板电脑的学习专注度监测方法及系统,解决相关技术中如何在无教师实时监督的学习平板自主学习场景中,克服现有专注度监测方案难以识别伪专注状态、难以在注视落点仍位于关键区域时及时捕捉认知脱离信号、以及通用模型难以适配个体差异并易造成干预误报的技术缺陷的技术问题

Benefits of technology

以内容管理接口输出的内容时序状态与激活事件流为基准,按视频讲解、图文阅读、互动习题三类场景构建动态更新的内容感知期望注视高斯混合模型,并从空间分布偏离与时序动态异常两个维度计算包括内容激活事件注视响应性、关键区域凝视占比、扫视方向对齐度在内的多维偏离度特征,可在注视落点仍位于关键区域时识别伪专注状态,弥补单纯依赖注视空间位置检测的识别盲区,使专注度判别与内容节奏保持一致;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122796579A_ABST
    Figure CN122796579A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent learning terminals, and discloses a learning concentration monitoring method and system for a learning tablet computer, wherein the learning concentration monitoring method for the learning tablet computer collects original behavior data from a front camera and a touch module of the learning tablet, obtains learning content time sequence states in combination with a content management interface, and forms a synchronous data stream aligned with a time axis; a content perception expected gaze model is constructed, a multi-dimensional deviation degree feature sequence is calculated from two dimensions of spatial distribution deviation and time sequence dynamic anomaly, and a touch response time delay is fused; four types of concentration state time sequence sequences are obtained through online parameter self-adaptation and Viterbi decoding of a personalized hidden Markov model, hierarchical content intervention is performed, and a learning process concentration analysis report is generated. The application can recognize a pseudo-concentration state and accurately locate a content position of attention loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent learning terminal technology, and more specifically, to a method and system for monitoring learning focus on a learning tablet computer. Background Technology

[0002] With the increasing popularity of learning tablets in K-12 self-learning scenarios, front-facing camera gaze tracking and touch behavior collection have been used to help judge students' learning focus. Related solutions usually infer focus based on whether the gaze falls within the screen content area or by statistically analyzing the proportion of gaze points staying in a preset area, and trigger reminders or record learning behaviors accordingly.

[0003] However, students often experience pseudo-focus during prolonged self-directed learning, where their eyes remain fixed on the screen but their cognitive processing has deviated from the learning content. At this time, the focus may still be located in key areas such as the subtitle bar, text line, or question stem. Existing solutions rely solely on the spatial location of the gaze, making it difficult to distinguish between normal focus and pseudo-focus. Furthermore, when abnormalities in the dynamics of the gaze sequence have already occurred in the early stages of pseudo-focus, the solutions fail to capture the signals in time, leading to distorted focus assessments and delayed intervention.

[0004] Furthermore, individual differences exist in students' learning pace and attention baselines, making it easy to cause systematic misjudgments when using fixed thresholds or generalized models. Existing interventions often rely on content-irrelevant pop-up windows or vibration prompts, which can easily disrupt the normal learning process. Therefore, a focus monitoring technology solution is needed that uses the temporal structure of learning content as a reference, takes into account spatial deviations and temporal dynamic anomalies, supports personalized state recognition, and implements content-oriented tiered interventions. Summary of the Invention

[0005] This invention provides a method and system for monitoring learning focus on a learning tablet computer, addressing the technical problems in related technologies of how to overcome the shortcomings of existing focus monitoring schemes in self-directed learning scenarios on learning tablet computers without real-time teacher supervision. These shortcomings include difficulty in identifying pseudo-focused states, difficulty in timely capturing cognitive disengagement signals when the gaze point is still in the key area, and the difficulty of adapting general models to individual differences, which can easily lead to false alarms during intervention.

[0006] This invention provides a method for monitoring learning focus on a learning tablet computer, comprising the following steps: S1 collects raw behavioral data from the front-facing camera and touch module of the learning tablet, and combines it with the content management interface to obtain the real-time temporal status of the learning content, resulting in a multimodal behavior and content synchronized data stream aligned with the timeline. S2, based on the temporal state of content in the synchronous data stream, constructs the gaze behavior expectation distribution according to three types of learning content scenarios, and uses the Gaussian mixture modeling method to obtain a dynamically updated content-aware gaze expectation model; S3, based on the content-aware expected gaze model, calculates gaze behavior deviation features from two dimensions: spatial distribution deviation and temporal dynamic anomaly, integrates touch response latency, and outputs a multi-dimensional deviation feature sequence. S4. Input the multidimensional deviation feature sequence into the personalized hidden Markov attention model, and obtain four types of attention time-series state sequences and state transition confidence through online parameter adaptation and Viterbi decoding. S5, based on the four types of attention time-series state sequences and state transition confidence, adopts a graded content intervention strategy and learning event recording mechanism to output real-time intervention instructions and attention analysis reports during the learning process.

[0007] Furthermore, in S1, the content management interface continuously outputs content type tags, key area coordinate sets, content progress markers, and paragraph summary fields, and generates activation event records when a substantial update occurs in the key area, thus forming a content activation event stream; The front-facing camera uses a lightweight facial landmark detection model with MobileNet as the backbone network to output the coordinates of 68 semantic landmarks in the binocular regions. When a new user starts the camera for the first time, a 5-point gaze calibration process is performed to establish a personalized eye geometric mapping parameter matrix. The gaze dwell event sequence and saccade motion event sequence are maintained simultaneously and stored independently with timestamps as indexes. They are not included in the 10Hz resampling matrix. After the three data streams are uniformly resampled to 10Hz, if the confidence level of key point detection at several consecutive sampling points is lower than the first confidence threshold, a gaze data missing marker is written for the corresponding time period; if the gaze point exceeds the screen boundary, a gaze off-screen marker is written. Both types of quality markers are transmitted to step S3 along with the synchronous data stream.

[0008] Furthermore, in S2, the expected attention GMM of the video explanation content includes a subtitle component, a whiteboard component, and a social attention component. The weight of the subtitle component is determined by the ratio of the current number of subtitle characters to the total number of visible text characters. The weight of the whiteboard component is increased by adding the weight increment according to the ratio of the area increment to the area of ​​the subtitle bar after the whiteboard area expansion event, and then linearly decays within a preset decay time. If there is no visible text in the current frame, the remaining weight is borne by the background browsing component. If the lecturer's avatar area coordinate field is empty, the social attention component is not set, and the corresponding weight is merged into the subtitle component. The expected attention GMM for interactive exercise content dynamically adjusts the weight of each component as the answering time progresses and as events occur in the touch option area: when the question appears, the weight of each area is relatively uniform, the weight of the question stem area gradually increases as time progresses, and the weight of the corresponding option component surges when events occur in the touch option area. When the content type label changes, the active expectation gaze GMM is immediately replaced with the model corresponding to the new content type, and the mean parameters of each component are initialized according to the coordinates of the key region of the new content.

[0009] Furthermore, in S2, the system maintains a reading frontier pointer, which is defined as the historical maximum value of the vertical coordinate of the gaze landing point within the current content item, and is only updated as reading progresses downwards; when the content identifier changes, the touch page jump crosses more than one page, or a new learning session is started, the pointer is reset to the vertical coordinate of the topmost text line of the current content display area; The main line is the paragraph where the reading front pointer is located. The image-text reading expected gaze GMM sets a high-weight main line component at the position of the main line, sets an auxiliary component above the main line, and sets a pre-read component below the main line. The mixed weight of the auxiliary component is one-third of the mixed weight of the main line component, and the mixed weight of the pre-read component is one-fifth of the mixed weight of the main line component. The mean x-axis of the main line component is determined in real time by the median x-axis of the gaze landing point in the main line area within the current analysis window. When the sample is less than the preset minimum sample threshold, it degenerates into a uniform calculation from the left end of the line according to the preset average reading speed. If there is an image area on the screen, an additional weight is added to the image component, which is calculated as the ratio of the image area to the total area of ​​the content area.

[0010] Furthermore, in step S3, at the beginning of each analysis window, the off-screen gaze marker and gaze data missing marker in the synchronization time matrix are checked first. If they exist, the state 4 trigger indication is directly output to step S4 without entering the feature calculation process. Spatial distribution deviation is measured by KL divergence, which calculates the KL divergence from the actual gaze distribution to the direction of the content-aware expected gaze model. The actual gaze distribution is constructed using a kernel density estimation method, which uses a Gaussian kernel function with fixed bandwidth to transform the gaze coordinate samples in the current analysis window into a two-dimensional continuous probability density function. The kernel bandwidth is adaptively set according to the key region size of the current content type. Feature calculation uses a fixed-length sliding time window, moving forward by one sampling point at a time to ensure that the deviation feature sequence is consistent with the temporal resolution of the original data stream.

[0011] Furthermore, in S3, the content activation event gaze responsiveness feature NRR takes the activation events that have completed response judgment within the current analysis window as the statistical objects. The response judgment of the activation event is completed within a preset response judgment time after the event timestamp. For each statistical object, the gaze landing point coordinates at the event timestamp are taken as the response starting point. The gaze displacement vector is obtained by subtracting the response starting point from the average gaze landing point in the response detection window. The content direction vector is obtained by subtracting the response starting point from the center coordinates of the newly activated area. When the magnitude of the gaze displacement vector is less than the preset effective gaze displacement threshold, it is marked as no response; when the magnitude meets the condition, if the direction cosine value of the gaze displacement vector and the content direction vector is lower than the preset response direction cosine threshold, it is also marked as no response; the NRR value is obtained by the ratio of the number of no response events in the current window to the total number of events that have been judged. If there are no completed activation events in the current analysis window, the NRR is taken from the NRR value of the window containing the most recent valid statistical events in the current session, and an NRR validity marker is added to the feature vector as a historical retention state.

[0012] Furthermore, in S3, the key area gaze proportion feature FS takes the gaze dwell events whose center coordinates fall within the key area of ​​the content in the current analysis window as the statistical object, and calculates the proportion of the total duration of gaze events whose dwell time exceeds the preset gaze judgment duration threshold; the scanning direction alignment feature SA is extracted from the scanning motion events of the key area of ​​text content, after filtering out line break back scan events with negative horizontal coordinate components and positive vertical coordinate components, the cosine mean of the direction is calculated for the remaining positive reading scanning, and the SA feature of non-text content areas is set as invalid; In interactive exercise scenarios, the touch response latency deviation index is calculated by segmented mapping from the time the question appears to the time the student makes the first valid click in the option area: the deviation index is 0 within the first latency threshold, and reaches its maximum value after exceeding the second latency threshold, increasing linearly within the interval; touch events falling in non-critical areas are marked with abnormal touch features and included in the deviation feature vector.

[0013] Furthermore, in S4, two independent HMMs are maintained for video / text / image scenarios and interactive exercise scenarios respectively: the observation vector of the video / text / image HMM is four-dimensional: KL divergence, NRR, FS, and SA; the observation vector of the interactive exercise HMM is four-dimensional: KL divergence, NRR, FS, and touch response delay deviation exponent. When switching content types, the posterior probability distribution of the state at the last moment of the previous type is used as the initial state prior of the new active HMM. HMM parameters are constructed in two stages: the initial stage is initialized with prior parameters of learners of the same age group, and the group prior is obtained by offline training from classroom learning session data; after each learning session, the corresponding type of HMM parameters are incrementally corrected using online Baum-Welch, and the update amount and historical parameters are integrated using exponential moving average. The EMA weight coefficient is decreased by a preset step size in each session, and remains unchanged after reaching the minimum EMA weight coefficient. When the Viterbi decoding result shows a shift from shallow focus to attentional drift, confirmation is only possible if the ratio of the posterior probabilities of state 3 to state 2 exceeds the third confirmation threshold and the duration of the shift exceeds the first duration threshold. After state 3 confirmation, when a valid confirmation touch is detected in the specified response area, the confirmation mark of state 3 is cleared, the current active HMM state is initialized to shallow focus, and Viterbi incremental decoding is restarted.

[0014] Furthermore, in S5, for shallow focus, the system enters a silent monitoring mode; when the continuous duration of shallow focus exceeds the first duration threshold, the first level of intervention is triggered, and a dynamic visual focusing effect is superimposed on the key area of ​​the current content to guide the learner to pay attention to the current content area. For confirmed attention drift states, if the duration does not exceed the second duration threshold, a second-level intervention is triggered: depending on the content type, a video pause prompt, a paragraph summary pop-up, or a prompt about the nature of the problem is executed; if the duration exceeds the second duration threshold, a third-level intervention is triggered: based on the content pause, a voice broadcast is used to wake up the student, waiting for the student to make an active touch operation in the designated confirmation area to trigger the state 3 recovery mechanism in step S4; After the learning session ends, a focus analysis report is generated, which includes a focus time series curve, statistics on the time ratio of each state, location of content where attention was lost, and longitudinal comparison of historical sessions. The location of content where attention was lost is mapped to the timestamps of attention drift and attention absence events with the content progress markers, accurately marking the content position when attention was lost. For video explanations, the playback progress is marked down to the second level; for text and image readings, it is marked down to the paragraph number; and for interactive exercises, it is marked down to the exercise number.

[0015] Furthermore, the learning attention monitoring system for the learning tablet computer performs the steps in the above-described learning attention monitoring method for the learning tablet computer, including: The data acquisition and synchronization module is used to collect raw behavioral data from the front-facing camera and touch module of the learning tablet, and combine it with the content management interface to obtain the real-time temporal status of the learning content, so as to obtain a multimodal behavioral and content synchronization data stream aligned with the time axis. The expected gaze distribution modeling module is used to construct the expected gaze distribution based on the temporal state of content in the synchronous data stream and three types of learning content scenarios. It adopts the Gaussian mixture modeling method to obtain a dynamically updated content-aware expected gaze model. The multidimensional deviation feature calculation module is used to calculate the gaze behavior deviation features from two dimensions, spatial distribution deviation and temporal dynamic anomaly, based on the content-aware expected gaze model, and integrates touch response latency to output a multidimensional deviation feature sequence. The personalized state recognition module is used to input the multidimensional deviation feature sequence into the personalized hidden Markov attention model. After online parameter adaptation and Viterbi decoding, four types of attention time-series state sequences and state transition confidence are obtained. The intervention and reporting output module is used to output real-time intervention instructions and learning process attention analysis reports based on four types of attention temporal state sequences and state transition confidence, using a graded content intervention strategy and learning event recording mechanism.

[0016] The beneficial effects of this invention are as follows: Based on the content time sequence status and activation event flow output by the content management interface, a dynamically updated content perception expectation gaze Gaussian mixture model is constructed according to three scenarios: video explanation, text and image reading, and interactive exercises. Multidimensional deviation features, including content activation event gaze responsiveness, key area gaze ratio, and saccade direction alignment, are calculated from two dimensions: spatial distribution deviation and temporal dynamic anomaly. This model can identify pseudo-focused state when the gaze point is still in the key area, making up for the blind spot of recognition that relies solely on gaze spatial position detection, and ensuring that focus judgment is consistent with the content rhythm. Personalized Hidden Markov Attention Models are maintained for video and text content and interactive exercises respectively. Combined with online parameter adaptation, a dual confirmation mechanism for shallow attention to attention drift, and a graded content intervention strategy, it can adapt to individual student differences and suppress intervention false alarms caused by short-term fluctuations. The learning report maps attention drift and attention loss events to content progress markers, which can trace back and mark the content location where attention loss occurred, providing a basis for teaching adjustments and review arrangements. Attached Figure Description

[0017] Figure 1 This is a flowchart of the learning focus monitoring method for a learning tablet computer according to the present invention; Figure 2 This is a flowchart of the learning attention monitoring method for a learning tablet computer according to the present invention. Figure 1 ; Figure 3 This is a flowchart of the learning attention monitoring method for a learning tablet computer according to the present invention. Figure 2 . Detailed Implementation

[0018] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.

[0019] At least one embodiment of this invention discloses a method for monitoring learning focus on a learning tablet computer. Using the temporal structure of learning content as a reference benchmark for behavioral analysis, a content rhythm-aware attention monitoring method is constructed. The entire method sequentially includes five main steps: multimodal behavior-content synchronous data acquisition and preprocessing, content-aware expected gaze distribution modeling, gaze-content multidimensional deviation feature calculation, personalized attention state temporal identification, and generation of tiered intervention strategies and output of learning reports. Figures 1 to 3 As shown, it includes the following steps: S1 collects raw behavioral data from the front-facing camera and touch module of the learning tablet, and combines it with the content management interface to obtain the real-time temporal status of the learning content, resulting in a multimodal behavior and content synchronized data stream aligned with the timeline. In one embodiment of the present invention, a real-time data set is established that synchronously integrates student behavior signals with screen content status. The acquisition logic of the three signals has different focuses, and the acquisition methods and specific contents of each signal are as follows: The acquisition of the temporal status of learning content relies on the content management interface built into the tablet learning application. This interface continuously outputs the current screen's content status information during application operation, including the following fields: content type tag (one of three categories: video explanation, text / image reading, or interactive exercises); content presentation timestamp, accurate to milliseconds, marking the moment the current content status occurred; a set of coordinates for key content areas, defined according to content type. In video explanation content, this includes the subtitle bar, the active writing area on the whiteboard, the illustration area, and the instructor's avatar area coordinates field pre-configured in the video metadata (recording the fixed rectangular coordinates of the instructor's face within the video screen display area; if this field is empty, it indicates that the video does not contain any instructor-related content, and no social attention component will be set during subsequent GMM modeling); in text / image reading content, this includes the rectangular bounding boxes of each paragraph's text lines and the outer rectangles of the accompanying illustrations; in interactive exercises content, this includes the question stem text area, the text areas for each option, and the submit / confirm button area.

[0020] In addition to coordinate information, the content management interface also outputs two types of supplementary metadata synchronously with the content status. The first is a content progress marker: for video tutorials, it records the current playback progress (accurate to the second); for text-based content, it records the current paragraph number (an integer starting from 1); and for interactive exercises, it records the current exercise number (an integer starting from 1). These three types of progress markers are written to the data stream with each content status update. Each sampling point corresponds to a specific content progress value, allowing step five to accurately locate attention drift events to the corresponding content position during report generation. The second is a paragraph summary field, which only applies to text-based content. The platform pre-enters the core points of each paragraph during content creation. If a paragraph summary field is empty, the first line of text in that paragraph is used as an approximate substitute for the summary content, which is used for displaying the pop-up content during the second-level intervention in step five.

[0021] The content management interface also continuously outputs a stream of activation events for key content areas. Whenever a key area undergoes a substantial update, including a video subtitle text switching to a new segment, an expansion of the whiteboard's active area, or a new exercise question, the interface generates an activation event record at the current timestamp. This record includes three fields: event type label, event timestamp, and the screen coordinate range of the newly activated area. These activation events form the time reference for calculating gaze responsiveness features in subsequent steps. The entire acquisition process is completed locally on the tablet, requiring no additional image recognition computation overhead.

[0022] Real-time extraction of gaze trajectories is performed using the tablet's front-facing camera. The camera continuously captures facial image frames of students at a fixed frame rate. A lightweight facial landmark detection model with MobileNet as its backbone is deployed on the device. This model undergoes specific quantization and compression for near-field tablet use, enabling real-time inference at 30 frames per second on the tablet processor's neural network acceleration unit. For each input image frame, the model outputs the coordinates of 68 semantic landmarks in the binocular regions, covering the upper and lower eyelid contours, inner and outer corners of the eyes, and the iris edge. Each landmark is accompanied by a confidence score, reflecting the detection reliability of the current frame.

[0023] Based on these 68 key points, the estimation of the gaze point coordinates from the key point coordinates to the screen coordinates is completed through eye geometry mapping: The two-dimensional offset vectors of the left and right iris center points (calculated from the geometric centers of the iris edge key points) relative to the corresponding eye opening geometric centers are used as the original input. Combined with the estimated distance from the device camera to the student's face (calculated from the ratio of the pixel length of the interpupillary distance to the known average interpupillary distance) as a depth reference parameter, and utilizing a pre-established perspective projection mapping relationship, the gaze offset vectors of both eyes are converted into gaze point coordinates in the screen coordinate system. Given the individual differences in eye anatomy among students, a 5-point gaze calibration process is performed when a new user first starts the learning application: the system displays calibration guide points sequentially at the four corners and the center of the screen, guiding the student to gaze at each guide point and hold for approximately 2 seconds. A personalized eye geometry mapping parameter matrix for the current user is then established through affine transformation fitting. This parameter matrix is ​​stored locally on the device and directly accessed in all subsequent learning sessions.

[0024] Based on the frame-by-frame gaze point coordinates, the system maintains two types of derived event sequences in real time. These sequences are stored independently using timestamps as indexes and are not included in the 10Hz resampling matrix. They are used directly in step three for calculating dwell and saccade features via time range lookup. The dwell event sequence is maintained as follows: a 28-pixel radius neighborhood is used as the static determination range. When consecutive frames of gaze points fall within this neighborhood, a dwell timer is initiated, recording the entry timestamp and the current gaze center coordinates. Dwell ends when the gaze point leaves the neighborhood, recording the end timestamp. The difference between the two timestamps is the dwell duration. The complete dwell event (start timestamp, end timestamp, center coordinates, dwell duration) is written into the dwell event sequence. The saccade motion event sequence is maintained as follows: whenever the gaze point shifts between two consecutive dwell events, a saccade event is recorded, including the occurrence timestamp, start coordinates (center of the previous dwell event), and end coordinates (center of the current dwell event). The sign information of both the horizontal and vertical coordinate components is retained for use in step three when calculating the saccade direction alignment features.

[0025] Touch interaction logs are collected asynchronously by the system's touch event listening module. Each touch log entry records the following fields: event type (click, long press, continuous swipe start / end point, page jump click, input confirmation button), event timestamp, the position of the touch point in the screen coordinate system, and the time interval from the current event timestamp to the previous event timestamp. In the touch event type field, continuous swipes and page jump clicks are clearly distinguished for use in step two when determining the conditions for resetting the reading frontier pointer. In interactive exercise-type content scenarios, touch response latency has special value as an indicator of attention; the length of time from when the question appears to when the student first makes a valid touch in the option area reflects the time the student actively participates in processing the question.

[0026] The three data streams (content state sequence and activation event stream, gaze coordinate sequence, and touch log sequence) are generated in parallel by their respective threads, each carrying an independent timestamp. To ensure accurate temporal correspondence among the three data streams, a unified system clock millisecond-level timestamp is used as the global alignment benchmark. Linear interpolation is used to uniformly resample the three data streams to a fixed frequency of 10 analysis sampling points per second. Gaze coordinates, content state, content progress markers, and touch records form a multi-channel synchronous timing matrix. Gaze persistence event sequences and saccade motion event sequences are maintained in parallel as independent event lists and do not participate in the 10Hz resampling. After alignment, a basic quality screening is performed on the gaze trajectory data: if the confidence score of eye keypoint detection at several consecutive sampling points is lower than the first confidence threshold (the default value of this threshold is 0.60), the time period is marked as a missing gaze data segment, and the missing marker is written to the corresponding sampling point in the synchronous timing matrix; if the gaze landing coordinates of a sampling point exceed the screen boundary, a gaze-off marker is written to the corresponding sampling point. The two types of quality markers mentioned above serve as trigger signals for the absence of attention and are transmitted to step three along with the synchronous data stream. Step three checks these markers first before the start of each analysis window.

[0027] The output dataset specifically includes: a multi-channel synchronous timing matrix (including gaze coordinates, content status, content progress markers, touch records, and quality markers), a content activation event stream, a gaze persistence event sequence, and a saccade motion event sequence; step two only reads the content status information and activation event stream for expected distribution modeling, while step three reads all of the above data to calculate multidimensional deviation features.

[0028] S2, based on the temporal state of content in the synchronous data stream, constructs the gaze behavior expectation distribution according to three types of learning content scenarios, and uses the Gaussian mixture modeling method to obtain a dynamically updated content-aware gaze expectation model; In one embodiment of the present invention, a corresponding expected distribution model of gaze behavior is established for each of the three types of learning content. This model uses a Gaussian mixture model to parameterize the spatial probability distribution of gaze points that learners in a cognitively engaged state should follow under these types of content. The Gaussian mixture model (GMM) is chosen as the representation of the expected distribution because visual attention in a cognitively engaged state is not uniformly distributed across the entire screen, but is highly concentrated in a few areas carrying key information, and switches between these areas in a predictable order. The multi-component structure of the Gaussian mixture model naturally expresses this spatial distribution characteristic of multi-area concentrated attention. Each Gaussian component corresponds to a key content area, the component mean vector corresponds to the center coordinates of that area, the covariance matrix reflects the allowable range of gaze point diffusion around that area, and the mixture weight coefficient reflects the proportion of gaze time that a cognitively engaged learner allocates to that area.

[0029] Taking video-based explanations as an example, this section explains the construction process of the Expected Fixation Distribution (GMM) and the determination of its component weights. During a video explanation, the fixation behavior of students who are actually following the content mainly consists of two alternating modes: horizontal scanning following the subtitles and fixation shifts when new content appears in the whiteboard or illustrated areas. Correspondingly, the Expected Fixation GMM for video-based explanations includes three categories: subtitle component, whiteboard component, and social attention component. The subtitle component, with its mean value calculated from the three typical fixation points on the left, center, and right of the subtitle bar, corresponds to the fixation behavior of following the subtitles. The whiteboard component, with its mean value calculated from the area where new content appears in the whiteboard in the current frame, corresponds to the fixation behavior of following the whiteboard. The social attention component, with its mean value calculated from the center of the rectangle in the instructor's portrait area coordinate field (provided by the content management interface in step one), corresponds to the normal fixation switching when students occasionally observe the instructor's facial expressions. If this field is empty, no social attention component is set, and its corresponding weight is merged into the subtitle component.

[0030] The weights of the three components are dynamically determined through calculations based on content management interface data. The subtitle component weight W_sub is calculated as the ratio of the number of characters in the current subtitle segment N_sub to the total number of visible text characters in the current frame (the sum of the number of subtitle characters and the number of whiteboard text characters, denoted as N_all), i.e., W_sub = N_sub / N_all. When N_all is 0 (there are no visible subtitles or whiteboard text in the current frame, such as in a pure animation segment), the weights of the subtitle component and the whiteboard component are both set to 0, and the background browsing component, which takes the average of the entire video screen display area, bears the remaining weight, corresponding to the gaze pattern of cognitive input learners observing the overall content of the screen in a textless frame. The whiteboard component weights employ a dynamic update mechanism: the default value of the whiteboard component's base weight, W_board_base, is 0.08, serving as a persistent weight when there are no whiteboard updates. Each time a whiteboard area expansion event occurs in the activation event stream, the ratio of the current whiteboard active area increment ΔA to the caption bar area A_sub is calculated and added as a weight increment to the current whiteboard component weight. This weight then linearly decays back to W_board_base over a duration T_decay (the default value of T_decay is 4 seconds). The weight of the social attention component is fixed at W_social (the default value of W_social is 0.05). The sum of the weights for the three components is always equal to one, and the covariance matrix of each component is set proportionally based on the corresponding pixel area.

[0031] The expected gaze GMM construction logic differs for text-based reading content. During text-based reading, learners with cognitive engagement begin their gaze from the left end of the first line of a paragraph, progressing line by line to the right and down, with extended lingering time when encountering bolded words, headings, or key terms. Due to significant differences in reading speed among students, determining the current main line in the expected distribution requires tracking the student's actual reading progress, rather than relying on a fixed temporal preset. Therefore, the system maintains a reading frontier pointer, defined as the historical maximum value of the gaze placement point within the current content item along the vertical axis of the content area. This pointer is updated only as reading progresses downwards and remains unchanged when the student rereads upwards, always marking the deepest position the student has reached. The reading frontier pointer is reset to the vertical coordinate of the topmost text line of the current content display area in the following three situations: the content management interface reports a change in the content identifier of the text-based reading content (switching to a new article); the touch event type is a page jump click and the page number spans more than one page; and a new learning session is started (the pointer is not retained across sessions). When a touch-based page-turning event occurs, the system synchronously corrects the overall offset of the content area coordinate system to ensure that the reading front pointer always remains consistent in the physical content coordinate system.

[0032] Using the paragraph where the reading pointer is located as the main line, the expected gaze GMM for text-image reading organizes its components as follows: A high-weight main line component is set at the core of the current pointer line. Its mean x-coordinate is determined in real-time by the median x-coordinate of the gaze points falling within the main line area in the current 3-second analysis window, to track the student's actual position within the line. If there are fewer than 3 valid gaze samples in the main line area within the current window, it degenerates into an approximate substitute by uniformly calculating the current position from the left end to the right end of the line at the average reading speed (the default average reading speed is 3.5 Chinese characters per second). An auxiliary component with a weight of approximately one-third of the main line is set above it, corresponding to the learner's normal forward-referencing reading behavior; a pre-reading component with a weight of approximately one-fifth of the main line is set below it, corresponding to the natural switching of gaze anticipation of the beginning of the next line. If there is an illustrated area on the current screen, an additional illustrated component is added with the center of the bounding rectangle of the illustrated area as the average value. Its weight is calculated as the ratio of the illustrated area to the total area of ​​the current screen content area, and the weight gradually decreases as the reading pointer moves beyond the illustrated area.

[0033] Modeling the expected gaze distribution for interactive exercise content requires considering the phased changes in cognitive activity during the answering process. After an exercise is presented, students with cognitive engagement typically go through three gaze behavior phases: full question browsing, in-depth understanding, and answer confirmation. The expected gaze distribution model (GMM) for this type of exercise dynamically adjusts the weights of each component within different time windows after the question appears: in the first few seconds after the question appears, the weights of the stem and each option area are relatively evenly distributed; as time progresses, the weight of the stem gradually increases; when a student touches the option area, the weight of the corresponding option component surges, corresponding to the gaze concentration during the answer confirmation phase. This mechanism of dynamically adjusting weights as the answering process progresses allows the expected gaze distribution model to closely match the actual cognitive rhythm of answering exercises, avoiding misjudging the full-screen scanning behavior before completing the full question browsing as attentional drift.

[0034] The three types of expected gaze models mentioned above are seamlessly replaced when content states change: when the content interface reports a change in the content type label, the system immediately replaces the currently active expected gaze GMM with the model corresponding to the new content type, and initializes the mean parameters of each component based on the key region coordinates of the new content. Layout changes within the same type of content also trigger local updates to the mean and weight of the corresponding components, ensuring that the expected gaze model always accurately corresponds to the current screen content.

[0035] S3, based on the content-aware expected gaze model, calculates gaze behavior deviation features from two dimensions: spatial distribution deviation and temporal dynamic anomaly, integrates touch response latency, and outputs a multi-dimensional deviation feature sequence. Within each 3-second analysis window, a State 4 quality marker check is performed first. If successful, feature calculations are performed on two dimensions: spatial distribution deviation and temporal dynamic anomaly. At the start of each analysis window, step three checks if there are any gaze-off markers or missing gaze data markers in the synchronization temporal matrix from step one within the current window. If so, the window does not calculate feature vectors and directly outputs a State 4 trigger indication to step four. Step four receives this and directly outputs an attention absence (State 4) label, without entering the HMM inference process. Only after the check passes does step three proceed with the normal feature calculation process, quantifying the gap between actual gaze behavior and content expectations from two dimensions: spatial distribution deviation and temporal dynamic anomaly. Spatial distribution deviation detects whether the gaze point deviates from the key content area, while temporal dynamic anomaly features detect pseudo-focus situations where the gaze point is located in the key area but the behavioral rhythm does not match the cognitive engagement state. These two features complement each other, covering different manifestations of attention loss.

[0036] The core metric for spatial distribution deviation is KL divergence. KL divergence measures the degree of difference between two probability distributions; it is zero when the two distributions are perfectly identical and increases monotonically as the difference increases, satisfying the measurement requirement where the degree of deviation is the core semantic. This method calculates the KL divergence from the actual gaze distribution to the desired gaze distribution. When a large number of actual gaze points fall in areas where the probability quality of the desired distribution is concentrated, the KL divergence approaches zero; conversely, when a large number of actual gaze points fall in non-critical areas, the KL divergence reaches a larger value. In the actual calculation, firstly, using a kernel density estimation method, a fixed-bandwidth Gaussian kernel function is used to transform several gaze coordinate samples within the current time window into a continuous probability density function on a two-dimensional plane, and then the divergence integral is calculated with the desired GMM. The kernel bandwidth parameter is adaptively set according to the size of the critical region of the current content type: the default kernel bandwidth for video subtitle regions is 40 pixels, for image / text paragraph regions it is 55 pixels, and for exercise option regions it is 25 pixels.

[0037] The design of the sliding time window is crucial to the sensitivity and stability of deviation calculation. A window that is too short includes insufficient fixation samples, leading to significant bias in kernel density estimation; a window that is too long reduces response speed, as early signals of attentional drift are diluted by historical normal fixation data, resulting in detection lag. Considering that the transition of human attention from a normal state to a significantly drifting state typically occurs on the order of several seconds to tens of seconds, and combined with a data frequency of 10 analysis sampling points per second, a fixed-length sliding window is used for analysis. The default value for the sliding window length is 3 seconds (i.e., 30 sampling points). The window slides forward in increments of 1 sampling point each time, ensuring that the temporal resolution of the deviation feature sequence is consistent with the original data stream.

[0038] Temporal dynamic anomaly features are extracted from two sources, both based on a 3-second time window identical to the KL divergence. The first type is the content activation event gaze responsiveness feature (NRR), which uses activation events that have completed response assessment within the current 3-second window as the statistical object (the response assessment of activation events is completed within T_resp milliseconds after the event timestamp t_e, therefore the statistical object is activation events whose t_e is earlier than the end time of the current window minus T_resp; the default value of T_resp is 500 milliseconds): For each statistical object, the system records the gaze landing point coordinates G_start and the center coordinates of the newly activated region P_target at time t_e, using the average gaze landing point coordinates within the response detection window as a reference. Subtracting G_start from the value yields the gaze displacement vector V_gaze, and subtracting G_start from P_target yields the content direction vector V_content. When the magnitude of V_gaze is less than the effective gaze displacement threshold (default is 25 pixels), it is marked as no response. If the magnitude meets the condition, and the direction cosine of V_gaze and V_content is lower than the response direction cosine threshold (default is 0.55), it is also marked as no response; otherwise, it is marked as a response. The NRR is calculated as the ratio of the number of no-response events in the current window to the total number of completed events. If there are no completed activation events in the current window, the NRR is taken from the NRR value of the window containing the most recent valid statistical events in the current session (keeping the most recent valid value), and an NRR validity marker is appended to the feature vector as a historical retention state, allowing step four to appropriately reduce the reference weight of this dimension.

[0039] The second type of temporal dynamic anomaly features are extracted from the two types of derived event sequences maintained in step one. The key area gaze proportion feature (FS) takes gaze persistence events whose center coordinates fall within the key area of ​​the content within the current 3-second window as the analysis object (the key area coordinates are taken from the content state snapshot corresponding to the start timestamp of each persistence event to ensure accurate time correspondence), and calculates the proportion of gaze events whose persistence duration exceeds the gaze judgment duration threshold (default value is 600 milliseconds) in the total duration. When cognitively engaged, a single gaze persistence is usually between 150 and 450 milliseconds. Long-term gazes exceeding 600 milliseconds often mean that the gaze has stopped progressing with the content rhythm, which is one of the temporal manifestations of pseudo-focus. The higher the FS value, the more severe the gaze. The scan alignment feature (SA) is extracted from scanning events whose starting coordinates fall within the key area of ​​text-based content in the current window. Before calculation, line break swipes are filtered out. When the horizontal coordinate component of a scanning event is negative (to the left) and the vertical coordinate component is positive (downward, in line break direction), it is identified as a normal reading swipe and excluded from SA statistics. After filtering, the direction cosine value (scanning horizontal coordinate component divided by scanning vector magnitude) is calculated for each of the remaining positive reading scans (the horizontal coordinate component is positive and the absolute value of the vertical coordinate component is less than half the current line height, which is obtained from the height of the line rectangle bounding box of the key content area). The average value is the SA. An SA close to 1 indicates that the scanning height advances along the text direction, while a low SA indicates that the direction is more random, which is one of the behavioral characteristics of pseudo-focused state. The SA feature of non-text-based content areas is marked as invalid and does not participate in subsequent calculations.

[0040] In interactive exercise scenarios, gaze-based features have certain blind spots: when a student's cognition has shifted, their gaze may remain mechanically fixed on the question stem area without generating effective touch responses. To address this, touch response latency features are introduced as a supplementary dimension: when the content status interface reports a new question, the system starts a timer to record the time elapsed from the question's appearance to the student's first effective touch in the option area. This time is then converted into a normalized deviation index through segmented mapping: the deviation index is 0 when the latency is within the first latency threshold (default value is 15 seconds), and increases linearly with increasing latency thereafter, reaching a maximum value of 1 after exceeding the second latency threshold (default value is 45 seconds). If a touch event is detected falling on a non-critical area of ​​the screen, this type of touch is marked as an abnormal touch feature and included in the deviation feature vector.

[0041] The output is divided into two paths based on the current content type, inputting the two sets of HMMs from step four respectively: In video-based and text-based content scenarios, the feature vector consists of four dimensions: KL divergence, NRR, FS, and SA, inputting the video / text HMM; In interactive exercise content scenarios, the feature vector consists of four dimensions: KL divergence, NRR, FS, and touch response latency deviation index. The SA feature is not applicable to exercise scenarios, inputting the interactive exercise HMM. The state 4 trigger indicator is output independently, without any HMM inference. All numerical features are subjected to min-max normalization. For the first learning session of a new user, the group normalization parameters obtained from the group prior data are used as the initial boundary (from the same source as the group prior HMM parameters in step four). From the second learning session onwards, the observation range of each dimension of the features in this session is merged with the historical normalization parameters using an exponential moving average method, gradually converging towards the student's own feature distribution range. This step outputs the feature vector (or state 4 trigger indicator) window by window, accumulating it into a temporal feature sequence, which serves as the core observation input for the temporal modeling of attention state in step four.

[0042] S4. Input the multidimensional deviation feature sequence into the personalized hidden Markov attention model, and obtain four types of attention time-series state sequences and state transition confidence through online parameter adaptation and Viterbi decoding. The system receives the output feature sequence and uses a Hidden Markov Model (HMM) to model and identify the temporal dynamics of attention states, outputting a temporally consistent sequence of attention states and state transition confidence scores. The HMM is chosen because students' attention states exhibit significant continuity over time; attention states tend to persist for a period before changing, and state transitions typically involve gradual transitions rather than abrupt changes. The hidden state transition matrix of the HMM can capture this temporally dependent structure, and its modeling framework, which separates the hidden states from the observed probability distribution, allows for decoupling the modeling of internal cognitive states and external observable behavioral features, aligning with the essential structure of the attention detection problem.

[0043] In one embodiment of the present invention, to address the differences in attention behavior between video / text / image scenarios and interactive exercise scenarios, two independent Hidden Markov Models (HMMs) are maintained: the observation vector of the video / text / image HMM is 4-dimensional (KL divergence, NRR, FS, SA), and the observation vector of the interactive exercise HMM is also 4-dimensional (KL divergence, NRR, FS, touch response latency deviation index). Both HMMs define the same four hidden states and the same state transition matrix constraints, maintain independent sets of observation distribution parameters and population prior initial parameters, and accept online Baum-Welch personalized updates. When the content type changes, the system activates the HMM for the corresponding content type, using the posterior probability distribution of the state at the last moment of the previous content type as the initial state prior of the newly active HMM, smoothing the state continuity when switching between content types. The parameters of the two HMMs are independent and updated separately. Four latent states are defined to correspond to the following attention levels. Deep focus (State 1) corresponds to a state where the student's cognition is fully engaged in the content. In this state, the KL divergence remains below the first deviation threshold (default value 0.30), the NRR is close to 0, the FS is within the normal range, and the SA (in text content scenarios) is close to 1. This is a state where the gaze location and temporal behavior rhythm are highly matched with the content rhythm. Shallow focus (State 2) corresponds to a state where attention remains on the content to a certain extent, but the deviation signal shows a slight increase. The judgment is mainly based on slight anomalies in temporal features. The NRR begins to rise (the gaze response to content activation events begins to slow down), and the FS begins to increase (the gaze dwell time begins to lengthen), usually before significant changes in the KL divergence. This is a typical manifestation of the pseudo-focus precursor stage and the core advantage of multi-dimensional feature vectors compared to simple spatial deviation detection. Attention drift (State 3) corresponds to a state where cognition has clearly deviated from the content. The KL divergence consistently exceeds the second deviation threshold (default value 0.60), the NRR remains high, the FS increases significantly, and the gaze trajectory drifts on the screen unrelated to the content rhythm. The absence of attention (state 4) corresponds to the state 4 trigger indication output in step 3, and the corresponding label is directly output in step 4 without HMM inference; the four states are arranged from high to low attention level, and the transition matrix design constrains the transition probability between adjacent states to be much higher than the transition probability across states, in order to conform to the gradual law of attention evolution.

[0044] Each hidden state corresponds to an observation probability distribution with respect to a multidimensional deviation feature vector, which is parameterized using a multivariate Gaussian distribution. The mean of the observation distribution corresponding to the deep focus state (state 1) is close to 0 in all dimensions of KL, NRR, and FS, and SA is close to 1, with a small covariance. The mean of the observation distribution of the shallow focus state (state 2) is slightly higher than that of state 1 in the NRR and FS dimensions, while the mean of the KL divergence dimension increases relatively little, reflecting the feature structure that temporal anomalies precede spatial anomalies in the early stage of pseudo-focus. The observation distribution of the attention drift state (state 3) has a high mean in all dimensions of KL, NRR, and FS, a low mean of SA, and a large covariance, reflecting the irregular changes of each feature in the distraction state. The two sets of HMMs each maintain independent observation distribution parameters, which are initialized by the population prior and then updated individually by the online Baum-Welch.

[0045] Because there are significant individual differences in the baseline attention characteristics of different students, using a general fixed parameter for the group can lead to systematic misjudgments for students with large individual differences. This method designs a two-stage personalized parameter construction strategy. The first stage is the initialization stage: when students use it for the first time, the group statistical parameters of learners of the same age (8 to 12 years old) are used as the prior initial values ​​for two sets of HMMs. The group prior parameters are obtained by collecting learning conversation data in a classroom learning environment with on-site teacher supervision. The teacher evaluates the students' attention status every 30 seconds as a weak supervision label, and records task completion indicators such as answer accuracy and reading completion as supplementary labeling basis. The combined data form a label sequence of attention status. The Baum-Welch algorithm is used to train the data of the two content scenarios offline on this batch of data to obtain two sets of group statistical prior parameters for video / text and interactive exercises.

[0046] The second stage is the personalized iterative update stage: starting from the student's second learning session, after each learning session, the system incrementally corrects the HMM parameters for the corresponding content type using an online Baum-Welch update mechanism, based on the feature observation sequence of that session. The two HMMs are updated independently according to the observation data of their respective scenarios. During each parameter update, the parameter update amount estimated for that session is fused with historical parameters using an exponential moving average. The initial EMA weight coefficient (in the second learning session) has a default value of 0.5, which decreases by 0.05 with each subsequent learning session until it reaches the minimum EMA weight coefficient (default value 0.10) and remains unchanged. For the parameter updates in the 10th and subsequent learning sessions, new session data accounts for 10%, and historical parameters account for 90%, allowing the personalized parameters to converge after several sessions, stably reflecting the student's individual attention behavior characteristics.

[0047] In real-time attention state decoding, step four, upon receiving an output from step three, first determines whether the output is a state 4 trigger indication: if so, it directly outputs the attention absence (state 4) label at the current moment without performing HMM inference; if it is a normal 4-dimensional feature vector, it activates the HMM corresponding to the current content type and uses the Viterbi algorithm to decode the state sequence of the observation sequence within the current analysis window using the maximum a posteriori probability. The Viterbi algorithm uses dynamic programming to find the state path that best matches the observation sequence among all possible state sequences. The algorithm runs incrementally on a sliding window, requiring only one recursive calculation update when a new sampling point arrives, ensuring real-time performance.

[0048] To address the uncertainty in detecting the transition from shallow focus (state 2) to attentional drift (state 3), a dual confirmation mechanism is added to the Viterbi decoding results: when the decoding results indicate a transition from state 2 to state 3, the system simultaneously calculates the ratio of the posterior probability of state 3 to the posterior probability of state 2 (which must exceed the third confirmation threshold, default value 3.0) and the continuous duration of state 3 (which must exceed the first duration threshold, default value 10 seconds). Only when both conditions are met is the state transition marked as a confirmed attentional drift event and passed to step five. This mechanism effectively distinguishes between frequent, brief attentional fluctuations (such as looking up to think or adjusting posture) that occur during normal student learning and genuine, continuous attentional drift, keeping the false alarm rate within an acceptable level.

[0049] After state 3 is confirmed, when step five intervenes and detects a valid confirmation touch from the student in the designated response area on the screen (this touch event is captured in real-time by the touch log from step one), step four performs the following recovery process: clears the confirmation marker for state 3, initializes the current active HMM state to shallow focus (state 2), restarts Viterbi incremental decoding from this starting point, and resets the duration timer of the double confirmation mechanism to zero. The recovery starting state is set to state 2 instead of state 1 because cognitive recovery after intervention is usually gradual; if the deviation feature does indeed decline after recovery, the HMM will naturally output state 1 within the following few seconds; if the student immediately deviates after clicking, the 10-second duration requirement of the double confirmation mechanism must be reset to zero, preventing immediate re-triggering of the intervention due to the click action. The focus state label for each sample point and the transition confidence value accompanying the state transition are output.

[0050] S5, based on the four types of attention time sequence and state transition confidence, adopts a graded content intervention strategy and learning event recording mechanism to output real-time intervention instructions and attention analysis reports during the learning process; The system receives the attention state sequence and state transition confidence scores, transforming the results into two types of outputs: real-time intervention instructions for the current learning process and a attention analysis report for post-learning review. The overall design principles of the intervention strategy are tiered response and content orientation: tiered response means adopting progressive interventions from mild to severe for different degrees of attention decline, avoiding over-intervention that could disrupt the normal learning experience; content orientation means that interventions should act as directly as possible on the learning content, redirecting learners' attention back to the content, rather than interrupting the learning process with external stimuli such as pop-ups or vibrations unrelated to the content.

[0051] For deep focus (State 1), the system does not perform any active intervention actions to maintain the stability of the learning environment. For shallow focus (State 2), the system enters silent monitoring mode, continuously accumulating the continuous duration count of shallow focus; if the continuous duration of shallow focus exceeds the first duration threshold (default value is 10 seconds), the system triggers the first level of content micro-intervention: a lightweight dynamic visual focusing effect is superimposed on the main key area of ​​the current content area (the current subtitle line in video explanation content, the current paragraph text area in text and image reading content), a brief halo pulse animation is generated at the edge of the subtitle line area, or a gentle brightness change occurs in the background color of the paragraph text line area, to awaken the learner's attention to the current content area through content-level visual guidance without creating a strong sense of interruption.

[0052] For confirmed attention drift (state 3), the system implements second-level and third-level interventions based on the current content type and the duration of the state. When the duration of attention drift after confirmation does not exceed the second duration threshold (default 30 seconds), the second-level intervention is triggered: video explanations are automatically paused briefly at the current playback position, while a semi-transparent information bar in the center of the screen prompts the student to click to continue; text-based content triggers a short content summary pop-up at the current reading position, displaying the current paragraph summary field provided by the content management interface (if this field is empty, the first line of the paragraph is used instead); interactive exercises trigger a prompt question related to the current exercise's knowledge point, guiding the student to refocus on the question. When attention drift is confirmed and its duration exceeds the second duration threshold, the third-level intervention is triggered: in addition to pausing the content, a short wake-up prompt is emitted via voice broadcast, lasting no more than 2 seconds. The system then waits for the student to actively touch a designated confirmation area on the screen to resume playback of the learning content. This touch operation also serves as the trigger signal for the state 3 recovery mechanism in step four.

[0053] For the attention deficit state (state 4), the system immediately implements mandatory intervention: pausing all content playback, issuing a voice prompt to wake up the student, and displaying a full-screen learning guidance page in the center of the screen, requiring the student to click on a designated area to confirm restarting. The full-screen guidance page also displays the progress information of the current learning task, helping the student re-establish contextual awareness of the current learning task. Regarding learning process recording, the system continuously maintains a time-series record of attention levels for each learning session in the background, based on the status label and timestamp of each sampling point. After the learning session ends, an attention analysis report is generated based on the complete time-series sequence of states. The report includes four core components: a session attention time-series curve, statistics on the time percentage of each state, location of content with weak attention, and a longitudinal comparison of historical sessions. These are presented to students and parents in chart form after the learning session, and can also be pushed to teachers for teaching effectiveness evaluation. Regarding the identification of content with weak attention spans, the report correlates the timestamps of attention drift and attention absence events with the corresponding content progress markers in the synchronized data stream of Step 1, accurately pinpointing the content location at the time of attention loss: video explanation content is accurate to the second of playback progress, text and image reading content is accurate to the paragraph number, and interactive exercise content is accurate to the exercise number, providing direct evidence for teachers to adjust their teaching design and for parents to guide students to focus on key review areas.

[0054] In one embodiment of the present invention, a scenario is illustrated by a fifth-grade (11-year-old) student using a learning tablet for independent math learning. The learning session includes approximately 18 minutes of video explanations on fraction multiplication (video-based content) followed by approximately 10 minutes of accompanying exercises (interactive practice content). The student learns independently at home without parental supervision, and the room has normal lighting. The following data illustrates the process based on the method in the above typical scenario, used to explain the data processing logic of each step. Specific values ​​may vary within a normal range due to individual differences, device variations, and learning content.

[0055] Table 1. Example of multimodal synchronous data stream sampling (video tutorial content, from minute 8 to minute 8.4) As shown in Table 1, the gaze points at 480.0 and 480.1 seconds were both located in the critical area of ​​the subtitle bar (y-coordinates within the range of 660 to 720), resulting in low KL divergence. From 480.2 seconds onwards, the y-coordinates of the gaze points significantly deviated from the subtitle bar area (423, 189, and 201 respectively), falling into the non-critical area at the top of the screen, leading to a significant increase in KL divergence and triggering anomaly markers for spatial deviation features. The content progress marker (at 480 seconds) was recorded at each sampling point. When the attention drift event was finally confirmed, the reporting module directly obtained the corresponding video playback progress from this field, pinpointing the attention loss to the narration segment at the 8-minute mark of the video, without requiring additional timestamp lookup.

[0056] Table 2, Examples of Attention State Recognition and Intervention Outputs (Status snapshot every 30 seconds from the 8th to the 10th minute) As shown in Table 2, at 08:30, the NRR feature (0.25) and KL divergence (0.31) both showed an abnormal increase. The increase in NRR indicates that the student's gaze response to the subtitle update event had begun to slow down during this period, which is a temporal manifestation of the loosening of cognitive input in the early stage of pseudo-focus. At 09:00, the KL divergence rose to 0.67 and the NRR reached 0.75. Both dimensions pointed to attentional drift, but due to the double confirmation mechanism, intervention was not triggered immediately. At 09:30, the duration of the attentional drift state exceeded the first duration threshold (default value 10 seconds), and the transfer confidence reached 0.89. The system confirmed the attentional drift and triggered the second level of intervention. After the student clicked to continue, step four cleared the confirmation mark of state 3, initialized the HMM state to shallow focus, and restarted decoding. At 10:00, the KL divergence dropped to 0.18 and the NRR fell back to 0.25. The focus state returned to shallow focus, and the intervention effect was significant. This example shows that the time from the occurrence of the deviation anomaly to the confirmation of intervention is about 30 seconds, which avoids false alarms caused by brief attentional fluctuations. At the same time, it demonstrates the rationality of the double confirmation mechanism in grasping the timing of intervention, even when the actual attentional drift continues.

[0057] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.

Claims

1. A method for monitoring learning focus on a learning tablet computer, characterized in that, Includes the following steps: S1 collects raw behavioral data from the front-facing camera and touch module of the learning tablet, and combines it with the content management interface to obtain the real-time temporal status of the learning content, resulting in a multimodal behavior and content synchronized data stream aligned with the timeline. S2, based on the temporal state of content in the synchronous data stream, constructs the gaze behavior expectation distribution according to three types of learning content scenarios, and uses the Gaussian mixture modeling method to obtain a dynamically updated content-aware gaze expectation model; S3, based on the content-aware expected gaze model, calculates gaze behavior deviation features from two dimensions: spatial distribution deviation and temporal dynamic anomaly, integrates touch response latency, and outputs a multi-dimensional deviation feature sequence. S4. Input the multidimensional deviation feature sequence into the personalized hidden Markov attention model, and obtain four types of attention time-series state sequences and state transition confidence through online parameter adaptation and Viterbi decoding. S5, based on the four types of attention time-series state sequences and state transition confidence, adopts a graded content intervention strategy and learning event recording mechanism to output real-time intervention instructions and attention analysis reports during the learning process.

2. The method for monitoring learning focus on a learning tablet computer according to claim 1, characterized in that, In S1, the content management interface continuously outputs content type tags, key area coordinate sets, content progress markers and paragraph summary fields, and generates activation event records when a key area is substantially updated, thus forming a content activation event stream. The front-facing camera uses a lightweight facial landmark detection model with MobileNet as the backbone network to output the coordinates of 68 semantic landmarks in the binocular regions. When a new user starts the camera for the first time, a 5-point gaze calibration process is performed to establish a personalized eye geometric mapping parameter matrix. The gaze dwell event sequence and saccade motion event sequence are maintained simultaneously and stored independently with timestamps as indexes. They are not included in the 10Hz resampling matrix. After the three data streams are uniformly resampled to 10Hz, if the confidence level of key point detection at several consecutive sampling points is lower than the first confidence threshold, a gaze data missing marker is written for the corresponding time period; if the gaze point exceeds the screen boundary, a gaze off-screen marker is written. Both types of quality markers are transmitted to step S3 along with the synchronous data stream.

3. The method for monitoring learning focus on a learning tablet computer according to claim 1, characterized in that, In S2, the expected attention GMM of video explanatory content includes subtitle components, whiteboard components, and social attention components. The weight of the subtitle component is determined by the ratio of the current number of subtitle characters to the total number of visible text characters. The weight of the whiteboard component is added to the ratio of the area increment to the area of ​​the subtitle bar after the whiteboard area expansion event, and decays linearly within the preset decay time. If there is no visible text in the current frame, the background browsing component will bear the remaining weight. If the lecturer's avatar area coordinate field is empty, the social attention component is not set, and the corresponding weight is merged into the subtitle component. The expected attention GMM for interactive exercise content dynamically adjusts the weight of each component as the answering time progresses and as events occur in the touch option area: when the question appears, the weight of each area is relatively uniform, the weight of the question stem area gradually increases as time progresses, and the weight of the corresponding option component surges when events occur in the touch option area. When the content type label changes, the active expectation gaze GMM is immediately replaced with the model corresponding to the new content type, and the mean parameters of each component are initialized according to the coordinates of the key region of the new content.

4. The method for monitoring learning focus on a learning tablet computer according to claim 1, characterized in that, In S2, the system maintains a reading frontier pointer, which is defined as the historical maximum value of the vertical coordinate of the gaze landing point within the current content item. It is only updated as reading progresses downwards. When the content identifier changes, the touch page jump crosses more than one page, or a new learning session is started, the pointer is reset to the vertical coordinate of the topmost text line in the current content display area. The main line is the paragraph where the reading front pointer is located. The image-text reading expected gaze GMM sets a high-weight main line component at the position of the main line, sets an auxiliary component above the main line, and sets a pre-read component below the main line. The mixed weight of the auxiliary component is one-third of the mixed weight of the main line component, and the mixed weight of the pre-read component is one-fifth of the mixed weight of the main line component. The mean x-axis of the main line component is determined in real time by the median x-axis of the gaze landing point in the main line area within the current analysis window. When the sample is less than the preset minimum sample threshold, it degenerates into a uniform calculation from the left end of the line according to the preset average reading speed. If there is an image area on the screen, an additional weight is added to the image component, which is calculated as the ratio of the image area to the total area of ​​the content area.

5. The method for monitoring learning focus on a learning tablet computer according to claim 1, characterized in that, In S3, at the beginning of each analysis window, the gaze off-screen marker and gaze data missing marker in the synchronization time matrix are checked first. If they exist, the state 4 trigger indication is directly output to step S4, and the feature calculation process is not entered. Spatial distribution deviation is measured by KL divergence, which calculates the KL divergence from the actual gaze distribution to the direction of the content-aware expected gaze model. The actual gaze distribution is constructed using a kernel density estimation method, which uses a Gaussian kernel function with fixed bandwidth to transform the gaze coordinate samples in the current analysis window into a two-dimensional continuous probability density function. The kernel bandwidth is adaptively set according to the key region size of the current content type. Feature calculation uses a fixed-length sliding time window, moving forward by one sampling point at a time to ensure that the deviation feature sequence is consistent with the temporal resolution of the original data stream.

6. The method for monitoring learning focus on a learning tablet computer according to claim 1, characterized in that, In S3, the content activation event gaze responsiveness feature NRR takes the activation events that have completed response judgment within the current analysis window as the statistical object, and the response judgment of the activation event is completed within a preset response judgment time after the event timestamp; For each statistical object, the gaze point coordinates at the event timestamp are used as the response start point. The gaze displacement vector is obtained by subtracting the response start point from the average gaze point within the response detection window. The content direction vector is obtained by subtracting the response start point from the center coordinates of the newly activated region. When the magnitude of the gaze displacement vector is less than the preset effective gaze displacement threshold, it is marked as no response; when the magnitude meets the condition, if the direction cosine value of the gaze displacement vector and the content direction vector is lower than the preset response direction cosine threshold, it is also marked as no response; the NRR value is obtained by the ratio of the number of no response events in the current window to the total number of events that have been judged. If there are no completed activation events in the current analysis window, the NRR is taken from the NRR value of the window containing the most recent valid statistical events in the current session, and an NRR validity marker is added to the feature vector as a historical retention state.

7. The method for monitoring learning focus on a learning tablet computer according to claim 1, characterized in that, In S3, the key area gaze proportion feature FS takes the gaze dwelling events whose center coordinates fall in the key area of ​​the content within the current analysis window as the statistical object, and calculates the proportion of the total duration of gaze events whose dwelling time exceeds the preset gaze judgment duration threshold. The scanning direction alignment feature (SA) is extracted from the scanning motion events of key areas of text-based content. After filtering out line break back scan events with negative horizontal coordinate components and positive vertical coordinate components, the mean cosine of the direction is calculated for the remaining forward reading scans. The SA features of non-text-based content areas are marked as invalid. In interactive exercise scenarios, the touch response latency deviation index is calculated by segmented mapping from the time the question appears to the time the student makes the first valid click in the option area: the deviation index is 0 within the first latency threshold, and reaches its maximum value after exceeding the second latency threshold, increasing linearly within the interval; touch events falling in non-critical areas are marked with abnormal touch features and included in the deviation feature vector.

8. The method for monitoring learning focus on a learning tablet computer according to claim 1, characterized in that, In S4, two independent HMMs are maintained for video / text / image scenarios and interactive exercise scenarios respectively: the observation vector of the video / text / image HMM is four-dimensional: KL divergence, NRR, FS, and SA; the observation vector of the interactive exercise HMM is four-dimensional: KL divergence, NRR, FS, and touch response latency deviation exponent. When switching content types, the posterior probability distribution of the state at the last moment of the previous type is used as the initial state prior of the new active HMM. HMM parameters are constructed using a two-stage personalized approach: the initial stage initializes the parameters with prior parameters from a group of learners of the same age, and the group prior is obtained through offline training using classroom learning conversation data. After each learning session, the corresponding type of HMM parameters are incrementally corrected using an online Baum-Welch method. The update amount and historical parameters are merged using an exponential moving average method. The EMA weight coefficient is reduced by a preset step size in each session until it reaches the minimum EMA weight coefficient and then remains unchanged. When the Viterbi decoding result shows a shift from shallow focus to attentional drift, confirmation is only possible if the ratio of the posterior probabilities of state 3 to state 2 exceeds the third confirmation threshold and the duration of the shift exceeds the first duration threshold. After state 3 confirmation, when a valid confirmation touch is detected in the specified response area, the confirmation mark of state 3 is cleared, the current active HMM state is initialized to shallow focus, and Viterbi incremental decoding is restarted.

9. The method for monitoring learning focus on a learning tablet computer according to claim 1, characterized in that, In S5, for shallow focus, the system enters silent monitoring mode; when the continuous duration of shallow focus exceeds the first duration threshold, the first level of intervention is triggered, and a dynamic visual focusing effect is superimposed on the key area of ​​the current content to guide the learner to pay attention to the current content area. For confirmed attention drift states, if the duration does not exceed the second duration threshold, a second-level intervention is triggered: depending on the content type, a video pause prompt, a paragraph summary pop-up, or a prompt about the nature of the problem is executed; if the duration exceeds the second duration threshold, a third-level intervention is triggered: based on the content pause, a voice broadcast is used to wake up the student, waiting for the student to make an active touch operation in the designated confirmation area to trigger the state 3 recovery mechanism in step S4; After the learning session ends, a focus analysis report is generated, which includes a focus time series curve, statistics on the time ratio of each state, location of content where attention was lost, and longitudinal comparison of historical sessions. The content location tracking for attention loss maps the timestamps of attention drift and attention absence events to the content progress markers, accurately marking the content location at the time of attention loss. For video explanations, the progress is marked down to the second; for text and image readings, it is marked down to the paragraph number; and for interactive exercises, it is marked down to the exercise number.

10. A learning attention monitoring system for a learning tablet computer, used to perform the steps of the learning attention monitoring method for a learning tablet computer as described in any one of claims 1-9, characterized in that, include: The data acquisition and synchronization module is used to collect raw behavioral data from the front-facing camera and touch module of the learning tablet, and combine it with the content management interface to obtain the real-time temporal status of the learning content, so as to obtain a multimodal behavioral and content synchronization data stream aligned with the time axis. The expected gaze distribution modeling module is used to construct the expected gaze distribution based on the temporal state of content in the synchronous data stream and three types of learning content scenarios. It adopts the Gaussian mixture modeling method to obtain a dynamically updated content-aware expected gaze model. The multidimensional deviation feature calculation module is used to calculate the gaze behavior deviation features from two dimensions, spatial distribution deviation and temporal dynamic anomaly, based on the content-aware expected gaze model, and integrates touch response latency to output a multidimensional deviation feature sequence. The personalized state recognition module is used to input the multidimensional deviation feature sequence into the personalized hidden Markov attention model. After online parameter adaptation and Viterbi decoding, four types of attention time-series state sequences and state transition confidence are obtained. The intervention and reporting output module is used to output real-time intervention instructions and learning process attention analysis reports based on four types of attention temporal state sequences and state transition confidence, using a graded content intervention strategy and learning event recording mechanism.