A learning state monitoring method and device based on multi-modal perception
By using multimodal perception technology, learners' multimodal data can be acquired in real time, a feature monitoring model can be constructed, and feature temporal judgment can be performed. This solves the problems of low compliance and misjudgment in existing technologies, and realizes seamless learning status monitoring and intelligent reminders.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 北京爱宾果科技有限公司
- Filing Date
- 2026-06-29
- Publication Date
- 2026-07-24
AI Technical Summary
Existing learning status monitoring methods suffer from low compliance and discomfort when worn, and pure software solutions have a single perception dimension, making them susceptible to momentary interference and leading to misjudgments.
Employing multimodal perception technology, it acquires multimodal monitoring data in real time, constructs a monitoring model for sitting posture, gaze, and movement features, calculates non-focus by determining the time sequence of features, and performs seamless monitoring and intelligent reminders locally on the terminal.
It achieves seamless end-to-end management, eliminates misjudgments due to momentary actions, provides intelligent recognition and reminders, and improves the accuracy of learning status monitoring and user experience.
Smart Images

Figure CN122451401A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of learning state monitoring technology, and specifically relates to a learning state monitoring method and device based on multimodal perception. Background Technology
[0002] Learning status monitoring allows for the analysis and evaluation of the monitored individual's learning process. This not only effectively prevents physical injuries such as myopia and scoliosis, but also helps them establish healthy self-directed learning habits, optimize their learning pace, and thus improve learning efficiency.
[0003] Currently, mainstream learning status monitoring methods fall into two categories: wearable solutions and pure software solutions. While wearable solutions can acquire relatively accurate data such as head posture, their mandatory wearing nature affects the learning experience, leading to low compliance, discomfort, and other issues, making them difficult to deploy continuously in daily learning scenarios. Pure software solutions rely on the device's built-in camera to collect user behavior data. Although they offer the advantage of convenient deployment, these solutions generally have a single perception dimension, feature blind spots, lack temporal evolution analysis, are susceptible to transient interference, and are prone to misjudgments.
[0004] Therefore, there is an urgent need for a technical solution to overcome or mitigate at least one of the aforementioned defects in the existing technology. Summary of the Invention
[0005] The purpose of this application is to provide a learning state monitoring method and apparatus based on multimodal perception to solve at least one problem existing in the prior art.
[0006] The technical solution of this application is:
[0007] The first aspect of this application provides a learning state monitoring method based on multimodal perception, comprising:
[0008] Step S1: Acquire multimodal monitoring data of the monitored subject in real time during the learning state monitoring period and extract feature information;
[0009] Step S2: Construct a feature monitoring model and calculate feature indicators based on the feature information;
[0010] Step S3: Determine whether the feature index meets the feature temporal determination condition;
[0011] Step S4: Calculate the non-focus degree based on the feature index that meets the feature time-series determination condition. When the non-focus degree is greater than the preset non-focus degree threshold, generate a reminder message and push it to the monitored person.
[0012] In at least one embodiment of this application, in step S1, the feature information includes sitting posture feature information, gaze feature information, and motion feature information.
[0013] In at least one embodiment of this application, step S2, which involves constructing a feature monitoring model and calculating feature indicators based on the feature information, includes:
[0014] Construct a sitting posture feature monitoring model and calculate sitting posture feature indicators based on the sitting posture feature information;
[0015] Construct a gaze feature monitoring model and calculate gaze feature indicators based on the gaze feature information;
[0016] A motion feature monitoring model is constructed, and motion feature indicators are calculated based on the motion feature information.
[0017] In at least one embodiment of this application, step S3, determining whether the feature index satisfies the feature temporalization determination condition, includes:
[0018] Calculate the time-series index of features:
[0019] ;
[0020] ;
[0021] in, Here, W is the time-series characteristic index, and W is the length of the time-series sliding window. Here, F is the timing indicator function, and F is the characteristic index.
[0022] like Then the feature index satisfies the feature temporal determination condition; if If the feature index does not meet the feature time-series determination condition, The threshold for the duration of the feature.
[0023] In at least one embodiment of this application, step S4, calculating the degree of non-focus based on the feature index that satisfies the feature temporal determination condition, includes:
[0024] The feature indicators that meet the feature time-series determination conditions are normalized.
[0025] Calculate the degree of non-focus based on the normalized feature indexes:
[0026] ;
[0027] in, For the degree of non-focus at time t, Let t be the characteristic index of sitting posture. Let t be the line-of-sight characteristic index. Let w1, w2, and w3 be the action characteristic indicators at time t, and w1, w2, and w3 be the weights of different characteristic indicators, respectively.
[0028] In at least one embodiment of this application, in step S4, when the degree of non-focus is greater than a preset degree of non-focus threshold, the non-focus learning state is divided into three levels of non-focus: slight distraction, moderate distraction, and severe distraction, according to the degree of non-focus range.
[0029] In at least one embodiment of this application, in step S4, the reminder information is pushed to the monitored person through a reminder box, and the visual presentation effect of the reminder box is adapted and adjusted according to the daytime mode or the nighttime mode.
[0030] A second aspect of this application provides a learning state monitoring device based on multimodal perception, which, based on the multimodal perception-based learning state monitoring method described above, includes:
[0031] The feature extraction module is used to acquire multimodal monitoring data of the monitored subjects in real time during the learning state monitoring period and extract feature information;
[0032] The indicator calculation module is used to construct a feature monitoring model and calculate feature indicators based on the feature information.
[0033] A time-series determination module is used to determine whether the feature index meets the feature time-series determination conditions;
[0034] The non-focus calculation module is used to calculate the non-focus based on the feature indicators that meet the feature time-series determination conditions. When the non-focus is greater than a preset non-focus threshold, a reminder message is generated and pushed to the monitored person.
[0035] In at least one embodiment of this application, the feature extraction module, the index calculation module, the temporal determination module, and the non-focus calculation module all run locally on the terminal.
[0036] The invention has at least the following beneficial technical effects:
[0037] The learning state monitoring method based on multimodal perception in this application performs feature temporal determination for each feature indicator, eliminates misjudgments of instantaneous actions, and calculates non-focus by weighting multiple feature indicators, forming a complete end-to-end management system of non-inattentive monitoring, intelligent recognition, and reminder intervention. Attached Figure Description
[0038] Figure 1 This is a flowchart of a learning state monitoring method based on multimodal perception according to one embodiment of this application;
[0039] Figure 2 This is a schematic diagram of a learning state monitoring device based on multimodal perception according to one embodiment of this application.
[0040] in:
[0041] 100 - Feature extraction module; 200 - Index calculation module; 300 - Temporal determination module; 400 - Non-focus calculation module. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be described in more detail below with reference to the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The described embodiments are some, but not all, embodiments of this application. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. The embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0043] The following is in conjunction with the appendix Figures 1 to 2 This application will be described in further detail.
[0044] The first aspect of this application provides a learning state monitoring method based on multimodal perception, such as... Figure 1 As shown, it includes the following steps:
[0045] Step S1: Acquire multimodal monitoring data of the monitored subject in real time during the learning state monitoring period and extract feature information;
[0046] Step S2: Construct a feature monitoring model and calculate feature indicators based on feature information;
[0047] Step S3: Determine whether the feature indicators meet the feature time-series determination conditions;
[0048] Step S4: Calculate the non-focus level based on the feature indicators that meet the feature time-series judgment conditions. When the non-focus level is greater than the preset non-focus level threshold, generate a reminder message and push it to the monitored person.
[0049] The learning state monitoring method based on multimodal perception in this application firstly involves, in step S1, a customized total monitoring duration, and real-time acquisition of multimodal monitoring data of the monitored subject during the learning state monitoring period. Preferably, wide-angle high-definition visual sensors and eye-tracking sensors are used to collect multimodal monitoring data, providing full coverage of the learning state monitoring area from multiple angles, achieving seamless acquisition of multimodal monitoring data without the need for mandatory device wearing or interference with learning. Clean, aligned, and standardized multimodal monitoring data is obtained through time synchronization and noise reduction processing. Posture features, gaze features, and motion features are extracted from the multimodal monitoring data.
[0050] Secondly, in step S2, a feature monitoring model is constructed, and feature indicators are calculated based on feature information, including the following processes:
[0051] Construct a sitting posture feature monitoring model and calculate sitting posture feature indicators based on sitting posture feature information;
[0052] Construct a gaze feature monitoring model and calculate gaze feature indicators based on gaze feature information;
[0053] Construct a motion feature monitoring model and calculate motion feature indicators based on motion feature information.
[0054] In this embodiment, a sitting posture feature monitoring model is constructed based on the three-dimensional spatial coordinates of the spine, neck, and shoulders. This model can identify abnormal sitting postures such as looking down, bending to the side, and slouching. The calculation formula for the sitting posture feature monitoring model is:
[0055] ;
[0056] in, Let t be the characteristic index of sitting posture. Let t be the three-dimensional coordinates of the spine. For the three-dimensional coordinates of the spine, Let t be the three-dimensional coordinates of the neck. For the neck reference three-dimensional coordinates, Let t be the three-dimensional coordinates of the left shoulder. The left shoulder is the reference three-dimensional coordinate. Let t be the three-dimensional coordinates of the right shoulder. The right shoulder reference three-dimensional coordinates, , , These represent sensitivity coefficients for different sitting postures.
[0057] The line-of-sight feature monitoring model consists of two parts: line-of-sight spatial accuracy and line-of-sight stability. Line-of-sight spatial accuracy measures whether the line of sight accurately falls on the target, while line-of-sight stability determines whether the line of sight is stable, rather than jumping or drifting. The specific formula is:
[0058] ;
[0059] in, Let t be the line-of-sight characteristic index. Let be the distance between the point where the line of sight falls at time t and the reference point where the line of sight falls. For space tolerance parameters, represents the steepness of the Sigmoid function. The standard deviation of the line of sight. The threshold value is the value of the Sigmoid function.
[0060] The motion feature detection model consists of two parts: mouth features and hand features, specifically:
[0061] ;
[0062] in, The action characteristic index at time t, Let t be the frequency of mouth opening and closing. Let t be the velocity of the hand key points, and β1 and β2 be the weights of different features.
[0063] Further, in step S3, it is determined whether the feature indicators calculated above meet the feature temporalization determination conditions. Specifically:
[0064] Calculate the time-series index of features:
[0065] ;
[0066] ;
[0067] in, Here, W is the time-series characteristic index, and W is the length of the time-series sliding window. Here, F is the timing indicator function, and F is the characteristic index.
[0068] like If the feature index satisfies the feature time series determination condition; if If the feature index does not meet the criteria for feature time-series determination, then the feature index does not meet the criteria for feature time-series determination. The threshold for the duration of the feature.
[0069] Finally, in step S4, the degree of non-focus is calculated based on the feature indicators that meet the feature temporalization judgment conditions, including:
[0070] The feature indicators that meet the criteria for feature temporalization are normalized.
[0071] Non-focus level is calculated based on the normalized feature indicators:
[0072] ;
[0073] in, For the degree of non-focus at time t, Let t be the characteristic index of sitting posture. Let t be the line-of-sight characteristic index. Let w1, w2, and w3 be the action characteristic indicators at time t, and w1, w2, and w3 be the weights of different characteristic indicators, respectively.
[0074] Using the above method, the degree of inattention is calculated in real time. The lower the degree of inattention, the higher the degree of focus in the learning state. When the degree of inattention is greater than the preset degree of inattention threshold, the inattention learning state is divided into three levels of inattention: slight distraction, moderate distraction and severe distraction. Corresponding reminder information is generated and pushed to the monitored person according to the degree of inattention.
[0075] This application's learning state monitoring method based on multimodal perception constructs a reminder information database. When a non-focused learning state is identified, reminder information is retrieved and pushed to the monitored individual via a reminder box. The visual presentation of the reminder box is adapted to either daytime or nighttime mode. In nighttime mode, the brightness of the reminder text is automatically reduced by 20%, and a warm-toned eye-protecting color temperature filter is applied to eliminate visual afterimages caused by strong light, ensuring eye health throughout the learning process. It is understood that all reminder information is pushed using a physical-grade projection eye-protection design. Through a semi-transparent, gradient soft-light reminder box, it avoids the core writing area, achieving non-invasive intervention without strong light stimulation or harsh noise, preventing visual fatigue caused by sudden pupil constriction, thus achieving the dual goals of reminders and eye protection.
[0076] Understandably, in the preferred embodiment of this application, after the learning status monitoring is completed, a learning status monitoring report is generated based on the monitoring results to analyze and evaluate the lack of focus during the monitoring period. Finally, the learning status monitoring report is pushed to the parent's device via end-to-end encrypted synchronization, achieving encrypted synchronization between the learning status monitoring report and the parent's mini-program, thus forming a closed-loop supervision system.
[0077] This application presents a learning state monitoring method based on multimodal perception. It collects multimodal monitoring data in real time, employs edge-side real-time inference, and uses feature-based temporal judgment to distinguish between short-term normal behavior and long-term abnormal behavior, eliminating misjudgments based on instantaneous actions. This achieves deep coupling between learning state monitoring and eye protection reminders, addressing industry pain points such as the inability of parents to monitor in real time, the lack of quantifiable learning states, and eye-harming intervention methods.
[0078] The second aspect of this application provides a learning state monitoring device based on multimodal perception, such as... Figure 2 As shown, it includes:
[0079] The feature extraction module 100 is used to acquire multimodal monitoring data of the monitored subject in real time during the learning state monitoring period and extract feature information.
[0080] The indicator calculation module 200 is used to build a feature monitoring model and calculate feature indicators based on feature information.
[0081] The time-series determination module 300 is used to determine whether the feature indicators meet the feature time-series determination conditions.
[0082] The non-focus calculation module 400 is used to calculate the non-focus based on the feature indicators that meet the feature time-series judgment conditions. When the non-focus is greater than the preset non-focus threshold, a reminder message is generated and pushed to the monitored person.
[0083] The learning state monitoring device based on multimodal perception in this application is installed on an eye-protecting smart learning terminal. The feature extraction module 100, the index calculation module 200, the temporal determination module 300, and the non-focus calculation module 400 all run locally on the terminal. The learning state monitoring algorithm runs locally, ensuring privacy and security and millisecond-level response, and it is available offline.
[0084] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A learning state monitoring method based on multimodal perception, characterized in that, include: Step S1: Acquire multimodal monitoring data of the monitored subject in real time during the learning state monitoring period and extract feature information; Step S2: Construct a feature monitoring model and calculate feature indicators based on the feature information; Step S3: Determine whether the feature index meets the feature temporal determination condition; Step S4: Calculate the non-focus degree based on the feature index that meets the feature time-series determination condition. When the non-focus degree is greater than the preset non-focus degree threshold, generate a reminder message and push it to the monitored person.
2. The learning state monitoring method based on multimodal perception according to claim 1, characterized in that, In step S1, the feature information includes sitting posture feature information, gaze feature information, and action feature information.
3. The learning state monitoring method based on multimodal perception according to claim 2, characterized in that, In step S2, a feature monitoring model is constructed, and feature indicators are calculated based on the feature information, including: Construct a sitting posture feature monitoring model and calculate sitting posture feature indicators based on the sitting posture feature information; Construct a gaze feature monitoring model and calculate gaze feature indicators based on the gaze feature information; A motion feature monitoring model is constructed, and motion feature indicators are calculated based on the motion feature information.
4. The learning state monitoring method based on multimodal perception according to claim 3, characterized in that, In step S3, determining whether the feature index meets the feature temporal determination condition includes: Calculate the time-series index of features: ; ; in, Here, W is the time-series characteristic index, and W is the length of the time-series sliding window. Here, F is the timing indicator function, and F is the characteristic index. like Then the feature index satisfies the feature temporal determination condition; if If the feature index does not meet the feature time-series determination condition, The threshold for the duration of the feature.
5. The learning state monitoring method based on multimodal perception according to claim 4, characterized in that, In step S4, the degree of non-focus is calculated based on the feature index that satisfies the feature temporal determination condition, including: The feature indicators that meet the feature time-series determination conditions are normalized. Calculate the degree of non-focus based on the normalized feature indexes: ; in, For the degree of non-focus at time t, Let t be the characteristic index of sitting posture. Let t be the line-of-sight characteristic index. Let w1, w2, and w3 be the action characteristic indicators at time t, and w1, w2, and w3 be the weights of different characteristic indicators, respectively.
6. The learning state monitoring method based on multimodal perception according to claim 5, characterized in that, In step S4, when the degree of non-focus is greater than the preset non-focus threshold, the non-focus learning state is divided into three levels of non-focus: slight distraction, moderate distraction, and severe distraction, according to the range of non-focus.
7. The learning state monitoring method based on multimodal perception according to claim 6, characterized in that, In step S4, the reminder information is pushed to the monitored person through a reminder box, and the visual presentation of the reminder box is adapted and adjusted according to the daytime mode or nighttime mode.
8. A learning state monitoring device based on multimodal perception, based on the learning state monitoring method based on multimodal perception according to any one of claims 1 to 7, characterized in that, include: The feature extraction module is used to acquire multimodal monitoring data of the monitored subjects in real time during the learning state monitoring period and extract feature information; The indicator calculation module is used to construct a feature monitoring model and calculate feature indicators based on the feature information. A time-series determination module is used to determine whether the feature index meets the feature time-series determination conditions; The non-focus calculation module is used to calculate the non-focus based on the feature indicators that meet the feature time-series determination conditions. When the non-focus is greater than a preset non-focus threshold, a reminder message is generated and pushed to the monitored person.
9. The learning state monitoring device based on multimodal perception according to claim 8, characterized in that, The feature extraction module, the index calculation module, the temporal determination module, and the non-focus calculation module all run locally on the terminal.