AI multi-mode fusion-based learning accompanying robot system and learning behavior intervention method thereof
By integrating the functions of desk lamp lighting, monitoring, and intervention through AI multimodal fusion technology, the problem of existing devices being unable to identify learning status and occupying space is solved, realizing efficient learning status monitoring and personalized intervention, and improving learning efficiency and quality.
Patent Information
- Application Number
- CN202511615117.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-01-30
Smart Images

Figure CN121436031A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of artificial intelligence and educational technology, specifically involving a learning companion lamp robot system that integrates AI multimodal fusion technology and robotic arm physical interaction function, as well as its learning behavior intervention method, which is applicable to the monitoring of learning status, behavior analysis and proactive intervention in the home learning scenario of K12 students. Background Technology
[0002] With the deepening development of educational informatization, home learning has become an important part of K-12 students' learning. However, in the absence of real-time supervision from teachers and parents, students generally have problems such as lack of concentration, poor posture, and pretending to study during home learning, which seriously affect learning efficiency and physical health.
[0003] The following problems exist in the existing technology: Limitations of traditional smart desk lamps: Smart desk lamps on the market mainly focus on the intelligence of lighting functions, such as automatic adjustment of brightness and color temperature, but lack the ability to actively perceive the learner's learning status, cannot identify whether the learner is really studying, and cannot actively intervene for different learning statuses. The shortcomings of standalone learning robots: Existing educational robots typically exist as standalone devices, which are large in size, expensive, and require additional desktop space. These robots mainly focus on content guidance and emotional support, with weak functions in monitoring behavior and correcting posture during the learning process, and lack deep integration with the learning environment; Limitations of posture correctors: Most posture correction products on the market are passive wearable devices or mechanical supports, which students generally resist, have poor wearing comfort, and can only monitor posture data in a single dimension, making it impossible to comprehensively judge the learning status. Pretending to study is difficult to identify: This is a major pain point for parents, but one that current technology struggles to address. Students may appear to be sitting at their desks, but their attention is completely elsewhere. Traditional surveillance cameras infringe on privacy and lack intelligent analysis capabilities, leaving parents unable to understand their child's true learning status. Lack of subject-specific support: The learning characteristics of different subjects vary significantly. Mathematics requires deep thinking, Chinese requires extensive reading, and English requires reading aloud practice, but existing technologies lack differentiated monitoring and intervention strategies tailored to the characteristics of different subjects.
[0004] Therefore, there is an urgent need for an intelligent system that can deeply integrate lighting, monitoring, analysis, and intervention functions, without increasing additional space occupation and learning burden, while accurately identifying learning status and implementing effective intervention, thereby truly improving the efficiency and quality of family learning. Summary of the Invention
[0005] Purpose of the invention The purpose of this invention is to overcome the shortcomings of the prior art and provide a learning companion robot system based on AI multimodal fusion and its learning behavior intervention method. By deeply integrating the desk lamp lighting function with multimodal perception, intelligent analysis, active intervention and intelligent tutoring functions, it can realize comprehensive monitoring and personalized intervention of learners' learning status, especially solving key problems such as "pretending to learn", fatigue prediction, posture correction and learning tutoring, and improving the efficiency of family learning.
[0006] Technical solution To achieve the above objectives, the present invention adopts the following technical solution:
[0007] A learning behavior intervention method includes the following steps: S1. Multidimensional learning state perception steps: Collect learner's posture image data through the visual perception module, including head position, body posture, hand movements, etc.; collect sound data of the learning environment through the audio perception module, including writing sounds, page turning sounds, sighing sounds, reading sounds, etc. S2. Intelligent Learning State Judgment Steps: The collected posture image data and sound data are input into a multimodal fusion model. This model, based on deep learning technology, can comprehensively analyze multi-dimensional data and identify the learner's current learning state. Learning states include: focused state (continuous reading or writing), fatigue state (frequent eye rubbing, yawning), distracted state (wandering eyes, prolonged stillness), poor posture state (forward tilt of the neck, crooked body), feigned learning state (book is still, no effective learning behavior), and normal rest state. S3. Steps for Generating Tiered Intervention Strategies: Based on the identified learning status, combined with the subject knowledge graph and the learner's historical behavioral data, personalized intervention strategies are generated. Intervention strategies include intervention levels (mild, moderate, severe), intervention methods (voice prompts, lighting adjustments, robotic arm interaction), and intervention content (specific prompts or action instructions). The knowledge graph is used to understand the learner's current learning content and determine whether pauses are due to normal thinking or difficulty encountered. S4. Multimodal Collaborative Execution Steps: Based on the generated intervention strategy, the intervention operation is executed collaboratively through the voice prompt module, lighting adjustment module, and robotic arm interaction module. Mild intervention only adjusts the lighting, moderate intervention adds voice prompts, and severe intervention activates the robotic arm for physical interaction, forming a progressive intervention mechanism. S5. Intelligent Learning Tutoring Steps: When a learner actively initiates a learning request, the system receives the learner's question through the speech recognition module, acquires images of learning materials through the visual perception module, inputs the question and image into the visual-language multimodal model for understanding and reasoning, generates tutoring content using a heuristic hierarchical tutoring method, and explains it in collaboration with the speech synthesis module and the robotic arm module. The heuristic hierarchical tutoring method includes four levels: knowledge point prompts, problem-solving framework, step-by-step explanation, and complete solution.
[0008] A learning companion robot system, comprising: The lamp body module includes an LED lighting unit. This unit uses a full-spectrum LED array with a color temperature dynamically adjustable from 2700K (warm light) to 6500K (cool light) and brightness adjustable from 100 to 2000 lumens, meeting the lighting needs of different study scenarios. The lamp head is adjustable at multiple angles via a robotic arm module.
[0009] Robotic arm module: It has 4 degrees of freedom, including a base rotation joint, a shoulder pitch joint, an elbow pitch joint, and a wrist rotation joint. The working radius is 60cm to 100cm, the load capacity is not less than 500g, and the positioning accuracy is ±5mm.
[0010] Multimodal sensing module: - First sensing submodule: Located at the lamp head of the desk lamp. This module is used to acquire image information and / or depth information within the illumination area below the lamp head, in order to achieve tracking illumination or content recognition of targets (such as books, hands); - Second sensing submodule: Located at the base connection position of the desk lamp. This module is used to acquire image information of the user in front of the desk lamp to realize user posture detection; - Sound source localization unit: used to collect ambient sound and determine the direction of the sound source. This unit can be integrated into the lamp head position or the non-lamp head fixed position.
[0011] The computing control module employs a two-tier architecture combining edge computing and cloud-based decision-making. The edge computing unit, based on an embedded AI chip, boasts a computing power of at least 4 TOPS, handling tasks such as real-time data processing, posture detection, and voice recognition, with a response latency of less than 100ms. The cloud-based decision-making unit deploys complex deep learning models for tasks such as fatigue prediction, knowledge graph matching, and learning report generation.
[0012] Data Management Module: This module stores learners' learning profile data, including daily study time, attention curves, abnormal posture records, and knowledge point mastery status. A strict privacy protection mechanism is implemented, saving only skeletal contour coordinates and statistical data; raw video images are not stored. All data is stored locally using AES-256 encryption.
[0013] The intelligent tutoring module includes a speech recognition unit, a visual-language multimodal model unit, a knowledge graph verification unit, and a speech synthesis unit. It is used to respond to learners' active requests for help and provide real-time learning tutoring services. The intelligent tutoring module adopts a heuristic hierarchical tutoring strategy, providing four levels of knowledge point prompts, problem-solving frameworks, step-by-step explanations, and complete solutions according to the learner's level of understanding.
[0014] Beneficial effects Compared with the prior art, the present invention has the following beneficial effects: Deep integration of functions: It integrates desk lamp lighting, learning monitoring, intelligent analysis and proactive intervention functions into one, without taking up additional desktop space, reducing the threshold and cost of use; Pretending to learn identification: This is the first time that a pretending to learn identification algorithm based on multimodal fusion has been proposed. It integrates multi-dimensional features such as the book being still, no writing action, and shifty eyes, and confirms the identification through a verification questioning mechanism. The accuracy rate reaches 92.7%, effectively solving the pain points of parents. Fatigue prediction in advance: Using a bidirectional LSTM time series prediction model, fatigue status can be predicted 3-5 minutes in advance with an accuracy of 89%, providing a time window for timely intervention and avoiding excessive fatigue from affecting learning efficiency and physical health. Subject-specific support: Differentiated monitoring thresholds and intervention strategies are set according to the learning characteristics of different subjects such as mathematics, Chinese, and English to avoid misjudgment and over-intervention; A progressive intervention mechanism: Based on the severity of the problem, three levels of intervention are implemented: mild, moderate, and severe. From light adjustment to voice prompts to physical interaction with robotic arms, the intervention is effective while avoiding excessive interference. Privacy protection: Only skeletal contour coordinate data is saved, and the original video images are not stored. Local encrypted storage is used to fully protect learners' privacy.
[0015] Experimental data shows that students using the system of this invention have seen an average increase in focus time from 23 minutes to 41 minutes, an increase of 78%; the rate of abnormal sitting posture has decreased from 4.2 times per hour to 1.3 times per hour, with a correction effectiveness rate of 68%; the accuracy rate of identifying fake learning behavior has reached 92.7%; the accuracy rate of fatigue prediction has reached 89%, with an early warning time of 3-5 minutes; the intelligent tutoring function is used an average of 2.3 times per day, and the completion rate of questions after tutoring has reached 87%. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the overall system architecture of the present invention; Figure 2 This is a schematic diagram of the hardware structure of the present invention; Figure 3 This is a flowchart of the multimodal data acquisition process of the present invention; Figure 4 This invention uses a learning state discrimination decision tree; Figure 5 This is a flowchart of the spoofing learning recognition algorithm of the present invention; Figure 6 This is a sequence diagram of the robotic arm's motion execution according to the present invention; Figure 7 This is a schematic diagram of knowledge graph matching in this invention; Figure 8 This is a curve showing the adjustment of the light color temperature in this invention. Figure 9 This is an example diagram of the learning report interface for this invention; Figure 10 This is a comparison chart of experimental data for the present invention. Detailed Implementation
[0017] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0018] Example 1: Overall System Hardware Structure like Figure 1 and Figure 2 As shown, the learning companion robot system of the present invention includes a desk lamp body module, a robotic arm module, a multimodal perception module, a computing control module, a data management module, and an intelligent tutoring module.
[0019] The lamp body features a modern minimalist design with a 20cm diameter lamp head. The LED lighting unit consists of 120 full-spectrum LED beads arranged in a ring, providing a luminous area of 200 square centimeters. The color temperature is adjustable from 2700K to 6500K in 100K increments, allowing for stepless smooth adjustment. Brightness is adjustable from 100 to 2000 lumens, meeting lighting needs under various ambient light conditions. The lamp head is adjustable at multiple angles via four degrees of freedom joints on a robotic arm, with a positioning accuracy of 1°.
[0020] The robotic arm module connects to the lamp head in a series structure, comprising four rotary joints. The base rotary joint connects to the top of the lamp post, with a rotation range of ±90°; the shoulder pitch joint has a rotation range of 0-120°; the elbow pitch joint has a rotation range of 0-135°; and the wrist pitch joint connects to the lamp head, with a rotation range of ±45°. Each joint uses a harmonic reducer with a transmission ratio of 50:1 to ensure smooth movement and positioning accuracy. The robotic arm weighs 1.2kg, has an arm span of 80cm, and an end effector load capacity of 500g (including the lamp head weight). The lamp head housing is made of flexible silicone and incorporates a 6-axis torque sensor to detect contact force and prevent injury to the learner.
[0021] Multimodal sensing module: The first sensing submodule (located in the lamp head) is specifically an RGB camera used to track and identify books. This camera can employ a structured light scheme, with its RGB sensor having a resolution of 1920×1080 and a frame rate of 30fps; its depth sensor has a resolution of 640×480, a frame rate of 30fps, a depth measurement range of 0.5-5 meters, and an accuracy of ±2mm. The camera's field of view is preferably 70°, covering the area of a standard desk (120cm×60cm). The second sensing submodule (located on the pole or base) can specifically be a wide-angle RGB camera for posture detection; The sound source localization unit (in this embodiment, for example, located at the lamp head) can specifically be a microphone array consisting of 6 MEMS microphones in a ring layout, with a diameter of 8cm and a microphone spacing of 4cm. This array supports beamforming technology and can achieve sound source localization within a range of ±30°, with a signal-to-noise ratio greater than 60dB.
[0022] The computing control module adopts a two-layer architecture. The edge computing unit is based on an embedded AI development board, equipped with a GPU acceleration unit, a computing power of no less than 4 TOPS, and a memory of no less than 4GB. It runs a Linux operating system and a deep learning inference engine, deploying lightweight pose detection models, sound classification models, and data fusion models. The cloud decision unit is deployed on a cloud server, configured with an 8-core CPU, 16GB of memory, and a GPU acceleration card. It runs a fatigue prediction model (Bi-LSTM), a knowledge graph matching engine, and a learning report generation module. The edge and cloud communicate via the MQTT protocol, and data transmission uses TLS encryption.
[0023] The data management module uses an SQLite database for local storage, with a database file size limit of 2GB. Stored content includes: learning session records (start time, end time, subject type), status statistics (focus duration, number of distractions, number of abnormal postures), skeletal keypoint coordinate sequences (1 frame per second), and intervention records (intervention time, intervention type, intervention content). All data is encrypted using AES-256 before being written, with the key stored in a secure chip. Users can view learning reports through the accompanying app and choose to sync data to the cloud for long-term archiving and multi-device synchronization.
[0024] Example 2: Sitting Posture Detection Algorithm like Figure 4 As shown, the specific implementation steps of the sitting posture detection algorithm are as follows: Step 1: Skeletal Keypoint Extraction. A skeletal keypoint detection algorithm is used to extract the learner's skeletal keypoints from the RGB image, including: head vertex, neck point, right shoulder point, right elbow point, right wrist point, left shoulder point, left elbow point, left wrist point, mid-hip point, right hip point, right knee point, right ankle point, left hip point, left knee point, left ankle point, right eye point, left eye point, right ear point, and left ear point. Combined with depth information, the 2D keypoint coordinates are converted to 3D coordinates (x, y, z), where x represents the horizontal direction, y represents the vertical direction, and z represents the depth direction.
[0025] Step 2: Calculate the cervical spine forward tilt angle. Define the midpoint of the shoulder as the midpoint between the left and right shoulder points, with the coordinates as follows: Shoulder midpoint = ((left shoulder point.x + right shoulder point.x) / 2, (left shoulder point.y + right shoulder point.y) / 2, (left shoulder point.z + right shoulder point.z) / 2) The formula for calculating the cervical spine forward tilt angle θ is: θ = arctan[(midpoint of shoulder.y - vertex of head.y) / (midpoint of shoulder.z - vertex of head.z)] × 180 / π In a normal sitting posture, θ should be within the range of 15°-25°. When θ > 30°, it indicates excessive forward tilting of the cervical spine, posing a risk of "smartphone addiction".
[0026] Step 3: Calculate the head tilt angle. The head tilt angle φ reflects whether the head is tilted to the left or right. The calculation formula is: φ = arctan[(left shoulder point x - right shoulder point x) / (left shoulder point y - right shoulder point y)] × 180 / π In a normal sitting posture, φ should be within the range of -5° to +5°. When |φ| > 15°, it indicates that the head is significantly tilted to the side.
[0027] Step 4: Calculate the spinal curvature. Using information from the cervical point, mid-gluteal point, and depth, calculate the radius of curvature R of the spinal curve: R = [(neck point.z - mid-glute point.z)² + (neck point.y - mid-glute point.y)²] / [2 × |neck point.z - mid-glute point.z|] In a normal sitting posture, the radius (R) should be greater than 50 cm. When R < 40 cm, it indicates excessive curvature of the spine.
[0028] Step 5: Anomaly Detection and Duration Statistics. Set the sliding time window T_window = 3 minutes, and count the percentage of frames that meet the anomaly conditions within the window. The judgment rule is: - If the percentage of frames with θ > 30° is > 80%, it is judged as "abnormal cervical lordosis"; - If the percentage of frames with |φ| > 15° is > 70%, it is judged as "abnormal head tilt". - If the percentage of frames with R < 40cm is greater than 75%, it is judged as "abnormal spinal curvature".
[0029] Step 6: Intervention Trigger. When abnormal sitting posture is detected and the duration exceeds the threshold, the corresponding level of intervention is triggered: - Lasting 3-5 minutes: Mild intervention, adjust the light color temperature to 5000K to increase alertness; - 5-10 minutes: Moderate intervention, with voice prompts such as "Pay attention to your posture, head up and chest out"; - If the condition persists for more than 10 minutes: Severe intervention. The robotic arm moves the light head to the learner's line of sight, and a voice prompt says, "You have been in poor posture for 10 minutes. Please adjust."
[0030] Actual test case: When a junior high school student was doing math homework, starting at 18:30, the forward tilt angle θ of the cervical spine gradually increased from 22° to 35°. At 18:33 (lasting for 3 minutes), the system adjusted the light color temperature and issued a voice prompt at 18:35 (lasting for 5 minutes). After the student adjusted his posture, θ dropped to 24°.
[0031] Example 3: Pretend Learning Recognition Algorithm like Figure 5 As shown, the spoofing learning recognition algorithm is one of the core innovations of this invention, and the specific implementation steps are as follows: Step 1: Book Static Detection. Use object detection algorithms to identify books, workbooks, test papers, and other learning materials on the desk, and extract their bounding box coordinates. Calculate the displacement distance d between the center points of the bounding boxes of adjacent frames: d = sqrt[(x_t - x_{t-1})² + (y_t - y_{t-1})²] Set the static threshold d_threshold = 5 pixels (corresponding to an actual distance of approximately 1 cm). Count the continuous static time T_static; when T_static > 5 minutes, record it as the first discriminative feature F1 = 1. Step 2: Writing Action Detection. The learner's hand movements are detected using an RGB camera to identify writing behavior. The characteristics of writing action are: periodic wrist movement above the book or paper, with a frequency of 2-5 Hz (corresponding to 2-5 strokes per second). The continuous no-writing time T_nowrite is recorded; when T_nowrite > 3 minutes, it is recorded as the second discriminant feature F2 = 1. Step 3: Eye gaze tracking. An eye-tracking algorithm (based on pupil detection and head pose estimation) is used to calculate the gaze direction vector v_gaze to determine if the gaze falls within the learning material area. The learning material area is defined as a 60cm × 40cm area in the center of the desk. The percentage of time P_distance spent with the gaze deviating from the learning material within a 10-minute time window T_window is calculated. P_off = T_offgaze / T_window × 100% When P_discretion > 60%, it is recorded as the third discriminant feature F3 = 1; Step 4: Comprehensive Judgment. When F1 = 1, F2 = 1, and F3 = 1 are all satisfied simultaneously, it is initially judged as "suspected pretense of learning." To avoid misjudgment (e.g., the student may be seriously thinking about a difficult problem), a verification questioning mechanism is activated; Step 5: Verify the question-asking mechanism. The system uses OCR to recognize the content of the current learning material and extract key information (such as page numbers, chapter titles, and question numbers). It then generates verification questions, for example: - "Which page are you looking at right now?" - "What is the question number?" - "What is the title of this section?" The system plays questions via speech synthesis and receives student responses via speech recognition. If a student can answer accurately (answer matching rate > 80%), they are considered to be in a "deep thinking state" and no intervention is triggered. If a student cannot answer or answers incorrectly, they are considered to be in a "pretending to learn state" and severe intervention is triggered.
[0032] Step 6: Intervention Implementation. After confirming the feigned learning, implement the following intervention actions: - The robotic arm moves the lamp head, changing the lighting angle to attract attention; - Voice prompt: "You haven't made any progress in your studies for 10 minutes. Do you need help?"; - Record instances of pretending to study and mark them in the study report.
[0033] Actual test case: Between 7:00 PM and 7:10 PM, a primary school student's book remained on page 15 without being turned (F1=1), the camera did not detect any writing activity (F2=1), and the student's gaze shifted between looking out the window and looking at their phone, deviating from the learning material for 75% of the time (F3=1). At 7:10 PM, the system triggered a verification question: "Which page are you looking at now?" The student answered "I don't know," confirming that they were pretending to study. The robotic arm moved the light head to attract attention and provided a voice prompt, after which the student resumed studying.
[0034] In a 30-day test with 200 students, the accuracy rate of the fake learning recognition was 92.7%, the false alarm rate was 5.3%, and the false negative rate was 2.0%.
[0035] Example 4: Fatigue Prediction Model like Figure 4 As shown, the fatigue prediction model uses a bidirectional long short-term memory network (Bi-LSTM) and can predict the learner's fatigue state 3-5 minutes in advance. Figure 4 The learning state decision tree is shown, which includes the logic for determining the fatigue state.
[0036] Step 1: Feature Vector Construction. Collect multidimensional physiological and behavioral features of learners to construct a feature vector X_t: X_t = [θ, f_blink, N_sigh, T_study, v_write, P_focus, N_yawn] in: θ: Cervical spine forward tilt angle (degrees), reflecting the degree of relaxation in body posture; f_blink: Blinking frequency (times / minute), normally 15-20 times / minute, increases to 25-35 times / minute when fatigued; N_sigh: Number of sighs (cumulative over the past 10 minutes), detected by audio recognition; T_study: Continuous study time (minutes); v_write: Writing speed (words per minute), estimated by the frequency of hand movements; P_focus: Focus level (0-100), calculated based on eye gaze stability; N_yawn: Number of yawns (cumulative over the past 10 minutes), identified by facial motion unit (AU).
[0037] Step 2: Time series data construction. Construct time series data with a 1-minute sampling interval: X = [X_{t-9}, X_{t-8}, ..., X_{t-1}, X_t] Even if the feature vectors of the past 10 minutes are used as model input, the fatigue state of the next 5 minutes can be predicted.
[0038] Step 3: Bi-LSTM Model Structure. The model consists of the following layers: - Input layer: Shape (10, 7), representing 10 time steps, with 7 features per step; - Forward LSTM layer: 128 hidden units, learning temporal dependencies from the past to the present; - Backward LSTM layer: 128 hidden units, learns the temporal dependencies from the future to the past; - Stitching layer: The outputs of the forward and backward LSTMs are stitched together in the shape (256,). - Fully connected layer 1: 128 neurons, activation function is ReLU; - Dropout layer: Dropout rate of 0.3 to prevent overfitting; - Fully connected layer 2: 64 neurons, activation function is ReLU; - Output layer: 1 neuron, activation function is sigmoid, output fatigue prediction value P_fatigue ∈ [0, 1].
[0039] Step 4: Model Training. The model is trained using learning data from 500 students, totaling 10,000 hours. Labeling method: Students self-assess their fatigue level (0-10 points) after the learning period, and this is adjusted based on learning efficiency metrics (number of questions completed per unit time). Training parameters: learning rate 0.001, batch size 32, training epochs 50, optimizer Adam, loss function mean squared error (MSE).
[0040] Step 5: Fatigue assessment. Map P_fatigue to the range of 0-100: Fatigue level = P_fatigue × 100 Judgment rules: Fatigue level < 40: Normal state, no intervention required; 40 ≤ Fatigue Level < 70: Mild fatigue. Adjust the light color temperature to 4000K (warm light) to relieve visual fatigue. 70 ≤ Fatigue Level < 85: Moderate fatigue, voice prompt "You have been studying for a long time, it is recommended to rest for 5 minutes"; Fatigue level ≥ 85: Severe fatigue. The robotic arm moves the lamp head to the resting position, and a voice prompt says, "Please rest immediately and have a drink of water to relax."
[0041] Step 6: Early Warning. The model predicts fatigue levels at time t+5, providing a 5-minute advance warning. This gives learners a buffer time, allowing them to rest after completing the current task and avoiding sudden interruptions to their train of thought.
[0042] Actual test case: A high school student started studying physics at 20:00. At 20:35, the model predicted that the student's fatigue level would reach 72 in 5 minutes (20:40). At 20:37 (when the current fatigue level was 65), the system issued a voice prompt "It is recommended to rest after completing the current question". The student completed the question at 20:39 and rested for 5 minutes, effectively avoiding excessive fatigue.
[0043] On the test set, the fatigue prediction accuracy was 89%, the average early warning time was 4.2 minutes, the false positive rate was 8%, and the false negative rate was 3%.
[0044] Example 5: Intelligent Learning Tutoring System like Figure 7 As shown, the intelligent learning tutoring module is one of the core functions of this invention, used to respond to learners' proactive requests for help and provide real-time personalized learning tutoring services. This module is deeply integrated with the knowledge graph matching function, realizing a complete closed loop from passive monitoring to proactive tutoring.
[0045] 5.1 Active Request Trigger Mechanism Learners can trigger intelligent tutoring in the following ways: Method 1: Voice wake-up. The learner says the wake-up word "Little Light, Little Light," and the system enters tutoring mode. Voice wake-up uses an offline wake-up engine with a response latency of <500ms and a false wake-up rate of <0.5 times / hour. Method 2: Physical Button. The lamp base is equipped with a tutoring button; the learner can activate the tutoring mode by pressing the button. The button uses a capacitive touch design and supports both short press (single request for help) and long press (continuous dialogue) modes. Method 3: Gesture Recognition. The system uses an RGB camera to recognize the learner's raised hand gestures. When the arm is raised at an angle greater than 120° and the duration is greater than 2 seconds, the tutoring mode is automatically triggered. Method 4: Passive Trigger. When the system detects that the learner's pause time on the same question exceeds 1.5 times the reasonable upper limit, it will proactively ask, "Do you need help?" The learner will enter tutoring mode if they answer "yes".
[0046] 5.2 Multimodal VLM Understanding and Tutoring Generation Step 1: Speech Recognition. A deep learning speech recognition model is used, supporting Mandarin and common dialects, with an accuracy rate greater than 95%. The model has been specifically optimized to address the speech characteristics of K-12 students (fast speaking speed, unclear pronunciation, and use of colloquial expressions). Step 2: Multimodal Input Construction. Learning material images are captured in real-time using an RGB camera, and visual information is combined with the learner's spoken questions to create a multimodal input. - Image input: A full-page image containing the questions (1920×1080 resolution); - Text input: Learners' spoken questions are converted into text; - Contextual information: Metadata such as the learner's grade, subject, and current chapter.
[0047] Step 3: Multimodal VLM Inference. Input the multimodal content into a large-scale vision-language model (VLM). The VLM model has the following capabilities: - Visual understanding: Directly recognize the content of the questions in images, including text, mathematical formulas, geometric figures, tables, etc.; - Semantic understanding: Understanding the learner's true intent in their question without requiring explicit intent categorization; - Knowledge Reasoning: Based on a massive amount of pre-trained knowledge, understand the knowledge points and problem-solving methods tested in the questions; - Multi-turn dialogue: Supports continuous dialogue and dynamically adjusts tutoring strategies based on learner feedback.
[0048] VLM model deployment methods: - Cloud-based reasoning: Calls the cloud API, with a response time of 2-5 seconds. To protect privacy, the question images are anonymized before uploading, retaining only the question content area and removing student personal information; - Edge inference: Deploy lightweight VLM models (model size 7B-13B parameters), with a response time of less than 1 second, but slightly weaker capabilities; - Hybrid mode: Simple problems (such as multiple choice and fill-in-the-blank questions) use the edge model, while complex problems (such as proof and application questions) call the cloud model.
[0049] Fault tolerance mechanism: - Network failure handling: When a cloud API call fails, automatically switch to the edge model or provide basic prompts; - Timeout handling: If the cloud response time exceeds 10 seconds, it will automatically degrade to the edge model; - Degradation strategy: When the edge model fails, the system extracts a preset tutoring template from the knowledge graph.
[0050] VLM model technology feasibility description: The VLM model of this invention has multiple implementation schemes and does not depend on any specific commercial service: Open source model solution: The system can use open source vision-language multimodal models, such as models trained based on open source frameworks such as LLaMA and CLIP, with a model parameter size of 7B-13B, which can run on edge devices; Self-training scheme: System developers can use public datasets (such as K12 question banks and scanned images of textbooks) to fine-tune the model and train a dedicated model suitable for K12 education scenarios; Hybrid deployment solution: For simple questions (multiple choice, fill-in-the-blank, calculation), a lightweight edge model can be used to achieve high-accuracy identification and guidance; for complex questions (proof, application), you can choose to call the cloud API or use a rule-based backup solution. Offline backup solution: When the network is unavailable and the edge model cannot handle the situation, the system automatically switches to a knowledge graph-based rule engine, matching preset tutoring templates according to the question type and knowledge points. Although the flexibility is reduced, basic tutoring functions can still be provided. Model intellectual property rights: All models used in the system are either open-source or self-trained, and there are no intellectual property dependencies. Training data comes from public datasets and legally authorized educational resources.
[0051] Step 4: Knowledge Point Location and Verification. The VLM model output includes: - Question type identification: multiple choice, fill-in-the-blank, calculation, proof, application, etc. - Knowledge point annotation: The core knowledge points involved in the question (such as "application of the Pythagorean theorem"); - Difficulty Assessment: The difficulty level (easy / medium / hard) is assessed based on the complexity of the problem. - Problem-solving approach: Step-by-step solution method and key hints.
[0052] The system cross-validates the knowledge points output by VLM with the local subject knowledge graph to ensure accuracy.
[0053] 5.3 VLM-Driven Heuristic Coaching Strategies This invention employs a VLM large-model-driven heuristic teaching method, using the Prompt project to implement hierarchical guidance.
[0054] Coaching on Prompt Design: The system sends a structured Prompt to the VLM model, which includes: Character setting: You are an experienced K-12 teacher who is good at heuristic teaching; Task objective: Help students of {grade level} understand the questions in {subject level}, but do not provide the answers directly; Problem Image: [Image Input] Student question: {Student speech to text} Coaching levels: {Current level 1-4} Historical Dialogue: [Records of the First 3 Rounds of Dialogue] Tutoring requirements: Level 1: Only the knowledge points are mentioned, without providing solutions. - Level 2: Provide a problem-solving framework and list the key steps. - Level 3: A detailed explanation of the first step, using Socratic questioning to guide thinking. - Level 4: Provides complete solutions and recommends similar practice problems. Output format: JSON { "response": "Tutoring content", "knowledge_point": This refers to a specific knowledge point. "difficulty": Level of difficulty "next_action": "Wait / Continue with the explanation / Recommended practice" } Level 1 tutoring (knowledge point tips): - VLM output example: "This question tests the application of the Pythagorean theorem in practical problems. Can you think of a way to transform the figure formed by the ladder, wall, and ground into a mathematical model?"; - Waiting time: 30 seconds, visual monitoring is used to determine whether the learner has started writing.
[0055] Second-level coaching (thinking framework): - VLM output example: "The solution involves three steps: First, draw a diagram and label the right triangle formed by the ladder, wall, and ground; second, extract the known conditions from the problem (ladder length, wall height, or distance from the ground); third, set the unknown quantity as x and use the Pythagorean theorem to set up an equation and solve it. You can try the first step first." - Waiting time: 60 seconds.
[0056] Third-level coaching (step-by-step explanation): - VLM output example: "Let's look at the first step together. The problem says 'A 5-meter-long ladder is leaning against a wall,' and this ladder is the hypotenuse of a right triangle. Which side is the wall? Which side is the ground? Can you draw this triangle on scratch paper?" - It uses a conversational approach, explaining only one step at a time; - After visual monitoring confirms that the learner has completed the current step, proceed to the next step.
[0057] Level 4 tutoring (complete solutions + exercises): - VLM outputs the complete solution process and generates similar problems; - Example: "Great! Now let's review everything... [complete solution]. I'll give you a similar problem: a 13-meter-long ladder... give it a try." - Adaptive adjustment: The VLM model dynamically adjusts its tutoring strategy based on learners' real-time feedback (such as "I still don't understand" or "Can you explain in more detail?"), without the need for preset fixed rules.
[0058] 5.4 Explanation of VLM Output and TTS Collaboration To enhance tutoring effectiveness, the system converts the text content generated by VLM into speech via TTS, and combines it with visual and motor input.
[0059] Text-to-Speech (TTS): Converts the tutoring text output by the VLM into speech using neural network TTS technology. Employing an advanced TTS model, it possesses the following characteristics: - Natural tone: Close to the voice of a real teacher, with a strong sense of approachability; - Prosody control: Automatically identifies key words and emphasizes them, pausing naturally at punctuation marks; - Adjustable speaking speed: 150 words / minute by default, which can be adjusted according to the learner's age (120 words / minute for primary school students, 180 words / minute for middle school students). - Emotional expression: Supports various tones such as encouragement, reminders, and questions to enhance the vividness of the explanation.
[0060] TTS processing flow: - The VLM output text is preprocessed, highlighting key words and pauses; - Generate speech by calling the TTS API or the local TTS engine; - Speech synthesis latency is less than 500ms, enabling a smooth conversation experience.
[0061] Visual assistance: A robotic arm points to key locations within the problem. For example: - When explaining "finding the right triangle", the robotic arm points to the geometric figure in the problem; - When explaining "labeling known side lengths", the robotic arm points to the data given in the problem in sequence; - When explaining "setting up equations", the robotic arm points to the learner's draft paper, prompting them to write here.
[0062] The robotic arm has a pointing accuracy of ±5mm, a pointing speed of 10cm / s, and maintains its pointing posture for 3-5 seconds.
[0063] Lighting adjustment: In tutoring mode, the light color temperature is automatically adjusted to 5000K (natural light) and the brightness is increased to 800 lux to ensure that learners can clearly see the details of the questions.
[0064] 5.5 Evaluation of Coaching Effectiveness Step 1: Real-time monitoring. During the tutoring process, the system continuously monitors the learner's behavior: - Writing motion: The learner's hand movements are detected by an RGB camera to determine whether they are writing; - Eye contact: Determine whether the learner is looking at the question through eye tracking; - Facial expressions: Through facial action unit (AU) recognition, determine whether the learner understands (smiles, nods) or is confused (frowns, shakes head).
[0065] This behavioral data can be used as feedback input to the VLM model to help the model adjust its coaching strategies.
[0066] Step 2: Comprehension Confirmation. After each tutoring level, the system asks "Did you understand?" or "Can you continue?" and determines whether to proceed to the next level based on the learner's response.
[0067] Step 3: Completion Verification. After the learner completes the questions, the system takes an image of the answer and inputs it along with the original questions into the VLM model for grading. The VLM model can: - Recognize handwritten answers; - Determine if the answer is correct; - Analyze error types (misunderstanding of concepts, calculation errors, omissions in problem-solving steps, etc.); - Generate targeted corrective suggestions.
[0068] Compared to traditional OCR + rule matching, the VLM model can understand the problem-solving process. Even if the final answer is wrong, it can still identify some correct steps and affirm them.
[0069] Step 4: Knowledge Point Mastery Assessment. Calculate the knowledge point mastery level based on tutoring level, completion time, and answer accuracy: Mastery = 100 - Coaching Level × 20 - Percentage of Time Exceeding Standard Time × 10 - Number of Errors × 15 Mastery level >80: Proficient; 60-80: Basic mastery; 40-60: Partial mastery; <40: No mastery.
[0070] Step 5: Data Recording. Record detailed data from each tutoring session into the learning portfolio: - Tutoring time, question content, knowledge points, tutoring level; - Learner questions and system answers; - Completion time, correctness of answers, and mastery of knowledge points; - Used to generate analysis of weak knowledge points and personalized learning suggestions.
[0071] Real-world test case: A 8th-grade student, while working on a Pythagorean theorem application problem, asked for help via voice, "How do I do this problem?" The system input the problem image and the student's question into the VLM model. The model outputs the first level of guidance: "This problem tests the application of the Pythagorean theorem in practical problems. Can you identify the right triangle in the problem?" After a 35-second pause with no writing action, the system requests the second level of guidance from VLM. The model outputs a detailed problem-solving framework. The student begins writing and completes the problem 3 minutes later. The system inputs the student's answer image into VLM for grading, confirms the answer is correct, assesses the student's mastery of the knowledge point as 75% (basic mastery), and records it in the student's learning portfolio.
[0072] In a 30-day test involving 200 students, the VLM-based intelligent tutoring function was used an average of 2.3 times per day, resulting in an 87% completion rate for questions after tutoring and a student satisfaction rating of 8.5 / 10. Compared to traditional methods, VLM can handle more complex question types (such as open-ended questions and essay correction), significantly improving tutoring quality.
[0073] Example 6: Knowledge Graph Matching and Passive Tutoring Trigger In addition to responding to learners' requests for help, the system can also proactively provide guidance when learners encounter difficulties through knowledge graph matching.
[0074] Step 1: Learning Material Recognition. Use OCR technology to recognize the text content in books or workbooks. The recognition range is all text on the currently opened page, with an accuracy rate greater than 95%. Extract key information: - Book title: Extracted by cover recognition or header information; - Chapter: Identified by page title; - Page number: Identified by numbers in the footer; - Question number: Matched by patterns such as "1.", "(1)", "Question 1".
[0075] Step 2: Knowledge Point Localization. The identified text content is matched against the subject knowledge graph. The knowledge graph is constructed using a graph database and contains the following node types: - Subject nodes: Mathematics, Chinese, English, Physics, Chemistry, etc.; - Grade milestones: Elementary school grades 1-6, Junior high school grades 1-3, Senior high school grades 1-3; - Chapter nodes: such as "Chapter 17 Pythagorean Theorem in the second semester of eighth grade mathematics"; - Key knowledge points: such as "applications of the Pythagorean theorem" and "converse of the Pythagorean theorem"; - Question type nodes: such as "proof questions", "calculation questions", "application questions".
[0076] The relationships between nodes include: belonging, containing, preceding, and associating. A graph traversal algorithm is used to locate the knowledge point node corresponding to the currently learning content.
[0077] Knowledge Graph Scale and Maintenance: - Scale of the map: It contains approximately 50,000 knowledge point nodes, covering mainstream textbooks such as People's Education Press, Beijing Normal University Press, and Jiangsu Education Press; - Number of nodes: 12 subject nodes, 15 grade nodes, approximately 2,000 chapter nodes, approximately 50,000 knowledge point nodes, and approximately 200 question type nodes; - Number of relationships: Approximately 150,000 relationship edges, including types such as "belongs to", "contains", "precedes", "associates", and "similar"; - Update mechanism: Updated once per semester, adjusting the knowledge point structure according to the latest textbook syllabus; - Textbook adaptation: The system automatically matches the knowledge graph branch of the corresponding version by recognizing the book title.
[0078] Step 3: Determine Thinking Time. The normal thinking time varies significantly across different knowledge points and question types. Query the "average thinking time" and "standard deviation of thinking time" for the current knowledge point from the knowledge graph, and calculate the reasonable range of thinking time: Reasonable range = [Average duration - Standard deviation, Average duration + 2 × Standard deviation] For example, if the average thinking time for Pythagorean theorem application problems is 10 minutes and the standard deviation is 3 minutes, then the reasonable range is [7, 16] minutes.
[0079] Step 4: Pause Analysis. Monitor the learner's pause duration T_pause on the current question in real time. Judgment Logic: If T_pause is within a reasonable range, it is considered "normal thinking" and no intervention is made. If T_pause exceeds the upper limit of a reasonable range, it is determined as "encountering difficulties" and intelligent tutoring is triggered.
[0080] Step 5: Generating Tutoring Content. Retrieve the relevant information for the current knowledge point from the knowledge graph: - Concept Explanation: Definitions and core points of the knowledge points; - Problem-solving approach: General problem-solving steps for this type of question; - Common Mistakes: Common mistakes students make and how to correct them; - Related examples: Solution examples for similar questions.
[0081] Generate personalized coaching scripts, for example: "This question tests the application of the Pythagorean theorem. The solution is as follows: First, identify the right triangle in the problem; second, label the known side lengths; third, let the unknown side length be x; fourth, set up an equation using the Pythagorean theorem. Do you need any hints?" Step 6: Coaching Implementation. Play the coaching script using speech synthesis. If the learner responds "needs," continue providing detailed prompts; if the learner responds "no need" or "let me think about it," stop coaching and continue monitoring.
[0082] Actual test case: A junior high school student paused for 18 minutes while working on question 5 on page 128 of the People's Education Press textbook for eighth-grade mathematics, exceeding the reasonable upper limit (16 minutes). The system identified the question as an application problem of the Pythagorean theorem, extracted tutoring content from the knowledge graph, and provided a voice prompt: "The key to this problem is to find the right triangle. You can try drawing auxiliary lines." After listening to the prompt, the student completed the solution within 3 minutes.
[0083] Example 7: Subject-Specific Differentiation Parameter Settings The learning characteristics of different subjects differ significantly. This invention sets differentiated monitoring parameters and intervention strategies for the three main subjects of mathematics, Chinese, and English.
[0084] Mathematics: - Threshold for pauses in thinking: 8-15 minutes (mathematics requires deep thinking, and the pause time is longer). - Writing frequency threshold: 5-20 times / minute (including calculations, drawing, and formulating); - Focus requirement: Sustained focus time > 80% (mathematics requires high concentration); - Intervention strategies: When the pause exceeds 15 minutes, provide hints on how to solve the problem; when frequent erasing is detected (eraser usage > 5 times / 10 minutes), prompt "You can skip if you encounter difficulties".
[0085] Language and Literature Subject: - Reading stillness threshold: 10-20 minutes (when the body is relatively still while reading); - Page turning frequency threshold: 0.5-3 times / minute (depending on reading speed); - Reading aloud detection: Reading aloud behavior is detected through audio recognition; reading aloud time accounting for >20% is considered normal. - Intervention strategy: When the page is not turned for a long time (> 20 minutes), prompt "Have you encountered a difficult paragraph to understand?"; when the reading volume is detected to be too low (< 40dB), prompt "You can read aloud a little louder".
[0086] English subject: - Reading aloud detection threshold: Volume > 50dB, reading time percentage > 30%; - Word lookup detection: based on the frequency of use of mobile phones or electronic dictionaries (based on hand gesture recognition); - Hearing pattern recognition: When headphones are detected being worn and audio is playing, the system automatically switches to hearing mode to reduce the frequency of intervention. - Intervention strategies: When the reading time is insufficient, prompt "English learning requires more reading"; when frequent word lookup is detected (>10 times / 10 minutes), prompt "You can read the whole text first, and then look up new words".
[0087] Adaptive parameter adjustment: The system automatically adjusts the parameter thresholds for each subject based on the learner's historical data. For example, if a student thinks about math quickly and has an average pause time of 6 minutes, the student's "thinking pause threshold" will be adjusted to 5-12 minutes.
[0088] Example 8: Robotic Arm Motion Planning and Execution like Figure 6 As shown, the robotic arm module can perform three types of interactive actions: pushing actions, pointing actions, and periodic prompting actions.
[0089] Push action: Application scenario: When a learner is detected to be severely fatigued (fatigue level ≥ 85) or has been studying for more than 45 minutes, the robotic arm will push a water cup or rest reminder card in front of the learner to remind them to take a break.
[0090] Motion planning: - Initial state: The robotic arm is in standby position (retracted, with joint angles of [0°, 30°, 45°, 0°]). - Target recognition: Identify the position of the water cup using an RGB camera and extract its 3D coordinates (x_cup, y_cup, z_cup); - Trajectory planning: Use fifth-order polynomial interpolation to generate smooth trajectories, avoiding sudden acceleration or deceleration; - Pushing motion: The robotic arm pushes the water cup to the learner's face at a speed of 5-15cm / s (10cm away from the edge of the book); - Return to standby: The robotic arm returns to the standby position along the original trajectory.
[0091] Security mechanisms: The robotic arm module of this invention adopts a multi-layered safety protection design, complies with the relevant requirements of GB / T 20867-2007 "Safety Implementation Specifications for Industrial Robots" and ISO 10218 international standard, and has been specifically optimized for K-12 student use scenarios. - Collision detection and force limitation: If the force sensor detects an abnormal contact force (greater than 500g), movement will immediately stop, triggering emergency stop protection. The maximum contact force during normal interaction will not exceed 200g, which is far below the human pain threshold (1000g). - Speed and acceleration limits: The robotic arm's movement speed shall not exceed 15cm / s and the acceleration shall not exceed 10cm / s². A fifth-order polynomial trajectory planning method shall be used to achieve smooth start and stop, avoiding sudden acceleration that may cause fright. - Workspace limitations: The robotic arm's range of motion is restricted to the desk area (a circular area with a radius of 80cm), and it will not enter the learner's personal safety space (within 30cm of the learner's body). This is achieved through dual protection via software and hardware limits. - Emergency Stop System: The lamp base is equipped with a physical emergency stop button (compliant with GB 16754-2008 standard). Pressing this button will stop the robotic arm's movements and cut off power within 100ms. Manual reset is required after the emergency stop to restart it. - Visual safety monitoring: During the movement of the robotic arm, the camera continuously monitors the learner's position and posture. If the learner is detected to have entered the safe zone of the movement trajectory (20cm buffer zone), the movement will be stopped immediately and an audio warning will be issued. - Child safety design: Designed with K-12 students in mind, all safety parameters of the robotic arm are set according to children's usage scenarios, increasing the safety factor by 50% compared to adult usage scenarios. The system includes a parental control mode to limit the robotic arm's range of motion and functions. - Fail-safe design: Adopting the "fail-safe" design principle, when the system detects sensor failure, controller abnormality or power failure, the robotic arm will automatically lock in the current position or slowly return to the standby position, and will not lose control or fall suddenly; - Safety Certification: The system has passed mandatory certifications such as GB 4706.1-2005 "Safety of Household and Similar Electrical Appliances" and GB / T9254-2008 "Electromagnetic Compatibility of Information Technology Equipment", and has applied for CE certification and FCC certification, meeting the domestic and international market access requirements; - User safety education: The product manual details the safety precautions for use. The system will play a safety education video when used for the first time to ensure that learners and parents understand the correct usage method.
[0092] Pointing action: Application scenario: When the robot arm detects that the learner is distracted, pretending to study, or pausing for a long time, it points to a clock, learning materials, or other cues to guide the learner's attention.
[0093] Motion planning: - Target localization: Identify the 3D coordinates (x_target, y_target, z_target) pointing towards a target (such as a clock); - Attitude calculation: Calculate the attitude that the end effector of the robotic arm should achieve so that the end effector points to the target; - Inverse kinematics solution: Based on the target position and orientation, solve for the joint angles [θ1, θ2, θ3, θ4]; - Motion execution: The robotic arm moves to the pointing position at a speed of 10cm / s and holds for 3 seconds; - Voice prompts: Play voice prompts such as "Pay attention to the time" or "Return to learning materials" while the pointing action is being performed; - Return to standby: The robotic arm returns to the standby position.
[0094] Pointing accuracy: ±5mm, ensuring clear pointing to the target.
[0095] Periodic prompt action: Application scenario: When learners maintain the same posture for a long time (> 20 minutes), the robotic arm performs a prompting action of lightly touching the table every 20-30 minutes to remind learners to move their bodies.
[0096] Motion planning: - Triggering conditions: The learner's posture change amplitude is detected to be < 5% (based on the variance of skeletal keypoint coordinates) and the duration is > 20 minutes; - Motion execution: The robotic arm moves to the edge of the desk, and the end effector lightly touches the tabletop (contact force 100-200g), producing a slight sound; - Frequency control: Execute every 20-30 minutes to avoid excessive interruptions; - Adaptive adjustment: If the learner does not respond to the prompt (posture remains unchanged), the intensity of the prompt is increased (contact force is increased to 300g, in conjunction with voice prompts).
[0097] Actual test case: A high school student was doing chemistry homework continuously from 21:00 to 21:45, with only a 2% change in posture. The system executed periodic prompts at 21:20 and 21:40, prompting the student to get up and move around, effectively relieving fatigue from prolonged sitting.
[0098] Example 9: Lighting Coordination Control like Figure 8As shown, the lighting adjustment module dynamically adjusts the color temperature and brightness according to the learning status and time period, in conjunction with the intervention strategy.
[0099] Color temperature adjustment strategy: - Focused state: Color temperature 5500-6500K (cool white light), improves alertness and concentration; - Mild fatigue: Color temperature 4500-5500K (natural light) to relieve visual fatigue; - Moderate fatigue: Color temperature 3500-4500K (warm white light) to create a relaxing atmosphere; - Severe fatigue: Color temperature 2700-3500K (warm yellow light), strongly suggests rest; - Resting state: Color temperature 2700K (warm light), protects eyesight.
[0100] Brightness adjustment strategy: - Automatic adjustment based on ambient light intensity: Uses a photosensor to detect ambient light intensity and maintains desktop illuminance within the range of 500-750 lux (compliant with national standard GB / T 9473-2017). - Adjust brightness according to the learning content: higher brightness (700 lux) when reading, moderate brightness (600 lux) when writing, and lower brightness (300 lux) when resting; - Anti-blue light mode: The anti-blue light mode is automatically activated when studying at night (after 9:00 PM), reducing the proportion of blue light to below 30% to protect eyesight and sleep quality.
[0101] Color temperature dynamic adjustment curve: The system dynamically adjusts the color temperature based on learning duration and fatigue level. The adjustment curve is as follows: Color temperature (K) = 6500 - fatigue level × 38 For example, when the fatigue level is 0, the color temperature is 6500K; when the fatigue level is 50, the color temperature is 4600K; and when the fatigue level is 100, the color temperature is 2700K.
[0102] Color temperature adjustment speed: To avoid discomfort caused by sudden changes, the color temperature adjustment speed is limited to 100K / minute. For example, adjusting from 6500K to 4500K takes 20 minutes. The adjustment process is smooth and natural, and learners hardly notice it.
[0103] Actual test case: A middle school student started studying at 7:00 PM with an initial color temperature of 6500K. As the study time increased, fatigue gradually rose. The system adjusted the color temperature to 5500K at 7:30 PM, 4500K at 8:00 PM, and 3500K at 8:30 PM. The student reported that the lighting changes were natural and comfortable, effectively relieving visual fatigue.
[0104] Example 10: Learning Portfolio Generation and Privacy Protection like Figure 9 As shown, the data management module generates multi-level learning reports and implements a strict privacy protection mechanism.
[0105] Daily Learning Report: Generation time: Automatically generated at 23:00 daily, or manually generated after the learning period ends.
[0106] Report content: - Total study time: includes effective study time and rest time; - Subject distribution: Percentage of learning time for each subject (pie chart); - Attention curve: The horizontal axis represents time, and the vertical axis represents attention level (0-100), showing the trend of attention level change over time; - Posture abnormality record: number of abnormalities, type of abnormality (forward tilt of the neck / lateral head tilt / spinal curvature), and time period of abnormality; - Pretend learning log: number of occurrences, time period, duration; - Fatigue warning record: number of warnings, warning period, peak fatigue level; - Intervention effectiveness evaluation: the number of times each type of intervention is triggered and its effectiveness rate (the proportion of learners responding to the intervention).
[0107] Weekly Study Report: Generation time: Automatically generated every Sunday at 23:00.
[0108] Report content: - Learning duration trend: Bar chart of daily learning duration over the past 7 days; - Attention Trend: Line chart of average attention over the past 7 days; - Heat map of peak fatigue periods: The horizontal axis represents the weekday, the vertical axis represents the time period (in 1-hour units), and the color depth represents the degree of fatigue; - Subject-specific learning time distribution: cumulative learning time and percentage for each subject; - Progress indicators: Compared to last week, changes in focus duration, abnormal posture rate, and number of times pretending to study; - Personalized suggestions: Improvement suggestions generated based on data analysis, such as "It is recommended to increase the frequency of rest between 19:00 and 20:00, as fatigue levels are higher during this period."
[0109] Monthly Learning Report: Generation time: Automatically generated at 23:00 on the last day of each month.
[0110] Report content: - Study habits assessment: Summarize the strengths and weaknesses of study habits, such as "high level of concentration, but posture needs improvement"; - Progress trend analysis: Key indicator trends over the past 4 weeks (focus duration, abnormal posture rate, number of times pretending to study); - Knowledge point mastery: Assess the mastery of each knowledge point based on the duration of pauses and tutoring requests; - Learning efficiency analysis: Indicators such as the number of questions completed per unit of time and the accuracy rate; - Long-term improvement recommendations: Long-term improvement plans generated based on monthly data.
[0111] Privacy protection mechanism: - Data minimization: Only the skeleton contour coordinates (x, y, z coordinates of 25 key points) are saved, and the original RGB image and depth image are not saved; - Local storage: All data is stored locally by default and is not uploaded to the cloud; - Encrypted storage: Database files are encrypted using AES-256, and the key is stored in a secure chip (TPM) to prevent data leakage; - Access control: Learning reports require password or fingerprint verification to view, preventing unauthorized access; - Data deletion: Users can delete historical data at any time; the deletion operation is irreversible. - Pause monitoring: Learners can pause monitoring via physical button or voice command. No data is collected during the pause, and the system only records the pause duration. - Transparency: When users use the system for the first time, clearly inform them of the scope of data collection, storage method, and purpose of use, and enable the system only after obtaining the user's consent.
[0112] Parental synchronization (optional): Parents can view learning reports through the accompanying app to understand their child's learning progress. Synchronization mechanism: - Learner authorization: Parents can only view the report if the learner actively authorizes it; - Data anonymization: The parent app only displays statistical data and charts, without showing specific monitoring screens or detailed records; - Synchronization frequency: Synchronize once a day to avoid the psychological pressure of real-time monitoring.
[0113] Example 11: Experimental Verification and Data Comparison like Figure 10 As shown, a comparative experiment was conducted over a period of 30 days to verify the effectiveness of the present invention.
[0114] Experimental Design: - Test subjects: 200 primary and junior high school students, aged 8-15, with a male-to-female ratio of 1:1; - Grouping: The participants were randomly divided into an experimental group (100 people, using the system of this invention) and a control group (100 people, using ordinary desk lamps). - Test duration: 30 consecutive days, with 2-3 hours of after-class study time each day; - Data collection: Data is collected through learning logs, parent questionnaires, learning efficiency tests, etc.
[0115] Evaluation indicators: - Average focus duration: The average duration of a single continuous focus session for learning; - Posture abnormality rate: The average number of abnormal sitting postures per hour; - Pretend learning detection rate: The proportion of pretend learning instances identified by the system out of the actual pretend learning instances; - Learning efficiency: Number of questions completed per unit of time; - User satisfaction: Student and parent satisfaction ratings for the system (1-10 points).
[0116] Experimental results: index control group experimental group Increase Statistical significance Average focus time 23±3.2 minutes 41 ± 4.5 minutes +78% p<0.001 abnormal sitting posture rate 4.2 ± 0.8 times / hour 1.3±0.4 times / hour -69% p<0.001 Pretending to learn detection rate not applicable 92.7% not applicable not applicable Learning efficiency 8.5 ± 1.2 questions / hour 12.3 ± 1.8 questions / hour +45% p<0.01 Student satisfaction 6.2 ± 1.1 points 8.1 ± 0.9 points +31% p<0.01 Parent satisfaction 5.8 ± 1.3 points 8.7 ± 1.0 points +50% p<0.001 Note: The control group used ordinary desk lamps, which do not have the ability to recognize fake learning; therefore, this indicator is marked as "not applicable." Data are mean ± standard deviation. The independent samples t-test was used for statistical analysis, with a significance level of p < 0.05. Sample size calculations were based on an effect size of 0.5, a power of 0.8, and a significance level of 0.05, resulting in a minimum of 64 participants per group. In practice, 100 participants per group met the statistical requirements.
[0117] Data Analysis: - Significantly improved focus duration: The average focus duration of students in the experimental group increased from 23 minutes to 41 minutes, an increase of 78%. This is mainly attributed to the fake learning detection and distraction reminder functions, which enable students to maintain a genuine learning state; - The posture correction effect was significant: the rate of abnormal posture among students in the experimental group decreased from 4.2 times per hour to 1.3 times per hour, with a correction effectiveness rate of 68%. The progressive intervention mechanism (light → voice → robotic arm) was proven to be more effective than a single reminder; - Accurate identification of spoofing: In the experimental group, the system identified 327 instances of spoofing, and after manual verification, the actual number of spoofing instances was 353, with an accuracy rate of 92.7%, a false alarm rate of 5.3%, and a false negative rate of 2.0%. - Improved learning efficiency: The learning efficiency (number of questions completed per unit of time) of the experimental group students increased by 45%. This was due to fatigue prediction and timely rest, which avoided inefficient learning; - High user satisfaction: Students and parents rated the system with 8.1 and 8.7 points respectively, significantly higher than the control group. Parents particularly appreciated the "pretend learning recognition" and "learning report" functions, believing they helped them better understand their children's learning status.
[0118] Typical Case: Case 1: A fifth-grade student in elementary school had an average attention span of 18 minutes before using the system and frequently pretended to study (2-3 times a day). After using the system, the number of times he pretended to study decreased to once a week, his average attention span increased to 35 minutes, and his monthly test score improved by 15 points. Case 2: A second-year junior high school student had a long-standing problem with poor posture (the forward tilt of the cervical spine often exceeded 40°). After using the system of this invention for 30 days, the rate of abnormal posture decreased from 5 times per hour to 1 time per hour, and the parents reported that the child's neck pain symptoms were significantly relieved. Case 3: A junior high school student spent long hours studying but had low efficiency and was often overly fatigued. After using the system of this invention, the fatigue prediction function helped him to reasonably arrange rest, improving his learning efficiency by 50% and reducing his daily study time from 4 hours to 3 hours, but the amount of homework he completed actually increased.
[0119] Industrial application This invention has broad industrial application prospects and can be applied to the following scenarios: - Home education scenario: As an intelligent auxiliary tool for K12 students' after-school learning, it helps students develop good study habits, improve learning efficiency, and reduce the burden of supervision on parents; - Online education institutions: Integrate into online education platforms to provide learning status monitoring and behavior analysis services for remote learning, helping teachers understand students' true learning status and optimize teaching strategies; - Library study rooms: Deployed as shared smart desk lamps in library study rooms to provide students with a personalized learning environment and behavioral reminders; - Special Education Area: Providing attention training and behavioral intervention services for children with ADHD (Attention Deficit Hyperactivity Disorder) to assist in rehabilitation treatment; - Corporate training scenario: Applied to corporate employee training and examination scenarios, it monitors learning status and improves training effectiveness.
[0120] Legal Compliance Statement: The data collection and processing of this invention system comply with the requirements of laws and regulations such as the "Personal Information Protection Law of the People's Republic of China," the "Data Security Law of the People's Republic of China," and the "Regulations on the Protection of Children's Personal Information Online." Specific measures include: - Informed consent mechanism: Before the system is used for the first time, learners and their guardians must be clearly informed of the scope of data collection, storage method and purpose of use, and explicit consent must be obtained before the monitoring function can be enabled; - Data minimization principle: Only collect and store the minimum data necessary to achieve the function, do not save the original video images, only save the skeletal outline coordinates and statistics; - Special protection for minors: Enhanced privacy protection measures are implemented for K-12 students (minors), including parental or guardian consent, data access control, and a regular data deletion mechanism; - Data security measures: AES-256 encryption algorithm is used for local storage, the key is stored in a security chip, and data transmission is encrypted with TLS to prevent data leakage and unauthorized access; - Data subject rights: Learners and their guardians have the right to query, copy, correct, and delete data. The system provides a convenient data management interface. - Pause monitoring mechanism: Learners can pause monitoring at any time via physical button or voice command. No data is collected during the pause, fully respecting personal privacy;
[0121] The industrialization prospects of this invention are broad, with an estimated market size of several billion yuan. With the deepening development of educational informatization and smart homes, multifunctional intelligent learning devices will become essential products.
[0122] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A learning behavior intervention method characterized by, The method comprises the following steps: S1. Multi-dimensional learning state perception step: collecting posture image data of the learner through a visual perception module and collecting sound data of the learning environment through an audio perception module; S2. Learning state intelligent discrimination step: inputting the posture image data and the sound data into a multi-modal fusion model to identify the current learning state of the learner, the learning state including at least one of a focused state, a fatigue state, a mind-wandering state, a poor posture state, a fake learning state and a normal rest state; S3. Hierarchical intervention strategy generation step: generating a corresponding intervention strategy according to the learning state, combining a subject knowledge graph and historical behavior data of the learner, the intervention strategy including an intervention level, an intervention method and an intervention content; S4. Multi-modal collaborative execution step: executing an intervention operation through at least one of a voice prompt module, a light adjustment module and a mechanical arm interaction module according to the intervention strategy; S5. Intelligent learning tutoring step: when the learner initiates a learning request, receiving a question of the learner through a voice recognition module, collecting learning material images through the visual perception module, inputting the question and the images into a visual-language multi-modal model for understanding and reasoning, generating tutoring content by using a heuristic hierarchical tutoring method including four levels of knowledge point prompt, problem solving thought framework, step-by-step explanation and complete solution, and explaining through a voice synthesis module and the mechanical arm module.
2. The learning behavior intervention method according to claim 1, characterized in that, In the learning state intelligent discrimination step: The identification method of the fake learning state includes: detecting, through the RGB camera, that the book is stationary for more than 5 minutes, there is no writing action for more than 3 minutes, and the eye gaze deviation from the learning material accounts for more than 60%, triggering a verification question mechanism when the above three features are met at the same time, and determining whether it is a fake learning state according to the accuracy of the learner's answer; The fatigue state prediction method includes: collecting a multi-dimensional feature vector including cervical lordosis angle, blink frequency, sigh frequency, learning time and writing speed, inputting the multi-dimensional feature vector into a bidirectional long short-term memory network model, and outputting a fatigue degree prediction value to predict the fatigue state 3 to 5 minutes in advance, with a prediction accuracy of not less than 85%; The poor posture state detection method includes: extracting the bone key point coordinates through the RGB camera, calculating the cervical lordosis angle θ and the head lateral deviation angle φ, and determining the poor posture state when θ is greater than 30° and lasts for more than 3 minutes, or φ is greater than 15° and lasts for more than 2 minutes.
3. The learning behavior intervention method of claim 1, wherein, The hierarchical intervention strategy includes: Mild intervention strategy: when a mild fatigue state or a short-term mind-wandering state is detected, only adjusting the color temperature to the range of 4000K to 6000K through the light adjustment module; Moderate intervention strategy: when a moderate fatigue state or a poor posture state lasting for more than 3 minutes is detected, issuing a gentle reminder through the voice prompt module and adjusting the brightness through the light adjustment module; Severe intervention strategy: when detecting that the severe fatigue state, pretending to learn state or bad posture state lasts more than 8 minutes, the physical interaction action is executed through the mechanical arm module, and the explicit reminder is issued through the voice prompt module.
4. The learning behavior intervention method of claim 1, wherein, The data fusion method of the multi-modal fusion model comprises: Feature extraction layer: respectively extracting feature vectors from visual data and audio data, visual features including posture key point coordinates, action frequency and eye gaze direction, audio features including sound type, volume and frequency; Feature fusion layer: splicing or weighted fusion of visual feature vectors and audio feature vectors, and the fusion weight is dynamically adjusted according to the current subject type; State classification layer: input the fused feature vector into the classifier, output the learning state category and confidence, and trigger the multi-modal cross verification mechanism when the confidence is lower than the threshold.
5. The learning behavior intervention method of claim 1, wherein, The heuristic hierarchical tutoring method in the intelligent learning tutoring step comprises: First level tutoring: only provide knowledge point name and related concept prompt, and wait for 30 to 60 seconds; Second level tutoring: provide problem solving idea framework, list key steps but do not explain in detail, and wait for 60 to 120 seconds; Third level tutoring: adopt socrates type questioning method, guide the learner to complete the first key step step by step, and continue after confirming the completion through visual monitoring; Fourth level tutoring: provide complete solution process, and generate similar questions for practice, and correct and feedback the answers of the learner through the visual-linguistic multi-modal model.
6. A companion robot system, characterized by It comprises: A desk lamp body module comprising an LED lighting unit, the color temperature of the LED lighting unit can be dynamically adjusted in the range of 2700K to 6500K; A mechanical arm module connected to the desk lamp head, having at least 4 degrees of freedom, and having a working radius of 60cm to 100cm, for moving the lamp head and executing physical interaction actions; A multi-modal sensing module: a first sensing sub-module: arranged at the lamp head position of the desk lamp. This module is used to obtain image information and / or depth information in the lighting area below the lamp head, to realize tracking lighting or content recognition of the target (such as a book, a hand); a second sensing sub-module: arranged at the base link position of the desk lamp. This module is used to obtain image information of the user in front of the desk lamp, to realize user posture detection; a sound source positioning unit: used to collect environmental sound and determine the sound source direction, which can be integrated in the lamp head position or the non-lamp head fixed position; A computing control module comprising an edge computing unit and a cloud decision unit, the edge computing unit is used for real-time data processing and rapid response, and the cloud decision unit is used for complex model reasoning and data analysis; A data management module for storing the learning record data of the learner, and implementing a privacy protection mechanism, the learning record data is stored locally by AES-256 encryption algorithm, only the skeletal outline coordinate data is saved, and the original video image is not stored; An intelligent tutoring module comprising a voice recognition unit, a visual-linguistic multi-modal model unit, a knowledge graph verification unit and a voice synthesis unit, for responding to the active help of the learner, and providing real-time learning tutoring service.
7. The tutorial robot system according to claim 6, wherein The knowledge graph matching method in the hierarchical intervention strategy generation step includes: Text content in learning materials is identified through optical character recognition technology, and book title, chapter, and page number information is extracted; The text content is matched with a subject knowledge graph to locate the current learning knowledge point node; According to the association relationship of the knowledge point node, relevant concept explanations, problem solving ideas, and common errors are obtained; According to the learner's pause length and historical learning data, it is judged whether the learner needs coaching prompts, and personalized coaching rhetoric is generated.
8. The tutorial robot system according to claim 6, wherein, The subject differentiation parameter setting method includes: For mathematics, set the thinking pause threshold to 8-15 minutes and the writing frequency threshold to 5-20 times per minute; For Chinese, set the reading stillness threshold to 10-20 minutes and the page turning frequency threshold to 0.5-3 times per minute; For English, set the reading detection threshold to a volume of not less than 50 dB and a reading duration ratio of not less than 30%; According to the subject type, automatically adjust the triggering conditions and intervention content of the intervention strategy.
9. The tutorial robot system according to claim 6, wherein, The motion planning method of the mechanical arm module includes: Lamp head movement: when the learner needs intervention, the mechanical arm module moves the lamp head to change the lighting angle and position, with a moving speed of 5-15 cm / s; Pointing action: when the learner's attention is distracted, the mechanical arm moves the lamp head to point to the learning materials or the clock, with a positioning accuracy of ±5 mm; Periodic prompting action: when the learner maintains the same posture for a long time, the mechanical arm slightly shakes the lamp head every 20-30 minutes.
10. The companion robot system of claim 6, wherein The data management module generates daily learning reports, weekly learning reports, and monthly learning reports, the daily learning report includes total learning time, focused time, abnormal posture times, pretend learning times, fatigue warning times, and help seeking times, the weekly learning report includes focus trend curve, fatigue high incidence time heat map, subject learning time distribution graph, and knowledge point mastery heat map, and the monthly learning report includes learning habit evaluation, progress trend analysis, weak knowledge point analysis, and personalized improvement suggestions.