Auxiliary learning method and system based on AI multi-mode interactive screen projection eye protection table lamp
By leveraging the desk lamp's omnidimensional environmental perception matrix and two-way interactive channel, user data is collected and analyzed in real time to optimize the learning environment. This solves the accuracy and comprehensiveness issues of traditional desk lamp-assisted learning systems, improving learning efficiency and eye health.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN SHENGPINYUAN IND CO LTD
- Filing Date
- 2025-12-24
- Publication Date
- 2026-05-05
AI Technical Summary
Traditional desk lamp-assisted learning systems rely on limited data sources and cannot fully integrate multi-dimensional state information such as visual fatigue and distraction, resulting in a lack of accuracy and comprehensiveness in lighting control and posture reminders.
The integrated desk lamp's full-dimensional environmental perception matrix collects high-dimensional perception feature streams in real time, analyzes user profile information, calculates visual fatigue index and focus, constructs a two-way interaction channel between the user and the desk lamp, and optimizes the auxiliary learning environment.
It achieves precise quantification of users' visual fatigue and learning status, provides personalized and seamless learning support, improves learning efficiency and eye health, and stimulates long-term learning motivation.
Smart Images

Figure CN121979381A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an auxiliary learning method and system for an AI-based multimodal interactive projection eye-protection desk lamp, belonging to the field of multimodal interaction technology. Background Technology
[0002] Desk lamp-assisted learning refers to an intelligent service model that utilizes desk lamps as smart terminals, going beyond their basic function of "illumination." Through proactive sensing, intelligent analysis, and personalized interaction, it creates a more efficient, healthier, and more immersive learning environment, thereby optimizing the learning process and improving learning outcomes. Desk lamp-assisted learning proactively intervenes before the user becomes distracted to maintain an efficient learning state by monitoring focus. When the user encounters difficulties, it instantly projects relevant auxiliary information next to the book to reduce interruptions. Simultaneously, it can monitor and correct poor posture and viewing distance in real time, and proactively reminds the user to take a break based on fatigue levels, thus providing effective health management.
[0003] Traditional desk lamp-assisted learning provides a static or semi-static, standardized physical environment through preset, fixed hardware functions. For example, existing pulse-width modulation (PWM) technology is used in desk lamp-assisted learning to control the light and adjust its brightness and color. A single distance sensor detects the user's posture and distance from the book during study to monitor the user's sitting posture and thus assist learning. However, this method is too rigid in its intervention, relying on limited data sources and failing to integrate multi-dimensional information such as visual fatigue and distraction. This results in a lack of comprehensive basis for intervention decisions, leading to inaccurate and precise light control and posture reminders. Summary of the Invention
[0004] This invention provides an auxiliary learning method and system based on an AI-powered multimodal interactive screen projection eye-protection desk lamp, the main purpose of which is to improve the adaptability and comprehensiveness of desk lamp-assisted learning.
[0005] To achieve the above objectives, the present invention provides an assisted learning method for an AI-based multimodal interactive screen-projecting eye-protection desk lamp, comprising: The integrated desk lamp has a full-dimensional environmental perception matrix to collect high-dimensional perception feature streams of the target user in real time during the learning process; Based on the high-dimensional perceptual feature stream, the user profile information of the target user is analyzed to determine the potential needs of the target user. The visual fatigue index and visual focus of the target user are calculated to analyze the current learning state of the target user. Based on the potential needs and the current learning state, an auxiliary learning environment for the target user is constructed. A two-way interaction channel is constructed between the target user and the desk lamp. Based on the two-way interaction channel, the learning environment characteristics of the auxiliary learning environment are optimized to execute the auxiliary learning of the desk lamp.
[0006] Optionally, calculating the visual fatigue index and visual focus of the target user includes: The eye-tracking features and facial features of the target user corresponding to the high-dimensional perceptual feature stream are extracted to construct a three-dimensional facial model of the target user. Based on the three-dimensional facial model, the head posture of the target user is analyzed, and the visual focus of the target user is determined based on the head posture. Based on the facial 3D model, the blinking behavior index, pupil and eye movement index, and head posture index of the target user are calculated to determine the visual fatigue index of the target user.
[0007] Optionally, constructing a 3D model of the target user's face includes: Construct a basic 3D facial model of the target user, and determine the facial feature points of the target user based on the facial features of the target user, so as to map the corresponding points of the 3D facial model; Based on the facial feature points and the corresponding points of the model, the rotation matrix and translation vector of the three-dimensional facial model are determined. According to the rotation matrix and the translation vector, the pose parameters and shape parameters of the three-dimensional facial model are calculated to map the three-dimensional facial model to a two-dimensional image space to obtain the mapped feature points. The error between the facial feature points and the mapped feature points is calculated to optimize the basic 3D model and obtain a facial 3D model.
[0008] Optionally, based on the head posture, determining the visual focus of the target user includes: Based on the head posture, determine the head center axis and eye center axis of the head posture; Using the rotation matrix and initial direction vector of the head posture, the head direction vector of the head central axis is calculated. Based on the pupil center of the head posture, the eye direction vector of the eye central axis is determined. The distance between the central axes of the head central axis and the eye central axis is calculated to determine the visual focus of the target user.
[0009] Optionally, based on the facial 3D model, the blink behavior index, pupil and eye movement index, and head posture index of the target user are calculated, including: Based on the eye feature sequence of the facial 3D model, the blink frequency, blink duration, and PERCLOS of the target user are calculated to determine the blink behavior index of the target user. The pupil diameter change, saccade velocity, saccade amplitude, and fixation point dispersion of the target user are calculated to calculate the pupil and eye movement index of the target user; Based on the head pose sequence of the facial 3D model, the head nodding frequency and head pose stability of the target user are calculated to determine the head pose index of the target user.
[0010] Optionally, calculating the gaze dispersion of the target user includes: Based on the eye feature sequence, a 3D gaze point cloud of the target user is generated to calculate the spatial probability density of the gaze point corresponding to the target user in 3D space. Based on the spatial probability density, the information entropy of the gaze point is analyzed to determine the gaze point dispersion of the target user.
[0011] Optionally, the integrated desk lamp's multi-dimensional environmental perception matrix includes: The central control unit and multi-dimensional subspace unit of the desk lamp are integrated. The physical layout diagram and communication protocol between the central control unit and the multi-dimensional subspace unit are defined. Based on the physical layout diagram and the communication protocol, the full-dimensional environmental perception matrix of the desk lamp is integrated.
[0012] Optionally, based on the high-dimensional perceptual feature stream, the user profile information of the target user is analyzed, including: Extract the high-level feature vector of the high-dimensional perceptual feature stream, and analyze the target user's attention endurance, knowledge gap map and learning preferences based on the high-level feature vector; Based on the focus endurance, the knowledge gap map, and the learning preferences, a user profile of the target user is constructed.
[0013] Optionally, a two-way interaction channel is constructed between the target user and the desk lamp, including: Define the output mode of the desk lamp and the input mode of the target user; Based on the output mode and the input mode, a three-layer intent negotiation protocol for the desk lamp is generated, wherein the three-layer intent negotiation protocol includes: an implicit intent proposal layer, an explicit confirmation and correction layer, and a negotiation learning layer; Based on the three-layer intent negotiation protocol, a two-way interaction channel is constructed between the target user and the desk lamp.
[0014] To address the aforementioned problems, this invention also provides an auxiliary learning system based on an AI-powered multimodal interactive screen-projecting eye-protecting desk lamp, the system comprising: The high-dimensional data acquisition module is used to integrate the full-dimensional environmental perception matrix of the desk lamp to collect the high-dimensional perception feature stream of the target user in real time during the learning process. The learning environment construction module is used to analyze the user profile information of the target user based on the high-dimensional perceptual feature stream to determine the potential needs of the target user, calculate the visual fatigue index and visual focus of the target user to analyze the current learning state of the target user, and construct the auxiliary learning environment of the target user based on the potential needs and the current learning state. The assisted learning execution module is used to construct a two-way interaction channel between the target user and the desk lamp, and based on the two-way interaction channel, optimize the learning environment characteristics of the assisted learning environment to execute the assisted learning of the desk lamp. Compared to the problems described in the background technology, this invention, by integrating a desk lamp to construct a full-dimensional environmental perception matrix, achieves a revolutionary leap from passive monitoring to proactive empowerment. It transforms scattered user data into a unified high-dimensional perceptual feature stream, enabling real-time analysis of the user's visual fatigue index and focus, accurately quantifying their current learning state, and deeply mining user profiles to proactively identify their potential needs. The assisted learning environment constructed by this invention is no longer a static set of rules, but a dynamically evolving intelligent ecosystem. By building a two-way interaction channel between the user and the desk lamp, intelligent decision-making can be transformed into physical world lighting adjustment, attention guidance, and emotional feedback, extending assisted behavior from the screen to real learning scenarios. Finally, this complete closed loop of perception, analysis, decision-making, and execution allows the assisted learning environment to continuously self-optimize, providing users with highly personalized, seamless, and empathetic learning support, thereby fundamentally improving learning efficiency, protecting eye health, and stimulating long-term learning motivation. Therefore, this invention can enhance the adaptability and comprehensiveness of desk lamp-assisted learning. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating an embodiment of the auxiliary learning method for an AI-based multimodal interactive screen projection eye-protection desk lamp provided by the present invention. Figure 2 This is a schematic diagram of a module for implementing the auxiliary learning method of the AI-based multimodal interactive screen projection eye-protection desk lamp according to an embodiment of the present invention; Figure 3 A schematic diagram of a computer device for an auxiliary learning method based on an AI multimodal interactive projection eye-protection desk lamp, according to an embodiment of the present invention; The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0016] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0017] This application provides an assisted learning method based on an AI-powered multimodal interactive screen-projecting eye-protection desk lamp. The execution entity of this assisted learning method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the assisted learning method based on an AI-powered multimodal interactive screen-projecting eye-protection desk lamp can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.
[0018] Reference Figure 1 The diagram shown is a flowchart illustrating an auxiliary learning method for an AI-based multimodal interactive screen-projecting eye-protection desk lamp according to an embodiment of the present invention. In this embodiment, the auxiliary learning method for an AI-based multimodal interactive screen-projecting eye-protection desk lamp includes: S1, the integrated desk lamp's full-dimensional environmental perception matrix, to collect high-dimensional perception feature streams of the target user in real time during the learning process.
[0019] This invention, through an integrated desk lamp's multi-dimensional environmental perception matrix, can capture data such as ambient light intensity, color temperature, distance, posture, and user physiological rhythms in real time, forming a dynamic perception feature stream and providing high-precision input for subsequent analysis. The multi-dimensional environmental perception matrix refers to a systematic perception architecture integrated into the smart desk lamp, based on a multi-source heterogeneous sensor array, synchronized and controlled through a unified spatiotemporal reference. It aims to capture all key information in the learning scenario in real time from three orthogonal and complementary dimensions: physical environment, user behavior, and learning content.
[0020] As an embodiment of the present invention, the integrated desk lamp's full-dimensional environmental perception matrix includes: The central control unit and multi-dimensional subspace unit of the table lamp are integrated; Define the physical layout diagram and communication protocol between the central control unit and the multi-dimensional subspace unit; Based on the physical layout diagram and the communication protocol, the full-dimensional environmental perception matrix of the desk lamp is integrated.
[0021] The central control unit refers to a high-performance microprocessor module integrating a CPU, GPU / NPU, memory, and multiple I / O interfaces. The multi-dimensional subspace unit refers to a functional module integrating specific types of sensors, categorized according to data source and functional domain, including vision subunits, depth subunits, spectral subunits, acoustic subunits, and interaction subunits. The physical layout diagram refers to an engineering blueprint defining the three-dimensional spatial positions, structural relationships, physical interfaces, and heat dissipation channels of each functional unit within the lamp. The communication protocol refers to a set of rules ensuring accurate, efficient, and orderly data exchange between the central control unit and each multi-dimensional subspace unit.
[0022] Optionally, the central control unit of the desk lamp can be integrated using semiconductor packaging technology, such as 2.5D / 3DIC packaging, fan-out wafer-level packaging, etc. The multi-dimensional subspace unit can be integrated using functional chip-level technology.
[0023] Optionally, the physical layout diagram between the central control unit and the multidimensional subspace unit can be defined by digital twins and generative design algorithms, such as topology optimization algorithms, evolutionary algorithms, gradient optimization algorithms, etc.
[0024] This invention, through real-time acquisition of high-dimensional perceptual feature streams of target users during the learning process, can transform subjective learning states into quantifiable indicators, thereby achieving precise quantification of user states and in-depth analysis of learning behaviors. The high-dimensional perceptual feature stream refers to a continuous data structure formed by real-time acquisition from multimodal sensors, followed by processing and fusion, including visual features, posture features, environmental features, and behavioral and interaction features.
[0025] S2. Based on the high-dimensional perceptual feature stream, analyze the user profile information of the target user to determine the potential needs of the target user, calculate the visual fatigue index and visual focus of the target user to analyze the current learning state of the target user, and construct an auxiliary learning environment for the target user based on the potential needs and the current learning state.
[0026] This invention, through analyzing the user profile information of the target user based on the high-dimensional perceptual feature stream, can achieve the identification of the user's cognitive patterns and the classification of learning styles, thereby establishing a quantitative model of their fatigue level and predicting the fatigue threshold. The user profile information refers to a set of digital models automatically constructed and updated through continuous analysis of the high-dimensional perceptual feature stream, used to accurately describe the target user's learning state, behavioral habits, and cognitive characteristics.
[0027] As an embodiment of the present invention, the step of analyzing the user profile information of the target user based on the high-dimensional perceptual feature stream includes: Extract the high-level feature vector of the high-dimensional perceptual feature stream; Based on the advanced feature vectors, analyze the target user's attention span, knowledge gap map, and learning preferences; Based on the focus endurance, the knowledge gap map, and the learning preferences, a user profile of the target user is constructed.
[0028] The high-level feature vector refers to a low-dimensional, high-information-density numerical representation strongly correlated with the specific analysis target, extracted from the original, timestamped, high-dimensional perceptual feature stream after data cleaning, noise reduction, and feature engineering. Examples include learning duration, focus duration, and learning material type. Focus endurance refers to a long-term model that quantitatively describes a user's ability to maintain a high level of focus. The knowledge gap graph refers to a graph-structured data model with knowledge concepts as nodes and association strength as edges. Learning preferences refer to a set of parameterized configurations describing the external conditions under which a user's learning efficiency and experience are highest, such as lighting, interaction methods, and learning types.
[0029] Optionally, the high-level feature vectors of the high-dimensional perceptual feature stream can be extracted using deep learning techniques, such as convolutional neural networks, recurrent neural networks, long short-term memory networks, etc.
[0030] Optionally, the target user's attention span, knowledge gap map, and learning preferences can be analyzed using causal inference models, such as difference-in-differences or causal forests.
[0031] This invention, by identifying the potential needs of the target user, achieves a fundamental shift from passive response to proactive intervention, ultimately constructing a personalized assisted learning environment with forward-looking, adaptive, and deeply collaborative capabilities. The potential needs refer to functions inferred from the fusion analysis of real-time user profile information and the current context, functions that the user has not yet actively expressed but are necessary for achieving better learning outcomes.
[0032] Optionally, the potential needs of the target users can be determined by a demand analysis model, such as a time series model or a causal analysis model.
[0033] This invention, through calculating the visual fatigue index and visual focus trajectory of the target user, can accurately quantify the user's visual health and attention state, providing precise triggering criteria for adaptive lighting adjustment, fatigue intervention, and gaze-based contextual assistance. The visual fatigue index is a composite indicator calculated through multimodal visual feature fusion, used to objectively quantify the user's eye muscle tension and visual function state. The visual focus refers to the precise point where the target user's gaze lands in three-dimensional physical space and the semantic object they are interested in.
[0034] As an embodiment of the present invention, calculating the visual fatigue index and visual focus of the target user includes: Extract eye movement features and facial features from the high-dimensional perceptual feature stream corresponding to the target user to construct a 3D facial model of the target user; Based on the facial 3D model, analyze the head posture of the target user; Based on the head posture, determine the visual focus of the target user; Based on the facial 3D model, the blinking behavior index, pupil and eye movement index, and head posture index of the target user are calculated to determine the visual fatigue index of the target user.
[0035] The eye-movement features refer to quantitative data describing eye movement states and visual attention patterns, such as gaze vectors, saccades, and pupil diameter. Facial features refer to key points, regions, and their geometric and textural information used to identify, reconstruct, and analyze the face, such as facial key points, facial regions, and texture information. The facial 3D model refers to a general, parameterized 3D facial mesh. Head posture refers to the rotation angle and displacement position of the head in 3D space. The blinking behavior index is a quantitative indicator used to assess the degree of fatigue reflected by blinking behavior. The pupil and eye-movement index is a quantitative indicator used to assess cognitive load and fatigue levels reflected by pupil changes and eye movement patterns. The head posture index is a quantitative indicator used to assess fatigue levels reflected by unconscious changes in head posture.
[0036] Optionally, constructing the 3D facial model of the target user includes: Construct a basic 3D facial model of the target user; Based on the facial features of the target user, determine the facial feature points of the target user to map the corresponding points of the facial 3D model; Based on the facial feature points and the corresponding points of the model, the rotation matrix and translation vector of the three-dimensional facial model are determined. Based on the rotation matrix and the translation vector, the pose parameters and shape parameters of the three-dimensional facial model are calculated to map the three-dimensional facial model to a two-dimensional image space to obtain the mapped feature points; The error between the facial feature points and the mapped feature points is calculated to optimize the basic 3D model and obtain a facial 3D model.
[0037] The basic 3D facial model refers to a general, parameterizable 3D facial mesh template. Facial feature points are specific points in a facial image used to identify key facial features, such as eyes, nose, and ears. The corresponding points in the model are the points on the 3D model that correspond to the feature points detected in the 2D image. The rotation matrix is a mathematical representation of an object's rotation in 3D space; it is a 3x3 matrix used to transform a point in 3D space from its original coordinate system to a target coordinate system. The translation vector is a vector describing the change in an object's position in 3D space; for pose estimation of the 3D facial model, the translation vector describes the movement of the basic 3D model relative to its initial position in 3D space. The pose parameters are a set of parameters describing the rotation and displacement of the human body in space. The shape parameters are parameters used to describe the shape of an object in computer vision applications such as 3D modeling and face recognition. The mapped feature points are points on the 2D image plane that are mapped from the corresponding points on the 3D basic facial model through projection transformation. The error refers to the difference between the mapped feature points of the 3D basic model on the 2D image and the feature points actually detected in the 2D image.
[0038] Optionally, the three-dimensional facial model can be obtained through optimization algorithms, such as gradient descent or particle swarm optimization.
[0039] In another implementation, the attitude parameters and the shape parameters are calculated using the following formula: ; ; in, Indicates attitude parameters, Indicates shape parameters, Indicates function parameters, Represents the function to be minimized. This indicates the number of facial feature points. Indicates the first Facial feature points, Indicates the first Each model corresponds to a point. Represents the rotation matrix. Represents the translation vector. Indicates projection transformation, Indicates the first The change in the corresponding points of each model.
[0040] Optionally, determining the visual focus of the target user based on the head posture includes: Based on the head posture, determine the head center axis and eye center axis of the head posture; The head orientation vector of the head center axis is calculated using the rotation matrix of the head posture and the initial orientation vector. Based on the pupil center of the head posture, determine the eye direction vector of the eye central axis; The distance between the central axes of the head and the eyes is calculated to determine the visual focus of the target user.
[0041] In this context, the head central axis refers to a hypothetical virtual axis passing through the center of the head in a face image or 3D model. The eye central axis refers to an axis passing through the center of the pupil and perpendicular to the surface of the eye in a face image. The initial direction vector refers to the starting direction used to determine the head and eye central axes during head pose estimation. The head direction vector is a vector describing the direction of the head central axis in a face image or 3D model; this vector typically points from the center of the head to the head central axis. The eye direction vector is a vector describing the direction of the eye central axis in a face image or 3D model; this vector typically points from the center of the eye to the eye central axis. The central axis distance refers to the distance between the head and eye central axes in a face image or 3D model; this distance can be used to describe the relative position between the head and eyes.
[0042] Optionally, the rotation matrix and translation vector of the basic 3D model can be determined by a nonlinear optimization method, wherein the nonlinear optimization refers to an algorithm that considers the nonlinear relationship of rotation and translation and can handle more complex geometric transformations, such as the Levenberg-Marquardt algorithm.
[0043] Optionally, the step of calculating the blink behavior index, pupil and eye movement index, and head posture index of the target user based on the facial 3D model includes: Based on the eye feature sequence of the facial 3D model, the blink frequency, blink duration, and PERCLOS of the target user are calculated to determine the blink behavior index of the target user. The pupil diameter change, saccade velocity, saccade amplitude, and fixation point dispersion of the target user are calculated to calculate the pupil and eye movement index of the target user; Based on the head pose sequence of the facial 3D model, the head nodding frequency and head pose stability of the target user are calculated to determine the head pose index of the target user.
[0044] The eye feature sequence refers to a series of key geometric and dynamic data related to the eyes extracted sequentially from the 3D facial model. The blink frequency refers to the number of complete blinks occurring per unit time. The blink duration refers to the duration of a single complete blink. PERCLOS refers to the percentage of time within a specific time window during which the eyelids remain closed for more than a certain threshold (e.g., 80%). The pupil diameter change refers to the dynamic characteristics of pupil diameter changes over time. The saccade velocity refers to the peak angular velocity of the eyeball during rapid movement between two fixations. The saccade amplitude refers to the range of visual angles covered by a single saccade, i.e., the angle between the starting and ending fixation points. The fixation point dispersion refers to the spatial dispersion of the user's fixation point within the target area when performing a task. The head posture sequence refers to a series of data describing the head's state in 3D space extracted sequentially from the 3D facial model. The head nodding frequency refers to the number of unconscious or low-frequency nodding movements along the pitch axis that conform to fatigue characteristics per unit time. Head posture stability refers to the degree of slight fluctuation in the rotation angle and displacement of the head while maintaining a basic posture.
[0045] Optionally, the blink frequency, blink duration, and PERCLOS of the target user can be calculated using statistical methods.
[0046] Optionally, calculating the gaze dispersion of the target user includes: Based on the eye feature sequence, a 3D gaze point cloud of the target user is generated to calculate the spatial probability density of the gaze point corresponding to the target user in 3D space. Based on the spatial probability density, the information entropy of the gaze point is analyzed to determine the gaze point dispersion of the target user.
[0047] The 3D gaze point cloud refers to a set of discrete 3D coordinate points formed by the intersection of all gaze directions of the target user with multiple virtual planes in three-dimensional space within a specific time window. The spatial probability density is a continuous function describing the density of the 3D gaze point cloud distribution in space. The information entropy is an indicator used to quantify the uncertainty of gaze point distribution.
[0048] This invention enables precise quantitative assessment of a user's current learning state by analyzing the user's current learning status, providing real-time decision-making support for adaptive content scheduling and personalized intervention strategies. The current learning state refers to a dynamic description of the user's cognitive engagement, focus, and effectiveness in a specific learning task, obtained through a comprehensive evaluation that integrates the user's visual focus, visual fatigue index, and interactive behavior data.
[0049] Optionally, the current learning state of the target user can be analyzed using causal discovery algorithms, such as PC algorithm, NOTEARS, etc.
[0050] This invention, through its embodiments, constructs an auxiliary learning environment for the target user based on the potential needs and the current learning state. This environment proactively configures learning resources and environmental parameters before the user explicitly expresses their needs, dynamically optimizing the lighting spectrum, content presentation methods, and interaction modalities based on real-time cognitive states. This forms an environment matrix that matches the user's physiological and cognitive rhythm. Simultaneously, based on a knowledge gap map and a focus endurance model, it achieves accurate content recommendation and learning path planning. The auxiliary learning environment refers to a personalized configuration set generated in the digital space and used to drive physical devices based on real-time analysis of the target user's potential needs and current learning state.
[0051] Optionally, the auxiliary learning environment for the target user can be constructed using graph neural networks, such as graph convolutional networks or graph attention networks.
[0052] S3. Construct a two-way interaction channel between the target user and the desk lamp. Based on the two-way interaction channel, optimize the learning environment characteristics of the auxiliary learning environment to execute the auxiliary learning of the desk lamp.
[0053] This invention, through the construction of a two-way interaction channel between the target user and the smart lamp, forms a closed-loop calibration mechanism of "proposal-user feedback-strategy optimization." This not only ensures the accuracy and acceptability of intervention strategies but also provides real-time data support for the continuous evolution of user profiles, ultimately constructing an adaptive learning space with self-optimization capabilities. The two-way interaction channel refers to a structured communication architecture established between the user and the smart lamp, supporting bidirectional flow of multimodal information and intent collaboration.
[0054] As an embodiment of the present invention, constructing a two-way interaction channel between the target user and the desk lamp includes: Define the output mode of the desk lamp and the input mode of the target user; Based on the output mode and the input mode, a three-layer intent negotiation protocol for the desk lamp is generated, wherein the three-layer intent negotiation protocol includes: an implicit intent proposal layer, an explicit confirmation and correction layer, and a negotiation learning layer; Based on the three-layer intent negotiation protocol, a two-way interaction channel is constructed between the target user and the desk lamp.
[0055] The output modality refers to the set of all sensory channels used by the desk lamp to convey information, express system status, and intent to the target user. The input modality refers to the set of all natural interaction channels used by the target user to convey instructions, feedback, and intent to the desk lamp. The three-layer intent negotiation protocol is a set of structured, progressively deepening interaction rules designed to achieve human-computer intent coordination and alignment through a progressive process from gentle probing to explicit confirmation. The implicit intent proposal layer is the first stage of the protocol, which gently hints at the detected potential needs to the user through non-verbal, low-interference sensory signals and probes the user's response intention. The explicit confirmation and correction layer is the second stage of the protocol, where specific options are provided through clear language and the interface when the user shows interest in the implicit layer or the system determines that a clear decision is needed, inviting the user to confirm, refuse, or modify. The negotiation learning layer is the third stage of the protocol, which is also a continuously running background process responsible for recording, analyzing, and learning the user's final choice in the explicit layer, and using this feedback data to dynamically optimize the triggering conditions and content of the implicit proposal and the option settings of the explicit layer.
[0056] Optionally, the three-layer intent negotiation protocol of the desk lamp can be generated through hierarchical reinforcement learning, such as Hindsight Experience Replay.
[0057] This invention, through its bidirectional interactive channel, optimizes the learning environment characteristics of the auxiliary learning environment. By recording user feedback data on various environmental configurations, a "feedback-adjustment-evaluation" optimization loop is established. This not only allows for real-time correction of lighting parameters and content presentation strategies to accurately match the user's physiological state and cognitive needs, but also drives online updates of the environmental strategy model through continuously accumulating labeled datasets, ultimately constructing an adaptive learning space with self-evolution capabilities. The learning environment characteristics refer to a quantifiable and controllable set of parameters that constitute and define the auxiliary learning environment, such as physical lighting characteristics, information presentation characteristics, and interaction strategy characteristics.
[0058] Optionally, the learning environment features can be obtained through Bayesian optimization, such as Gaussian Process Upper Confidence Bound, Expected Improvement, etc.
[0059] The embodiments of the present invention, by executing the assisted learning of the desk lamp, can achieve dynamic planning of personalized learning paths and non-intrusive intelligent intervention through continuous quantitative tracking of the user's cognitive state and learning process.
[0060] like Figure 2 The diagram shown is a functional module diagram of the auxiliary learning system of the AI-based multimodal interactive screen projection eye-protection desk lamp of the present invention.
[0061] The AI-based multimodal interactive screen-projection eye-protection desk lamp auxiliary learning system 200 described in this invention can be installed in an electronic device. Depending on the functions implemented, the AI-based multimodal interactive screen-projection eye-protection desk lamp auxiliary learning system may include a high-dimensional data acquisition module 201, a learning environment construction module 202, and an auxiliary learning execution module 203. The module described in this invention can also be called a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, stored in the memory of the electronic device.
[0062] In this embodiment of the invention, the functions of each module / unit are as follows: The high-dimensional data acquisition module 201 is used to integrate the full-dimensional environmental perception matrix of the desk lamp to collect the high-dimensional perception feature stream of the target user in the learning process in real time. The learning environment construction module 202 is used to analyze the user profile information of the target user based on the high-dimensional perceptual feature stream to determine the potential needs of the target user, calculate the visual fatigue index and visual focus of the target user to analyze the current learning state of the target user, and construct the auxiliary learning environment of the target user based on the potential needs and the current learning state. The assisted learning execution module 203 is used to construct a two-way interaction channel between the target user and the desk lamp, and based on the two-way interaction channel, optimize the learning environment characteristics of the assisted learning environment to execute the assisted learning of the desk lamp.
[0063] In detail, the modules in the AI-based multimodal interactive screen-projecting eye-protecting desk lamp auxiliary learning system 200 described in this embodiment of the invention employ the same methods as described above during use. Figure 1 The method used is the same as the auxiliary learning method of the AI-based multimodal interactive projection eye-protection desk lamp described in the article, and can produce the same technical effect, so it will not be repeated here.
[0064] In one embodiment, a computer device is provided, which may be a server or a client, and its internal structure diagram may be as follows: Figure 3As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of an AI-based multimodal interactive screen-projecting eye-protecting desk lamp-assisted learning method on the server or client side.
[0065] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: The integrated desk lamp has a full-dimensional environmental perception matrix to collect high-dimensional perception feature streams of the target user in real time during the learning process; Based on the high-dimensional perceptual feature stream, the user profile information of the target user is analyzed to determine the potential needs of the target user. The visual fatigue index and visual focus of the target user are calculated to analyze the current learning state of the target user. Based on the potential needs and the current learning state, an auxiliary learning environment for the target user is constructed. A two-way interaction channel is constructed between the target user and the desk lamp. Based on the two-way interaction channel, the learning environment characteristics of the auxiliary learning environment are optimized to execute the auxiliary learning of the desk lamp.
[0066] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: The integrated desk lamp has a full-dimensional environmental perception matrix to collect high-dimensional perception feature streams of the target user in real time during the learning process; Based on the high-dimensional perceptual feature stream, the user profile information of the target user is analyzed to determine the potential needs of the target user. The visual fatigue index and visual focus of the target user are calculated to analyze the current learning state of the target user. Based on the potential needs and the current learning state, an auxiliary learning environment for the target user is constructed. A two-way interaction channel is constructed between the target user and the desk lamp. Based on the two-way interaction channel, the learning environment characteristics of the auxiliary learning environment are optimized to execute the auxiliary learning of the desk lamp.
[0067] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0068] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0069] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0070] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0071] Finally, it should be noted that in the above embodiments, each embodiment can be combined with each other or independent. Deleting any one of them will not affect the technical implementation of other embodiments. The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. An auxiliary learning method for an AI-based multimodal interactive screen-projecting eye-protection desk lamp, characterized in that, The method includes: The integrated desk lamp has a full-dimensional environmental perception matrix to collect high-dimensional perception feature streams of the target user in real time during the learning process; Based on the high-dimensional perceptual feature stream, the user profile information of the target user is analyzed to determine the potential needs of the target user. The visual fatigue index and visual focus of the target user are calculated to analyze the current learning state of the target user. Based on the potential needs and the current learning state, an auxiliary learning environment for the target user is constructed. A two-way interaction channel is constructed between the target user and the desk lamp. Based on the two-way interaction channel, the learning environment characteristics of the auxiliary learning environment are optimized to execute the auxiliary learning of the desk lamp.
2. The assisted learning method of the AI-based multimodal interactive screen-projecting eye-protecting desk lamp as described in claim 1, characterized in that, Calculating the visual fatigue index and visual focus of the target user includes: The eye-tracking features and facial features of the target user corresponding to the high-dimensional perceptual feature stream are extracted to construct a three-dimensional facial model of the target user. Based on the three-dimensional facial model, the head posture of the target user is analyzed, and the visual focus of the target user is determined based on the head posture. Based on the facial 3D model, the blinking behavior index, pupil and eye movement index, and head posture index of the target user are calculated to determine the visual fatigue index of the target user.
3. The assisted learning method of the AI-based multimodal interactive screen-projecting eye-protecting desk lamp as described in claim 2, characterized in that, Constructing a 3D facial model of the target user includes: Construct a basic 3D facial model of the target user, and determine the facial feature points of the target user based on the facial features of the target user, so as to map the corresponding points of the 3D facial model; Based on the facial feature points and the corresponding points of the model, the rotation matrix and translation vector of the three-dimensional facial model are determined. According to the rotation matrix and the translation vector, the pose parameters and shape parameters of the three-dimensional facial model are calculated to map the three-dimensional facial model to a two-dimensional image space to obtain the mapped feature points. The error between the facial feature points and the mapped feature points is calculated to optimize the basic 3D model and obtain a facial 3D model.
4. The assisted learning method of the AI-based multimodal interactive screen-projecting eye-protecting desk lamp as described in claim 2, characterized in that, Based on the head posture, determining the visual focus of the target user includes: Based on the head posture, determine the head center axis and eye center axis of the head posture; Using the rotation matrix and initial direction vector of the head posture, the head direction vector of the head central axis is calculated. Based on the pupil center of the head posture, the eye direction vector of the eye central axis is determined. The distance between the central axes of the head central axis and the eye central axis is calculated to determine the visual focus of the target user.
5. The assisted learning method of the AI-based multimodal interactive screen-projecting eye-protecting desk lamp as described in claim 2, characterized in that, Based on the aforementioned 3D facial model, the blink behavior index, pupil and eye movement index, and head posture index of the target user are calculated, including: Based on the eye feature sequence of the facial 3D model, the blink frequency, blink duration, and PERCLOS of the target user are calculated to determine the blink behavior index of the target user. The pupil diameter change, saccade velocity, saccade amplitude, and fixation point dispersion of the target user are calculated to calculate the pupil and eye movement index of the target user; Based on the head pose sequence of the facial 3D model, the head nodding frequency and head pose stability of the target user are calculated to determine the head pose index of the target user.
6. The assisted learning method of the AI-based multimodal interactive screen-projecting eye-protecting desk lamp as described in claim 5, characterized in that, Calculating the gaze dispersion of the target user includes: Based on the eye feature sequence, a 3D gaze point cloud of the target user is generated to calculate the spatial probability density of the gaze point corresponding to the target user in 3D space. Based on the spatial probability density, the information entropy of the gaze point is analyzed to determine the gaze point dispersion of the target user.
7. The assisted learning method of the AI-based multimodal interactive screen-projecting eye-protecting desk lamp as described in claim 1, characterized in that, The integrated desk lamp's multi-dimensional environmental perception matrix includes: The central control unit and multi-dimensional subspace unit of the desk lamp are integrated. The physical layout diagram and communication protocol between the central control unit and the multi-dimensional subspace unit are defined. Based on the physical layout diagram and the communication protocol, the full-dimensional environmental perception matrix of the desk lamp is integrated.
8. The assisted learning method of the AI-based multimodal interactive screen-projecting eye-protecting desk lamp as described in claim 1, characterized in that, Based on the high-dimensional perceptual feature stream, the user profile information of the target user is analyzed, including: Extract the high-level feature vector of the high-dimensional perceptual feature stream, and analyze the target user's attention endurance, knowledge gap map and learning preferences based on the high-level feature vector; Based on the focus endurance, the knowledge gap map, and the learning preferences, a user profile of the target user is constructed.
9. The assisted learning method of the AI-based multimodal interactive screen-projecting eye-protecting desk lamp as described in claim 1, characterized in that, Constructing a two-way interaction channel between the target user and the desk lamp includes: Define the output mode of the desk lamp and the input mode of the target user; Based on the output mode and the input mode, a three-layer intent negotiation protocol for the desk lamp is generated, wherein the three-layer intent negotiation protocol includes: an implicit intent proposal layer, an explicit confirmation and correction layer, and a negotiation learning layer; Based on the three-layer intent negotiation protocol, a two-way interaction channel is constructed between the target user and the desk lamp.
10. An auxiliary learning system based on an AI-powered multimodal interactive projection eye-protection desk lamp, characterized in that: The system includes: The high-dimensional data acquisition module is used to integrate the full-dimensional environmental perception matrix of the desk lamp to collect the high-dimensional perception feature stream of the target user in real time during the learning process. The learning environment construction module is used to analyze the user profile information of the target user based on the high-dimensional perceptual feature stream to determine the potential needs of the target user, calculate the visual fatigue index and visual focus of the target user to analyze the current learning state of the target user, and construct the auxiliary learning environment of the target user based on the potential needs and the current learning state. The assisted learning execution module is used to construct a two-way interaction channel between the target user and the desk lamp, and based on the two-way interaction channel, optimize the learning environment characteristics of the assisted learning environment to execute the assisted learning of the desk lamp.