Intelligent table lamp based on AI large model and control method thereof

By embedding AI large models in smart desk lamps, using multimodal perception modules and actuators to generate anthropomorphic user feedback, the problem of single interactive functions of existing desk lamps is solved, and emotional interaction and dynamic feedback are achieved.

CN120282340APending Publication Date: 2025-07-08CHENGDU HUMANOID ROBOT INNOVATION CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510700758.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing smart desk lamps lack emotional interaction capabilities and cannot generate anthropomorphic feedback based on environmental information. The interaction function is single and depends on preset programs.

Method used

The AI large model is used to deeply embedded into the desk lamp control system, and environmental data is collected through the multi-modal perception module to generate a comprehensive decision vector containing user intentions, emotional states and environmental characteristics, and to control the coordinated actions of multiple degrees of freedom actuators.

Benefits of technology

It realizes anthropomorphic interaction between smart desk lamps and users, and can generate dynamic feedback based on multimodal data, improving the interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120282340A_ABST
    Figure CN120282340A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent table lamp based on an AI large model and a control method of the intelligent table lamp, and relates to the technical field of artificial intelligence control. And multi-modal data information is imported into an AI large model, feature fusion is carried out, a comprehensive decision vector containing user intentions, emotional states and environment features is generated, a control instruction set is generated, and cooperative control of a multi-degree-of-freedom execution mechanism is realized. And thus, more personified feedback can be made in user interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence control technology, and more particularly, to an intelligent desk lamp based on an AI large model and its control method. Background Art

[0002] A desk lamp is a commonly used daily necessity. With the progress of technology, the desk lamp has been given more intelligent functions. However, existing intelligent desk lamps often have a single interaction function, only supporting basic voice control or preset scene switching (such as dimming, color adjustment), lacking the ability of emotional interaction. And they often rely on preset programs and cannot generate anthropomorphic feedback by combining environmental information (such as user actions, voice content). Summary of the Invention

[0003] The purpose of the present invention is to provide an intelligent desk lamp based on an AI large model and its control method, which deeply embeds the AI large model into the desk lamp control system to achieve multi-modal perception, dynamic decision-making and emotional interaction.

[0004] The embodiments of the present invention are implemented as follows: An intelligent desk lamp based on an AI large model, comprising: A multi-modal perception module for collecting environmental data to obtain multi-modal data; A calculation and control module for processing the multi-modal data through the AI large model to generate a control instruction set and controlling the execution of each specified action in the control instruction set; An actuator module for completing the specified action under the control of the calculation and control module.

[0005] A control method for an intelligent desk lamp based on an AI large model, comprising: S1. Collect environmental data through the multi-modal perception module to obtain multi-modal data, and the multi-modal data includes visual, voice and physical sensing information; S2. Input the multi-modal data into the AI large model for feature fusion to generate a comprehensive decision vector including user intent, emotional state and environmental features; S3. Dynamically generate a control instruction set according to the comprehensive decision vector, and the instruction set includes a spatio-temporal synchronization scheme of mechanical action parameters, optical parameters and audio output instructions; S4. Execute the control instruction set to achieve the coordinated control of the multi-degree-of-freedom actuator.

[0006] The beneficial effects of the embodiments of the present invention are: An embodiment of the present invention provides an intelligent table lamp based on an AI large model and its control method. It can collect environmental data through a multimodal perception module to obtain multimodal data information; then import the multimodal data information into the AI large model for feature fusion to generate a comprehensive decision vector including user intentions, emotional states, and environmental features, and generate a control instruction set to achieve the coordinated control of a multi-degree-of-freedom actuator. Thus, more anthropomorphic feedback can be made in the interaction with users. Brief Description of the Drawings

[0007] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, so they should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0008] Figure 1 It is a gesture processing flowchart of a control method for an intelligent table lamp based on an AI large model provided by an embodiment of the present invention; Figure 2 It is an audio processing flowchart of a control method for an intelligent table lamp based on an AI large model provided by an embodiment of the present invention. Detailed Embodiments

[0009] To make the purpose, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the present invention to be protected, but only represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0010] An embodiment of the present invention provides an intelligent table lamp based on an AI large model, which includes: A multimodal perception module for collecting environmental data to obtain multimodal data; A computing and control module that processes multimodal data through an AI large model to generate a control instruction set and controls the execution of various specified actions in the control instruction set; An actuator module that completes specified actions under the control of the computing and control module.

[0011] Furthermore, the multi-modal perception module includes: An RGB-D camera for identifying and collecting the user's gesture commands; A lidar for ranging and three-dimensional space perception; An array microphone for identifying and collecting the user's voice commands; A pressure sensor is disposed on the base of the smart table lamp for identifying the position and weight information of the items on the base.

[0012] Optionally, the RGB-D camera has a resolution of 1080P and a frame rate of 30fps; the ranging range of the lidar is 0.1 - 5m, and the accuracy is ±2mm. The array microphone uses a six-microphone circular array.

[0013] The actuator module includes: A three-axis gimbal carrying an RGB-D camera, a lidar, and an array microphone, which can adjust the angle according to the angle parameters in the control instruction set; An LED matrix that can adjust the brightness according to the lighting parameters in the control instruction set; A servo motor that can adjust the operating trajectory of the robotic arm according to the angle parameters in the control instruction set; A speaker that can play audio according to the audio parameters in the control instruction set.

[0014] The embodiment of the present invention also provides a control method for a smart table lamp based on an AI large model, which includes: S1. Collect environmental data through the multi-modal perception module to obtain multi-modal data, where the multi-modal data includes visual, voice, and physical sensing information; S2. Input the multi-modal data into the AI large model for feature fusion to generate a comprehensive decision vector including the user's intention, emotional state, and environmental characteristics; S3. Dynamically generate a control instruction set according to the comprehensive decision vector, and the instruction set includes the spatio-temporal synchronization scheme of mechanical action parameters, optical parameters, and audio output instructions; S4. Execute the control instruction set to achieve the coordinated control of the multi-degree-of-freedom actuator.

[0015] Furthermore, in step S1, the process of collecting environmental data through the multi-modal perception module includes: S1-1. Obtain the three-dimensional coordinates of the user's skeletal joints through the RGB-D camera; S1-2. Collect voice commands through the array microphone; S1-3. Real-time monitor the pressure distribution on the base of the table lamp through the pressure sensor matrix.

[0016] Optionally, the AI large model is improved based on an existing model, such as the DEEPSSK architecture, which includes a cross-modal attention fusion layer, an emotional state calculation layer, and a dynamic decision-making layer. The S2 step includes: S2-1. The cross-modal attention fusion layer uses the multi-head attention mechanism to associate visual and speech features; S2-2. The emotional state calculation layer analyzes the fundamental frequency change rate of speech based on the GRU network to generate an emotional intensity value; S2-3. The dynamic decision-making layer generates a control instruction queue containing timestamps through an improved Transformer decoder.

[0017] It can be implemented through the following commands:

[0018] Furthermore, the S1-1 step includes: S1-11. Detect the hand bounding box through an improved YOLOv5 model; S1-12. Extract hand key points based on the OpenPose algorithm and establish a three-dimensional space coordinate system.

[0019] Figure 1 A specific and feasible gesture processing process is shown, including a gesture command generation process and a gesture command processing process.

[0020] The gesture command generation process specifically includes: 1) System initialization Step 1: Initialize the camera device (configure parameters such as resolution and frame rate); Step 2: Create a HandCtrlArm robotic arm control object and establish a communication connection; Step 3: Reset the robotic arm to the initial position (zero pose).

[0021] 2) Gesture recognition and command generation Step 4: Enter the main loop and continuously read camera frames (RGB image data); Step 5: Call the ctrl_arm.process method to process the current frame and perform the following operations: Image preprocessing: Horizontally flip the image (to solve the camera mirror problem); Hand detection: Detect whether there is a hand in the picture based on the YOLOv5 model; Action parsing: If a hand is detected, calculate the three-dimensional coordinates (X / Y / Z) of the hand key points.

[0022] Step 6: Generate control commands (such as target position and movement speed), and add the commands to the control command queue.

[0023] 3) Queue Management Step 7: If no hand is detected or the control is not activated, directly add an empty command to the queue (to maintain queue continuity). Step 8: Real-time display of the processing frame rate (FPS) for performance monitoring.

[0024] Step 9: Detect whether the user presses the 'q' key to trigger the termination signal. If so, release the camera and robotic arm resources and exit the program.

[0025] The specific process of gesture command processing includes: 1) Command acquisition and parsing Step 1: Obtain the latest command from the control command queue (the queue follows the FIFO principle). Step 2: Parse the position control information (3D coordinates, Euler angles) in the command.

[0026] 2) Motion planning and execution Step 3: Calculate the robotic arm joint angles based on the target position (inverse kinematics solution based on D-H parameters). Step 4: Perform motion filtering processing (Kalman filter to eliminate jitter, path tracking error ≤ 1mm). Step 5: Drive the robotic arm to move along the planned path and real-time monitor the status of the end effector.

[0027] Furthermore, in other preferred embodiments of the present invention, the S1-2 step includes: S1-21: Perform adaptive beamforming through an array microphone. S1-22: Use the Conformer model for speech recognition.

[0028] Figure 2 Displays a specific and feasible gesture processing process, including: 1) System initialization Step 1: Initialize the robotic arm and set the servo action range (such as horizontal ±180°, pitch ±90°). Step 2: Call init_arm_position to reset the robotic arm to the initial position. Step 3: Initialize the PID controller (set the proportional coefficient Kp = 0.5, integral time Ti = 0.1s).

[0029] Step 4: Set the audio sampling rate (48kHz), bit depth (16bit). Step 5: Create and start the audio stream, and bind the audio_callback callback function.

[0030] 2) Audio data acquisition and preprocessing Step 6: In the audio_callback function, convert audio data in real time (PCM format → floating-point array); Step 7: Calculate the audio volume (RMS energy value) and determine whether it exceeds the noise threshold (default 30dB): Condition judgment: Yes → Calculate the pitch (fundamental frequency extraction, using the YIN algorithm); No → Set the volume and pitch to 0 (mute state).

[0031] 3) Servo control logic Step 8: If the volume < 30dB, keep the robotic arm in its initial position; Step 9: If the volume ≥ 30dB, perform the following operations: Target angle calculation: Generate the target angle of the servo according to the volume size (linear mapping of 0 - 100dB) and pitch (200 - 800Hz); PID smooth output: Calculate the smooth command through the PID controller (suppress overshoot, overshoot amount ≤ 5%); Robotic arm control: Drive the servo to move according to the target angle (response time ≤ 50ms).

[0032] 4) Loop termination condition Step 10: Detect whether the program continues to run (default loop condition is True); Step 11: If a termination signal is triggered (such as an external interrupt), clean up the audio stream resources and exit the program.

[0033] Furthermore, in other preferred embodiments of the present invention, establish the mapping relationship between the emotion label and the control instruction set; perform emotion voiceprint analysis on the recognized speech content and generate emotion labels, such as happy, sad, angry, disappointed, fearful, anxious, etc.; according to the emotion label, call the preset control instruction set. For example, for the emotion label of happy, the robotic arm can swing at a high frequency, combined with cold-colored lights to match the user's mood. Specifically, the swing frequency of the servo can be set to 4 - 6Hz and the LED color temperature > 5000K. For the emotion label of sad, the pitch angle of the three-axis gimbal can be lowered, combined with lowering the brightness to match the user's mood. Specifically, the pitch angle of the three-axis gimbal can be adjusted to -20° to -30°, and the LED brightness is reduced by 30% - 50%.

[0034] The generation of emotion labels combines the conversation content and the conclusion of emotion voiceprint analysis. Emotion voiceprint analysis can generate an emotion intensity value based on the fundamental frequency change rate (ΔF0 / Δt ≥ 5Hz / s). For example, when the emotion intensity value E ∈ [0.7, 1.0], a happy emotion label is generated, and when E ∈ [0, 0.3], a sad emotion label is generated.

[0035] In addition, a mapping relationship between the running duration and the control instruction set can be established; when the running duration of the smart table lamp reaches a preset value, a preset control instruction set is called. At the same time, a RGB-D camera is used to detect the user's posture. When it is detected that the user maintains a sitting posture for a long time, a preset control instruction set is called for active interaction to implement operations such as pushing a water cup, dimming the light, and voice reminder.

[0036] In addition, this control method can also adaptively learn according to user feedback. For example, collect the user's correction instructions for device behavior, including the frequency statistics of voice negative words and the record of manually adjusted parameters; and update the model parameters using an online fine-tuning algorithm (weight coefficient λ = 0.5 ± 0.2) constrained by KL divergence. In addition, a user preference database can also be established, such as storing the association matrix between music types and light color temperatures (dimension ≥ 16×16).

[0037] The following is further described through specific application scenarios. Embodiment 1

[0038] Application scenario: Transferring items Control method: S1. Detect the user's gesture (pointing to the water cup) through a RGB-D camera, collect voice commands (I want to drink water, thirsty) through an array microphone, and detect the specific position of the water cup through a pressure sensor; S2. Input the multi-modal data into the AI large model for feature fusion, calculate the three-dimensional coordinate relationship of the user, the robotic arm, and the water cup, and generate a comprehensive decision vector; S3. Dynamically generate a control instruction set according to the comprehensive decision vector, control the servo motor and the robotic arm to move, and push the water cup in front of the user. Embodiment 2

[0039] Application scenario: Emotional interaction Control method: S1. Detect the user's facial expression (sad) through a RGB-D camera, and collect voice commands (unhappy tone) through an array microphone; S2. Input the multi-modal data into the AI large model for feature fusion, generate an emotion label through emotion analysis, and call a preset control instruction set; S3. According to the control instruction set, control the pitch angle of the three-axis gimbal to be adjusted to -20°, and the LED brightness is reduced by 30%. Embodiment 3

[0040] Application scenario: Active interaction Control method: S1. Detect the user's posture (sitting posture maintained) through a RGB-D camera, and detect the running duration (reaching the threshold) through a timer; S2. Input the multi-modal data into the large AI model for feature fusion, and call the preset control instruction set according to the running duration. S3. According to the control instruction set, the speaker emits synthesized audio for sedentary reminder, the LED brightness is reduced by 20%, and the light source is adjusted to soft warm light. Embodiment 4

[0041] Application scenario: Entertainment interaction Control method: S1. Detect the user's posture (non-working state) through the RGB-D camera, collect music information through the array microphone, and convert the audio signal into a Mel spectrogram. S2. Input the multi-modal data into the large AI model for feature fusion, identify the rhythm type (such as fast-paced EDM, slow-paced jazz) according to the Mel spectrogram; generate a comprehensive decision vector according to the music rhythm. S3. Dynamically generate a control instruction set according to the comprehensive decision vector, control the servo and the robotic arm to swing according to the music rhythm, and the lights flash with the beats.

[0042] In summary, the embodiments of the present invention provide an intelligent table lamp based on a large AI model and its control method. It can collect environmental data through a multi-modal perception module to obtain multi-modal data information; then import the multi-modal data information into the large AI model for feature fusion, generate a comprehensive decision vector including user intention, emotional state and environmental characteristics, generate a control instruction set, and realize the coordinated control of a multi-degree-of-freedom actuator. Thus, more anthropomorphic feedback can be made in the interaction with the user.

[0043] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An intelligent desk lamp based on an AI large model, characterized in that, It includes: A multi-modal perception module for collecting environmental data to obtain multi-modal data; A calculation and control module that processes the multi-modal data through an AI large model to generate a control instruction set, and controls the execution of each specified action in the control instruction set; An actuator module that, under the control of the calculation and control module, completes the specified action.

2. The intelligent table lamp according to claim 1, wherein The multi-modal perception module includes: An RGB-D camera for identifying and collecting gesture instructions of the user; A lidar for ranging and three-dimensional space perception; An array microphone for identifying and collecting voice instructions of the user; A pressure sensor disposed on the base of the smart table lamp for identifying the position and weight information of items on the base.

3. The intelligent table lamp according to claim 1, characterized in that, The actuator module includes: A three-axis gimbal that carries the RGB-D camera, the lidar, and the array microphone, and can adjust the angle according to the angle parameters in the control instruction set; An LED matrix that can adjust the brightness according to the lighting parameters in the control instruction set; A servo motor that can adjust the operating trajectory of the robotic arm according to the angle parameters in the control instruction set; A speaker that can play audio according to the audio parameters in the control instruction set.

4. A control method for an intelligent table lamp based on a large AI model, characterized in that, It includes: S1. Collect environmental data through the multi-modal perception module to obtain multi-modal data, and the multi-modal data includes visual, voice, and physical sensing information; S2. Input the multi-modal data into the AI large model for feature fusion to generate a comprehensive decision vector including user intent, emotional state, and environmental characteristics; S3. Dynamically generate a control instruction set according to the comprehensive decision vector, and the instruction set includes a spatio-temporal synchronization scheme of mechanical action parameters, optical parameters, and audio output instructions; S4. Execute the control instruction set to achieve coordinated control of the multi-degree-of-freedom actuator.

5. The control method according to claim 4, wherein In the step S1, the process of collecting environmental data through the multi-modal perception module includes: S1-1. Obtain the three-dimensional coordinates of the user's skeletal joint points through the RGB-D camera; S1-2. Collect voice instructions through the array microphone; S1-3. Real-time monitor the pressure distribution of the table lamp base through the pressure sensor matrix.

6. The control method according to claim 4, wherein The AI large model includes a cross-modal attention fusion layer, an emotional state calculation layer, and a dynamic decision layer, and the step S2 includes: S2-1. The cross-modal attention fusion layer uses a multi-head attention mechanism to associate visual and voice features; S2-2. The emotional state calculation layer analyzes the fundamental frequency change rate of the voice based on the GRU network to generate an emotional intensity value; S2-3. The dynamic decision layer generates a control instruction queue including timestamps through an improved Transformer decoder.

7. The control method according to claim 5, characterized in that The step S1-1 includes: S1-11. Detect the hand bounding box through an improved YOLOv5 model; S1-12. Extract hand key points based on the OpenPose algorithm and establish a three-dimensional space coordinate system.

8. The control method according to claim 5, characterized in that, The step S1-2 includes: S1-21. Perform adaptive beamforming through the array microphone; S1-22. Use the Conformer model for speech recognition.

9. The control method according to claim 8, characterized in that, Establish a mapping relationship between the emotion tags and the control instruction set; perform emotion voiceprint analysis on the recognized speech content and generate emotion tags; call the preset control instruction set according to the emotion tags.

10. The control method according to claim 4, characterized in that Establish a mapping relationship between the running duration and the control instruction set; when the running duration of the smart table lamp reaches a preset value, call the preset control instruction set.