Lamplight illumination angle intelligent control system and method
By constructing a dynamic three-dimensional semantic scene graph and parameterized lighting bodies, the problem of insufficient dynamic perception of the environment and users in existing lighting systems is solved, precise lighting control and energy optimization are achieved, and the user experience is improved.
Patent Information
- Application Number
- CN202511130563.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing lighting systems lack the ability to perceive the environment and user dynamics in real time and adapt to changes, resulting in wasted lighting resources and poor user experience. This makes it difficult to meet refined lighting needs, especially in complex scenarios where multiple users or multiple tasks coexist.
Through the multimodal 3D perception module, depth information and acoustic data are obtained to construct a dynamic 3D semantic scene graph. The temporal graph convolutional network is combined to identify the user's task intention and generate a parameterized 3D illumination volume. The multi-degree-of-freedom beam projection mechanism is used for precise lighting control, and the closed-loop correction mechanism is combined to optimize the lighting effect.
It achieves a deep understanding of the environment and accurate identification of user tasks. The light beam can be accurately projected to the area required for the task, reducing ineffective lighting, improving energy utilization efficiency, optimizing the user's light environment experience, and ensuring the stability and accuracy of the lighting effect.
Smart Images

Figure CN120640491A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent control technology, and in particular relates to an intelligent control system and method for lighting angles. Background Art
[0002] The field of lighting control technology includes multiple levels such as traditional manual control, automatic control, centralized management systems and intelligent lighting systems. The core content of this technology field is to meet the visual needs in specific scenarios, improve human comfort and achieve energy conservation through effective management of light sources. Intelligent lighting refers to an advanced control form in which the system can sense the environmental status through sensors, automatically adjust parameters such as brightness, color temperature and lighting angle, and interact with users or other devices. With the development of Internet of Things technology and sensor technology, intelligent lighting technology has been widely used in many scenarios such as homes, commercial buildings, urban landscapes and industrial production. The development of this field has not only improved energy utilization efficiency, but also brought significant value in improving user experience and ensuring safety.
[0003] Among them, the lighting angle control method refers to the technology of accurately projecting the light beam to a specified area by adjusting the physical orientation of the lamp. This technology aims to dynamically adjust the lighting coverage according to changes in the user's position, work plane or specific target object, thereby achieving precise lighting. In this method, the control system changes the pitch angle and azimuth angle of the light source by driving the mechanical structure so that the light can be focused on the effective area and avoid ineffective illumination of non-target areas. This directional lighting capability plays a key role in maximizing light energy utilization, avoiding glare interference, and creating a specific light environment. This method also improves the flexibility and adaptability of lighting by optimizing the angle adjustment strategy.
[0004] Existing technologies for lighting angle control mostly rely on preset modes or manual adjustment, lacking real-time perception and adaptive adjustment capabilities for environmental and user dynamics, making it difficult to meet personalized and refined lighting needs. For example, when a user moves around a room, fixed-angle lighting cannot follow the user's movement, resulting in localized areas being too bright or too dark. Angle control logic is often based on a simple trigger mechanism, unable to predict user behavior and intention, resulting in delayed adjustment response and a poor user experience. The lack of comprehensive analysis of ambient lighting prevents angle adjustment from being effectively linked to brightness control, resulting in energy waste. Control systems are typically open-loop or simple closed-loop, unable to dynamically calibrate based on actual lighting feedback, resulting in low control accuracy. In complex scenarios with multiple users or multiple tasks, existing technologies struggle to effectively divide lighting areas and coordinate angles. This problem is particularly prominent in smart homes or modern office spaces that require high flexibility and comfort, directly impacting the intelligence level and energy efficiency of the lighting system. Summary of the Invention
[0005] One purpose of the present invention is to provide an intelligent control system and method for lighting angles, aiming to solve the technical problem in the prior art that lighting control strategies are out of touch with users' actual task requirements, resulting in waste of lighting resources and poor user experience.
[0006] In order to achieve the above object, the present invention provides a method for intelligently controlling lighting angles, the method comprising the following steps:
[0007] Acquire scene point cloud data containing depth and color information collected in real time by the multimodal 3D perception module, as well as scene acoustic event data collected synchronously by the microphone array;
[0008] Based on the scene point cloud data, a 3D instance segmentation network model is called to segment and identify independent physical entities in the scene, and combined with acoustic event data to construct a dynamic 3D semantic scene graph containing object categories, 3D poses, and state attributes;
[0009] Based on the dynamic three-dimensional semantic scene graph, a temporal graph convolutional network is called to perform temporal analysis on the user's skeleton key point sequence and the object interaction relationship, identify the task category currently performed by the user, and generate a task intention label;
[0010] According to the task intention label, query the target object set associated with the current task in the dynamic three-dimensional semantic scene graph, call a preset illumination volume generation rule library, and calculate and generate a parameterized three-dimensional illumination volume representing the target lighting area;
[0011] Converting the geometric and optical parameters of the parameterized three-dimensional illumination object into underlying hardware control instructions for a multi-degree-of-freedom light beam projection mechanism, driving the mechanism to adjust the projection angle, convergence, and color properties of the light beam to precisely match the three-dimensional illumination object;
[0012] The actual lighting effect after adjustment of the multi-degree-of-freedom beam projection mechanism is fed back to the multimodal three-dimensional perception module, the deviation between the actual lighting area and the three-dimensional lighting body is calculated through image difference and geometric comparison, and the hardware control instructions are closed-loop corrected.
[0013] As an embodiment of the present invention, the steps of constructing the dynamic three-dimensional semantic scene graph specifically include:
[0014] A time-of-flight (Time-of-Flight) sensor is fused with a red, green, and blue (RGB) image sensor to generate a three-dimensional point cloud data stream with color information. A sparse convolutional neural network is applied to the point cloud data stream to segment the point cloud into independent instance clusters corresponding to each physical entity in the scene. For each instance cluster, a three-dimensional classification network is used to identify its object category, and an iterative closest point algorithm is used to accurately calculate its six-degree-of-freedom pose information in the world coordinate system. The pose information includes three-dimensional coordinates (x, y, z) and attitude angles (roll, pitch, yaw). Simultaneously, a sound source localization algorithm based on the time difference of arrival (TDOA) is used to process the microphone array data to determine the three-dimensional spatial location of acoustic events such as tapping and page turning. All identified object categories, pose information, and acoustic events with spatial locations are integrated into a unified data structure indexed by timestamps, namely the dynamic three-dimensional semantic scene graph.
[0015] As an embodiment of the present invention, the step of identifying the category of the task currently being performed by the user specifically includes:
[0016] A three-dimensional coordinate sequence of skeletal key points associated with the user is extracted from the dynamic three-dimensional semantic scene graph; a time sequence graph is constructed with the skeletal key points as nodes and limb connection relationships as edges; at the same time, the Euclidean distance and relative posture between the user's skeletal key points (such as hands) and other objects in the scene (such as keyboards and books) are extracted as dynamic attributes of the graph; this time sequence graph with dynamic attributes is input into a pre-trained time sequence graph convolutional network; the network outputs a probability distribution representing the current task category by learning the evolution pattern of the user's posture and object interaction in the time dimension, and selects the category with the highest probability as the task intention label. The task categories include reading, writing, computer operation or environmental interaction.
[0017] As an embodiment of the present invention, the step of calculating and generating the parameterized three-dimensional illumination volume specifically includes:
[0018] The parameterized three-dimensional illumination volume is uniquely determined by a set of parameters V=(P, D, S, Φ, C), where P is the three-dimensional position vector of the center point of the illumination volume, D is the unit vector of the illumination direction, S is the shape and size parameters of the illumination volume, Φ is the target luminous flux, and C is the target correlated color temperature. According to the task intention label, the corresponding illumination volume generation function G_model is activated; the function takes the current task category, user posture, and the posture and size of the associated target object as input, that is, ; For example, when the task category is "reading", the function determines the surface center of the target object "book" as P, the surface normal vector as D, sets the shape S to an elliptical cone proportional to the size of the book, and sets the luminous flux Φ and color temperature C to the preset 500 lux and 3000 Kelvin.
[0019] As an embodiment of the present invention, the method further includes a task intention prediction step:
[0020] After identifying the current task category, the task intent label sequence of the past N time steps is input into a long short-term memory (LSTM) network; the LSTM network is used to learn the typical patterns and probabilities of users switching between different tasks; the network outputs a probability of each task category occurring within the next M time steps; when the predicted probability of a future task exceeds a preset switching threshold, the system calculates and caches the parameterized three-dimensional illumination volume corresponding to the future task in advance, and generates a smooth transition path between the illumination volumes of the two tasks to achieve seamless switching of lighting effects and avoid visual mutations.
[0021] As an embodiment of the present invention, the steps of converting and driving the hardware control instructions specifically include:
[0022] The position vector P and direction vector D of the parameterized three-dimensional illumination body are solved through an inverse kinematics model and converted into target rotation angles for driving the horizontal stepper motor and the vertical stepper motor in the multi-degree-of-freedom beam projection mechanism; the data representing the beam divergence angle in the illumination body shape parameter S is converted into a control voltage value applied to the electrically controlled tunable lens, or the number of step pulses for driving the micromotor in the mechanical zoom component; and the target luminous flux Φ and the target correlated color temperature C are decomposed into the duty ratios of the pulse width modulation (PWM) signals for the warm white light LED array and the cool white light LED array by consulting a preset chromaticity-luminance lookup table.
[0023] As an embodiment of the present invention, the closed-loop correction step specifically includes:
[0024] After the multi-degree-of-freedom beam projection mechanism completes adjustment, a frame of target scene image is captured; by performing pixel-level differential operation on the frame image and the reference image captured immediately before the adjustment, the area with significant brightness change is extracted, which is the actual illumination area; the centroid, circumscribed rectangle and main axis direction of the actual illumination area are calculated; the geometric parameters of the three-dimensional illumination body V are projected onto a two-dimensional plane under the current camera viewing angle to form a target illumination area; the intersection over union (IoU) indicator is used to quantify the geometric deviation between the actual illumination area and the target illumination area. ; Using a proportional-integral-derivative (PID) controller, based on the deviation Generate a set of corrections , and its calculation formula is
[0025]
[0026] Wherein Kp, Ki, and Kd are proportional, integral, and differential coefficients, respectively; the correction amount is superimposed on the original hardware control instruction to form a new control instruction and sent to the multi-degree-of-freedom beam projection mechanism again, and iteratively executed until the deviation is less than the preset tolerance.
[0027] The present invention also provides an intelligent lighting angle control system, which is used to execute any of the above methods and includes:
[0028] a multimodal 3D perception module integrating a time-of-flight depth sensor, a red, green, and blue image sensor, and a MEMS microphone array, the module being configured to simultaneously output a 3D point cloud data stream of a scene with color information and multi-channel acoustic data;
[0029] a dynamic scene semantic parsing module, coupled to the multimodal 3D perception module, configured to receive point cloud and acoustic data and, by running a 3D instance segmentation algorithm and a sound source localization algorithm, construct and maintain a dynamic 3D semantic scene graph representing objects, users, and their spatial relationships within the scene;
[0030] a task intention recognition and prediction module, coupled to the dynamic scene semantic parsing module, configured to analyze the temporal evolution of the user's skeletal posture and object interaction in the scene graph to identify the current task category and predict the probability of future task switching using historical task sequences;
[0031] an illumination volume parameter generation module, coupled to the task intention recognition and prediction module, configured to retrieve associated objects from the scene graph according to a current or predicted task category, and calculate a parameterized three-dimensional illumination volume of a target lighting area according to a preset mathematical model;
[0032] a multi-degree-of-freedom beam projection module comprising a dual-axis stepper motor-driven pan / tilt stage, an electrically controlled tunable zoom lens system, and an LED light source array with adjustable color temperature, the module being configured to receive illumination object parameters and convert them into precise physical drives for various hardware components;
[0033] A closed-loop performance correction module is coupled to the multimodal three-dimensional perception module and the multi-degree-of-freedom beam projection module. The module is configured to evaluate the deviation between the actual lighting effect and the target lighting object through image comparison, and generate correction instructions using a PID control algorithm to dynamically fine-tune the output of the beam projection module.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] By constructing a dynamic three-dimensional semantic scene graph, this invention enables lighting systems to, for the first time, possess a deep understanding of their environment, surpassing traditional sensors that merely sense presence. Rather than simply tracking a moving target, the system can accurately identify the user's specific task intent, such as "reading" or "writing," thereby achieving a fundamental shift from "illuminating people" to "illuminating things." Furthermore, by introducing the concept of a parameterized three-dimensional illumination volume, this invention transforms vague lighting requirements into precise, calculable geometric and optical targets, enabling unprecedented precision in lighting control. Light beams can be precisely shaped and projected into the three-dimensional spatial region required for the task, significantly reducing lighting ineffective areas and significantly improving energy efficiency. Furthermore, the task intent prediction capability based on a long-short-term memory network enables the system to anticipate changes in user behavior and smoothly transition the lighting layout in advance, avoiding visual discomfort caused by sudden changes in light intensity and optimizing the user's light environment experience. Finally, a closed-loop correction mechanism ensures high consistency between theoretical calculations and physical reality. The system can autonomously compensate for deviations caused by mechanical errors or environmental changes, ensuring consistent, stable, and precise lighting effects. This completes an intelligent closed loop of perception, understanding, decision-making, execution, and feedback. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 A schematic diagram of the overall process of a lighting angle intelligent control method provided by an embodiment of the present invention;
[0037] Figure 2 A schematic diagram of the functional modules of an intelligent lighting angle control system provided by an embodiment of the present invention;
[0038] Figure 3 for Figure 1 A detailed flowchart of the steps for constructing a dynamic 3D semantic scene graph;
[0039] Figure 4 for Figure 1 Detailed flow chart of the closed-loop calibration steps;
[0040] Figure 5 A schematic diagram of an application scenario provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solutions and advantages of the present invention clearer, the specific implementation methods of the present invention will be described in detail with reference to the accompanying drawings. It should be noted that the described embodiments are some embodiments of the present invention, rather than all embodiments. In the description of the present invention, the orientations or positional relationships indicated by the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc. are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, in the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.
[0042] Reference Figure 1 The embodiment of the present invention provides a method for intelligently controlling lighting angles, the overall process of which includes the following steps:
[0043] Step S1: Acquire scene point cloud data containing depth information and color information collected in real time by a multimodal three-dimensional perception module, and scene acoustic event data collected synchronously by a microphone array.
[0044] Step S2: Based on the scene point cloud data, a three-dimensional instance segmentation network model is called to segment and identify independent physical entities in the scene, and combined with acoustic event data to construct a dynamic three-dimensional semantic scene graph containing object categories, three-dimensional poses and state attributes.
[0045] Step S3: Based on the dynamic three-dimensional semantic scene graph, a temporal graph convolutional network is called to perform temporal analysis on the user's skeleton key point sequence and the object interaction relationship, identify the task category currently performed by the user, and generate a task intention label.
[0046] Step S4: according to the task intention label, query the target object set associated with the current task in the dynamic three-dimensional semantic scene graph, call the preset illumination body generation rule library, and calculate and generate a parameterized three-dimensional illumination body representing the target lighting area.
[0047] Step S5, converting the geometric and optical parameters of the parameterized three-dimensional illumination object into underlying hardware control instructions of a multi-degree-of-freedom light beam projection mechanism, driving the mechanism to adjust the projection angle, convergence degree and light color properties of the light beam to accurately match the three-dimensional illumination object.
[0048] In step S6, the actual illumination effect after adjustment of the multi-degree-of-freedom beam projection mechanism is fed back to the multimodal three-dimensional perception module, the deviation between the actual illumination area and the three-dimensional illumination body is calculated through image difference and geometric comparison, and the hardware control instruction is closed-loop corrected.
[0049] Reference Figure 2 An embodiment of the present invention provides an intelligent lighting angle control system, whose functional modules include a multimodal three-dimensional perception module, a dynamic scene semantic analysis module, a task intent recognition and prediction module, an illuminated body parameter generation module, a multi-degree-of-freedom beam projection module, and a closed-loop performance correction module. The system is configured to perform the aforementioned method.
[0050] Each step of the method is described in detail below.
[0051] In a specific embodiment, step S1, i.e., the step of obtaining scene point cloud data and acoustic event data, can be specifically decomposed into the following sub-steps:
[0052] Step S101 activates the Time-of-Flight (ToF) depth sensor and red, green, and blue (RGB) image sensor in the multimodal 3D perception module to generate a 3D point cloud data stream that incorporates color information. Specifically, the ToF sensor emits near-infrared light (e.g., a vertical cavity surface emitting laser array with a central wavelength of 940 nanometers) and measures the round-trip time of the light to calculate the depth value of each pixel. The ToF sensor has a depth resolution of 640x480 pixels, a depth measurement accuracy of ±5 mm, a maximum range of 5 meters, and a data output frame rate of 30 Hz. At the same time, the RGB image sensor, such as a 1 / 2.3-inch complementary metal oxide semiconductor (CMOS) sensor, captures the color information of the scene at a resolution of 1920x1080 pixels, and its output frame rate is also set to 30 Hz to match the ToF sensor.
[0053] Furthermore, in order to generate a 3D point cloud with color information, the ToF sensor and RGB sensor need to be calibrated with internal and external parameters. Intrinsic calibration is used to correct the lens distortion of each sensor, while extrinsic calibration is used to determine the rigid transformation relationship between the two sensor coordinate systems (a rotation matrix R and a translation vector T). Extrinsic calibration can be completed by taking image pairs of a checkerboard calibration plate at multiple different poses. After calibration, for each depth pixel output by the ToF sensor And its depth value d, we can first use the intrinsic parameter matrix K_d of the ToF sensor to back-project it into the three-dimensional coordinate system of the ToF sensor to obtain the three-dimensional point Then, using the external reference (R, T) Transformed to the coordinate system of the RGB sensor, we get Finally, using the intrinsic parameter matrix of the RGB sensor Will Project to the RGB image plane to get the corresponding color pixel coordinates , and extract the (r, g, b) value at that coordinate. Ultimately, the data structure of each point is defined as Point = {x, y, z, r, g, b, timestamp}, where (x, y, z) are the three-dimensional coordinates in the world coordinate system (obtained by aligning the sensor coordinate system with the predefined world coordinate system) and are 32-bit floating-point numbers; (r, g, b) are the color components of 8-bit unsigned integers; and timestamp is a 64-bit integer representing the acquisition timestamp generated by a system clock synchronized with a Network Time Protocol (NTP) server, with microsecond accuracy.
[0054] Step S102 synchronously activates the microelectromechanical system (MEMS) microphone array in the multimodal 3D perception module to collect multi-channel acoustic data. In one specific embodiment, the microphone array consists of four omnidirectional MEMS microphones arranged in a square planar layout with a side length of 5 cm. The sampling rate of each microphone is set to 48 kHz, and the quantization bit depth is 24 bits. The collected four-channel audio data stream is assigned the same synchronization timestamp as the corresponding point cloud data to ensure precise temporal alignment of acoustic events with the visual scene.
[0055] As an exception handling mechanism, the system continuously monitors the status of the two sensors. If no new point cloud frame or audio frame is received within a preset timeout threshold (for example, 100 milliseconds), the system will determine that the corresponding sensor has failed and enter a degraded working mode or issue a maintenance alert to the user. For example, if the ToF sensor fails, the system can attempt to perform two-dimensional target detection and tracking based only on RGB images to provide basic lighting tracking capabilities. If a microphone fails, the sound source localization algorithm can automatically switch to a positioning model using the remaining three microphones, although the positioning accuracy will be reduced.
[0056] Reference Figure 3 Step S2, i.e., constructing a dynamic 3D semantic scene graph, is the core of scene understanding. This step receives the output of step S1 as input and generates a structured scene representation that is dynamically updated over time. Specifically, this step includes:
[0057] Step S201: pre-process the original point cloud data stream fused with color information. Since the original point cloud data is huge and contains noise, pre-processing is necessary.
[0058] First, voxel downsampling is performed. The entire 3D space is divided into a cubic voxel grid with a side length of 1.0 cm. For all points falling within the same voxel, the average of their 3D coordinates and color components is calculated, and this average point is used as the representative point of that voxel. This process significantly reduces the density of the point cloud while preserving the macroscopic structure of the scene. For example, a frame of the original point cloud containing 2,073,600 points (not all pixels have valid depth at a 1920x1080 resolution) can be reduced to 350,000 points after voxel downsampling, greatly reducing the computational burden of subsequent processing.
[0059] Next, statistical outlier removal is performed. For each point in the downsampled point cloud, the average distance to its 50 nearest neighbors is calculated. The global mean μ and standard deviation σ of these average distances are then calculated. Any point with an average neighbor distance greater than μ + 1.0*σ is considered an outlier and removed. This method effectively removes isolated noise points caused by sensor measurement errors or multipath reflections.
[0060] Step S202: Perform 3D instance segmentation on the preprocessed point cloud. This step aims to segment the point cloud into multiple clusters, with each cluster corresponding to an independent physical entity in the scene. This embodiment uses a sparse convolutional neural network (SCNN)-based model, pre-trained on a dataset containing a large amount of annotated indoor scene data (e.g., ScanNet). The input is the preprocessed point cloud, and the output is the instance label and semantic label for each point.
[0061] For example, for a scene containing a desk, a chair, a user, and a book, the segmentation network will output four main instance clusters. All points belonging to the desk are assigned instance ID 1 and the semantic label "desk"; all points belonging to the chair are assigned instance ID 2 and the semantic label "chair"; all points belonging to the user are assigned instance ID 3 and the semantic label "person"; and all points belonging to the book are assigned instance ID 4 and the semantic label "book".
[0062] As an alternative, 3D instance segmentation can also be achieved using non-deep learning methods. For example, a segmentation method based on region growing is used. This method first calculates local geometric features, such as the normal vector and curvature, for each point in the point cloud. Then, a randomly selected point in the point cloud with low curvature that has not yet been assigned to any instance is used as a seed point, and a new instance region is created. Next, the neighborhood of the boundary point of this region is iteratively checked. If the points in the neighborhood are similar to points in the region in terms of normal vector direction (normal vector dot product greater than 0.95) and color (Euclidean distance in CIELAB color space less than 15.0), the neighboring points are merged into the current instance region. This process is repeated until no more points can be added to any existing regions. This method has the advantage of not requiring large amounts of training data, but its segmentation performance is very sensitive to the settings of parameters (such as normal vector and color similarity thresholds) and is prone to under-segmentation when dealing with physically touching or closely adjacent objects.
[0063] As another alternative, density-based clustering algorithms can be used.
[0064] For example, DBSCAN (Density-Based Spatial Clustering of Applications with Noise). This algorithm performs clustering based on two parameters: the neighborhood radius ε and the minimum number of neighbors required to form a core point, MinPts. For example, set ε to 3.0 cm and MinPts to 50. The algorithm starts from any point. If the number of points in the ε neighborhood of the point is not less than MinPts, the point is marked as a core point and a new cluster is created. Then, all points (including other core points) that are density-reachable from the core point are recursively added to the cluster. The algorithm can discover clusters of arbitrary shapes and is insensitive to noise. Its main limitation is that for different objects with significant density differences in the scene, it is difficult to achieve ideal segmentation results with a single ε and MinPts parameter.
[0065] In step S203 , for each segmented instance cluster, the precise six-degree-of-freedom (6-DoF) pose information and other attributes thereof are extracted.
[0066] Specifically, for each instance cluster identified as a specific category (such as "book", "keyboard", "cup"), the system retrieves the corresponding standard model from a preset digital 3D model library. The model library pre-stores standardized CAD models or high-precision scanned models of various common objects, each of which is defined in its own local coordinate system. Subsequently, the Iterative Closest Point (ICP) algorithm is used to calculate the optimal rigid transformation required to align the standard model to the observed point cloud instance cluster. The ICP algorithm solves a 4x4 homogeneous transformation matrix by iteratively finding the nearest corresponding points between points on the standard model and the instance cluster, and minimizing the average square distance between these corresponding point pairs.
[0067] Taking instance ID4 (semantic label "book") as an example, its point cloud cluster contains about 1500 points. The system loads the "standard_book.ply" model from the model library. The ICP algorithm converges after 25 iterations, and the resulting transformation matrix is From this matrix, the translation vector can be decomposed =[1.352,-0.411,0.850] (unit: meter), which represents the position of the center of the book in the world coordinate system; and a rotation matrix, which can be further converted to Euler angles =[0.05,0.0,-1.57] (unit: radian), which represents the roll angle, pitch angle and yaw angle around the X, Y and Z axes respectively. In addition, by calculating the axis-aligned bounding box (AABB) of the instance cluster point cloud, the size information of the object can be obtained, such as =[0.210,0.297,0.030] m.
[0068] Step S204 : Processing the multi-channel acoustic data collected in step S102 in parallel to locate transient acoustic events in the scene.
[0069] First, a sound event detection model based on a convolutional recurrent neural network (CRNN) continuously analyzes the input four-channel audio stream. The model is trained to recognize specific task-related sounds, such as keyboard tapping, mouse clicking, turning pages of a book, or the sound of an object being placed on a table. When the model detects an event, such as at a timestamp When a "knocking" sound is detected and its confidence level is higher than 0.8, the system triggers the sound source localization process.
[0070] Next, the time difference of arrival (TDOA) between microphone pairs is calculated using a generalized cross-correlation-phase transform (GCC-PHAT) algorithm. For an array of four microphones, multiple sets of pairwise TDOA values can be calculated, such as , , , Etc. Each TDOA value A hyperboloid is defined with microphones i and j as the foci. The sound source must lie on this hyperboloid. By solving the intersection of these hyperboloid equations, the 3D spatial position of the sound source can be estimated.
[0071] For example, for the detected "knocking" sound, the calculated TDOA value is,
[0072] =-0.042ms, =0.088ms, =0.047ms. The overdetermined equations are solved by the least squares method to obtain the sound source position =[0.955,-0.251,0.848] m, and its positioning error is estimated to be ±2 cm.
[0073] In complex acoustic environments with strong echoes or reverberation, the GCC-PHAT algorithm's cross-correlation peaks can become blurred, resulting in reduced positioning accuracy. In such cases, the system calculates the signal-to-noise ratio of the cross-correlation peaks. If the signal-to-noise ratio falls below a preset threshold (e.g., 5dB), the localization result is considered unreliable and the system discards the acoustic event to avoid introducing erroneous sound source information into the scene graph.
[0074] Step S205: Integrate and update all extracted visual and acoustic information into a unified dynamic three-dimensional semantic scene graph. The scene graph is a directed graph data structure G=(N, E), where N is a node set and E is an edge set.
[0075] A node n∈N represents an entity in the scene. Each node contains the following fields:
[0076] node_id: A globally unique string identifier, such as object-uuid-1234.
[0077] node_type: enumeration type, optional values are STATIC_OBJECT (such as wall, floor), DYNAMIC_OBJECT (such as book, cup), USER, ACOUSTIC_EVENT.
[0078] timestamp_created: The timestamp when the node was first detected.
[0079] timestamp_updated: The timestamp when the node attribute was last updated.
[0080] attributes: A key-value pair collection that stores the specific attributes of the node.
[0081] For nodes of type DYNAMIC_OBJECT, the attributes include:
[0082] {semantic_label:'book',pose_6dof:Pose_object,dimensions:[w,h,d],point_cluster_ref:...}.
[0083] For nodes of type USER, the attributes include {skeleton_3d:Skeleton_data,...}.
[0084] For nodes of type ACOUSTIC_EVENT, the properties include:
[0085] {event_label:'knock',location_3d:[x,y,z],confidence:0.92}. This type of node is transient and will be automatically removed after a short life cycle (such as 1 second) after creation.
[0086] An edge e∈E represents a relationship between nodes. Each edge contains the following fields:
[0087] edge_id: A globally unique string identifier.
[0088] source_node_id: The ID of the starting node.
[0089] target_node_id: The ID of the target node.
[0090] relationship_type: enumeration type, optional value is,
[0091] SPATIAL_PROXIMITY,INTERACTION,SUPPORTED_BY.
[0092] attributes: The attributes of the relationship. For example, for the SPATIAL_PROXIMITY relationship, the attributes can be {distance:0.15}, which means that the Euclidean distance between the centroids of the two objects is 0.15 meters.
[0093] The scene graph is updated at a frequency of 10 Hz. In each update cycle, the system associates the perception results of the new frame with the existing nodes in the graph. For existing objects, if they are still detected in the current frame, their attributes such as pose and timestamp_updated are updated. If a DYNAMIC_OBJECT is not detected for 5 consecutive frames, its status is marked as LOST. If it does not appear again for 30 consecutive frames (3 seconds), the node is removed from the graph. For newly detected objects, new nodes are created in the graph. At the same time, relationship edges such as SPATIAL_PROXIMITY are dynamically created or updated based on the spatial position relationship between nodes.
[0094] Step S3, identifying the user's current task category, aims to infer high-level semantic intent from underlying geometric and motion information. This step performs temporal analysis based on the dynamic 3D semantic scene graph generated in step S2.
[0095] Step S301 extracts a 3D coordinate sequence of user-related skeletal key points from a dynamic 3D semantic scene graph. Specifically, the system first locates a node with a node_type of USER in the scene graph. Then, for the point cloud cluster associated with this node, a pre-trained 3D human pose estimation algorithm model is invoked. This model inputs the user's point cloud data and outputs a set (e.g., 17) of 3D skeletal key points, including the head, neck, shoulders, elbows, wrists, hips, knees, and ankles.
[0096] For example, at time step t, the extracted user skeleton data structure is , All coordinates are expressed in the world coordinate system. The system performs this operation continuously at a frequency of 10 Hz and stores the skeleton data sequence of the past T time steps (for example, T = 30, that is, 3 seconds) in a fixed-size ring buffer.
[0097] Step S302: Construct a spatio-temporal graph. This graph uses the user's skeletal keypoints as nodes and the natural limb connections of the human body as internal edges. For example, in each frame, there is an edge between the left shoulder keypoint and the left elbow keypoint, and an edge between the left elbow and the left wrist. These internal edges are topologically fixed.
[0098] Furthermore, to capture user interaction with the environment, dynamic inter-body edges are introduced. The system calculates the Euclidean distance between the user's hand keypoints (such as the left and right wrists) and the center of mass of other dynamic objects in the scene graph (such as "book" and "keyboard"). When the distance between a hand keypoint and an object is less than an interaction threshold d_interact (for example, 15 cm), a temporary inter-body edge is established between the hand keypoint node and the virtual node representing the object.
[0099] The graph's node and edge attributes also change dynamically over time. The attributes of each skeletal keypoint node are its 3D coordinates (x, y, z). The attributes of interaction edges can be the precise distance and relative pose vector between the hand and the object.
[0100] In step S303, the constructed temporal spatial graph sequence is input into a pre-trained temporal graph convolutional network (TGCN). The core idea of this network is to aggregate the spatial information within each node's neighborhood through graph convolution operations, and to learn the evolution pattern of these aggregated features in the temporal dimension through a gated recurrent unit (GRU) or long short-term memory network (LSTM) unit.
[0101] Specifically, at each time step, the graph convolutional layer updates the feature representation of each skeletal keypoint, so that it not only contains its own position information, but also incorporates the relative position information of other connected limbs and the objects being interacted with. This updated whole-image feature is then fed into the GRU unit and combined with the hidden state of the previous time step to capture the dynamic changes of the action.
[0102] For example, when a user performs a "writing" task, their hand key points will remain near the "book" object for a long time, accompanied by subtle, rhythmic movements. TGCN recognizes the "writing" task by learning the persistence of this "hand-book" interaction edge and the specific motion trajectory pattern of the hand nodes. When a user performs "computer operations", the key points of their hands will alternately interact with the "keyboard" and "mouse" objects at close range, and TGCN will learn these different spatiotemporal interaction patterns.
[0103] The last layer of the network is a fully connected layer, followed by a Softmax activation function, which outputs a probability distribution covering all predefined task categories. The predefined task categories may include:
[0104] READING, WRITING, PC_OPERATION, ENVIRONMENTAL_INTERACTION (e.g., turning on and off a light, getting a water cup), IDLE. The system selects the category with the highest probability as the task intent label A_t at the current moment. For example, at a certain moment, if the network output is {READING: 0.1, WRITING: 0.85, PC_OPERATION: 0.04, ...}, the current task intent label is determined to be WRITING. To smooth the output and avoid label jitter caused by transient noise, the system can use a majority voting strategy, selecting the label with the highest number of occurrences from the most recent K (e.g., 5) output labels as the final stable output.
[0105] Step S4, i.e., the step of calculating and generating a parameterized three-dimensional illumination volume based on the task intent label, is the core of converting abstract task requirements into precise and computable geometric and optical targets.
[0106] In step S401, based on the task intent label A_t generated in step S303, the target object set associated with the current task is searched in the dynamic 3D semantic scene graph. This query process is implemented through a predefined task-object association rule library. For example, the rule library is stored in the form of key-value pairs:
[0107] READING: Query nodes whose semantic_label is book, e-reader, or document.
[0108] WRITING: Query nodes whose semantic_label is notebook or paper, and also query nodes whose semantic_label is pen or pencil and have an interaction relationship with the key points of the user's hand.
[0109] PC_OPERATION: Query nodes whose semantic_label is keyboard and mouse.
[0110] ENVIRONMENTAL_INTERACTION: Query the DYNAMIC_OBJECT node that is closest to the user's hand key point and whose distance is less than the interaction threshold d_interact.
[0111] The query result is a set of node IDs {node_id_1, node_id_2,...} of one or more target objects. The system then extracts the 3D pose Pose_obj and size Size_obj attributes of these nodes from the scene graph.
[0112] In step S402, the corresponding illumination volume generation function G_model is activated based on the task intent label A_t. Each predefined task category is mapped to a unique parameterized function. This function takes as input the current task category, the user's pose Pose_user, and the pose and size {Pose_obj, Size_obj} of one or more target objects retrieved from step S401, and outputs a 3D illumination volume uniquely defined by the parameter set V = (P, D, S, Φ, C).
[0113] Step S403: Execute specific calculation of the illuminated body parameters.
[0114] In a specific embodiment, when the task intention label A_t is READING and the target object is a book, the calculation logic of the generated function G_model_reading is as follows:
[0115] Calculate the illuminated center point P: Use the center of mass of the target object "book," P_book, as the initial center point. Furthermore, to ensure that the light is perpendicular to the book surface, translate this center point a small distance (e.g., 2 cm) in the negative direction of the book surface normal to ensure that the light projection mechanism is not blocked by the book itself.
[0116] Calculate the illumination direction D: Use the surface normal vector of the target object "book" point cloud cluster as the illumination direction. This normal vector can be obtained by plane fitting (for example, using the RANSAC algorithm) on the point cloud of the book instance cluster. The resulting unit normal vector is D.
[0117] Calculate the shape and size S of the illumination volume: Set the illumination volume shape to an elliptical cone. The major and minor axes of the base ellipse of the elliptical cone are proportional to the width and height of the book. The proportional coefficient k_size (for example, 1.1) ensures that the light spot completely covers the book with a small margin. The cone's angle is determined by the distance from the light source to the book and the spot size. Therefore, the parameter S is defined as a structure containing information such as the shape type, the major and minor axes of the base, and the cone's angle.
[0118] Setting the target luminous flux Φ and target correlated color temperature C: Based on ergonomic standards, the recommended illuminance for reading tasks is 300 to 500 lux. The system consults a preset parameter table and sets the target luminous flux Φ to a lumen value that produces 450 lux at the target distance. Simultaneously, the target correlated color temperature C is set to 4000 Kelvin, providing a neutral white light environment that facilitates concentration.
[0119] In another specific embodiment, when the task intention label A_t is PC_OPERATION, the target object set includes keyboard and mouse. The calculation logic of its generation function G_model_pc is as follows:
[0120] Calculate the center point P of the light source: Calculate the weighted average position of the centroid P_keyboard of the keyboard and the centroid P_mouse of the mouse as the center of the illumination area. The weights can be set according to the object size or usage frequency.
[0121] Calculate the illumination direction D: Usually, the desktop is horizontal. Therefore, the illumination direction D can be set as a slightly inclined vector pointing in the user direction, such as [0.1, 0.1, -0.98], to avoid glare on the monitor.
[0122] Calculate the shape and size S of the light source: Calculate the smallest bounding box that can contain both the keyboard and the mouse. The shape of the light source is set as a rectangular pyramid, and the bottom size is the size of the bounding box multiplied by an expansion coefficient (such as 1.2).
[0123] Set the target luminous flux Φ and the target correlated color temperature C: Computer operations usually require a softer and wider ambient light. Therefore, the target luminous flux Φ is set to produce a level of 150 lux, and the target correlated color temperature C is set to 3500 Kelvin.
[0124] Step S404, implement the prediction of the task intention and the smooth transition of the illumination. After identifying the current task A_t, the system inputs the sequence of task intention labels {A_{t - N + 1},..., A_t} in the past N time steps (for example, N = 50, that is, 5 seconds) into a long short-term memory network (LSTM). This network is trained offline to learn the typical patterns of the user switching between different tasks. For example, the network may learn that after the READING task, there is a relatively high probability of switching to the WRITING or PC_OPERATION task.
[0125] The LSTM network outputs a probability matrix of the occurrence of each task category within the next M time steps (for example, M = 20, that is, 2 seconds). When the predicted probability of a certain future task A_future continuously exceeds a preset switching threshold (such as 0.7) within the next m steps (m < M), the system triggers the preparatory switching process.
[0126] The system calls the corresponding illumination volume generation function G_model_future in advance to calculate the parameterized three-dimensional illumination volume V_future corresponding to the future task. Subsequently, the system calculates a smooth transition path from the current illumination volume V_current to the future illumination volume V_future. This path is implemented by linear or spline interpolation of all parameters (P, D, S, Φ, C) of the two illumination volumes. The interpolation process is designed to be completed within m time steps. In this way, when the user actually completes the task switch, the lighting effect has already smoothly and seamlessly transitioned to the new state, greatly avoiding visual discomfort and distraction caused by sudden changes in light.
[0127] Step S5, which converts the parameters of the parameterized three-dimensional illumination body into underlying hardware control instructions, is a bridge connecting software decision-making and physical execution.
[0128] Step S501: Perform inverse kinematics to solve the beam projection angle. The core of the multi-DOF beam projection mechanism is a dual-axis (pan and tilt) stepper motor-driven gimbal. The target position vector P and direction vector D in the world coordinate system need to be converted into the gimbal's translational rotation angle. and pitch rotation angle .
[0129] First, the target point P in the world coordinate system is transformed to the gimbal base coordinate system through a fixed transformation matrix T_world_to_base to obtain P_base. Available through The projection calculation on the XY plane is: .
[0130] Then, the direction vector D is also converted to the gimbal base coordinate system to obtain D_base. is the direction vector The angle with the XY plane can be determined by Calculated.
[0131] The calculated target angles θ_pan and θ_tilt are converted into the number of pulses required for the stepper motors. For example, if the stepper motor has a 1.8-degree step angle and uses 16 microstepping, 360 / 1.8*16 = 3200 pulses are required per revolution. This allows the precise number of pulses required to drive both motors to the target angles to be calculated.
[0132] Step S502: Convert control instructions for the degree of beam convergence. The degree of beam convergence is determined by the data representing the beam divergence angle or size in the illuminated object shape parameter S. In one embodiment, the beam projection mechanism utilizes an electrically tunable lens (ETL) system. The focal length of this lens can be varied by applying different control voltages.
[0133] The system maintains a precalibrated lookup table (LUT) or polynomial function f_lens, which maps the desired beam spread angle α (calculated from the parameter S and the distance from the light source to the target) to the required control voltage V_lens. That is, V_lens = f_lens(α). For example, when a narrow beam with a 10-degree spread angle is desired, the table lookup or calculation shows the required voltage is 3.2 volts; when a wide beam with a 35-degree spread angle is desired, the required voltage is 4.5 volts. This voltage is generated by a digital-to-analog converter (DAC) and applied to the ETL's driver circuit.
[0134] As an alternative, if a mechanical zoom assembly is used, the parameter S will be converted into the number of step pulses or rotation duration of the micro DC motor or stepper motor driving the zoom lens group to achieve physical adjustment of the focal length.
[0135] Step S503: Convert control instructions for light color attributes. Light color attributes are defined by target luminous flux Φ and target correlated color temperature C. The light source module consists of a warm white LED array (e.g., with a color temperature of 2700K) and a cool white LED array (e.g., with a color temperature of 6500K). By independently adjusting the brightness of the two arrays, any color temperature and total brightness between the two can be mixed.
[0136] Brightness adjustment is achieved through pulse-width modulation (PWM). The system maintains a two-dimensional chromaticity-luminance lookup table. This table is generated through offline calibration: a spectrometer and illuminance meter are used to measure the actual output luminous flux and color temperature under different combinations of warm light PWM duty cycle Duty_warm and cool light PWM duty cycle Duty_cool. The corresponding relationship between (Duty_warm, Duty_cool) and (Φ, C) is stored.
[0137] When the target (Φ, C) is received, the system performs a reverse lookup or bilinear interpolation on the two-dimensional lookup table to find the PWM duty cycle combination (Duty_warm_target, Duty_cool_target) that produces the light color closest to the target. For example, for a target (Φ = 800lm, C = 4000K), interpolation yields Duty_warm_target = 75% and Duty_cool_target = 42%. These two duty cycle values are then loaded into the microcontroller's PWM generator registers to drive the corresponding LED array.
[0138] In step S504, all calculated underlying hardware control instructions are packaged and transmitted. All instructions, including the stepper motor pulse count, ETL control voltage, and LED PWM duty cycle, are encapsulated into a defined data packet. This data packet structure may be: [Header (2 bytes), Instruction Length (1 byte), Pan Pulse Count (4 bytes), Tilt Pulse Count (4 bytes), ETL Voltage (2 bytes), PWM Warm Light (1 byte), PWM Cold Light (1 byte), Checksum (1 byte)]. This data packet is sent to the underlying hardware controller of the multi-degree-of-freedom beam projection mechanism via a high-speed serial interface (e.g., UART, baud rate 115200). The controller parses the instructions and drives the various execution components.
[0139] Reference Figure 4 Step S6 is the step of performing closed-loop correction on the lighting effect. Its purpose is to ensure that the actual lighting effect in the physical world is highly consistent with the theoretical target calculated in step S4, and to compensate for deviations introduced by factors such as mechanical tolerances, assembly errors, ambient light changes, or component aging.
[0140] In step S601, after the multi-DOF beam projection mechanism completes adjustment according to the instructions in step S5, it captures a reference image and a target scene image. Specifically, the system first briefly turns off the light source to capture a reference image I_base completely in ambient light. Then, the light source is immediately turned on and the adjustment is completed, capturing a target image I_adj with the active illumination spot. This operation is completed within a very short time (e.g., within 50 milliseconds), assuming that other objects in the scene are not moving.
[0141] Step S602: Extract the actual illuminated area through image differencing. Calculate the pixel-level difference image I_diff = abs(I_adj - I_base) between the two images. To improve robustness, this calculation is performed in grayscale space. Regions with higher brightness values in the difference image I_diff correspond to the light spots generated by active illumination. Apply an adaptive threshold (e.g., Otsu's method) to I_diff for binarization, generating a binary mask image M_actual, where regions with pixel values of 1 represent the actual illuminated areas, and regions with pixel values of 0 represent the background.
[0142] Step S603: Quantitatively analyze the geometric properties of the actual illuminated area and process the binary mask M_actual using an image moment algorithm.
[0143] The zeroth-order moment m00 gives the total area of the region (number of pixels).
[0144] The first-order moments m10 and m01 are used to calculate the coordinates of the centroid (center of mass) of the region (cx, cy), where cx = m10 / m00 and cy = m01 / m00.
[0145] The second-order central moments μ20, μ02, and μ11 are used to calculate the principal axis direction (orientation angle) θ of the region, θ=0.5*atan2(2*μ11,μ20-μ02).
[0146] At the same time, the minimum bounding rectangle of the binary mask is calculated to obtain its width and height as an evaluation of the spot size.
[0147] Step S604: Generate a 2D projection of the target illuminated area. The geometric parameters of the parameterized 3D illuminated volume V calculated in step S4 are projected onto the 2D image plane under the current camera's view angle using a calibrated camera projection model (including the camera's intrinsic parameter matrix K_c and extrinsic parameter matrix [R|T]). For example, the intersection of a 3D elliptical cone illuminated volume (a 3D ellipse) with the target object plane will be projected as a 2D ellipse on the image. This 2D elliptical region constitutes the mask M_target of the target illuminated area.
[0148] Step S605 , calculating the deviation between the actual illumination area and the target illumination area.
[0149] Geometric shape deviation E_geo: quantified using the Intersection over Union (IoU) metric. ,in
[0150]
[0151] The value range of is [0,1], and the larger the value, the greater the shape and size deviation.
[0152] Position deviation : Calculates the Euclidean distance between the centroids of two regions.
[0153]
[0154] Directional deviation : Calculate the angular difference between the main axis directions of two regions. .
[0155] These deviations are combined into a comprehensive error vector
[0156]
[0157] Step S606 : Based on the deviation vector E, a proportional-integral-derivative (PID) controller is used to generate a correction command ΔCmd. The controller includes multiple parallel PID loops, each corresponding to a different deviation component.
[0158] Position correction: and Depend on and drive.
[0159]
[0160] in , , It is the coefficient tuned for the pan / tilt position control loop.
[0161] Size / Shape Correction: Depend on drive.
[0162]
[0163] This correction amount will be converted into a fine-tuning value for the ETL voltage.
[0164] Direction correction (in some applications where a non-circular light spot is required): Depend on If the projection mechanism contains a rotatable optical element (such as an anamorphic lens or a grating), the light spot can be rotationally corrected.
[0165] Step S607, iteratively apply the correction instruction. Add the generated correction amount ΔCmd to the original hardware control instruction sent last time. On, form a new control instruction .Will The whole closed-loop correction process (S601 to S607) is iteratively executed until the norm of the integrated error vector E is less than a preset tolerance threshold (for example, IoU>0.98 and <3 pixels). Typically, this iterative process converges quickly within 2-3 cycles.
[0166] Reference Figure 5 The complete workflow of the method provided by the present invention is further illustrated below through a specific application scenario schematic diagram.
[0167] Figure 5 Shows a typical home study or office environment, consisting of a user, a desk, a laptop, a book, and a water cup.
[0168] Initially, the user is sitting in a chair, not performing any specific tasks, and is in the IDLE state. The Task Intent Recognition Module (S3) outputs an IDLE tag. Based on this, the Illumination Volume Parameter Generation Module (S4) generates a wide-angle ambient illumination volume V_idle covering the entire desk area, with low illumination (e.g., 100 lux) and a warm color temperature (3000K). The Multi-DoF Beam Projection Module (S5) executes the instructions to provide basic lighting.
[0169] The user then turns their gaze to the book on the table and reaches for it. The multimodal 3D perception module (S1) captures this sequence of actions. The dynamic scene semantic parsing module (S2) continuously updates the user's skeletal pose and detects that the SPATIAL_PROXIMITY relationship between the user's hand key points and the book object has evolved into an INTERACTION relationship. The temporal graph convolutional network in the task intention recognition module (S3) analyzes this spatiotemporal interaction pattern and switches the task intention label from IDLE to READING.
[0170] The illumination volume parameter generation module (S4) immediately responds to the READING tag and finds that the associated object is a book. The module obtains the 3D pose and size of the book and calls the G_model_reading function to calculate a new, focused illumination volume. The center P of the illumination body is located at the center of the book surface, the direction D is perpendicular to the book pages, the shape S is an elliptical cone slightly larger than the book, and the optical parameters Φ and C are set to 500 lux and 4000K, which are suitable for reading.
[0171] The multi-degree-of-freedom beam projection module (S5) receives The parameters are then transformed through inverse kinematics and parameter conversion to drive the pan / tilt motor, adjust the voltage of the tunable lens, and update the PWM duty cycle of the LED to the new target value. The light beam smoothly transitions from wide-angle ambient light to task light focused on the book in about 0.5 seconds.
[0172] At this point, the closed-loop performance correction module (S6) is activated. The module captures the adjusted scene image and, through image differencing, discovers that the center of the actual light spot is offset by 8 pixels to the right of the target center, with an Intersection Over Union (IoU) of only 0.91. This is likely due to motor microstep error. The PID controller calculates ΔCmd based on this deviation, applies a slight negative correction to the pan and pitch angles, and slightly increases the ETL voltage to expand the light spot. The correction command is sent, and the light spot position and size are precisely corrected, improving the IoU to 0.99.
[0173] After reading for a while, the user closes the book and turns to face the laptop. The LSTM network in the task intention prediction module (S404) predicts with high probability that the upcoming task is PC_OPERATION based on the typical action sequence of the user turning from the book area to the computer area. The system immediately pre-calculates the wide-angle illumination volume covering the keyboard and mouse area. , and plan well from arrive When the user's hands rest on the keyboard and the task intent is confirmed as PC_OPERATION, the system executes a pre-planned transition animation, seamlessly transforming the lighting from focused book light to soft computer work light. The entire process is smooth and natural, providing users with a highly intelligent, adaptive, and comfortable lighting experience.
[0174] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for intelligently controlling lighting angles, characterized in that: The following steps are involved: S1: Acquire scene point cloud data and acoustic event data, call the 3D instance segmentation network and sound source localization algorithm to segment, locate, and identify physical entities and acoustic events in the scene, and construct a dynamic 3D semantic scene graph containing object categories, 3D poses, and state attributes; S2: Based on the dynamic three-dimensional semantic scene graph, extract the user's skeleton key point sequence and the object interaction relationship, construct a temporal spatial graph and input it into the temporal graph convolutional network for temporal analysis, identify the task category currently performed by the user, and generate a task intention label; S3: According to the task intention label, a target object set associated with the current task is searched in the dynamic 3D semantic scene graph, a light volume generation rule library is called, and a parameterized 3D light volume representing the target lighting area is calculated and generated based on the pose and size of the target object; S4: The geometric and optical parameters of the parameterized three-dimensional illumination body are converted into hardware control instructions of a multi-degree-of-freedom light beam projection mechanism through an inverse kinematics model and a chromaticity-luminance lookup table, thereby driving the mechanism to adjust the projection angle, convergence degree, and light color properties of the light beam.
2. The intelligent control method for lighting angle according to claim 1, characterized in that: The dynamic three-dimensional semantic scene graph includes entity nodes and relationship edges indexed by timestamps. The task intention label includes task category, execution confidence and associated object identifier. The parameterized three-dimensional illumination body includes the illumination body center position, projection direction vector, shape and size parameters, target luminous flux and target correlated color temperature. The hardware control instructions include the target rotation angle of the stepper motor, the control voltage of the tunable lens and the duty cycle of the pulse width modulation signal of the light source array.
3. The intelligent control method for lighting angle according to claim 1, characterized in that: The steps for constructing the dynamic three-dimensional semantic scene graph are specifically as follows: S111: Fusion of the time-of-flight sensor and the red, green, and blue image sensor outputs to generate a 3D point cloud data stream with color information, while simultaneously collecting multi-channel acoustic data from the microphone array. S112: Based on the point cloud data stream, a sparse convolutional neural network is applied to process the point cloud to segment the point cloud into independent instance clusters corresponding to each physical entity in the scene, and an iterative closest point algorithm is used to calculate the six-degree-of-freedom pose and size information of each instance cluster to form a scene entity set; S113: Using a positioning algorithm based on arrival time difference to process the multi-channel acoustic data, determine the three-dimensional spatial position of the acoustic event, and integrate the scene entity set and the acoustic event position into a unified data structure indexed by timestamp to generate the dynamic three-dimensional semantic scene graph.
4. The intelligent control method for lighting angle according to claim 3, characterized in that: The steps for obtaining the task intention label are as follows: S211: Extracting a three-dimensional coordinate sequence of skeletal key points associated with the user from the dynamic three-dimensional semantic scene graph, and calculating the Euclidean distance between the user's hand key points and other objects in the scene, to construct a temporal spatial graph with skeletal key points as nodes and limb connections and object interaction relationships as edges; S212: Input the temporal spatial graph sequence into a pre-trained temporal graph convolutional network, aggregate the spatial information of the node neighborhood and learn its temporal evolution pattern, output a probability distribution covering predefined task categories, and select the category with the highest probability as the task intention label.
5. The intelligent control method for lighting angle according to claim 4, characterized in that: The calculation and generation steps of the parameterized three-dimensional illumination volume are specifically as follows: S311: activating a corresponding illuminated volume generation function according to the task intention label, and querying one or more target objects associated with the current task in the dynamic three-dimensional semantic scene graph to extract their pose and size information; S312: Using the current task category, user posture, and the posture and size of the target object as inputs to the illumination body generation function, a set of parameters that uniquely determine the center position, direction, shape, luminous flux, and color temperature of the illumination body are calculated and generated to form the parameterized three-dimensional illumination body.
6. The intelligent control method for lighting angle according to claim 5, characterized in that: The conversion steps of the hardware control instructions are specifically as follows: S411: Solving the position vector and direction vector of the parameterized three-dimensional illumination object through an inverse kinematics model to convert them into target rotation angles of the horizontal and vertical stepping motors in the light beam projection mechanism; S412: Converting the data representing the beam diffusion angle in the shape parameters of the illuminated object into a control voltage value applied to the electrically controlled tunable lens or a step pulse number for driving a mechanical zoom component; S413: Decomposing the target luminous flux and target correlated color temperature into duty cycles of pulse width modulation signals for the warm white light and cool white light LED arrays by referring to a preset chromaticity-luminance lookup table.
7. The intelligent control method for lighting angle according to claim 1, characterized in that: The method further comprises step S5: S5: Feeding back the actual illumination effect after adjustment by the multi-degree-of-freedom beam projection mechanism to the multimodal three-dimensional perception module, calculating the geometric deviation between the actual illumination area and the three-dimensional illumination object through image difference and geometric comparison, and performing closed-loop correction on the hardware control instructions; The geometric deviation includes the difference between the actual illumination area and the target illumination area in the centroid, the circumscribed rectangle and the main axis direction.
8. The intelligent control method for lighting angle according to claim 7, characterized in that: The steps of the closed-loop correction are specifically as follows: S511: capturing a target scene image after the mechanism is adjusted, performing pixel-level difference calculation with the reference image before adjustment, and extracting areas with significant brightness changes as actual illuminated areas; S512: Projecting the geometric parameters of the three-dimensional illumination volume onto a two-dimensional plane under the current camera viewing angle to form a target illumination area, and quantifying the geometric deviation between the actual illumination area and the target illumination area using an intersection-over-union (IoU) metric; S513: Using a proportional-integral-differential controller, a set of correction values is generated based on the geometric deviation, and the correction values are superimposed on the original hardware control instructions to form new control instructions and sent to the multi-degree-of-freedom beam projection mechanism again.
9. The intelligent control method for lighting angle according to claim 1, characterized in that: The method further comprises step S6: S6: After identifying the current task category, the task intention label sequence of the past N time steps is input into the long short-term memory network to learn the user's task switching pattern, output the occurrence probability of each task in the next M time steps, and when the predicted probability exceeds the switching threshold, pre-calculate and cache the parameterized three-dimensional illumination volume corresponding to the future task.
10. An intelligent control system for lighting angle, characterized in that: The system is used to implement the lighting angle intelligent control method according to any one of claims 1 to 9, and the system includes: The dynamic scene parsing module is used to obtain scene point cloud data and acoustic event data, and through 3D instance segmentation and sound source localization, construct and maintain a dynamic 3D semantic scene graph that represents objects, users and their spatial relationships in the scene; A task intention recognition module is used to extract the temporal information of the interaction between the user skeleton and the object based on the dynamic three-dimensional semantic scene graph, identify the user's current task category through a temporal graph convolutional network, generate a task intention label, and predict future tasks through a long short-term memory network; an illumination volume generation module, configured to retrieve associated objects from the scene graph according to the task intent label, and calculate and generate a parameterized three-dimensional illumination volume representing the target lighting area according to a preset mathematical model; A hardware instruction conversion module, configured to convert the geometric and optical parameters of the parameterized three-dimensional illumination body into underlying hardware control instructions for a multi-degree-of-freedom beam projection mechanism; The closed-loop correction module is used to calculate the geometric deviation between the actual illuminated area and the three-dimensional illuminated body through image comparison, and generate a correction value using a proportional-integral-differential control algorithm to dynamically fine-tune the hardware control instructions.
Citation Information
Cited By
Multi-scene-oriented intelligent light adaptive control method and system
CN121038059A
Multi-scene-oriented intelligent light adaptive control method and system
CN121038059B