Multi-scene decorative lighting control system based on artificial intelligence
By introducing computer vision technology into the lighting control system, using intelligent recognition and feature extraction methods, the problem of traditional lighting control methods lacking intelligence and dynamic adaptability is solved, and intelligent dynamic control of lighting and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202510347852.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional lighting control methods are based on preset scenes or manual operations, lack intelligence and dynamic adaptability, and cannot accurately identify indoor scenes and character behaviors, resulting in the lighting mode being unable to be automatically adjusted.
The idea of computer vision is introduced, and intelligently recognize indoor scenes and character behaviors, and a cross-view feature extraction method based on text guidance and a cross-modal bidirectional state space feature extraction method are adopted to realize intelligent dynamic control of lighting.
It realizes intelligent dynamic control of lighting, enhances adaptability to indoor scene recognition, has higher accuracy, and can automatically adjust the lighting mode according to different scenarios.
Smart Images

Figure CN120220237A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of smart home control, and specifically relates to a multi-scenario lighting control system based on artificial intelligence. Background Art
[0002] Lighting control systems play an important role in smart homes and commercial lighting. Traditional lighting control methods are usually based on preset scenarios or manual operations, lacking intelligence and dynamic adaptability. Summary of the Invention
[0003] In view of the above situation, to overcome the defects of the prior art, aiming at the problems that traditional lighting control methods are usually based on preset scenarios or manual operations and lack intelligence and dynamic adaptability, the present invention introduces the idea of computer vision. By intelligently identifying the indoor scene and human behavior, the lighting mode of the lamps is automatically adjusted according to the results of the intelligent identification, realizing the intelligent dynamic control of the lamps. The present invention creatively designs a text-guided cross-view feature extraction method. Through text-guided multi-view fusion, the model preferentially considers the views more relevant to the text content, and adaptively fuses the back-projected image features and point cloud features into a unified scene visual representation. Combined with context guidance, cross-modal context interaction is realized, and then accurate 3D scene recognition is achieved. The present invention creatively designs a cross-modal bidirectional state space feature extraction method, extracts the RGB features and skeleton features of the human body images from multiple time series and multiple perspectives, performs cross-fusion of the two features based on the attention mechanism, considers the multi-view attributes and multi-moment attributes of the fused features, rearranges and further fuses the fused features, and finally introduces a bidirectional state space model for prediction output to obtain the action classification and recognition results from multiple perspectives. Compared with traditional action recognition models, the accuracy is higher. By combining accurate 3D scene recognition and action recognition, the adaptability to indoor scene recognition is increased, and the mode conversion and control of different scene lamps are more accurate.
[0004] The multi-scenario lighting control system based on artificial intelligence provided by the present invention includes a multi-view camera, a point cloud construction module, a context word-level input module, an indoor scene perception module, an indoor human state recognition module, and a lighting control module;
[0005] The multi-view camera collects multi-view images of M views in the room where the lamps are located;
[0006] The point cloud construction module constructs an indoor point cloud through the multi-view images;
[0007] The context word-level input module inputs a label description of the indoor scene;
[0008] The indoor scene perception module performs text-guided cross-view feature extraction on multi-view images, point clouds, and label descriptions of the scene indoors to obtain scene description features;
[0009] The indoor human state recognition module uses a cross-modal bidirectional state space feature extraction method to recognize the actions of people in multi-view images and obtain human state description features;
[0010] The lighting control module assigns different control strategies to the combination of scene description features and human state description features, and the control strategies control the lighting to switch states.
[0011] Furthermore, in the indoor scene perception module, a text-guided cross-view feature extraction method specifically includes the following steps:
[0012] Step S1: Cross-modal feature extraction, extracting 3D features , multi-view image features , context word-level features and back-projected features , specifically including the following steps:
[0013] Step S11: 3D point cloud feature extraction, using a pre-trained Pointnet++ model to extract the 3D features of the point cloud ;
[0014] Step S12: Multi-view image feature extraction, using a pre-trained Swin Transformer model to extract the multi-view image features of multi-view images ;
[0015] Step S13: Context word-level feature extraction, using the context word-level features of the label description of the indoor scene extracted by the Sentence-BERT model ;
[0016] Step S14: Feature back-projection, back-projecting the multi-view image features into the three-dimensional coordinate space of the 3D features , and obtaining the back-projected features based on the correspondence between the points in the point cloud and the pixels in the multi-view image ;
[0017] Step S2: Multi-view feature aggregation, using the attention mechanism to learn the importance weights based on the context word-level, and aggregating the multi-view features according to the importance weights to obtain the multi-view aggregation features: ; ;
[0018] where is the global pooling feature of multi-view image features ; is the global pooling feature of context word-level features ; and are learnable weights , and are the obtained features of the attention mechanism represents the activation function processing represents the multi-view aggregation feature;
[0019] Step S3: Adaptive dual-vision perception, concatenate the multi-view aggregation feature with the 3D feature, and use a multi-layer perceptron and the Sigmoid function to perform feature reduction, and use a fully connected layer for feature mapping to obtain the point-level vision feature: Z h =σ(MLP( Z i , Z p ]))⨀[ Z i , Z p ], Z v =FC( Z h ) ;
[0020] In the formula, [ Z i , Z p ] represents concatenating the multi-view aggregation feature with the 3D feature represents the multi-layer perceptron processing represents the Sigmoid function represents element-wise multiplication represents the fully connected layer processing is an intermediate parameter is the point-level vision feature;
[0021] Step S4: Cross-modal context-guided reasoning, perform cross-modal context-guided reasoning on the context word-level feature and the point-level vision feature to obtain the scene description feature, which specifically includes the following steps:
[0022] Step S41: Farthest point sampling, use the farthest point sampling method to sample the point-level vision feature to obtain the point-level sparse feature ;
[0023] Step S42: Position embedding. The coordinates corresponding to the points in the point-level visual features and the point-level sparse features are transformed through a learnable multi-layer perceptron to obtain the position embedding;
[0024] Step S43: Visual embedding feature calculation. The position embedding is added to the point-level visual features and the point-level sparse features respectively to obtain the dense visual embedding and the sparse visual embedding: ; ;
[0025] In the formula, , respectively represent the coordinates corresponding to the points in the point-level visual features and the point-level sparse features, , represent the position embedding transformation by the multi-layer perceptron, and represent the dense visual embedding and the sparse visual embedding;
[0026] Step S44: Cross-attention processing. The dense visual embedding and the sparse visual embedding are processed by cross-attention-Transformer to obtain the scene description features. The formula used for cross-attention-Transformer processing is as follows: F C =Transformer([CrossAtt( E c , E v ), E c ]) ;
[0027] In the formula, represents the cross-attention processing, represents the Transformer mechanism processing, represents the scene description features.
[0028] Furthermore, in the indoor person status recognition module, a cross-modal bidirectional state space feature extraction method specifically includes the following steps:
[0029] Step Q1: RGB feature extraction. The convolutional network is used to extract the RGB features from the real-time multi-view images to obtain the temporal RGB features: z rgb =[ z 1 rgb , z 2 rgb ,..., z t rgb ] ;
[0030] In the formula, represents the RGB feature of the multi-view image at the first moment, represents the RGB feature of the multi-view image at the t-th moment, represents the temporal RGB feature;
[0031] Step Q2: Temporal skeleton feature extraction. Use the OpenPose algorithm to detect and extract the human skeleton key points in the real-time multi-view image to obtain the temporal skeleton feature: z sk =[ z 1 sk , z 2 sk ,..., z t sk ] ;
[0032] In the formula, represents the skeleton feature of the multi-view image at the first moment, represents the skeleton feature of the multi-view image at the t-th moment, represents the temporal skeleton feature;
[0033] Step Q3: Feature fusion. Use the attention mechanism to fuse the temporal RGB feature and the temporal skeleton feature to obtain the temporal fusion feature: ; ;
[0034] In the formula, , and represent learnable weights, , , are intermediate parameters, represents the attention weight, represents the temporal fusion feature;
[0035] Step Q4: Multi-permutation flattening. Rearrange the temporal fusion feature according to the view forward priority sequence, view reverse priority sequence, time forward priority sequence, and time reverse priority sequence respectively to obtain four rearranged features: P f v =[ w 1 1 , w 1 2 ,..., w t M ] ; P b v =[ w t M , w t M-1 ,..., w 1 1 ] ; P f t =[ w 1 1 , w 2 1 ,..., w t M ] ; P b t =[ w t M , w t M-1 ,..., w 1 1 ] ;
[0036] In the formula, represents the temporal fusion feature of the Mth perspective at the tth moment, where M is the total number of perspectives, , , and are four rearrangement features respectively;
[0037] Step Q5: One-dimensional convolution and state space model processing, fusing four rearrangement features and performing one-dimensional convolution and state space model prediction processing to obtain the character state description feature: ; ;
[0038] In the formula, represents the way of feature fusion by feature addition, represents the one-dimensional convolution weight, represents one-dimensional convolution, is the activation function, is the fully connected layer processing, represents the convolutional layer processing, represents the state space model prediction processing, represents the character state description feature.
[0039] The beneficial effects achieved by the present invention using the above solution are as follows:
[0040] (1)In view of the problem that traditional lighting control methods are usually based on preset scenarios or manual operations and lack intelligence and dynamic adaptation capabilities, the present invention introduces the idea of computer vision. By intelligently identifying indoor scenes and human behaviors, the lighting mode of the lighting fixtures is automatically adjusted according to the results of the intelligent identification, realizing the intelligent dynamic control of the lighting fixtures;
[0041] (2)The present invention creatively designs a text-guided cross-view feature extraction method. Through text-guided multi-view fusion, the model preferentially considers the views more relevant to the text content and adaptively fuses the back-projected image features and point cloud features into a unified scene visual representation. Combined with context guidance, cross-modal context interaction is realized, and then accurate 3D scene recognition is achieved;
[0042] (3)The present invention creatively designs a cross-modal bidirectional state space feature extraction method. The RGB features and bone features of time-series multi-view human body images are extracted, and cross-fusion based on the attention mechanism is performed on the two features. Considering the multi-view attributes and multi-moment attributes of the fusion features, the fusion features are rearranged and further fused. Finally, a bidirectional state space model is introduced for prediction output to obtain the multi-view action classification and recognition results, which has higher accuracy compared with traditional action recognition models;
[0043] (4)Combining accurate 3D scene recognition and action recognition increases the adaptability to indoor scene recognition and more accurately performs mode conversion and control on different scene lighting fixtures. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a module diagram of the multi-scene lighting control system based on artificial intelligence provided by the present invention;
[0045] Figure 2 It is a schematic flow diagram of a text-guided cross-view feature extraction method;
[0046] Figure 3 It is a schematic flow diagram of a cross-modal bidirectional state space feature extraction method.
[0047] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0049] Embodiment 1. Refer to Figure 1 , the multi-scenario lighting control system based on artificial intelligence provided by the present invention includes a multi-view camera, a point cloud construction module, a context word-level input module, an indoor scene perception module, an indoor human state recognition module, and a lighting control module;
[0050] The multi-view camera collects multi-view images of 10 views in the room where the lighting is located;
[0051] The point cloud construction module constructs an indoor point cloud through the multi-view images;
[0052] The context word-level input module inputs a label description of the indoor scene;
[0053] The indoor scene perception module performs a text-guided cross-view feature extraction on the multi-view images, point cloud, and label description of the scene in the room to obtain scene description features;
[0054] The indoor human state recognition module uses a cross-modal bidirectional state space feature extraction method to perform action recognition on the people in the multi-view images to obtain human state description features;
[0055] The lighting control module assigns different control strategies to the combination of the scene description features and the human state description features, and the control strategies control the lighting to perform state switching.
[0056] By performing the above operations, for the problem that traditional lighting control methods are usually based on preset scenarios or manual operations and lack intelligence and dynamic adaptation capabilities, the present invention introduces the idea of computer vision, intelligently recognizes the indoor scene and human behavior, and automatically adjusts the lighting mode of the lighting according to the results of the intelligent recognition, realizing the intelligent dynamic control of the lighting.
[0057] Embodiment 2. Refer to Figure 2 , based on the above embodiment, in the indoor scene perception module, a text-guided cross-view feature extraction method specifically includes the following steps:
[0058] Step S1: Cross-modal feature extraction, extracting 3D features , multi-view image features , context word-level features and back-projection features ;
[0059] Step S2: Multi-view feature aggregation. An attention mechanism is used to learn the importance weights at the context word level, and multi-view features are aggregated according to the importance weights to obtain multi-view aggregated features: ; ;
[0060] In the formula, is the global pooling feature of the multi-view image feature , is the global pooling feature of the context word-level feature , and are learnable weights, , and are the obtained features of the attention mechanism, represents the activation function processing, represents the multi-view aggregated feature;
[0061] Step S3: Adaptive dual-vision perception. The multi-view aggregated feature is concatenated with the 3D feature, and a multi-layer perceptron and a Sigmoid function are used for feature reduction, and a fully connected layer is used for feature mapping to obtain point-level visual features:
[0062] Z h =σ(MLP( Z i , Z p ]))⨀[ Z i , Z p ], Z v =FC( Z h ) ;
[0063] In the formula, [ Z i , Z p ] represents concatenating the multi-view aggregated feature with the 3D feature, represents the multi-layer perceptron processing, represents the Sigmoid function, represents element-wise multiplication, represents the fully connected layer processing, is an intermediate parameter, Is a point-level visual feature;
[0064] Step S4: Cross-modal context-guided reasoning is performed on the context word-level feature And the point-level visual feature To perform cross-modal context-guided reasoning to obtain scene description features.
[0065] By performing the above operations, through text-guided multi-view fusion, the model gives priority to the views more relevant to the text content and adaptively fuses the back-projected image features and point cloud features into a unified scene visual representation. Combining context guidance, cross-modal context interaction is achieved, and thus accurate 3D scene recognition is realized.
[0066] Embodiment 3. Based on the above embodiment, step S1 specifically includes the following steps:
[0067] Step S11: 3D point cloud feature extraction. A pre-trained Pointnet++ model is used to extract the 3D features of the point cloud ;
[0068] Step S12: Multi-view image feature extraction. A pre-trained Swin Transformer model is used to extract the multi-view image features of the multi-view images ;
[0069] Step S13: Context word-level feature extraction. The context word-level features of the label descriptions of the indoor scenes extracted by the Sentence-BERT model ;
[0070] Step S14: Feature back-projection. The multi-view image features Are back-projected into the three-dimensional coordinate space of the 3D features Based on the correspondence between the points in the point cloud and the pixels in the multi-view images, the back-projected features .
[0071] Embodiment 4. Based on the above embodiment, step S4 specifically includes the following steps:
[0072] Step S41: Farthest point sampling. The farthest point sampling method is used to sample the point-level visual feature To obtain point-level sparse features ;
[0073] Step S42: Position embedding. The coordinates corresponding to the points in the point-level visual feature and the point-level sparse features are transformed through a learnable multi-layer perceptron to obtain position embeddings;
[0074] Step S43: Visual embedding feature calculation. Add the position embedding to the point-level visual feature and the point-level sparse feature respectively to obtain the dense visual embedding and the sparse visual embedding: ; ;
[0075] In the formula, , respectively represent the coordinates corresponding to the points in the point-level visual feature and the point-level sparse feature, , represent the position embedding transformation by the multi-layer perceptron, and represent the dense visual embedding and the sparse visual embedding;
[0076] Step S44: Cross-attention processing. Perform cross-attention - Transformer processing on the dense visual embedding and the sparse visual embedding to obtain the scene description feature. The formula used for cross-attention - Transformer processing is as follows: F C =Transformer([CrossAtt( E c , E v ), E c ]) ;
[0077] In the formula, represents cross-attention processing, represents Transformer mechanism processing, represents the scene description feature.
[0078] Example 5. Refer to Figure 3 , this example is based on the above example. In the indoor human state recognition module, a cross-modal bidirectional state space feature extraction method specifically includes the following steps:
[0079] Step Q1: RGB feature extraction. Use a convolutional network to extract RGB features from real-time multi-view images to obtain temporal RGB features: z rgb =[ z 1 rgb , z 2 rgb ,..., z t rgb ] ;
[0080] In the formula, represents the RGB feature of the multi-view image at the first moment, represents the RGB features of the multi-view image at the t-th moment, representing the temporal RGB features;
[0081] Step Q2: Temporal skeleton feature extraction. Use the OpenPose algorithm to detect and extract the human skeleton key points in the real-time multi-view image to obtain the temporal skeleton features: z sk =[ z 1 sk , z 2 sk ,..., z t sk ] ;
[0082] wherein, represents the skeleton features of the multi-view image at the first moment, represents the skeleton features of the multi-view image at the t-th moment, representing the temporal skeleton features;
[0083] Step Q3: Feature fusion. Adopt the attention mechanism to fuse the temporal RGB features and the temporal skeleton features to obtain the temporal fusion features: ; ;
[0084] wherein, , and represent learnable weights, , , are intermediate parameters, represents the attention weight, representing the temporal fusion features;
[0085] Step Q4: Multi-permutation flattening. Rearrange the temporal fusion features according to the view forward priority sequence, view reverse priority sequence, time forward priority sequence, and time reverse priority sequence respectively to obtain four rearranged features: P f v =[ w 1 1 , w 1 2 ,..., w t 10 ] ; P b v =[ w t 10 , w t 9 ,..., w 1 1 ] ; P f t =[ w 1 1 , w 2 1 ,..., w t 10 ] ; P b t =[ w t 10 , w t 9 ,..., w 1 1 ] ;
[0086] In the formula, represents the temporal fusion feature of the Mth perspective at the tth moment, 10 is the total number of perspectives, , , and are four kinds of rearrangement features respectively;
[0087] Step Q5: One-dimensional convolution and state space model processing, fuse the four rearrangement features and perform one-dimensional convolution and state space model prediction processing to obtain the human state description features: ; ;
[0088] In the formula, represents the way of feature fusion by feature addition, represents the one-dimensional convolution weight, represents one-dimensional convolution, is the activation function, is the fully connected layer processing, represents the convolutional layer processing, represents the state space model prediction processing, represents the human state description feature.
[0089] By performing the above operations, the RGB features and bone features of the human body image with multiple temporal perspectives are extracted, cross-fused based on the attention mechanism for the two features, considering the multi-perspective attribute and multi-moment attribute of the fused features, rearranging and further fusing the fused features, and finally introducing a bidirectional state space model for prediction output to obtain the multi-perspective action classification and recognition results. Compared with the traditional action recognition model, the accuracy is higher.
[0090] Embodiment 6: This embodiment is based on the above embodiment and is aimed at indoor scenes described by different tags. The control strategy details are as follows:
[0091] Multi-scene classification
[0092] Reading scenario
[0093] Features: The characters in the picture are reading and the scene is a study. The light brightness is moderate and the color tone is gentle, mainly to provide good reading light;
[0094] Details: The light is concentrated in the reading area to avoid glare and shadows, providing a comfortable reading environment;
[0095] Rest scene
[0096] Features: The characters in the picture are in a resting position, and the lighting is soft and warm, creating a comfortable and relaxing atmosphere;
[0097] Details: Lights are scattered throughout the room, creating a warm and relaxing resting environment;
[0098] Party scene
[0099] Features: There are many characters in the picture with happy expressions, bright lighting and rich colors, creating a cheerful and lively atmosphere;
[0100] Details: The lights can be brightened and the colors can be changed to match the music and atmosphere to create a warm atmosphere;
[0101] Work Scene
[0102] Features: The characters in the picture are working, the lighting is bright and the color temperature is moderate, which improves work efficiency;
[0103] Details: The light is evenly distributed in the working area, avoiding glare and reflection, providing a clear and bright working environment;
[0104] Dinner scene
[0105] Features: The characters in the picture are eating and the scene is a restaurant, with soft lighting, creating a romantic dining atmosphere;
[0106] Details: The lighting can be slightly dimmed, with warm tones, to create a suitable atmosphere for dining;
[0107] Sleeping scene
[0108] Features: The characters in the picture are in sleeping postures, and the lighting is soft and dim, which helps to relax and fall asleep;
[0109] Details: The lights gradually dim and the tones become warm, creating a quiet and comfortable sleeping environment.
[0110] Example 7. Based on the above example, the detailed rules for 3D point cloud feature extraction are as follows:
[0111] Point cloud input:
[0112] Input the unprocessed point cloud;
[0113] Stratified sampling and grouping:
[0114] Sample the point cloud to reduce the number of points and ensure uniform coverage of the point cloud;
[0115] Spherical neighborhood query: Search for neighborhood points around the sampled points according to the radius to form a subset of local points;
[0116] Local feature extraction:
[0117] Use the PointNet sub-network to extract local features:
[0118] Perform multi-layer perceptron mapping on the geometric relationships of neighborhood points;
[0119] Use a symmetric function to aggregate the features of neighborhood points to obtain the local features of the sampled points;
[0120] After each layer of stratification, pass the local features upward to form a hierarchical feature representation;
[0121] Multi-layer feature extraction:
[0122] Repeat the stratified sampling and grouping operations, extract higher-level local geometric features layer by layer, and extract global features through global pooling;
[0123] Feature fusion:
[0124] Combine local and global features to obtain 3D point cloud features.
[0125] Example 8. Based on the above example, the detailed rules for multi-view image feature extraction are as follows:
[0126] Multi-view image input: Input multi-view images;
[0127] Chunking and embedding: Divide the input image into small chunks of a certain size, convert each small chunk into a vector representation, and use linear projection to embed each small chunk into a high-dimensional feature space;
[0128] Local attention calculation: Divide the image into multiple windows, apply the self-attention mechanism in each window to extract local features, and capture cross-window context information by moving the position of the window;
[0129] Hierarchical feature extraction: Use multiple layers of Swin Transformer Block. Each layer contains multi-head self-attention, multi-layer perceptron, and residual connection. After each layer, spatial downsampling is performed on the features to extract higher-level features layer by layer;
[0130] Global feature fusion: The last layer outputs multi-view image features.
[0131] Example 9. This example is based on the above example, and the details of context word-level feature extraction are as follows:
[0132] Input sentence:
[0133] Input the label description of the indoor scene;
[0134] Embedding layer processing: Convert the input sequence into an initial embedding vector, including word embedding, embedding for distinguishing sentence pairs, and position embedding;
[0135] Multi-layer Transformer encoding to extract context-related word-level features:
[0136] Multi-head self-attention captures the relationship between words, and the feed-forward network performs a non-linear mapping on the attention result. Each layer uses residual connection and LayerNorm;
[0137] Context feature extraction:
[0138] In each layer, the feature representation of the word is gradually updated to include its context information. The output of the last layer is the context-related feature of each word;
[0139] Sentence feature extraction:
[0140] Pool the word features of the entire sentence to obtain the sentence-level feature, that is, the context word-level feature.
[0141] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising", or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article, or device.
[0142] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention. The scope of the invention is defined by the appended claims and their equivalents.
[0143] The above description of the present invention and its embodiments is not restrictive. What is shown in the drawings is only one of the embodiments of the present invention, and the actual structure is not limited thereto. In general, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, design similar structural modes and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention.
Claims
1. Multi-scene lighting control system based on artificial intelligence, characterized by: It includes multi-view camera, point cloud construction module, context word-level input module, indoor scene perception module, indoor person status recognition module and lighting control module; The multi-view camera collects multi-view images of M viewpoints in the room where the lighting is located in real time; The point cloud construction module constructs indoor point clouds through multi-view images; The context word level input module inputs a label description of the indoor scene; The indoor scene perception module performs a text-guided cross-view feature extraction on the indoor multi-view images, point clouds and label descriptions of the scene to obtain scene description features; The indoor person state recognition module uses a cross-modal bidirectional state space feature extraction method to perform action recognition on the person in the multi-view image to obtain the person state description feature; The lighting control module allocates different control strategies to the combination of scene description features and character status description features, and the control strategies control the lighting to switch states.
2. The artificial intelligence-based multi-scene lighting control system according to claim 1, characterized in that: In the indoor scene perception module, a cross-view feature extraction method based on text guidance specifically includes the following steps: Step S1: Cross-modal feature extraction, extracting 3D features , multi-view image features , contextual word-level features and back-projection features ; Step S2: Multi-view feature aggregation, using the attention mechanism to learn the importance weights based on the context word level, and aggregate multi-view features according to the importance weights to obtain multi-view aggregated features ; Step S3: Adaptive dual-vision perception, concatenate multi-view aggregation features with 3D features, use multi-layer perceptron and Sigmoid function to simplify features, use fully connected layer for feature mapping, and obtain point-level visual features: ; In the formula, It represents the combination of multi-view aggregation features and 3D features. represents multi-layer perceptron processing, represents the Sigmoid function, represents element-wise multiplication, represents the fully connected layer processing, is the intermediate parameter, is a point-level visual feature; Step S4: Cross-modal context-guided reasoning, context word-level features With point-level visual features Perform cross-modal context-guided reasoning to obtain scene description features.
3. The artificial intelligence-based multi-scene lighting control system according to claim 2 is characterized in that: Step S1 specifically includes the following steps: Step S11: 3D point cloud feature extraction, using the pre-trained Pointnet++ model to extract 3D features of the point cloud ; Step S12: Multi-view image feature extraction, using the pre-trained Swin Transformer model to extract multi-view image features of the multi-view image ; Step S13: Extract contextual word-level features, using the Sentence-BERT model to extract contextual word-level features of the label description of the indoor scene ; Step S14: Feature back-projection, multi-view image features Back-projection to 3D features The back-projection feature is obtained based on the correspondence between the points in the point cloud and the pixels in the multi-view image. .
4. The artificial intelligence-based multi-scene lighting control system according to claim 2 is characterized in that: Step S4 specifically includes the following steps: Step S41: farthest point sampling, using the farthest point sampling method to sample point-level visual features Sampling is performed to obtain point-level sparse features ; Step S42: Position embedding, converting the coordinates corresponding to the midpoints of the point-level visual features and the point-level sparse features through a learnable multi-layer perceptron to obtain position embedding; Step S43: Calculate visual embedding features, add the position embedding to the point-level visual features and the point-level sparse features to obtain dense visual embedding and sparse visual embedding and ; Step S44: Cross-attention processing, cross-attention-Transformer processing is performed on dense visual migration and sparse visual embedding to obtain scene description features, wherein the formula used for cross-attention-Transformer processing is as follows: ; In the formula, stands for crossed attention processing, Represents the Transformer mechanism processing, Represents the scene description feature.
5. The multi-scene lighting control system based on artificial intelligence according to claim 1 is characterized in that: In the indoor person state recognition module, a cross-modal bidirectional state space feature extraction method specifically includes the following steps: Step Q1: RGB feature extraction: a convolutional network is used to extract RGB features from real-time multi-view images to obtain temporal RGB features. Step Q2: Time-series bone feature extraction: Use the OpenPose algorithm to detect and extract features of the human skeleton key points in the real-time multi-view images to obtain time-series bone features; Step Q3: Feature fusion, using the attention mechanism to fuse the temporal RGB features with the temporal skeleton features to obtain the temporal fusion features; Step Q4: Multi-arrangement flattening, rearrange the temporal fusion features according to the perspective forward priority sequence, perspective reverse priority sequence, time forward priority sequence and time reverse priority sequence, and obtain four rearranged features: Step Q5: One-dimensional convolution and state space model processing, fusing the four rearranged features and performing one-dimensional convolution and state space model prediction processing to obtain the character state description feature.