Intelligent lamp group control system based on artificial intelligence
By using YOLO in the intelligent lamp group control system and improving the deep learning model of the Unet network architecture, combined with the Kmeans clustering algorithm, fast and high-precision lamp personnel positioning and intelligent lighting control are achieved, solving the problems of inaccurate lighting control and high energy consumption in the existing technology, improving lighting comfort and reducing energy consumption.
Patent Information
- Application Number
- CN202510248443.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The prior art is difficult to realize lamp personnel positioning and multi-luminance control through image recognition methods based on artificial intelligence, resulting in inaccurate lighting control and high energy consumption.
YOLO's single-stage multi-objective detection model is used to quickly and highly accurate lamp personnel position recognition, combined with the monocular depth estimation model of the improved Unet network architecture, avoiding the high hardware cost of depth cameras, and implementing the energy-saving turn-on control of intelligent lamps through the Kmeans clustering algorithm.
Intelligent indoor lighting control is realized, the lighting comfort of indoor personnel is improved, lighting energy consumption is reduced, and the accuracy and generalization ability of the detection model are improved through transfer learning and improved intercomparison ratio.
Smart Images

Figure CN120224523A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of lamp group control, and in particular to an intelligent lamp group control system based on artificial intelligence. Background Art
[0003] Currently, the unified and non-distributed control of lamps through non-intrusive indoor personnel positioning technology is a research hotspot. For example, personnel positioning is carried out through infrared rays. Although the positioning accuracy is high, the penetration ability for obstacles is poor, the effective transmission distance is short, and the installation and deployment cost is high; personnel positioning through ultrasonic waves is easily affected by the Doppler effect and temperature. Therefore, researchers believe that with the popularization and progress of artificial intelligence technology, the personnel positioning method based on video image recognition has the advantages of low cost, strong applicability, and easy operation, and should become the mainstream development direction of the lighting industry.
[0004] For example, the invention patent with the application number 201680066164.1 discloses lighting control based on images, but it does not use an image recognition method based on artificial intelligence technology for lamp control.
[0005] Therefore, how to perform personnel positioning of lamps through image recognition based on artificial intelligence and control the brightness of a lamp group with multiple lamps is a technical problem to be solved at present. Summary of the Invention
[0006] To this end, the present invention provides an intelligent lamp group control system based on artificial intelligence, which realizes fast and high-precision recognition of the positions of lamps and personnel through a single-stage multi-object detection model of YOLO, avoids the high hardware cost generated by using a depth camera for depth detection through a monocular depth estimation model that improves the Unet network architecture, and realizes intelligent energy-saving turning-on control of lamps through the Kmeans clustering algorithm, thereby realizing intelligent indoor lighting control, improving the lighting comfort of indoor personnel, and reducing lighting energy consumption.
[0007] To achieve the above object, the present invention proposes an intelligent lamp group control system based on artificial intelligence. A plurality of lamps and a vision sensor form an intelligent lamp group, including:
[0008] A lamp personnel recognition module for identifying the positions of lamps, the positions of personnel, and the states of personnel in the image of the irradiated area of the lamp group collected by the vision sensor through a multi-object detection model based on YOLO, and generating a planar position map of lamp personnel according to the positions of the lamps and the positions of the personnel;
[0009] The personnel depth recognition module, connected to the lamp personnel position recognition module, is used to recognize the vertical distance between the personnel and the lamp through a monocular depth estimation model based on an improved Unet network architecture for the image of the illuminated area of the lamp group, and label the vertical distance on the lamp personnel plane position map;
[0010] The lamp selection module, connected to the personnel depth recognition module, is used to determine the selected lamp through a lamp selection model based on the lamp personnel plane position map and the vertical distance, where the lamp selection model is constructed based on the Kmeans clustering algorithm;
[0011] The lamp brightness control module, connected to the lamp selection module and the lamp personnel position recognition module, is used to set the desired brightness according to the personnel status and control the brightness of the selected lamp according to the desired brightness.
[0012] Further, the lamp personnel position recognition module includes a pre-training unit, a transfer learning unit, and a recognition unit, and the multi-object detection model includes an initial detection model, a pre-trained detection model, and a transfer model;
[0013] The pre-training unit is used to pre-train the initial detection model through a first data set to generate the pre-trained detection model, where the first data set is an unoccluded personnel data set and a lamp data set;
[0014] The transfer learning unit is used to perform fine-tuning operations on the network layer and output layer neurons of the pre-trained detection model through a second data set to generate the transfer model, where the second data set includes a personnel occlusion data set, a personnel action data set, and a personnel brightness data set;
[0015] The recognition unit is used to recognize the lamp position, personnel position, and personnel status for the image of the illuminated area of the lamp group through the transfer model, and generate a lamp personnel plane position map according to the lamp position and personnel position.
[0016] Further, the recognition unit includes an improved intersection over union calculation sub-unit;
[0017] The improved intersection over union calculation sub-unit is used to train and test the pre-trained detection model through an improved intersection over union to generate the transfer model, where the improved intersection over union is generated by performing scaling operations and averaging operations on the original intersection over union according to a set upper limit value and a set lower limit value.
[0018] Further, the improved intersection over union calculation sub-unit includes a scaled intersection over union node, an average value node, and a loss term iteration node;
[0019] The scaled intersection over union node is used to divide the difference between the original intersection over union and the middle value of the set limit value by the difference between the set upper limit value and the set lower limit value to obtain the scaled intersection over union; the average value node is used to calculate the average value of the scaled intersection over union and the original intersection over union, and use the average value as the adjusted loss term; the loss term iteration node is used to iterate the output loss term according to the adjusted loss term.
[0020] In the above solution, by using the YOLO-based multi-object detection model with transfer learning, the demand for its dataset is reduced, its training speed is accelerated, its accuracy and generalization ability are improved, and more accurate and restricted model penalties are provided by improving the intersection over union, enabling the prediction box to better regress to the true box, thereby achieving better guidance for the multi-object detection model to learn more accurate prediction boxes and bounding boxes.
[0021] Further, the monocular depth estimation model includes an encoder, a decoder, and an improved residual network, and the person depth recognition module includes a residual connection generation unit;
[0022] The residual connection generation unit is used to connect each layer of the encoder to each layer of the decoder through the improved residual network, and prune the number of connections of the improved residual network to generate the monocular depth estimation model for identifying the vertical distance.
[0023] Further, the encoder is sequentially provided with a first-level encoding convolutional layer to a fourth-level encoding convolutional layer along the encoding order, and the decoder is sequentially provided with a first-level decoding convolutional layer to a fourth-level decoding convolutional layer along the decoding order;
[0024] The output of the first-level encoding convolutional layer is connected to the inputs of the second-level decoding convolutional layer, the third-level decoding convolutional layer, and the fourth-level decoding convolutional layer through the improved residual network;
[0025] The output of the second-level encoding convolutional layer is connected to the inputs of the second-level decoding convolutional layer and the third-level decoding convolutional layer through the improved residual network;
[0026] The output of the third-level encoding convolutional layer is connected to the inputs of the first-level decoding convolutional layer and the second-level decoding convolutional layer through the improved residual network;
[0027] The fourth-level encoding convolutional layer is connected to the input of the first-level decoding convolutional layer through the improved residual network.
[0028] Further, the improved residual network has three-level residual networks connected in sequence, and each level of residual network respectively adds the features of the output of the large convolutional pooling layer and the output of the small convolutional layer.
[0029] In the above solution, by improving the monocular depth estimation model of the Unet network architecture, the effect of feature fusion of the model is improved, the number of parameters is reduced, and the model can make full use of feature information of different scales.
[0030] Furthermore, the lamp selection module includes a planar distance calculation unit, a planar clustering unit, and a vertical nearest lamp selection unit;
[0031] The planar distance calculation unit is used to calculate a set of planar distances reflecting the distance between the personnel and the lamps according to the personnel's planar position map;
[0032] The planar clustering unit is used to perform Kmeans clustering operation on the set of planar distances to generate a clustering set reflecting the planar classification of the personnel and the lamps;
[0033] The vertical nearest lamp selection unit is used to perform an operation on the clustering set to select the lamp with the nearest vertical distance to the personnel and determine it as the selected lamp.
[0034] In the above solution, the method of selecting lamps through two Kmeans clustering operations realizes fast, easy-to-implement lamp turning-on selection that meets the lighting comfort requirements, and further realizes intelligent energy-saving turning-on of lamps.
[0035] Furthermore, the personnel states include sleeping, lowering the head, standing, and discussing, and their corresponding desired brightnesses increase in sequence;
[0036] The lamp brightness control module is used to perform fuzzy control on the selected lamp according to the desired brightness and the actual brightness collected by the vision sensor.
[0037] Furthermore, multiple lamps are networked and communicated through NBIoT with one vision sensor with the OneNET cloud platform as the main node. The OneNET cloud platform stores and displays the personnel's planar position map of the lamps and the vertical distance.
[0038] Compared with the prior art, the beneficial effects of the present invention are as follows.
[0039] 1. The rapid and high-precision identification of the positions of lamps and personnel is realized through the single-stage multi-object detection model of YOLO. By improving the monocular depth estimation model of the Unet network architecture, the high hardware cost caused by using a depth camera for depth detection is avoided. The intelligent energy-saving turning-on control of lamps is realized through the Kmeans clustering algorithm, and thus the intelligent indoor lighting control is realized, the lighting comfort of indoor personnel is improved, and the lighting energy consumption is reduced.
[0040] 2. By using a YOLO-based multi-object detection model with transfer learning, the demand for its dataset is reduced, its training speed is accelerated, its accuracy and generalization ability are improved, and more accurate and restricted model penalties are provided by improving the intersection over union, enabling the prediction boxes to better regress to the ground truth boxes, thereby achieving better guidance for the multi-object detection model to learn more accurate prediction boxes and bounding boxes.
[0041] 3. By improving the monocular depth estimation model of the Unet network architecture, the effect of feature fusion of the model is improved, the number of parameters is reduced, and the model can make full use of feature information at different scales.
[0042] 4. By performing two Kmeans clustering operations to select lamps, a fast, easy-to-implement lamp turning-on selection that meets the lighting comfort requirements is achieved, and thus intelligent energy-saving lamp turning-on is realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a schematic structural diagram of the intelligent lamp group control system based on artificial intelligence according to an embodiment of the present invention;
[0044] Figure 2 is a schematic flow diagram of the intelligent lamp group control system based on artificial intelligence according to an embodiment of the present invention;
[0045] Figure 3 is a schematic structural diagram of the monocular depth estimation model with an improved Unet network architecture of the intelligent lamp group control system based on artificial intelligence according to an embodiment of the present invention;
[0046] Figure 4 is a schematic diagram of the improved residual network structure of the monocular depth estimation model with an improved Unet network architecture of the intelligent lamp group control system based on artificial intelligence according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] In order to make the objectives and advantages of the present invention clearer and more understandable, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0048] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.
[0049] It should be noted that in the description of the present invention, the terms indicating the direction or positional relationship such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the direction or positional relationship shown in the drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.
[0050] In addition, it should also be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected to" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0051] As Figures 1 to 4 shown, the present invention provides an intelligent lamp group control system based on artificial intelligence. Through the single-stage multi-object detection model of YOLO, it realizes fast and high-precision identification of the positions of lamps and personnel. By improving the monocular depth estimation model of the Unet network architecture, it avoids the high hardware cost caused by using a depth camera for depth detection. Through the Kmeans clustering algorithm, it realizes intelligent energy-saving control of lamps, thereby realizing intelligent indoor lighting control, improving the lighting comfort of indoor personnel, and reducing lighting energy consumption.
[0052] As Figures 1 to 4 shown, this embodiment proposes an intelligent lamp group control system based on artificial intelligence. A plurality of lamps and visual sensors form an intelligent lamp group, including:
[0053] A lamp and personnel identification module, which is used to identify the positions of lamps, personnel positions and personnel status in the image of the irradiated area of the lamp group collected by the visual sensor through a multi-object detection model based on YOLO, and generate a planar position map of lamps and personnel according to the lamp positions and personnel positions;
[0054] A personnel depth identification module, connected to the lamp and personnel identification module, which is used to identify the vertical distance between personnel and lamps in the image of the irradiated area of the lamp group through a monocular depth estimation model based on an improved Unet network architecture, and mark the vertical distance on the planar position map of lamps and personnel;
[0055] A lamp selection module, connected to the personnel depth identification module, which is used to determine the selected lamps through a lamp selection model based on the planar position map of lamps and personnel and the vertical distance, wherein the lamp selection model is constructed based on the Kmeans clustering algorithm;
[0056] The lamp brightness control module is connected to the lamp selection module and the lamp personnel recognition module, and is used to set the desired brightness according to the personnel status and control the brightness of the selected lamp according to the desired brightness.
[0057] It can be understood that YOLO (You Only Look Once) is a single-stage image recognition model with relatively high speed and accuracy. Its core idea is to convert the object detection task into a regression problem, so that it can directly predict the bounding box and its class probability at multiple positions in the image. Therefore, YOLO is applied to multi-object detection of personnel status and lamps. The monocular depth estimation task is to predict the depth value of each pixel from a single RGB image to the detection point. When the visual sensor and the lamp are integrated on the same horizontal plane or above the lamp horizontal plane, this depth value can reflect the vertical distance. The network architecture of UNet has the characteristics of multi-scale feature fusion, end-to-end training, and strong adaptability, so it is very suitable for the monocular depth estimation task, and thus demonstrates powerful performance and high generalization ability. KMeans is a clustering algorithm based on iterative optimization. Its core idea is to divide the data set into multiple clusters, so that each data point belongs to the cluster corresponding to the nearest cluster center, so as to minimize the sum of squared errors within the cluster and achieve simple and efficient clustering calculation, making the turning on of the lamp more intelligent. YOLO and UNet network architecture, as a deep learning algorithm, and KMeans, as an unsupervised machine learning algorithm, both belong to the field of artificial intelligence.
[0058] Further, as Figure 2 shown, the lamp personnel recognition module includes a pre-training unit, a transfer learning unit, and an identification unit, and the multi-object detection model includes an initial detection model, a pre-trained detection model, and a transfer model;
[0059] The pre-training unit is used to pre-train the initial detection model through the first data set to generate the pre-trained detection model, where the first data set is an unoccluded personnel data set and a lamp data set;
[0060] The transfer learning unit is used to fine-tune the network layer and output layer neurons of the pre-trained detection model through the second data set to generate the transfer model, where the second data set includes a personnel occlusion data set, a personnel action data set, and a personnel brightness data set;
[0061] The identification unit is used to identify the lamp position, personnel position, and personnel status in the image of the illuminated area of the lamp group through the transfer model, and generate a lamp personnel plane position map according to the lamp position and personnel position.
[0062] It is understandable that the training of object detection algorithms relies heavily on a large amount of data. The collection and production cost of large-scale datasets is expensive, while transfer learning can use an existing dataset to fine-tune on related new tasks to reduce the scale of the dataset and ensure model performance.
[0063] Specifically, the first dataset is the MS COCO dataset, which is a large object detection dataset released by Microsoft and contains categories of personnel materials and lighting fixture materials. However, its materials are relatively simple and conventional, and thus it is not applicable to the recognition of personnel and lighting fixtures under different brightness and environmental conditions of lighting fixtures. The second dataset is a dataset of the actual environment of the lighting fixtures collected and produced. For example, in a specific implementation, personnel images in a classroom scene are collected, considering various postures of personnel such as sleeping, standing, discussing, and lowering the head, as well as various occlusion information, and lighting fixture images of different types, angles, and sizes. The personnel images and lighting fixture images are labeled using the LabelImg tool to generate a corresponding second dataset of xml configuration files. The second dataset of 3600 labeled images is divided into two subsets, with 70% used to train the multi-object detection model and 30% used to test the training effect of the multi-object detection model.
[0064] Specifically, the visual sensor uses a distortion-free industrial camera CMOS sensor with a resolution of 1920×1080.
[0065] Furthermore, the recognition unit includes an improved intersection over union calculation sub-unit; the improved intersection over union calculation sub-unit is used to train and test the pre-trained detection model through the improved intersection over union to generate the transfer model, where the improved intersection over union is generated by performing a scaling operation and an average value operation on the original intersection over union according to a set upper limit value and a set lower limit value.
[0066] Furthermore, the improved intersection over union calculation sub-unit includes a scaled intersection over union node, an average value node, and a loss term iteration node;
[0067] The scaled intersection over union node is used to divide the difference between the original intersection over union and the middle value of the set limit value by the difference between the set upper limit value and the set lower limit value to obtain the scaled intersection over union; the average value node is used to calculate the average value of the scaled intersection over union and the original intersection over union, and use the average value as the adjusted loss term; the loss term iteration node is used to iterate the output loss term according to the adjusted loss term.
[0068] Specifically, the middle value of the set limit value is one-half of the difference between the set upper limit value and the set lower limit value.
[0069] Specifically, the calculation process of the improved intersection over union is as follows:
[0070]
[0071] L′ iou = 1, IOU > b
[0072]
[0073] loss IOUnew = L iou + loss IOUold
[0074] In the formula, IOU is the original intersection over union, S1 and S2 are the areas of the intersection part and the union part of the Ground Truth Box and the Prediction Box respectively, and L′ iou is the scaled intersection over union obtained by scaling the original intersection over union. a and b are the set upper limit value and the set lower limit value, preferably taking the values of 0.05 and 0.97. L iou is the adjustment loss term obtained by taking the average of the scaled intersection over union and the original intersection over union and performing linear adjustment. loss IOUnew and loss IOUold are the current output loss term and the output loss term of the previous prediction box respectively.
[0075] It can be understood that according to the calculation formula of the original intersection over union, the original intersection over union is between 0 and 1. The scaled intersection over union scales the original intersection over union to the region of the middle value of the set limit value, enhancing its value, and further enhancing the loss term. When the gap between the Ground Truth Box and the Prediction Box is large, the value of the output loss term is larger and the penalty effect is stronger. Furthermore, it can better constrain the positional relationship between the model prediction box and the Ground Truth Box, avoid overfitting or underfitting, and by linearly adjusting the loss term, it helps to improve the accuracy of the object detection model and the missed detection situation.
[0076] Specifically, considering the characteristics that the target size of the personnel image collected by the lamp group is small and distributed densely, and the problems of the YOLOv4 algorithm such as lack of resolution, weak detail extraction ability, and poor prediction effect for dense targets, this embodiment preferably adopts a multi-object detection model based on YOLOv5.
[0077] In the above solution, by using a YOLO-based multi-object detection model with transfer learning, the demand for its dataset is reduced, its training speed is accelerated, its accuracy and generalization ability are improved, and by improving the intersection over union, a more accurate and restricted model penalty is provided, enabling the prediction box to better regress to the Ground Truth Box, and thus better guiding the multi-object detection model to learn more accurate prediction boxes and bounding boxes.
[0078] Furthermore, as Figure 3 and 4As shown, the monocular depth estimation model includes an encoder, a decoder, and an improved residual network. The person depth recognition module includes a residual connection generation unit. The residual connection generation unit is used to connect each layer of the encoder to each layer of the decoder through the improved residual network, and prune the number of connections of the improved residual network to generate the monocular depth estimation model for identifying the vertical distance.
[0079] Further, as Figure 3 shown, the encoder is sequentially provided with a first-level encoding convolutional layer to a fourth-level encoding convolutional layer along the encoding order, and the decoder is sequentially provided with a first-level decoding convolutional layer to a fourth-level decoding convolutional layer along the decoding order.
[0080] The output of the first-level encoding convolutional layer is connected to the inputs of the second-level decoding convolutional layer, the third-level decoding convolutional layer, and the fourth-level decoding convolutional layer through the improved residual network.
[0081] The output of the second-level encoding convolutional layer is connected to the inputs of the second-level decoding convolutional layer and the third-level decoding convolutional layer through the improved residual network.
[0082] The output of the third-level encoding convolutional layer is connected to the inputs of the first-level decoding convolutional layer and the second-level decoding convolutional layer through the improved residual network.
[0083] The fourth-level encoding convolutional layer is connected to the input of the first-level decoding convolutional layer through the improved residual network.
[0084] It can be understood that the UNet network architecture only performs multi-scale feature extraction at the encoder end, and then gradually up-samples at the decoder end to restore to the size of the original image. Therefore, by re-designing the connection method of the network, the model can make full use of feature information at different scales. By allowing the decoder end to learn which layer of features of the encoder is useful, the network structure as Figure 3 shown is obtained through pruning operations.
[0085] Specifically, for each level from the first-level encoding convolutional layer to the fourth-level encoding convolutional layer of the encoder, the first layer of convolution uses a 3×3 convolution kernel, the stride is 1, and the number of output channels is set to 64. The activation function uses LeakyReLU, which increases the non-linear characteristics and allows a certain amount of negative values to pass through. The number of output channels of the second layer of convolution for each level is set to 128, using the same 3×3 convolution kernel and LeakyReLU activation. Therefore, through two convolutions, the underlying features of the image are extracted layer by layer, including edges, textures, and shapes, etc., ensuring rich feature information. The LeakyReLU activation function enables the network to capture more complex feature patterns, avoids the problem of neuron death, and improves the learning ability of the network.
[0086] Subsequently, the multi-layer perceptron (MLP) generates dynamic convolution kernels according to the input feature map, and uses 1×1 convolution to generate multiple convolution kernels. Convolution is performed on the input feature map based on the generated convolution kernels to adapt to different input features, thereby enhancing the network's adaptability to features. The max pooling layer uses a 2×2 pooling window and a stride of 2 to reduce the spatial resolution of the feature map. The most significant feature information is retained through the maximum value in the pooling window, enhancing the representativeness of the features.
[0087] The first-level to fourth-level encoding convolution layers are designed to further enhance the model's learning ability. As the number of parameters increases, the convolution layers can learn more abstract and complex features. These features may involve a wider range of visual patterns and can even be understood as high-level descriptions of the entire object or scene, retaining the most significant feature information and laying a foundation for the processing of the feature enhancement path.
[0088] After passing through the first-level to fourth-level encoding convolution layers, the obtained feature map enhances the important information in the feature map through an improved residual network, thereby enhancing the network's attention to key features. Through the combination of the improved residual network and the decoder, the network can better understand the key information in the input image, which is crucial for improving the accuracy and quality of the model.
[0089] The upsampling step is usually to restore or increase the spatial resolution of the feature map so that feature fusion and object detection can be better performed. During downsampling, the spatial size of the feature map shrinks, and some detailed information is lost. Upsampling can help restore these details and improve the spatial resolution of the feature map.
[0090] Furthermore, as Figure 4 shown, the improved residual network has three levels of improved residual networks connected in sequence, and each level of the improved residual network adds the outputs of the large convolution pooling layer and the small convolution layer in terms of features.
[0091] Specifically, as Figure 4 shown, the large convolution pooling layer uses a 3X3 convolution kernel, the activation function uses ReLU, and the small convolution layer uses a 1X1 convolution kernel.
[0092] It can be understood that since the conventional UNet network architecture only uses one convolution layer to implement residual connection and cannot make good use of spatial information and semantic information, the residual connection method is redesigned. In this way, the improved residual network can not only improve the effect of feature fusion but also further reduce the number of parameters.
[0093] In the above solution, by improving the monocular depth estimation model of the Unet network architecture, the effect of feature fusion of the model is improved, the number of parameters is reduced, and the model can make full use of feature information of different scales.
[0094] Further, the lamp selection module includes a planar distance calculation unit, a planar clustering unit, and a vertical nearest lamp selection unit;
[0095] The planar distance calculation unit is used to calculate a set of planar distances reflecting the distance between the personnel and the lamps according to the personnel planar position map;
[0096] The planar clustering unit is used to perform Kmeans clustering operation on the set of planar distances to generate a clustering set reflecting the planar classification of the personnel and the lamps;
[0097] The vertical nearest lamp selection unit is used to perform an operation on the clustering set to select the lamp with the shortest vertical distance from the personnel and determine it as the selected lamp.
[0098] Specifically, the planar distance calculation unit calculates the distance from the personnel to each lamp through a conventional distance calculation formula to generate the set of planar distances. Input the set of preset lamps, the set of personnel, and the set of planar distances into the Kmeans clustering algorithm, set the number of clustering clusters to 5, the maximum number of iterations to 2, and the iteration termination threshold to 10. Through the KMeans clustering algorithm, the set of planar distances clusters the people and lamps into 5 categories; then, in each element of the clustering set, find the lamp with the shortest vertical distance from the personnel, determine it as the selected lamp and output it.
[0099] More specifically, the loss function of the Kmeans clustering algorithm is set as:
[0100]
[0101] In the formula, J is the iterative output of the Kmeans clustering algorithm, the summation times of i is 5 which is the number of clustering clusters, N i represents the number of samples included in the i-th Kmeans clustering algorithm cluster, d j is the planar distance between the j-th personnel and the lamp, u j is the clustering center of the j-th clustering algorithm cluster in the Kmeans clustering algorithm. It can be understood that when the difference between the iterative outputs J of two iterations is less than the iteration termination threshold 10, the algorithm stops iterating, and the clusters at this time are used as the final clustering result.
[0102] The selected lamp is determined from the final clustering result through the following formula:
[0103] C = min i∈ { 1,2,3......k} dist(d j, u j )
[0104] where C is the selected lamp, k is the number of clusters of the Kmeans clustering algorithm, d j is the planar distance between the j-th person and the lamp, u j is the clustering center of the j-th clustering algorithm cluster in the Kmeans clustering algorithm, and dist() represents a function that calls the vertical distance corresponding to the planar distance and the clustering center. Therefore, the lamp closest to the person assigned to the same group of clustering algorithm clusters can be determined, and then the closest lamp is determined as the selected lamp.
[0105] Furthermore, the person states include sleeping, bowing the head, standing, and discussing, and their corresponding desired brightnesses increase in sequence; the lamp brightness control module is used to perform fuzzy control on the selected lamp according to the desired brightness and the actual brightness collected by the vision sensor.
[0106] Specifically, the indoor illuminance is basically between 50 and 300 lux (Lux). Therefore, in this embodiment, the basic domain of illuminance is taken as [70, 350]. When the person state is sleeping, bowing the head, standing, or discussing, the corresponding desired brightnesses are 70, 100, 280, and 350 lux (Lux) respectively. The basic domain of the error e between the actual brightness and the desired brightness can be obtained as [-230, 300], the basic domain of the error change rate is [-30, 30], and the basic domain of the output control quantity is [10%, 100%]. The error and the error change rate are fuzzified on the corresponding basic domains to obtain the corresponding fuzzy language variables. The fuzzy language variables and the output control quantity are subjected to fuzzy control reasoning through fuzzy control rules to obtain the fuzzy control quantity, and the fuzzy control quantity is defuzzified to obtain the output control quantity, realizing different control precisions and control speeds for the lamp brightness according to different person states.
[0107] It can be understood that the multi-object detection model based on YOLO outputs an image of the illumination area of the lamp group with multiple bounding boxes labeled. Each bounding box labels a single person and a single lamp, classifies the single person into one of the person states of sleeping, bowing the head, standing, or discussing, and determines the brightness of the single lamp using the brightness conversion formula for the RGB image of the single lamp, and uses the brightness for the control of the lamp. Therefore, the lamp group control system of this embodiment only needs to set one sensor, namely the vision sensor, to realize the intelligent control of turning on the lamps and the brightness for multiple people simultaneously, reducing its implementation cost.
[0108] Furthermore, multiple lamps are networked and communicated with one said vision sensor through NBIoT with the OneNET cloud platform as the main node, and the OneNET cloud platform stores and displays the planar position map of the lamp and the person and the vertical distance.
[0109] Specifically, as terminal devices, the lamps and visual sensors exchange data with the cloud platform through the NB-IoT wireless communication technology. In this design, NB-IoT adopts the CoAP transparent transmission mode to establish a connection with the OneNET cloud platform that supports the CoAP protocol, realizing a stable network connection.
[0110] It can be understood that the single-stage multi-object detection model based on YOLO in this embodiment realizes fast and high-precision identification of the positions of lamps and personnel. The monocular depth estimation model with an improved Unet network architecture avoids the high hardware cost caused by using a depth camera for depth detection. The intelligent energy-saving turning-on control of lamps is realized through the Kmeans clustering algorithm, thus realizing intelligent indoor lighting control, improving the lighting comfort of indoor personnel, and reducing lighting energy consumption. By using the multi-object detection model based on YOLO with transfer learning, the demand for its dataset is reduced, its training speed is accelerated, its accuracy and generalization ability are improved. By improving the intersection over union, more accurate and restricted model penalties are provided, enabling the prediction boxes to better regress to the true boxes, and thus better guiding the multi-object detection model to learn more accurate prediction boxes and bounding boxes. By improving the monocular depth estimation model with the Unet network architecture, the effect of feature fusion of the model is improved, the number of parameters is reduced, and the model can make full use of feature information of different scales. The method of selecting lamps through two Kmeans clustering operations realizes fast, easy-to-implement lamp turning-on selection that meets the lighting comfort requirements, and thus realizes the intelligent energy-saving turning-on of lamps.
[0111] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
[0112] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An intelligent light group control system based on artificial intelligence, characterized in that: Multiple lamps and visual sensors form a smart lighting group, including: A lamp and personnel recognition module is used to identify the lamp position, personnel position and personnel status of the lamp group illumination area image collected by the visual sensor through a YOLO-based multi-target detection model, and generate a lamp and personnel plane position map according to the lamp position and personnel position; A personnel depth recognition module is connected to the lamp personnel recognition module, and is used to identify the vertical distance between the personnel and the lamp through the image of the lamp group illumination area through a monocular depth estimation model based on an improved Unet network architecture, and mark the vertical distance on the lamp personnel plane position map; A lamp selection module, connected to the personnel depth recognition module, for determining the selected lamp by using the lamp personnel plane position map and the vertical distance through a lamp selection model, wherein the lamp selection model is constructed based on a Kmeans clustering algorithm; The lamp brightness control module is connected to the lamp selection module and the lamp personnel identification module, and is used to set the expected brightness according to the personnel status, and to control the brightness of the selected lamp according to the expected brightness.
2. The intelligent light group control system based on artificial intelligence according to claim 1 is characterized in that: The lamp personnel recognition module includes a pre-training unit, a transfer learning unit and a recognition unit, and the multi-target detection model includes an initial detection model, a pre-training detection model and a transfer model; The pre-training unit is used to pre-train the initial detection model through a first data set to generate the pre-trained detection model, wherein the first data set is an unobstructed person data set and a lamp data set; The transfer learning unit is used to generate the transfer model by fine-tuning the neurons of the network layer and the output layer of the pre-trained detection model through a second data set, wherein the second data set includes a personnel occlusion data set, a personnel action data set, and a personnel brightness data set; The recognition unit is used to recognize the lamp position, personnel position and personnel status of the lamp group illumination area image through the migration model, and generate a lamp and personnel plane position map according to the lamp position and personnel position.
3. The intelligent light group control system based on artificial intelligence according to claim 2 is characterized in that: The recognition unit includes an improved intersection-over-union calculation subunit; The improved intersection-and-union ratio calculation subunit is used to train and test the pre-trained detection model through the improved intersection-and-union ratio to generate the migration model, wherein the improved intersection-and-union ratio is generated by scaling and averaging the original intersection-and-union ratio according to a set upper limit value and a set lower limit value.
4. The intelligent light group control system based on artificial intelligence according to claim 3 is characterized in that: The improved intersection-over-union ratio calculation subunit includes a scaling intersection-over-union ratio node, an average value node, and a loss term iteration node; The scaled IoU node is used to obtain the scaled IoU by dividing the difference between the original IoU and the middle value of the set limit by the difference between the set upper limit and the set lower limit; The average value node is used to calculate the average value of the scaled intersection-and-union ratio and the original intersection-and-union ratio, and use the average value as the adjusted loss term; the loss term iteration node is used to iterate the output loss term according to the adjusted loss term.
5. The intelligent light group control system based on artificial intelligence according to claim 1 is characterized in that: The monocular depth estimation model includes an encoder, a decoder and an improved residual network, and the person depth recognition module includes a residual connection generation unit; The residual connection generation unit is used to connect each level of the encoder with each level of the decoder through the improved residual network, and to reduce the number of connections of the improved residual network through a pruning operation to generate the monocular depth estimation model for identifying the vertical distance.
6. The intelligent light group control system based on artificial intelligence according to claim 5 is characterized in that: The encoder is provided with a first-level encoding convolution layer to a fourth-level encoding convolution layer in sequence along the encoding order, and the decoder is provided with a first-level decoding convolution layer to a fourth-level decoding convolution layer in sequence along the decoding order; The output of the first-level encoding convolutional layer is connected to the input of the second-level decoding convolutional layer, the input of the third-level decoding convolutional layer, and the input of the fourth-level decoding convolutional layer through the improved residual network; The output of the second-level encoding convolutional layer is connected to the input of the second-level decoding convolutional layer and the input of the third-level decoding convolutional layer through the improved residual network; The output of the third-level encoding convolutional layer is connected to the input of the first-level decoding convolutional layer and the input of the second-level decoding convolutional layer through the improved residual network; The fourth-level encoding convolutional layer is connected to the input of the first-level decoding convolutional layer through the improved residual network.
7. The intelligent light group control system based on artificial intelligence according to claim 5 is characterized in that: The improved residual network has three levels of residual networks connected in sequence, and each level of residual network respectively adds the features of the output of the large convolution pooling layer and the output of the small convolution layer.
8. The intelligent light group control system based on artificial intelligence according to claim 1 is characterized in that: The lamp selection module includes a plane distance calculation unit, a plane clustering unit and a vertical nearest lamp selection unit; The plane distance calculation unit is used to calculate a plane distance set reflecting the distance between the personnel and the lamp according to the personnel plane position diagram; The plane clustering unit is used to perform Kmeans clustering operation on the plane distance set to generate a cluster set reflecting the plane classification of personnel and lamps; The vertically closest lamp selection unit is used to perform calculations on the cluster set to select the lamp that is closest to the person in vertical distance, and determine it as the selected lamp.
9. The intelligent light group control system based on artificial intelligence according to any one of claims 1 to 8, characterized in that: The personnel states include sleeping, bowing, standing and discussing, and the corresponding expected brightness increases in sequence; The lamp brightness control module is used to perform fuzzy control on the selected lamp according to the expected brightness and the actual brightness collected by the visual sensor.
10. The intelligent light group control system based on artificial intelligence according to any one of claims 1 to 8, characterized in that: Multiple lamps communicate with a visual sensor through NBIoT in an Internet of Things network with the OneNET cloud platform as the main node. The OneNET cloud platform stores and displays the planar position map of the lamps and personnel and the vertical distances.
Citation Information
Patent Citations
Image based lighting control
CN108293280A
Cascaded cavity convolutional network brain tumor segmentation method with attention mechanism
CN112215850A
X-ray security inspection article identification method and system based on hyper-parameter residual convolution and clustering fusion
CN114926785A
Intelligent classroom light control system based on artificial intelligence
CN116321617A
Pedestrian mask detection and safe distance early warning method based on deep learning
CN116682156A