An Adaptive Energy-saving Control Method and System for Smart Street Lights Based on Deep Reinforcement Learning

Through the smart street light system with deep reinforcement learning, combined with deep reinforcement learning algorithms and user needs, the adaptive energy-saving control of lighting equipment is realized, and the problems of waste of electricity and insufficient light in traditional control methods are solved, and efficient and user-friendly lighting solutions are provided.

CN113869482BActive Publication Date: 2025-07-18BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110816003.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-19
Publication Date
2025-07-18
Estimated Expiration
2041-07-19

AI Technical Summary

Technical Problem

Existing lighting systems often lead to eternal light when there is sufficient light or no one, causing waste of electricity, and traditional fuzzy control methods lack control accuracy and stability in large-scale lighting systems.

Method used

The smart street light system based on deep reinforcement learning is adopted, through the data processing of the perception control layer, edge computing layer and data service layer, combined with the deep reinforcement learning Deep Q-Network algorithm, the adaptive energy-saving control of the lighting equipment is realized, and the switching status of the lighting equipment is adjusted according to the user's needs priority.

Benefits of technology

It realizes efficient energy-saving control of lighting equipment, ensuring that the light intensity is within the comfort range of the human body. Users can adjust according to their needs, reduce power waste, and provide an economical, comfortable and environmentally friendly lighting environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113869482B_ABST
    Figure CN113869482B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent lighting adaptive energy-saving control method and system based on deep reinforcement learning. In the perception control layer, sensors are used to collect environmental state data, and the environmental state data is sent to the gateway in the edge computing layer; the gateway in the edge computing layer collects the environmental state data, caches and processes the data, and then sends the data to the data service layer for storage of the environmental state data; in the data service layer, control instructions are sent to the gateway in the edge computing layer; after receiving the on / off control instructions, the gateway in the edge computing layer performs adaptive regulation on the lighting equipment; the terminal application layer is directly connected to the data service layer; the server in the data service layer has set the user instruction priority to be greater than the algorithm output instruction, and determines the final action instruction according to the received instructions. By adopting the deep reinforcement learning algorithm, efficient adaptive energy-saving control of lighting equipment is achieved while providing appropriate lighting intensity for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of lighting energy conservation, and particularly to an intelligent street lamp adaptive energy conservation control method and system based on environmental perception. Background Art

[0002] In modern cities, most lighting systems rely on manual management of lighting. Especially in places such as parks or landscape areas, there are many lighting areas, long lighting hours, and the phenomenon of "constant lighting" often occurs when the light is sufficient or there are few people, resulting in a large amount of electrical energy waste. Due to the large number of lighting lamps, even if LED lamps are used, it will still cause a certain load on the power grid. Therefore, from the perspectives of energy conservation and environmental protection, efficient lighting adaptive energy conservation control is a necessary way to improve people's quality of life and create an economical, comfortable and environmentally friendly lighting environment for people.

[0003] In related technologies, the fuzzy control theory is also applicable to adaptive system control. However, since its fuzzy rules and membership functions are completely based on experience, the control accuracy is low and the problems of robustness and stability remain to be solved. Therefore, it is not very suitable for a large-scale adaptive lighting energy conservation system. Reinforcement learning is an advanced intelligent learning algorithm. It makes decisions through the interaction between an intelligent agent and a dynamic environment, continuously learns, accumulates experience, improves the action strategy, and finally obtains the optimal action plan. The behavior of its strategy is similar to the operation of a controller in a control system. Therefore, a deep neural network trained by reinforcement learning is more suitable for realizing the adaptive energy conservation control of such a lighting system. Summary of the Invention

[0004] The purpose of the present invention is to provide an intelligent street lamp adaptive energy conservation control method and system based on deep reinforcement learning. From the two perspectives of energy consumption reduction and human comfort consideration, the method of reinforcement learning is used to better realize the adaptive energy conservation control of the intelligent street lamp system.

[0005] The present invention proposes an intelligent street lamp system for adaptive energy conservation control based on deep reinforcement learning, including a perception control layer, an edge computing layer, a data service layer, and a terminal application layer.

[0006] The intelligent street lamp control system is divided into four layers according to its functional attributes, namely the perception and control layer, the edge computing layer, the data service layer, and the terminal application layer from bottom to top. The perception and control layer is mainly responsible for collecting environmental status data and controlling lighting devices, including lighting devices, human detection sensors, light sensors, and device controllers, etc. The edge computing layer mainly provides format conversion, caching, processing, and transmission services for monitoring data, reducing the load on the server and improving data processing efficiency at the same time. The data service layer is the core layer of the intelligent street lamp control system, mainly responsible for data storage and data processing. Through the Deep Q-Network algorithm of deep reinforcement learning, the optimal lighting on / off decision is obtained, minimizing the device energy consumption while providing appropriate light intensity for users. The terminal application layer is the layer where the client for the interaction between the intelligent street lamp control system and users is located. In the default mode, the system will adaptively regulate lighting devices by analyzing environmental status data through algorithms. In addition, users can also regulate lighting devices through the application program according to their actual needs.

[0007] The present invention proposes an intelligent street lamp adaptive energy-saving control method based on deep reinforcement learning, including:

[0008] S1. Use sensors in the perception and control layer to collect environmental status data, and send the environmental status data to the gateway in the edge computing layer;

[0009] S2. After the gateway in the edge computing layer collects the environmental status data, the data is cached and processed by the micro server, and then sent to the data service layer through the gateway;

[0010] S3. Store the environmental status data in the data service layer, and input the environmental status data into the pre-trained deep reinforcement learning Deep Q-Network model to obtain the optimal action decision sequence (1 on / 0 off) of the first quantity of lighting devices;

[0011] S4. After converting the 1 / 0 digital signal output by the model into the corresponding on / off control instruction in the data service layer, send the control instruction to the gateway in the edge computing layer, and send the regulation information analyzed by the system to the application program in the terminal application layer;

[0012] S5. After the gateway in the edge computing layer receives the on / off control instruction, send the instruction to the device controller in the perception and control layer, and the device controller realizes the adaptive regulation of the lighting device;

[0013] S6. The terminal application layer is directly connected to the data service layer. If the result of the adaptive regulation is not sufficient to meet the user's own lighting needs, the user adjusts the switch through the application program, and the application program sends the user instruction to the server in the data service layer;

[0014] S7. The server in the data service layer has set the user instruction priority to be greater than the algorithm output instruction, determines the final action instruction according to the received instruction, and executes steps S4 - S6 again.

[0015] In step S2 described in the present invention, the construction and training of deep reinforcement learning are included, including:

[0016] S21. Construct an environment for interacting with the agent, including:

[0017] Step A. Determine the environmental state characteristics. The environmental state characteristics are composed of multiple parameters and are represented by State;

[0018] According to the environmental monitoring information, obtain the first parameter, represented by person; according to the position of the moving target, obtain the lighting target closest to the lighting lamp in the traveling direction and obtain the second parameter, represented by Distance; according to the illuminance sensor, obtain the third parameter, represented by Light_Intensity; according to the intelligent street lamp system, obtain the fourth parameter, represented by Light_State.

[0019] Step B. Determine the switch action characteristic state, represented by Action;

[0020] Each of the N lighting devices has two states of on / off (1 / 0), so there are 2 N types.

[0021] Step C. Design a reward function, represented by Reward;

[0022] The reward function is mainly affected by three parts: the energy consumption of the lighting device, the energy consumption generated by continuous switching, and the appropriate light intensity. The energy consumption of the lighting device is related to the number of lit lights (represented by Light_number). The energy consumption generated by continuous switching is related to the change in the action Action for two consecutive actions (represented by change). The appropriate light intensity is only meaningful under the condition that people are present. This part is the sum of the products of the light intensity after the execution of the action Action and the first parameter of the environmental state. The maximum - minimum normalization method is used to perform data normalization operations on these three parts respectively. As shown in formula (1), the reward generated by a certain action is the weighted sum of these three parts.

[0023] Reward = ω1(Light_number × per_consumption)+ω2(change × ρ)+ω3(Light_Intensity × person) (1)

[0025] Among them, ω1, ω2, and ω3 are all weight coefficients, and ω1 < 0, ω2 < 0, ω3 > 0. per_consumption is the energy consumption generated by a single lighting device within the unit time granularity under standard conditions. ρ is the energy consumption generated by the continuous switching of a single lighting device under standard conditions.

[0026] S22. Initialize the experience pool, set the number of training rounds, and randomly initialize the observation state, denoted by S;

[0027] S23. Input the (n, m)-dimensional observation state S into the prediction neural network, and output the Q value corresponding to the action in the current state, denoted by Q(S, a);

[0028] S24. Select an action. Randomly select an action for exploration with probability ε, or select the action with the largest Q value calculated from the results of the neural network using the greedy strategy as the optimal action, denoted by a;

[0029] S25. The agent executes the action and obtains the reward signal (denoted by R) and the next state (denoted by S') feedback from the environment through the formula 1 reward function;

[0030] S26. Update the state S, and store the state, the next state generated after the execution of the action, the corresponding action, the corresponding reward signal, and the action completion flag in the experience pool;

[0031] S27. The agent randomly selects k mini-batch sample-related information from the experience pool, and calculates the target Q value of each state as shown in formula 2, denoted by y j The agent updates the Q value through the reward after executing the action by the target network Target Q as shown in formula 3.

[0032]

[0033] Q * Q(s,a)←Q(s,a)+α(TargetQ - Q(s,a)) (3)

[0034] Among them, γ is the decay coefficient of the maximum future reward obtained after adopting the action a, and θ is the weight coefficient of the Target Q neural network.

[0035] S28. Update the weight parameter θ in the prediction neural network using the stochastic gradient descent algorithm based on the mini-batch samples. Define the loss function as shown in formula 4.

[0036] L(θ)=E[(TargetQ - Q(s,a,θ)) 2 (4)

[0037] S29. Repeat all the steps from S22 to S28 until the end of the round. After training, the agent can automatically generate a set of optimal lighting device on / off decisions based on the actual environmental state data, minimizing energy consumption and ensuring appropriate light intensity.

[0038] In steps S23 and S27 described in the present invention, the neural network part is included, which includes:

[0039] In the above invention, the neural network consists of an input layer, two fully connected layers and an output layer, including a predictive neural network and a target network Target Q. The predictive neural network is used to obtain the Q value of the corresponding action according to the environmental state. The target network Target Q is used to stabilize the model learning process. It is a neural network with the same structure as the predictive neural network but different parameters. The parameters θ in the network are updated after a certain number of iterations.

[0040] The advantages of the present invention are as follows:

[0041] 1. In the present invention, the Deep Q-Network method of deep reinforcement learning is adopted to enable intelligent interaction between lighting devices and the environment, realizing adaptive energy-saving control of smart street lights based on environmental perception. The smart street lights can be automatically adjusted according to the real-time environmental state, getting rid of the traditional manual management method and being able to solve the energy-saving control problem of street lights in most cases.

[0042] 2. In the present invention, not only is the energy consumption minimized during the process of realizing the adaptive adjustment of smart street lights, but also a user-friendly design is adopted. By controlling the number of lighting devices, the light intensity generated by the lighting devices is within the comfortable range of the human body. In addition, users can also use the application program to adjust the lighting devices in real time according to their own needs. This method is convenient and efficient, creating an economical, comfortable and environmentally friendly lighting environment for people.

[0043] 3. The deep reinforcement learning algorithm of the present invention adopts the method with an experience pool. At the same time, two neural networks with the same structure and different parameters, namely a predictive neural network and a target network, are added, making the algorithm efficient and relatively stable. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 It is a flowchart of a method and system for adaptive energy-saving control of smart street lights based on deep reinforcement learning provided by the present invention.

[0046] Figure 2 This is a schematic structural diagram of the intelligent street lamp adaptive energy-saving control system provided in Embodiment 1 of the present invention.

[0047] Figure 3 This is a flowchart of the adaptive energy-saving control model based on Deep Q-Network of deep reinforcement learning provided by the present invention.

[0048] Figure 4 This is a schematic structural diagram of the adaptive energy-saving control model based on Deep Q-Network of deep reinforcement learning in Embodiment 2 of the present invention.

[0049] Figure 5 This is a schematic structural diagram of the neural network module in the adaptive energy-saving control model of Embodiment 2 of the present invention. Detailed implementation manners

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise.

[0052] The present invention will be further described in detail below with reference to the accompanying drawings:

[0053] Embodiment 1

[0054] As Figure 1 shown, the present invention provides an intelligent street lamp adaptive energy-saving control method and system based on deep reinforcement learning. Figure 2 This is a schematic structural diagram of the intelligent street lamp adaptive energy-saving control system provided in Embodiment 1 of the present invention. Taking the indoor lighting environment as an example, the specific steps are as follows:

[0055] S1. Use sensors in the perception control layer to collect environmental state data and send the environmental state data to the gateway in the edge computing layer;

[0056] Deploy an illuminance sensor to detect the current indoor light intensity, and a human detection sensor to detect the presence of people and their positions in the current lighting environment. Consider the person closest to the lighting device as the lighting object, and obtain the distance between this object and the lighting device. Devices such as lighting devices, human detection sensors, and illuminance sensors communicate with the edge layer gateway device through network protocols such as Zigbee or Echonet Lite.

[0057] S2. After the edge computing layer gateway collects the environmental status data, it caches and processes the data through a micro server, and then sends the data to the data service layer through the gateway.

[0058] The gateway is responsible for collecting data from the illuminance sensor and the human detection sensor, and sending the data to the data service layer through the MQTT protocol (a well-known Internet of Things protocol for data collection).

[0059] S3. Store the environmental status data in the data service layer, and input the environmental status data into a pre-trained deep reinforcement learning Deep Q-Network model to obtain an optimal action decision sequence (1 on / 0 off) for a first quantity of lighting devices.

[0060] Extract four state environmental features of light intensity, presence of people, distance between people and lighting devices, and device operating status according to the data collected in step S1, and adopt an adaptive energy-saving control model of deep reinforcement learning Deep Q-Network, such as Figure 3 The flowchart of the adaptive energy-saving control model based on deep reinforcement learning Deep Q-Network provided by the present invention is shown. According to this process, a stable decision model is trained, and after a certain number of rounds, an optimal action execution sequence (1, 0, 0,..., 0, 1) for N lighting devices can be obtained.

[0061] S4. After converting the 1 / 0 digital signal output by the model into a corresponding on / off control instruction in the data service layer, send the control instruction to the gateway of the edge computing layer, and send the regulation information analyzed by the system to the application program in the user layer.

[0062] S5. After the edge computing layer gateway receives the on / off control instruction, it sends the instruction to the device controller in the perception control layer, and the device controller realizes the adaptive regulation of the lighting device.

[0063] S6. The terminal application layer is directly connected to the data service layer. If the adaptive regulation result is not sufficient to meet the user's own lighting needs, the user adjusts the switch through the application program, and the application program sends the user instruction to the server in the data service layer.

[0064] S7. The server in the data service layer has set the priority of the user instruction to be higher than that of the algorithm output instruction. Determine the final action instruction according to the received instruction, and execute steps S4 - S6 again.

[0065] Embodiment 2

[0066] As Figure 4 shown is a schematic structural diagram of the adaptive energy-saving control model based on Deep Q-Network of the present invention. Taking a lighting device with 4 lights as an example, the numbers of each device are Light1, Light2, Light3, and Light4 respectively. The distances Distance between people and devices are represented as D1, D2, D3, and D4 respectively. Then, each time the distance D = min{D1, D2, D3, D4} is selected as the lighting target.

[0067] When the environmental state State is input into the Deep Q-Network network with an experience pool, the intelligent agent will select an action using the greedy strategy according to the output result of the neural network or explore unknown situations using random actions. After the intelligent agent executes the action Action, through dynamic interaction with the environment, the environment will feedback a reward corresponding to this action and the next state to the intelligent agent. After repeating the above process for a certain number of rounds, the optimal on / off decision of the lamp device can be obtained through training of this model. Figure 3 The specific process of this method includes:

[0068] S21. Construct an environment for interacting with the intelligent agent, including:

[0069] Step A. Determine the environmental state feature State. The current environmental state feature consists of 4 parameters and is represented by State;

[0070] According to the environmental monitoring information, obtain whether there is anyone in the current area. The value range of person is {1, 0}; according to the position of the moving target, obtain the lighting target that is closest to the lighting lamp in the traveling direction, and obtain the distance between people and the lighting device. D1 - D4 are all less than 3 meters; according to the illuminance sensor, obtain the light intensity of the current environmental area. Since the comfortable indoor light intensity is 150 - 300 lx, the value range of Light_Intensity is between 150 - 300; according to the intelligent street lamp system, obtain the operating state of the lighting devices in the lighting system. For example, Light_State = {1, 0, 1, 0} corresponds to the operating conditions of devices in a certain state.

[0071] Step B. Determine the on / off action feature state, represented by Action;

[0072] There are two on / off (1 / 0) states for all 4 lighting devices, so the number of states is 24 = 16 kinds.

[0073] Step C: Design a reward function, denoted by Reward;

[0074] The reward function is mainly affected by three parts: the energy consumption of the lighting equipment, the energy consumption generated by continuous switching, and the appropriate light intensity. The energy consumption of the lighting equipment is related to the number of lit lights (denoted by Light_number). The energy consumption generated by continuous switching is related to the change in the action Action for two consecutive actions (denoted by change). The appropriate light intensity is only meaningful when people are present. This part is the sum of the product of the light intensity after the execution of the action Action and the first parameter person of the environmental state. The maximum-minimum normalization method is used to perform data normalization operations on these three parts respectively. Among them, the maximum-minimum normalization method is to subtract the minimum value min(X) of the attribute X from the attribute value xi and then divide it by the difference between the maximum value max(X) and the minimum value min(X) of the attribute. As shown in Formula 1, the reward generated by a certain action is the weighted sum of these three parts.

[0075] Reward = ω1(Light_number × per_consumption) + ω2(change × ρ) + ω3(Light_Intensity × person) (1)

[0077] Where ω1, ω2, and ω3 are all weight coefficients, and ω1 < 0, ω2 < 0, ω3 > 0. One hour is divided into 12 small time periods. per_consumption is the energy consumption generated by a single lighting equipment in every 5 minutes under standard conditions. ρ is the energy consumption generated by continuous switching of a single lighting equipment under standard conditions.

[0078] Here, ω1 = -0.7, ω2 = -0.1, ω3 = 0.4, per_consumption = 50W, ρ = 80W.

[0079] S22. Initialize the experience pool, set the capacity of the experience pool to memory_size = 500, which is used to store training samples, randomly initialize the observation state S, and use the random function to initialize each parameter;

[0080] S23. The 4 environmental characteristic states of 4 lighting equipment can be represented as (4, 4). Input the observation state into the prediction neural network, and output the Q value corresponding to the action of the current state, denoted by Q(S, a);

[0081] Such as Figure 5The figure shows a schematic diagram of the structure of the neural network module in the adaptive energy-saving control model. The neural network includes an input layer, two fully connected layers, and an output layer. A set of observed states S is input into the Linear layer of the first layer (4×128), and then the output result of the first layer is input into the Linear layer of the second layer (128×128). After the output result of the second layer is input into the output layer (128×4), the Q value of the selected action corresponding to the state is output, denoted as Q(S,a). The activation function used is the rectified linear unit relu, and the optimizer used is Adam.

[0082] S24. Select an action. Randomly select an action for exploration with probability ε, or use the greedy strategy to select the action with the largest Q value calculated from the neural network as the optimal action, denoted by a. Initially, the value of ε is relatively large and gradually decreases as the number of training rounds increases.

[0083] S25. The agent executes the action and obtains the reward signal (denoted by R) and the next state (denoted by S’) feedback from the environment through the reward function in Equation 1.

[0084] S26. Update the state S, i.e., S = S’, and store the state, the next state generated after the action is executed, the corresponding action, the corresponding reward signal, and the action completion flag (S, S’, a, R, done) in the experience pool.

[0085] S27. The agent randomly selects m = 32 pieces of information of mini-batch samples from the experience pool and calculates the target Q value of each state as shown in Equation 2, denoted by y. j The agent updates the Q value according to the reward after executing the action through the target network Target Q as shown in Equation 3. If S’ is the terminal state, the corresponding reward is R. If S’ is not the terminal state, it is calculated according to the target network Target Q.

[0086]

[0087] Q * (s,a)←Q(s,a)+α(TargetQ-Q(s,a)) (3)

[0088] Among them, γ is the decay coefficient of the maximum future reward obtained after adopting action a, and θ is the weight coefficient of the Target Q neural network.

[0089] S28. Update the weight parameter θ in the prediction neural network using the stochastic gradient descent algorithm based on the mini-batch samples. Define the loss function as shown in Equation 4.

[0090] L(θ)=E[(TargetQ-Q(s,a,θ))2 (4)

[0091] S29. Repeat all steps from S22 to S28 until the end of the round. After training is completed, the agent can automatically generate a set of optimal lighting device on / off decisions according to the actual environmental state data, minimizing energy consumption and ensuring appropriate light intensity.

[0092] The optimal device on / off decisions output by the model training are in the form of

[0093] where each row of the matrix represents the (1 / on, 0 / off) probability of the corresponding device light. For example, L1(28.2%, 71.8%) means that the probability of the system recommending the first light to be on is 28.2%, and the probability of it being off is 71.8%. Then, the first light L1 is selected to be turned off, and so on for the others.

[0094] In the embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0095] The above integrated units implemented in the form of software functional units can be stored in a computer-readable storage medium. The above software functional units stored in a storage medium include several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of the present invention. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes.

[0096] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An intelligent street lamp system for adaptive energy-saving control based on deep reinforcement learning, characterized in that, Including: The intelligent street lamp control system is divided into four layers according to functional attributes, namely the perception control layer, the edge computing layer, the data service layer, and the terminal application layer from bottom to top; The perception control layer is mainly responsible for collecting environmental status data and controlling lighting devices, including lighting devices, human detection sensors, light sensors, and device controllers; The edge computing layer provides format conversion, caching, processing, and transmission services for monitoring data, reducing the load on the server and improving data processing efficiency at the same time; The data service layer is the core layer of the intelligent street lamp control system, responsible for data storage and data processing. Through the Deep Q-Network algorithm of deep reinforcement learning, the optimal lighting on / off decision is obtained, and the device energy consumption is minimized while providing appropriate light intensity for users; The terminal application layer is the layer where the client for the interaction between the intelligent street lamp control system and users is located; in the default mode, the system will adaptively regulate lighting devices by analyzing environmental status data through algorithms.

2. A method for self-adaptive energy-saving control of intelligent street lights based on deep reinforcement learning using the intelligent street light system described in claim 1, characterized in that, Including: S1. Use sensors in the perception control layer to collect environmental status data and send the environmental status data to the gateway in the edge computing layer; S2. After the gateway in the edge computing layer collects the environmental status data, cache and process the data through a micro server, and then send the data to the data service layer through the gateway; S3. Store the environmental status data in the data service layer, and input the environmental status data into a pre-trained Deep Q-Network model of deep reinforcement learning to obtain the optimal action decision sequence of the first number of lighting devices; Extract four state environmental features of light intensity, presence or absence of people, distance between people and lighting devices, and device operating status from the data collected in step S1. Adopt the adaptive energy-saving control model of Deep Q-Network of deep reinforcement learning. Based on the process of the adaptive energy-saving control model of Deep Q-Network of deep reinforcement learning, train a stable decision model according to this process. After completing a certain number of rounds, obtain the optimal action execution sequence (1, 0, 0,..., 0, 1) of N lighting devices; S4. After converting the 1 / 0 digital signal output by the model into the corresponding on / off control instruction in the data service layer, send the control instruction to the gateway in the edge computing layer, and send the regulation information analyzed by the system to the application program in the user layer; S5. After the gateway in the edge computing layer receives the on / off control instruction, send the instruction to the device controller in the perception control layer, and the device controller realizes the adaptive regulation of the lighting device; S6. The terminal application layer is directly connected to the data service layer. If the adaptive regulation result is not sufficient to meet the user's own lighting needs, the user adjusts the switch through the application program, and the application program sends the user instruction to the server in the data service layer; S7. The server in the data service layer has set the user instruction priority to be higher than the algorithm output instruction. Determine the final action instruction according to the received instruction, and execute steps S4 - S6 again.

3. The adaptive energy-saving control method for intelligent street lamps based on deep reinforcement learning according to claim 2, wherein Step S1 includes: Deploy an illuminance sensor to detect the current indoor light intensity, and a human detection sensor to detect the presence of people and their positions in the current lighting environment. Consider the person closest to the lighting device as the lighting target, and obtain the distance between this target and the lighting device. The lighting device, human detection sensor, and illuminance sensor communicate with the edge layer gateway device through the Zigbee or Echonet Lite network protocol.

4. The adaptive energy-saving control method for intelligent street lights based on deep reinforcement learning according to claim 2, characterized in that, Step S2 includes: The gateway is responsible for collecting data from the illuminance sensor and human detection sensor, and sending the data to the data service layer through the MQTT protocol.

5. The adaptive energy-saving control method for intelligent street lights based on deep reinforcement learning according to claim 2, wherein, Step S3 includes: S21. Construct an environment for interacting with the agent. S22. Initialize the experience pool, set the number of training rounds, and randomly initialize the observation state, denoted as S. S23. Input the (n, m)-dimensional observation state S into the prediction neural network, and output the Q value corresponding to the action in the current state, denoted as Q(S, a). S24. Select an action. Randomly select an action for exploration with probability ε, or use the greedy strategy to select the action with the largest Q value from the results calculated by the neural network as the optimal action, denoted as a. S25. The agent executes the action and obtains the reward signal R feedback from the environment and the next state S' through the reward function. S26. Update the state S, and store the state, the next state generated after the action is executed, the corresponding action, the corresponding reward signal, and the action completion flag in the experience pool. S27. The agent randomly selects information related to m mini-batch samples from the experience pool, calculates the target value for each state, denoted by y j and updates the Q value with the reward after the agent executes an action through the target network Target Q; S28. Update the weight parameter θ in the prediction neural network using the stochastic gradient descent algorithm based on a small batch of samples. S29. Repeat all steps from S22 to S28 until the round ends. After training, the agent automatically generates a set of optimal lighting device on / off decisions based on the actual environmental state data, minimizing energy consumption and ensuring appropriate light intensity.

6. The adaptive energy-saving control method for intelligent street lights based on deep reinforcement learning according to claim 5, characterized in that, Step S21 includes: Step A. Determine the environmental state characteristics, which are composed of multiple parameters, denoted as State. Step B. Determine the switch action characteristic state, denoted as Action. Step C. Design the reward function, denoted as Reward.

7. The adaptive energy-saving control method for intelligent street lights based on deep reinforcement learning according to claim 5, characterized in that, Step S27 includes: Each sampled sample status target value y j The calculation method is as follows: The way for the agent to update the Q value with the reward after executing the action through the target network Target Q is as follows: Among them, γ is the decay coefficient of the maximum future reward to be obtained after adopting the action a, θ is the weight coefficient of the Target Q neural network, and it will be updated to the θ of the prediction neural network after every C iterations.

8. The adaptive energy-saving control method for smart street lights based on deep reinforcement learning according to claim 5, characterized in that, Step S28 includes: Its loss function is defined as: .

9. The adaptive energy-saving control method for intelligent street lamps based on deep reinforcement learning according to claim 6, characterized in that, Step A includes: Obtain the first parameter, denoted as person, based on the environmental monitoring information. Based on the position of the moving target, obtain the lighting target closest to the lighting lamp in the traveling direction, and obtain the second parameter, denoted as Distance. Obtain the third parameter, denoted as Light_Intensity, according to the illuminance sensor. Obtain the fourth parameter, denoted as Light_State, according to the intelligent street lamp system.

10. The adaptive energy-saving control method for intelligent street lamps based on deep reinforcement learning according to claim 5, wherein, Step B includes: There are two states, on / off (1 / 0), for each of the N lighting devices, so there are 2 N kinds; The reward function is mainly affected by three parts: the energy consumption of lighting equipment, the energy consumption generated by continuous switching, and the appropriate light intensity. The energy consumption of lighting equipment is related to the number of lights represented by Light_number. The energy consumption generated by continuous switching is related to the change in the action represented by change, which is the change in the action for two consecutive times. The appropriate light intensity is only meaningful when people are present. This part is the sum of the product of the light intensity after the action Action is executed and the first parameter of the environmental state. The maximum-minimum normalization method is used to perform data normalization operations on these three parts respectively. As shown in Formula 1, the reward generated by a certain action is the weighted sum of these three parts. (1) Among them, ω1, ω2, and ω3 are all weight coefficients, and ω1 < 0, ω2 < 0, ω3 > 0; per_consumption is the energy consumption generated by a single lighting equipment within the unit time granularity under standard conditions; ρ is the energy consumption generated by continuous switching of a single lighting equipment under standard conditions.

Citation Information

Patent Citations

  • Mobile edge computing system task scheduling method based on migration and reinforcement learning

    CN111858009A

  • Radio frequency spectrum dynamic allocation method and device based on reinforcement learning algorithm

    CN112512121A