A stepped tunnel lighting control method based on deep Q-network

Through the tunnel lighting control method based on the deep Q network, the problem that traditional tunnel lighting systems cannot be dynamically adjusted is solved, and intelligent, energy-saving, safe and comfortable lighting effects are achieved.

CN117939754BActive Publication Date: 2025-05-27BEIJING CHINACOMMUNICATIONS UNISPLENDOUR TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410106419.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-05-27
Estimated Expiration
2044-01-25

AI Technical Summary

Technical Problem

Traditional tunnel lighting systems cannot dynamically adjust brightness according to changes in the external environment and traffic flow, resulting in waste of energy and may affect driving safety.

Method used

The step-type tunnel lighting control method based on deep Q network is adopted, and the data of high-definition bayonets and geomagnetic sensors are collected, the DQN model is constructed, and the lighting level in the tunnel is adjusted in real time.

Benefits of technology

It realizes intelligent lighting adjustments that respond to changes in the external environment and traffic flow in real time, reduces energy consumption, improves driving safety and comfort, and has self-learning and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117939754B_ABST
    Figure CN117939754B_ABST
Patent Text Reader

Abstract

The present invention discloses a stepped tunnel lighting control method based on a deep Q-network. This method utilizes intelligent sensing and control technologies to achieve real-time response and optimization of the tunnel lighting environment, aiming to improve driving safety and comfort while achieving energy-saving effects. The method collects the color temperature, brightness, weather conditions, and vehicle passing data inside and outside the tunnel in real time through intelligent sensing technologies, forming a high-dimensional state space of the tunnel traffic environment. By defining a clear reward mechanism, the system can automatically learn and obtain the optimal stepped lighting adjustment decision to achieve dynamic adjustment of the light environment inside the tunnel. The present invention brings innovation to the tunnel lighting system through intelligent and adaptive control technologies, providing a lighting solution that is both energy-saving and can enhance the driving experience. By implementing the present invention, the intelligent level of tunnel lighting can be effectively improved, providing a solid guarantee for road traffic safety and comfort.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tunnel lighting, and in particular to a stepped tunnel lighting control method based on a deep Q-network. Background Art

[0002] The cost of tunnel lighting is a heavy burden for many tunnel operation units and is the key to controlling the operation cost of expressways. For many newly opened mountain expressways, the tolls are not enough to cover the electricity expenses for tunnel lighting. Even in developed countries such as Europe and the United States, high-pressure sodium lamps with backward high energy consumption and low luminous efficiency are still used for lighting in tunnels, and the control method has not yet achieved intelligence. The lighting level in tunnels is crucial for ensuring traffic safety and driving comfort. Traditional tunnel lighting systems usually rely on preset brightness levels and cannot dynamically adjust according to external environmental changes and traffic flow, resulting in energy waste and potentially affecting driving safety in some cases.

[0003] Therefore, a vehicle-mounted lighting control method that can respond in real time and intelligently adjust the lighting level is developed. Aiming at realizing the optimal allocation of lighting resources on the basis of ensuring driving safety through intelligent sensing and control technologies. This control method dynamically adjusts the brightness of tunnel internal lighting equipment according to the real-time traffic flow of vehicles, realizing an intelligent response of "lights on when the vehicle comes, lights off when the vehicle leaves, and lights follow the vehicle", which not only greatly improves driving safety and the visual comfort of drivers, but also realizes the precise on-demand distribution of lighting energy consumption.

[0004] The information disclosed in this background art section is only intended to deepen the understanding of the overall background art of the present invention and should not be regarded as an admission or any form of suggestion that this information constitutes the prior art already known to those skilled in the art. Summary of the Invention

[0005] The purpose of the present invention is to provide a stepped tunnel lighting control method based on a deep Q-network, which uses intelligent sensing and control technologies to achieve real-time response and optimization of the tunnel lighting environment, aiming to improve driving safety and comfort while achieving energy-saving effects.

[0006] In order to achieve the above purpose, the present invention adopts the following technical solutions:

[0007] The present invention provides a stepped tunnel lighting control method based on a deep Q-network, including the following steps:

[0008] S1: Collect traffic flow data of high-definition checkpoints and geomagnetic sensors, outdoor brightness data of the tunnel, and add the longitude and latitude of the tunnel and the designed maximum speed limit value, and perform accuracy analysis and preprocessing on the above data;

[0009] S2: Convert high-order parameters into two indicators of illuminance and traffic flow state;

[0010] S3: Define the states including illuminance and AC vector respectively, adjust the actions of the supplementary light source, and design a reward function to reflect the quality of the dimming result, and construct a DQN model;

[0011] S4: Substitute the obtained data into the DQN model for training, output actions and improve the strategy;

[0012] S5: Monitor the actual dimming effect, and conduct performance evaluation in combination with lighting efficiency, energy efficiency ratio, energy saving rate, and safety improvement index.

[0013] Furthermore, step S1 specifically includes:

[0014] S11: Collect the data of the high-definition checkpoint and the geomagnetic sensor respectively, record and accumulate the data to form a data set M;

[0015] M = {(L t , T t , C t , G t )}, where L t is the outdoor brightness, T t is the current time, C t is the incoming vehicle flow state, and G t is the weather state;

[0016] Brightness L: Collect the light intensity data under different times, weather conditions and road conditions;

[0017] Time T: Record the exact time of data collection, including time and date;

[0018] Flow state C: Record the traffic flow on the road, including heavy flow, medium flow, low flow, no flow, and pedestrians;

[0019] Weather state G: Collect data on the weather, including sunny, cloudy, rainy, and snowy;

[0020] S12: Clean the data in the data set M, process missing values, outliers and error data to ensure the quality and consistency of the data, and form standardized data to adapt to the input of the model;

[0021] Missing value processing: Use the mean value to fill in the missing values; when the data x i is missing, then x i = A, where A is the mean value of all non-missing values;

[0022] Outlier processing: Use quartiles to define outliers. Values below Q1 - 1.5IQR or above Q3 + 1.5IQR are outliers;

[0023] x i If <Q1 - 1.5·IQR or xi>Q3 + 1.5·IQR, then x i is an outlier, where Q1 and Q3 are the 25th and 75th percentiles respectively.

[0024] Furthermore, step S2 specifically includes:

[0025] S21: Set the required illuminance set SI(Co, Iux) for the light environment of tunnel operation, where Co represents the color temperature value and Iux represents the luminance value;

[0026] SI = α·L + β·f(T) + γ·h(G) + δ·k(C);

[0027] where L is the luminance value measured in real time;

[0028] f(T) is a function of time, including the time of day and seasonal factors;

[0029] h(G) is a function of weather conditions, adjusted according to clear, cloudy, rainy, and snow-covered conditions;

[0030] k(C) is a function of road operation status, including heavy traffic, medium traffic, light traffic, no traffic, and pedestrians;

[0031] α, β, γ, and δ are coefficients determined through data analysis, representing the relative importance of their respective independent variables in illuminance;

[0032] S22: Set the traffic flow state Ts within a certain period of time in the tunnel. The formula is:

[0033] Ts = η·V + θ·g(T) + ι·l(G) + κ·m(C);

[0034] V is the number of vehicles passing through counted in real time;

[0035] g(T) is a function of time, including peak and off-peak hours of the day;

[0036] l(G) is a function of weather conditions, taking into account the impact of different weather on traffic flow;

[0037] m(S) is a function of road conditions, reflecting the traffic flow state under different road surface conditions;

[0038] η, θ, ι, and κ are coefficients determined through data analysis, indicating the influence of each variable on the traffic flow state.

[0039] Furthermore, step S3 specifically includes:

[0040] S31: Define the network structure:

[0041] Input layer: Accepts the current state S;

[0042] Hidden layer: Multiple fully connected layers, using non-linear activation functions; Captures the complex relationship between the state and action values through multiple fully connected layers;

[0043] Output layer: Provides a predicted Q-value for each possible action; In the case of a discretized action space, each node in the output layer corresponds to the Q-value of an action;

[0044] S32: Define the target and loss functions:

[0045] Target Q-value: Use the Bellman equation to calculate the target Q-value Yt, which is the immediate reward plus the discounted present value of the predicted future return of the best action for the next state;

[0046] Y t = r t + γ Maxα' Q(s', a'; θ ~ );

[0047] where r t is the immediate reward obtained from the environment after taking action a at time step t;

[0048] γ is the discount factor, and its value is between 0 and 1; This factor determines the current value of future returns, and the value closer to 1 indicates that future returns are more important for the current;

[0049] Q(s', a'; θ ~ ) is the q-value for all possible actions a' in the next state s', and here the parameters θ of the target network are used ~ ;

[0050] Maxα' q(s', a'; θ ~ ) represents the q-value of the action that maximizes the Q-value in the next state, that is, the predicted Q-value of the "best" action;

[0051] The target Q-value Y t reflects the expected return after taking action a t in state s t and following the optimal policy;

[0052] Loss function: The loss function is the mean squared error of the difference between the predicted Q-value and the target Q-value;

[0053] L(θ) = E[Y t - Q(s, a; θ) 2 ";

[0054] where L(θ) is the value of the loss function, and the dependence on the network parameters θ indicates that the loss will change with the change of the network parameters;

[0055] E is the sum of the number of samples in the batch;

[0056] Y t is the target Q value;

[0057] Q(s, a; θ) is the Q value predicted by the model for the current state-action pair, using the parameters θ of the main network.

[0058] Furthermore, step S4 specifically includes:

[0059] S41: Set the randomly initialized network parameters θ and the parameters θ of the target network ~ , and at the same time create an experience replay pool D;

[0060] S42: Take a dimming execution action from the environment according to the greedy strategy, and observe the light environment state s' and the reward parameter r after dimming, and convert it to (s, a, r, s') and put it into the experience replay pool and then into D; where s represents the current environment, a represents the dimming action, and s' represents the environment state after dimming execution;

[0061] S43: Randomly draw a batch of conversions from the experience replay pool D, and for each sample, calculate the target Q value Y t , and update the network parameters θ using the gradient descent method;

[0062] S44: Regularly copy the parameters θ of the main network to the target network θ ~ , to achieve a stable learning process;

[0063] S45: Use the trained model to predict the Q value of each action in the given state, and select the action with the highest predicted Q value to adjust the supplementary light source.

[0064] Furthermore, step S5 specifically includes:

[0065] S51: Calculate the energy saving rate, where E / 01 and E 234 are the energy consumptions of the old system and the new system respectively;

[0066] S52: Calculate the energy efficiency ratio, EER = E 567 / E, where E 567 is the average illuminance and E is the energy consumption; the higher the value, the higher the energy utilization rate;

[0067] S53: Calculate the system reliability, the mean time between failures, MTBF = total operating time / number of failures;

[0068] S54: Calculate the safety improvement index, calculate the percentage change in the accident rate before and after lighting improvement, Where A b3 and A 5@ are the accident rates before and after the improvement respectively; the higher the index, the more obvious the improvement effect of lighting improvement on safety is;

[0069] S55: By setting weights and scoring criteria, comprehensively consider the scores reflecting the overall performance of the lighting system such as lighting efficacy, energy efficiency ratio, energy saving rate, and safety improvement index, which is convenient for comparison, evaluation, and decision-making.

[0070] Adopting the above technical solutions, the present invention has the following beneficial effects:

[0071] Real-time response: It can capture and respond to external environmental changes and traffic flow in real time, and intelligently predict and adjust the tunnel lighting level.

[0072] Energy-saving effect: By precisely controlling the lighting level, reduce unnecessary energy consumption and achieve the energy-saving goal.

[0073] Safety and comfort: By optimizing the lighting conditions, enhance driving safety and comfort in the tunnel and improve the driving experience.

[0074] Self-learning ability: The deep Q-network enables the system to have self-learning and adaptation abilities and continuously optimize the lighting adjustment strategy.

[0075] The present invention brings innovation to the tunnel lighting system through intelligent and adaptive control technologies, and provides a lighting solution that is both energy-saving and can improve the driving experience. By implementing the present invention, the intelligent level of tunnel lighting can be effectively improved, providing a solid guarantee for road traffic safety and comfort. Description of the Drawings

[0076] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0077] Figure 1 It is a flowchart of the stepped tunnel lighting control method based on the deep Q-network provided by the embodiment of the present invention. Detailed Embodiments

[0078] The following will clearly and completely describe the technical solutions of the present invention with reference to the drawings. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0079] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not intended to limit the present invention.

[0080] Combined with Figure 1 As shown, this embodiment provides a stepped tunnel lighting control method based on a deep Q-network, which specifically includes the following steps:

[0081] Step 1: Collect the traffic flow data of high-definition checkpoints and geomagnetic sensors, the outdoor brightness data of the tunnel, and add the longitude and latitude of the tunnel and the designed maximum speed limit value, and perform accuracy analysis and preprocessing on the above data.

[0082] S11: Collect various data respectively, record and accumulate a large amount of data.

[0083] M = {(L t , T t , C t , G t )}, where L t is the outdoor brightness, T t is the current time, C t is the oncoming vehicle state, and G t is the weather state.

[0084] Brightness (L): Collect the light intensity data under different times, weathers and road conditions.

[0085] Time (T): Record the exact time of data collection, including the time (moment of the day) and the date (season may affect the light).

[0086] Traffic flow state (C): Record the traffic flow situation on the road, such as heavy traffic, medium traffic, low traffic, no traffic, pedestrians, etc.

[0087] Weather state (G): Collect data on the weather, which may include sunny, cloudy, rainy, snowy.

[0088] S12: Clean the above data, process missing values, outliers and errors, ensure the quality and consistency of the data, and form standardized data to adapt to the input of the model.

[0089] Missing value processing: Use the mean value to fill in the missing values. When x i is missing, then x i = A, where A is the mean value of all non-missing values.

[0090] Outlier processing: Use quartiles (the 25th and 75th percentiles) to define outliers. Values below Q1 - 1.5IQR or above Q3 + 1.5IQR are outliers. xi <If \(x_i>Q_3 + 1.5\cdot IQR\) or \(x_i<Q_1 - 1.5\cdot IQR\), then \(x\) i is an outlier, where \(Q_1\) and \(Q_3\) are the 25th and 75th percentiles respectively.

[0091] Step 2: Convert the high - order parameters into two key indicators: illuminance and traffic flow state.

[0092] S21: Set the required illuminance set \(SI(Co, Iux)\) for the light environment of tunnel operation, where \(Co\) represents the color temperature value and \(Iux\) represents the luminance value.

[0093] \(SI=\alpha\cdot L+\beta\cdot f(T)+\gamma\cdot h(G)+\delta\cdot k(C)\)

[0094] \(L\) is the luminance value measured in real - time.

[0095] \(f(T)\) is a function of time, which may include factors such as the time of day and seasons.

[0096] \(h(G)\) is a function of weather conditions, which is adjusted according to clear, cloudy, rainy, snowy and other situations.

[0097] \(k(C)\) is a function of road operation status, such as high traffic flow, medium traffic flow, low traffic flow, no traffic flow, presence of pedestrians, etc.

[0098] \(\alpha,\beta,\gamma,\) and \(\delta\) are coefficients determined through data analysis, representing the relative importance of their respective independent variables in illuminance.

[0099] S22: Set the traffic flow state (\(Traffic State, Ts\)) within a certain period of the tunnel. The formula is:

[0100] \(Ts=\eta\cdot V+\theta\cdot g(T)+\iota\cdot l(G)+\kappa\cdot m(C)\)

[0101] \(V\) is the number of vehicles passing through counted in real - time.

[0102] \(g(T)\) is a function of time, which may include peak and off - peak hours of the day.

[0103] \(l(G)\) is a function of weather conditions, considering the impact of different weather on traffic flow.

[0104] \(m(S)\) is a function of road conditions, reflecting the traffic flow state under different pavement conditions.

[0105] \(\eta,\theta,\iota,\) and \(\kappa\) are coefficients determined through data analysis, indicating the influence of each variable on traffic flow state.

[0106] Step 3: Define elements such as the state including illuminance and AC vector, the action of adjusting the supplementary light source (color temperature Co and brightness Lux), and design a reward function to reflect the quality of the dimming result, and construct a DQN model.

[0107] S31: Define the network structure

[0108] Input layer: Receive the current state S.

[0109] Hidden layer: Multiple fully connected layers, using non-linear activation functions such as ReLU. These layers can capture the complex relationships between states and action values.

[0110] Output layer: Provide a predicted Q-value for each possible action. In the case of a discretized action space, each node in the output layer corresponds to the Q-value of an action.

[0111] S32: Define the target and loss functions

[0112] Target Q-value: Use the Bellman equation to calculate the target Q-value Yt, which is the immediate reward plus the discounted present value of the predicted future return of the best action for the next state. Y t = r t + γMaxα‘Q(s’, a‘; θ ~ )

[0113] Loss function: The loss function is the mean squared error of the difference between the predicted Q-value and the target Q-value.

[0114] L(θ) = E[Y t - Q(s, a; θ) 2 "

[0115] Step 4: Substitute the obtained data into the model for training, output actions, and improve the policy.

[0116] S41: Set the randomly initialized network parameters θ and the parameters θ of the target network ~ , and at the same time create an experience replay pool D.

[0117] S42: Take the dimming execution action from the environment according to the greedy strategy, observe the light environment state s′ and the reward parameter r after dimming, and convert it to (s, a, r, s′) and put it into the experience replay pool and then into D.

[0118] S43: Randomly sample a batch of transitions from the experience replay pool D. For each sample, calculate the target Q-value Y t , and update the network parameters θ using the gradient descent method.

[0119] S44: Regularly copy the parameters θ of the main network to the target network θ ~ , to achieve a stable learning process.

[0120] S45: Predict the Q-value of each action in a given state using the trained model, and select the action with the highest predicted Q-value to adjust the supplementary light source.

[0121] Step Five: Monitor the actual dimming effect, and conduct performance evaluation considering multiple indicators such as lighting efficacy, energy efficiency ratio, energy saving rate, and safety improvement index.

[0122] S51: Calculate the energy saving rate, where E / 01 and E 234 are the energy consumptions of the old system and the new system respectively.

[0123] S52: Calculate the energy efficiency ratio, EER = E 567 / E, where E 567 is the average illuminance and E is the energy consumption. The higher the value, the higher the energy utilization rate.

[0124] S53: Calculate the system reliability, mean time between failures (MTBF) = total operating time / number of failures.

[0125] S54: Calculate the safety improvement index, calculate the percentage change in the accident rate before and after the lighting improvement, where A b3 and A 5@ are the accident rates before and after the improvement respectively. The higher the index, the more obvious the improvement effect of the lighting improvement on safety.

[0126] S55: By setting weights and scoring criteria, comprehensively consider the scores reflecting the overall performance of the lighting system such as lighting efficacy, energy efficiency ratio, energy saving rate, and safety improvement index, which is convenient for comparison, evaluation, and decision-making.

[0127] The present invention can be summarized as the following parts:

[0128] State space definition: The state space consists of the color temperature, brightness, weather conditions, basic tunnel design conditions, and traffic flow inside and outside the tunnel, reflecting the comprehensive lighting environment inside the tunnel.

[0129] Reward mechanism: Define the reward function according to the impact of tunnel lighting on driving safety and comfort, and reward the performance of the system in improving lighting effects and energy conservation.

[0130] Deep Q-network: Adopt the deep Q-network model, input the current state space, and output the expected reward values of each possible action to guide the stepped lighting adjustment strategy.

[0131] Experience replay: Through the experience replay mechanism, store the historical state transition data and perform random sampling during training to improve learning efficiency and model stability.

[0132] Stepped adjustment decision: Based on the optimal strategy output by the model, the lighting level in the tunnel is dynamically adjusted in a stepped manner to adapt to external environmental changes and internal traffic flow.

[0133] Currently, there are also the following two ways to achieve the invention purpose of the present invention:

[0134] One is to use a fuzzy logic controller, which simulates the human decision-making process based on fuzzy sets and fuzzy rules without the need to deeply understand the exact model of the system. It can handle uncertainties and ambiguities and dynamically adjust the lighting by defining a series of "if-then" rules.

[0135] The other is a traditional PID controller, which is a classic feedback control method that adjusts the lighting level through three parameters: proportional (P), integral (I), and derivative (D) to achieve the desired lighting effect.

[0136] These methods may be simpler and have lower computational costs compared to the deep Q-network, but they may lack the ability to handle complex, high-dimensional environments. The choice of which method depends on the requirements, complexity, and available resources of a specific scenario.

[0137] In summary, the present invention has the following advantages:

[0138] 1) Real-time response: By collecting color temperature, brightness, weather conditions, and passing vehicle data in real time, the system can immediately understand the lighting requirements and traffic conditions inside and outside the tunnel. The deep Q-network can quickly process this high-dimensional data and provide the best lighting adjustment strategy based on the current environmental state. This means that as external conditions change (such as weather changes, day-night alternation, and traffic flow increase or decrease), the system can timely adjust the lighting settings to ensure that the lighting inside the tunnel is always in an optimal state.

[0139] 2) Energy-saving effect: Traditional tunnel lighting systems may remain at high brightness when not needed or may not provide sufficient brightness when needed. The present invention intelligently predicts and adjusts the lighting level, enhancing the lighting only when necessary to avoid unnecessary energy waste. In addition, by finely regulating the brightness and color temperature of the light source, the situation of over-illumination or under-illumination can be reduced, thereby achieving higher energy efficiency.

[0140] 3) Safety and comfort: The safety of drivers in the tunnel highly depends on appropriate lighting conditions. Insufficient or excessive lighting may both lead to visual discomfort or dangerous situations. This system dynamically adjusts the lighting by considering the real-time traffic conditions and external light environment to ensure that drivers have a clear line of sight and a comfortable driving environment inside the tunnel. This reduces the risk of accidents caused by improper lighting and improves the overall comfort of drivers.

[0141] 4) Self-learning ability: As a reinforcement learning algorithm, the deep Q-network has the ability to learn from experience and continuously optimize strategies. Over time and with data accumulation, the network can better understand the relationship between lighting adjustment and rewards, thus providing more precise lighting adjustment strategies. Even when traffic and environmental conditions change, the system can self-adapt and update to maintain the effectiveness and precision of its dimming strategy.

[0142] Through this intelligent control method based on the deep Q-network, the tunnel lighting system can not only respond to environmental changes in real time, but also achieve energy conservation on the premise of ensuring safety and comfort. At the same time, it has the ability of continuous learning and adaptation, improving the intelligent level of the whole system.

[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A step-type tunnel lighting control method based on deep Q network, characterized in that: The steps include: S1: Collect the traffic flow data from high-definition checkpoints and geomagnetic sensors, the brightness data outside the tunnel, add the longitude and latitude of the tunnel and the designed maximum speed limit, and perform accuracy analysis and preprocessing on the above data; S2: Convert high-order parameters into two indicators: illuminance and traffic status; S3: Define the states including the illuminance and the AC quantity vector, adjust the action of the supplementary light source, and design the reward function to reflect the quality of the dimming result, and build the DQN model; S4: Substitute the obtained data into the DQN model for training, output actions and improve the strategy; S5: Monitor the actual dimming effect and conduct performance evaluation based on lighting efficiency, energy efficiency ratio, energy saving rate, and safety improvement index indicators; Step S1 specifically includes: S11: Collect data from the high-definition camera and the geomagnetic sensor respectively, record and accumulate the data to form a data set M; M={(L t ,T t ,C t ,G t )}, where L t is the brightness outside the cave, T t is the current time, C t is the traffic flow status, G t It is the weather conditions; Brightness L: collects light intensity data under different time, weather and road conditions; Time T: record the exact time of data collection, including time and date; Traffic status C: records the traffic conditions of the road, including heavy traffic, medium traffic, low traffic, no traffic, and pedestrians; Weather status G: collects data about the weather, including sunny, cloudy, rainy, and snowy; S12: Clean the data in the dataset M, process missing values, outliers and erroneous data to ensure the quality and consistency of the data, and form standardized data to adapt to the input of the model; Missing value processing: use the mean to fill in missing values; when the data x i If it is missing, then x i =A, A is the mean of all non-missing values; Outlier processing: Quartiles are used to define outliers. Values ​​below Q1-1.5IQR or above Q3+1.5IQR are outliers. x i <If Q1 - 1.5·IQR or xi > Q3 + 1.5·IQR, then x i is an outlier, where Q1 and Q3 are the 25th and 75th percentiles respectively; Step S3 specifically includes: S31: Define the network structure: Input layer: accept the current state S; Hidden layers: multiple fully connected layers, using nonlinear activation functions; multiple fully connected layers capture the complex relationship between state and action values; Output layer: Provides a predicted Q value for each possible action; in the case of discretized action space, each node in the output layer corresponds to the Q value of an action; S32: Define the objective and loss function: Target Q-value: Use the Bellman equation to calculate the target Q-value Yt, which is the immediate reward plus the discounted value of the predicted future reward for the best action in the next state; Y t =r t +γMaxα'Q(s',a';θ~); Among them, r t is the immediate reward obtained from the environment after taking action a at time step t; γ is the discount factor, which has a value between 0 and 1. This factor determines the current value of future returns. The closer the value is to 1, the more important the future returns are to the present. Q(s',a';θ~) is the Q value of all possible actions a' in the next state s', where the parameters θ~ of the target network are used; Maxα'Q(s',a';θ~) represents the Q value of the action that maximizes the Q value in the next state, that is, the predicted Q value of the "best" action; Target Q value Y t Reflects the state s t Take action a t and the expected return after following the optimal strategy; Loss function: The loss function is the mean squared error of the difference between the predicted Q value and the target Q value; L(θ)=E[Y t -Q(s, a;θ) 2 "; Among them, L(θ) is the value of the loss function, and its dependence on the network parameters θ indicates that the loss will change with the changes in the network parameters; E is the sum of the number of samples in the batch; Y t is the target Q value; Q(s, a; θ) is the Q value predicted by the model for the current state-action pair, using the parameters θ of the main network.

2. The step-type tunnel lighting control method based on deep Q network according to claim 1 is characterized in that: Step S2 specifically includes: S21: Set the required illuminance set SI (Co, Iux) for the light environment of the tunnel operation, where Co represents the color temperature value and Iux represents the brightness value; SI=α·L+β·f(T)+γ·h(G)+δ·k(C); Where L is the brightness value measured in real time; f(T) is a function of time, including the time of day and seasonal factors; h(G) is a function of weather conditions and is adjusted according to sunny, cloudy, rainy, and snowy conditions; k(C) is a function of the road operation status, including high traffic, medium traffic, low traffic, no traffic, and pedestrians; α, β, γ, and δ are coefficients determined through data analysis, representing the relative importance of each variable in the illuminance available; S22: Set the traffic flow state Ts in a certain period of time in the tunnel, the formula is: Ts=η·V+θ·g(T)+ι·l(G)+κ·m(C); V is the real-time count of the number of vehicles passing by; g(T) is a function of time, including peak and off-peak hours of the day; l(G) is a function of weather conditions, combining the impact of different weather conditions on traffic flow; m(S) is a function of the road state, reflecting the traffic flow state under different road conditions; η, θ, ι, and κ are coefficients determined through data analysis, indicating the influence of each variable on the traffic flow state.

3. The step-type tunnel lighting control method based on deep Q network according to claim 1 is characterized in that: Step S4 specifically includes: S41: Set the randomly initialized network parameters θ and the target network parameters θ~, and create an experience replay pool D; S42: Take dimming execution action from the environment according to the greedy dream strategy, and observe the light environment state s′ and reward parameter r after dimming, convert them into (s, a, r, s′), put them into the experience replay pool, and then put them into D; where s represents the current environment, a represents the dimming action, and s′ represents the environment state after dimming execution; S43: Randomly extract a batch of transformations from the experience replay pool D, and for each sample, calculate the target Q value Y t , update the network parameters θ using the gradient descent method; S44: Periodically copy the parameters θ of the main network to the target network θ~ to achieve a stable learning process; S45: Use the trained model to predict the Q value of each action in a given state, and select the action with the highest predicted Q value to adjust the supplementary light source.

4. The step-type tunnel lighting control method based on deep Q network according to claim 1 is characterized in that: Step S5 specifically includes: S51: Calculate the energy saving rate, Where E old and E new are the energy consumption of the old system and the new system respectively; S52: Calculate the energy efficiency ratio, EER = E avg / E, where E avg is the average illumination, and E is the energy consumption; the higher the value, the higher the energy utilization rate; S53: Calculate system reliability, mean time between failures, MTBF = total operating time / number of failures; S54: Calculate the safety improvement index and the percentage change in the accident rate before and after the lighting improvement. Among them A be and A af They are the accident rates before and after the improvement. The higher the index, the more obvious the improvement in lighting safety is. S55: By setting weights and scoring criteria, the lighting efficiency, energy efficiency ratio, energy saving rate, and safety improvement index are comprehensively considered to reflect the score of the overall performance of the lighting system, which is convenient for comparison, evaluation and decision-making.

Citation Information

Patent Citations

  • Multi-dimensional feature intelligent illumination adaptive control method and system based on deep reinforcement learning

    CN116887490A

  • Tunnel energy-saving dimming deep learning intelligent control method based on attention mechanism

    CN117279173A