Traffic signal lamp control method and device, computer device and storage medium
By acquiring the current phase of traffic lights and road condition data, and using phase prediction models and time prediction models to intelligently adjust the phase and duration of traffic lights, the problem of traffic lights not being able to be dynamically adjusted in existing technologies is solved, thereby improving the traffic efficiency of intersections.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 苏州万集车联网技术有限公司
- Filing Date
- 2022-12-23
- Publication Date
- 2026-05-29
AI Technical Summary
Existing urban traffic signal control methods cannot be dynamically adjusted according to actual traffic conditions, which can easily lead to traffic congestion and low traffic efficiency in special circumstances.
By acquiring the current phase type of the traffic lights and traffic condition data, the target phase type and control duration are generated using the trained phase prediction model and time prediction model, and the phase and duration of the traffic lights are intelligently adjusted.
It has improved the traffic efficiency of intersections, stimulated the potential communication capabilities of intersections, and enabled intelligent control of traffic lights.
Smart Images

Figure CN116186535B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of road traffic technology, and in particular to a traffic signal control method, device, computer equipment, storage medium, and computer program product. Background Technology
[0002] With the progress of social development, automobiles have become an indispensable part of human life, including private cars and buses that meet travel needs, and commercial vehicles that transport goods. As living standards improve, car ownership continues to increase, and urban traffic congestion has become an unavoidable problem in daily life. At road intersections, different traffic flows are given corresponding release times. Each control state of a traffic light, that is, the combination of different light colors displayed for different directions at various intersection entrances, is called a traffic light phase. Currently, the phase transition sequence and corresponding duration of urban traffic lights are preset without considering actual traffic conditions. In special circumstances, such as rainy weather, existing control methods cannot meet the traffic flow needs of different traffic volumes in real-world scenarios, easily causing traffic congestion and low traffic efficiency. Summary of the Invention
[0003] Therefore, it is necessary to provide a traffic signal control method, device, computer equipment, computer-readable storage medium, and computer program product that can efficiently and accurately adjust traffic signals, addressing the aforementioned technical problems.
[0004] Firstly, this application provides a traffic signal light control method. The method includes:
[0005] For the traffic lights at the target intersection, obtain the current phase type of the traffic lights and the current traffic conditions at the target intersection;
[0006] Based on the current phase type and current road condition data, generate the current traffic status of the target intersection;
[0007] The current traffic status is input into the trained phase prediction model and the trained time prediction model, respectively, and the target phase type and target control duration are output.
[0008] The control signal lights are switched to the target phase corresponding to the target phase type, and the duration of the target phase is controlled according to the target control duration.
[0009] In one embodiment, the current traffic data includes the current number of vehicles in each lane of the target intersection and the average travel distance of vehicles in each lane after the current phase type is put into use; based on the current phase and the current traffic data, the current traffic state of the target intersection is generated, including:
[0010] Based on the current phase type and the number of all phase types, generate a first feature vector to characterize the current phase type;
[0011] A second feature vector is generated based on the current number of vehicles in each lane and the total number of lanes. A third feature vector is generated based on the average driving distance in each lane and the total number of lanes.
[0012] The current traffic status of the target intersection is generated based on the first feature vector, the second feature vector, and the third feature vector.
[0013] In one embodiment, the training process of the time prediction model includes:
[0014] Obtain a first-time training target model with the same structure as the time prediction model, and obtain a time evaluation model and a second-time training target model with the same structure as the time evaluation model;
[0015] Obtain training samples, and train the second time training target model based on the training samples, the time prediction model, and the time evaluation model;
[0016] The first-time training target model is trained based on the training samples and the second-time training target model.
[0017] Based on the parameters in the training target model and the parameters in the time evaluation model after training, the parameters in the time evaluation model are updated;
[0018] The parameters in the time prediction model are updated based on the parameters in the target model and the parameters in the time prediction model at the first moment after training.
[0019] In one embodiment, the training process of the phase prediction model includes:
[0020] Obtain multiple training samples and a phase training target model with the same structure as the phase prediction model;
[0021] The phase training target model is trained based on multiple training samples and a phase prediction model;
[0022] The parameters in the phase prediction model are updated based on the parameters in the trained phase training target model.
[0023] In one embodiment, the process of obtaining training samples includes:
[0024] Obtain sample traffic status, input the sample traffic status into the phase prediction model and the time prediction model respectively, and output the corresponding sample phase type and sample control duration;
[0025] Traffic light simulation control is performed based on the sample traffic state, the corresponding sample phase type, and the sample control duration. The next sample traffic state and the corresponding sample reward value are obtained after the traffic light simulation control. The sample reward value is used to characterize the degree of traffic condition improvement in the next sample traffic state.
[0026] The sample traffic state is taken as the previous sample traffic state, and the next sample traffic state is taken as the subsequent sample traffic state. The previous traffic state, the corresponding sample phase type, the corresponding sample control duration, the corresponding sample reward value, and the subsequent traffic state constitute the training sample.
[0027] In one embodiment, the traffic light simulation control process is implemented through a simulation model; the simulation model construction process includes:
[0028] For the target intersection, real-time data is collected on the new traffic state formed after traffic light control is performed under different old traffic conditions according to different phase types and control durations, as well as the reward value brought by the new traffic state.
[0029] A simulation model is built based on the collected data. The simulation model is used to simulate traffic light control according to the old traffic conditions, phase types and control duration, and to obtain the new traffic conditions and corresponding reward values after the traffic light simulation control.
[0030] In one embodiment, the training samples include the previous traffic state and the corresponding sample control duration; based on the training samples and the time evaluation model, the first-time training target model is trained, including:
[0031] Based on the previous traffic conditions, the control duration of the corresponding samples under the previous traffic conditions is evaluated through a time evaluation model to obtain the corresponding evaluation value.
[0032] A first loss function is constructed based on the evaluation value, and the parameters in the target model are trained using the first loss function.
[0033] In one embodiment, the training samples include the preceding traffic state, the corresponding sample control duration, the corresponding sample reward value, and the subsequent traffic state; based on the training samples, the time prediction model, and the time evaluation model, a second time training target model is trained, including:
[0034] The post-traffic state will be input into the time prediction model, and the corresponding sample control duration in the post-traffic state will be output.
[0035] Based on the subsequent traffic conditions, the control duration of the corresponding samples in the subsequent traffic conditions is evaluated using a time evaluation model to obtain the first score value;
[0036] Based on the first score and the corresponding sample reward value in the previous traffic state, the training labels of the training samples are determined for the training target model in the second time.
[0037] Based on the previous traffic conditions, the control duration of the corresponding samples under the previous traffic conditions is evaluated using a time evaluation model to obtain a second score value.
[0038] Based on the difference between the second score and the training label, a second loss function is constructed, and the parameters in the target model are trained using the second loss function at the second time step.
[0039] In one embodiment, the training samples include the preceding traffic state, the corresponding sample phase type, the corresponding sample reward value, and the following traffic state; based on multiple training samples and the phase prediction model, the phase training target model is trained, including:
[0040] For the current training sample among multiple training samples, the subsequent traffic state of the current training sample is input into the phase prediction model, and the corresponding sample phase type and strategy value prediction value of the subsequent traffic state are output.
[0041] Based on the predicted value of the strategy corresponding to the subsequent traffic state and the sample reward value corresponding to the preceding traffic state, the training label of the current training sample for the phase training target model is determined.
[0042] The preceding traffic conditions are input into the phase prediction model, and the corresponding policy value prediction value is output in the preceding traffic conditions.
[0043] A third loss function is constructed based on the differences between the predicted policy value values corresponding to the previous traffic states and the corresponding training labels in each training sample.
[0044] The parameters in the phase training target model are trained using a third loss function.
[0045] Secondly, this application also provides a traffic signal light control device. The device includes:
[0046] The data acquisition module is used to acquire the current phase type of the traffic lights at the target intersection and the current traffic conditions at the target intersection.
[0047] The status generation module is used to generate the current traffic status of the target intersection based on the current phase type and current traffic condition data.
[0048] The action prediction module is used to input the current traffic state into the trained phase prediction model and the trained time prediction model, respectively, and output the target phase type and target control duration.
[0049] The action execution module is used to control the traffic lights to switch to the target phase corresponding to the target phase type, and to control the duration of the target phase according to the target control duration.
[0050] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0051] For the traffic lights at the target intersection, obtain the current phase type of the traffic lights and the current traffic conditions at the target intersection;
[0052] Based on the current phase type and current road condition data, generate the current traffic status of the target intersection;
[0053] The current traffic status is input into the trained phase prediction model and the trained time prediction model, respectively, and the target phase type and target control duration are output.
[0054] The control signal lights are switched to the target phase corresponding to the target phase type, and the duration of the target phase is controlled according to the target control duration.
[0055] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0056] For the traffic lights at the target intersection, obtain the current phase type of the traffic lights and the current traffic conditions at the target intersection;
[0057] Based on the current phase type and current road condition data, generate the current traffic status of the target intersection;
[0058] The current traffic status is input into the trained phase prediction model and the trained time prediction model, respectively, and the target phase type and target control duration are output.
[0059] The control signal lights are switched to the target phase corresponding to the target phase type, and the duration of the target phase is controlled according to the target control duration.
[0060] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0061] For the traffic lights at the target intersection, obtain the current phase type of the traffic lights and the current traffic conditions at the target intersection;
[0062] Based on the current phase type and current road condition data, generate the current traffic status of the target intersection;
[0063] The current traffic status is input into the trained phase prediction model and the trained time prediction model, respectively, and the target phase type and target control duration are output.
[0064] The control signal lights are switched to the target phase corresponding to the target phase type, and the duration of the target phase is controlled according to the target control duration.
[0065] The aforementioned traffic signal control method, device, computer equipment, storage medium, and computer program product, for a traffic light at a target intersection, acquire the current phase type of the traffic light being used and the current traffic conditions of the target intersection; generate the current traffic state of the target intersection based on the current phase type and current traffic conditions; input the current traffic state into a trained phase prediction model and a trained time prediction model, respectively, and output the target phase type and target control duration; control the traffic light to switch to the target phase corresponding to the target phase type, and control the duration of the target phase according to the target control duration. By simultaneously predicting the state of the traffic light through both the phase control agent and the time control agent, joint control of the traffic light is achieved in terms of both phase and duration. Furthermore, by combining actual intersection traffic conditions data, intelligent setting of the intersection traffic light phase and optimization of signal timing are realized, stimulating the intersection's potential capabilities and improving intersection communication efficiency. Attached Figure Description
[0066] Figure 1 This is an application environment diagram of a traffic signal control method in one embodiment;
[0067] Figure 2 This is a flowchart illustrating a traffic light control method in one embodiment;
[0068] Figure 3 This is a schematic diagram of the traffic light phase reference at an intersection in one embodiment;
[0069] Figure 4 This is a schematic diagram of the model training process in a traffic signal control method in one embodiment;
[0070] Figure 5 This is a flowchart illustrating a traffic light control method in another embodiment;
[0071] Figure 6 This is a flowchart illustrating the traffic light control method in yet another embodiment;
[0072] Figure 7 This is a flowchart illustrating the training process of an agent in one embodiment;
[0073] Figure 8 This is a structural block diagram of a traffic signal control device in one embodiment;
[0074] Figure 9 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0075] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0076] The traffic light control method provided in this application embodiment can be applied to, for example, Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. Specifically, terminal 102 acquires real-time traffic data at the intersection and sends the data to server 104. Server 104 processes the traffic data to obtain the current traffic light phase switching and time settings. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located in the cloud or on other network servers.
[0077] The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle systems. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0078] In one embodiment, such as Figure 2 As shown, a traffic signal light control method is provided, which can be applied to... Figure 1 Taking server 104 as an example, the following steps are included:
[0079] Step 202: For the traffic lights at the target intersection, obtain the current phase type of the traffic lights and the current traffic conditions of the target intersection;
[0080] In this context, "phase type" refers to the phase type of the traffic lights at the target intersection. The phase of a traffic light determines the permitted passage time for traffic flows in different directions. Specifically, at a signalized intersection, each control state of the traffic lights—that is, the combination of different light colors displayed for different directions at each approach lane—constitutes a signal phase. For example, for a crossroads, phase types could include east-west straight, east-west left turn, north-south straight, and north-south left turn.
[0081] Traffic data refers to data describing the distribution of vehicles at intersections. This data is acquired through roadside equipment such as millimeter-wave radar, lidar, and cameras. Traffic data can include the number of vehicles, the lane numbers of vehicles entering and exiting the intersection, the arrival and departure times of vehicles at the intersection, and the length of the vehicle queue in each lane.
[0082] It should be noted that the phase type of the target intersection can be obtained through roadside equipment, such as cameras at each approach lane of the target intersection, or directly from the control data of the traffic signal control system. This application does not specifically limit this. Specifically, in one embodiment, by acquiring multiple frames of images captured by multiple cameras at the target intersection at the current moment, and performing image processing on the multiple frames of images, the current phase type of the traffic lights at the target intersection, the number of vehicles, the numbers of the vehicle entry and exit lanes, and the time information of vehicle arrival and departure from the intersection can be determined.
[0083] Step 204: Generate the current traffic status of the target intersection based on the current phase type and current road condition data;
[0084] The current traffic state refers to the quantities that can be recognized by the intelligent agent and describe the current traffic state at the intersection. It is understandable that the acquisition of current phase type and current road condition data through roadside equipment in step 202 is directly obtained by physical devices. However, during the process of using the intelligent agent to process the data and control traffic lights, the agent cannot directly recognize this type of data and needs to process the acquired data, for example, by encoding the phase type using one-hot encoding to determine the one-hot vector corresponding to the current phase type. For current road condition data, for example, if the current traffic state includes lane queue length, then when m lane lines are identified, an m-dimensional feature vector is used to represent the lane queue length in the current traffic state, with the value of each dimension being the identified vehicle queue length. It should be noted that when there are n parameter types in the current traffic state, an m*n feature vector can be used for representation. For example, the current traffic state also includes the distance traveled by vehicles at the intersection, with the value of each dimension representing the average travel distance.
[0085] Specifically, for the current phase type, the corresponding feature vector is determined through encoding; for the current road condition data, the physical value data is defined and characterized to obtain the feature vector of the vehicle distribution at the intersection; the two together constitute the current traffic state s.
[0086] Step 206: Input the current traffic state into the trained phase prediction model and the trained time prediction model respectively, and output the target phase type and target control duration;
[0087] The phase prediction model is a discrete phase selection agent, and the time prediction model is a continuous time allocation agent. Both agents are trained based on reinforcement learning. After inputting the vector representing the current traffic state into the two trained agents, the discrete phase selection agent determines the target phase type corresponding to the next time step, while the continuous time allocation agent can obtain the target control duration corresponding to the target phase type.
[0088] It should be noted that there is a certain correlation between the phase prediction model and the time prediction model. When training both models through reinforcement learning, the same batch of data from the same training data pool is used. That is, during training, for a training sample, it needs to be input into both the discrete phase selection agent and the continuous time allocation agent to be trained. The phase prediction model and the time prediction model are trained at the same time to ensure the consistency of the learning data distribution, ensure joint optimization, and improve traffic efficiency.
[0089] Step 208: Control the traffic lights to switch to the target phase corresponding to the target phase type, and control the duration of the target phase according to the target control duration.
[0090] After obtaining the target phase type and target control duration, the traffic light phase at the target intersection is adjusted to the target phase type and maintained for the target control duration. For example, at a crossroads, if the current phase type of the target intersection is east-west straight for 30 seconds, after processing the traffic data and phase type of the current intersection using the algorithm in this embodiment, the target phase type is determined to be east-west left turn, and the target control duration is 45 seconds. Therefore, the traffic light at the target intersection is adjusted to east-west left turn for 45 seconds.
[0091] It should be noted that the target phase type can be any of the phase types at the target intersection. That is, the target phase type can be the same as the current phase type or different. If they are the same, it indicates that there are a large number of vehicles waiting to cross in the high-traffic direction, so the traffic light phase at the target intersection remains unchanged. Additionally, although the prediction of the target duration uses a continuous time allocation agent to allocate continuous time to obtain a duration, based on the intersection control situation, the traffic lights are already controlled in seconds. Therefore, the target control duration is generally taken as an integer number of seconds for easier traffic light control.
[0092] In the method provided in the above embodiments, for the traffic lights at the target intersection, the current phase type of the traffic lights being used and the current traffic conditions of the target intersection are obtained; based on the current phase type and current traffic conditions, the current traffic state of the target intersection is generated; the current traffic state is input into the trained phase prediction model and the trained time prediction model respectively, and the target phase type and target control duration are output; the traffic lights are controlled to switch to the target phase corresponding to the target phase type, and the duration of the target phase is controlled according to the target control duration. By simultaneously predicting the state of the traffic lights through the phase control agent and the time control agent, joint control of the traffic lights in terms of both phase and duration is achieved. Furthermore, by combining actual intersection traffic conditions data, the traffic light phases at the intersection are intelligently set and the signal timing is optimized, stimulating the potential capabilities of the intersection and improving the communication efficiency of the intersection.
[0093] In one embodiment, the current traffic data includes the current number of vehicles in each lane of the target intersection and the average travel distance of vehicles in each lane after the current phase type is put into use; based on the current phase and the current traffic data, the current traffic state of the target intersection is generated, including:
[0094] Based on the current phase type and the number of all phase types, generate a first feature vector to characterize the current phase type;
[0095] A second feature vector is generated based on the current number of vehicles in each lane and the total number of lanes. A third feature vector is generated based on the average driving distance in each lane and the total number of lanes.
[0096] The current traffic status of the target intersection is generated based on the first feature vector, the second feature vector, and the third feature vector.
[0097] In this context, "all phase types" refers to all the phase types of traffic lights that can be implemented at the target intersection. For different intersections, the phase types will vary due to differences in the actual conditions of the intersections where the traffic lights are located. For example, the total number of traffic lights at a crossroads and a T-junction differs, resulting in different phase types for the traffic lights at the two intersections. Similarly, at intersections of the same type, such as Crossroads 1 and Crossroads 2, although both have four traffic lights, their traffic rules differ. For instance, at Crossroads 1, left turns are not allowed in the east-west lanes. Furthermore, different combinations of phase types can exist for all traffic lights at the same intersection. For example, at a crossroads, vehicles in the same lane may be allowed to go straight and turn left simultaneously. In summary, the phase type mentioned in this embodiment should be one of all phase types determined under the same rules at the target intersection. That is, the traffic light control method mentioned in this embodiment is implemented with a fixed number of phase types. If the rules determining all phase types need to be changed, the traffic light phase prediction model and time prediction model for the target intersection need to be retrained.
[0098] The total number of lanes is a preset data for the target intersection. The number of vehicles in each lane can be obtained in real time through roadside equipment. Determining the number of vehicles in each lane at the current moment forms the second feature vector, which represents the vehicle distribution in each lane of the target intersection at that moment. The dimension of the second feature vector is determined by the total number of lanes at the target intersection, and the value of each dimension represents the number of vehicles in the corresponding lane. Similarly, the dimension of the third feature vector is also determined by the total number of lanes at the target intersection, and the value of each dimension represents the average travel distance in the corresponding lane.
[0099] The average travel distance of a lane is related to the total number of vehicles entering and exiting that lane. For the current moment, the first, second, and third feature vectors are integrated to obtain a vector matrix representing the current traffic state of the target intersection.
[0100] Understandably, for an intersection, only one of the phase types can be implemented at any given time; otherwise, traffic chaos will occur. Therefore, the first feature vector can be determined by encoding the current phase type using the total number of all phase types. See also Figure 3 The target intersections shown below have corresponding coding results for each phase type (unfilled fields indicate a red light status):
[0101] Table 1. Intersection Phase Types
[0102]
[0103] In the method provided in the above embodiments, the current phase type and current vehicle condition data of the target intersection are converted by a unified rule to generate data that the intelligent agent can recognize, thereby realizing the control of the traffic lights and improving the accuracy and robustness of the traffic light control system.
[0104] In one embodiment, the training process of the time prediction model includes:
[0105] Obtain a first-time training target model with the same structure as the time prediction model, and obtain a time evaluation model and a second-time training target model with the same structure as the time evaluation model;
[0106] Obtain training samples, and train the second time training target model based on the training samples, the time prediction model, and the time evaluation model;
[0107] The first-time training target model is trained based on the training samples and the second-time training target model.
[0108] Based on the parameters in the training target model and the parameters in the time evaluation model after training, the parameters in the time evaluation model are updated;
[0109] The parameters in the time prediction model are updated based on the parameters in the target model and the parameters in the time prediction model at the first moment after training.
[0110] The training samples are data obtained through simulation control of vehicle operation at the target intersection, or historical vehicle operation data and corresponding phase data acquired in real-world scenarios. It should be noted that the data included in the training samples describes the interaction experience between the target intersection and the agent controlling it, including the agent's input to the target intersection and its subsequent state (feedback). The acquired data can be used directly as sample data, or it can be processed through filtering, recombination, deduplication, and validation before being encoded and defined for each sample.
[0111] The time prediction model includes a sample duration strategy that selects the corresponding output control action based on the input state in the samples. The purpose of training the time prediction model is to determine the optimal sample duration strategy. However, if the prediction model is trained directly, changes in its parameters will affect the output, leading to significant training errors. Therefore, a first-time training target model with the same structure as the time prediction model is constructed. This first-time training target model is used to update parameters and select actions, while the time prediction model calculates the corresponding score value for each action according to a fixed method. Thus, the first-time training target model can be trained, and its parameters can then be directly copied to the time prediction model.
[0112] It is understood that the purpose of the first-time training target model is to update the parameters of the time prediction model. Therefore, the structure of the first-time training target model must be consistent with that of the time prediction model; otherwise, its parameters cannot be used to update the parameters of the time prediction model. Similarly, the structure of the second-time training target model should be consistent with that of the time evaluation model. The time prediction model is used to select the time corresponding to the phase action at the next moment based on the current traffic state. Therefore, the time prediction model should be a neural network model with a continuous action space, such as the Target Critic network in the Deep Deterministic Policy Gradient (DDPG) algorithm. The model outputs a data value representing a specific action; in this embodiment, the output action is the time allocation action.
[0113] In one embodiment, see Figure 4 The Target Actor network is the time prediction model, the Eval Actor network is the first-time training target model, the Target Critic network is the time evaluation model, and the Eval Critic network is the second-time training target model. Training samples are input into the Target Actor network, for example, sample (s, a, r, s′), where s represents the traffic state, a = (a... phase ,a time ) refers to a joint action, including phase assignment action a. phase and time allocation action a time Let r be the reward that the simulator at the target intersection gives to the time prediction model after executing the joint action a. Input (s,a,r,s′) into the Eval Actor network and output a new time-allocation action a. new =π(s;μ), where π(s;μ) is the time allocation strategy for training the target model Eval Actor network in the first time step, and μ is the weight of the Eval Actor network. Let s and a new The action a input to the Eval Critic network and its output to the Eval Actor network. new An evaluation is performed to obtain the evaluation value of the policy π(s;μ) of the Eval Actor network, and the parameters of the target model are adjusted based on the evaluation value.
[0114] In one embodiment, to deepen the exploration of the environment and obtain the globally optimal solution, we can... timeAdd random disturbances and control them between the minimum and maximum control durations, for example, the minimum control duration is 10s and the maximum control duration is 60s.
[0115] For the second-time training target model, the sample (s,a,r,s′) is input into the time prediction model to predict the action in the next state s′. Then, the difference between the target value and the expected value of the action is optimized, thereby updating the parameters of the second-time training target model Eval Critic network.
[0116] It should be noted that each training sample corresponds to a time step. When training the time prediction model, the parameters of the time prediction model need to be adjusted after each sample is input into the model. That is, at each time step, the parameter amplitudes of the trained second-time target model are transferred to the time prediction model to complete the update of the time prediction model. In addition, when updating the parameters of the time evaluation model and the time prediction model based on the first-time target model and the second-time target model, the weight parameters can be directly copied, or the weight parameters can be transformed through a specific method, which is not specifically limited here.
[0117] In the method provided in the above embodiments, the time prediction model is trained through a four-network model structure to complete the training of the prediction model for continuous time actions, thereby obtaining more accurate prediction results and making traffic light control more precise.
[0118] In one embodiment, the training process of the phase prediction model includes:
[0119] Obtain multiple training samples and a phase training target model with the same structure as the phase prediction model;
[0120] The phase training target model is trained based on multiple training samples and a phase prediction model;
[0121] The parameters in the phase prediction model are updated based on the parameters in the trained phase training target model.
[0122] As described above, if the phase prediction model is trained directly, its output will be affected by the training process, leading to significant errors. Therefore, a phase training target model with the same structure as the phase prediction model is obtained, and the parameters of the phase prediction model are updated through the phase training target model. Specifically, the role of the phase training target model is to update the parameters for action selection. During training, the phase prediction model calculates its reward evaluation value, which serves as the learning target. After training samples are input into the phase training target model, the model estimates the action value based on the current state of the training samples, and updates the parameters using the evaluation value and the value obtained from the phase training target model.
[0123] It should be noted that the method provided in this application uses multi-sample data to perform soft updates on the phase prediction model. That is, the training of the phase prediction model requires sample accumulation. In terms of sample acquisition, the samples acquired at each time step are used to train the phase training target model. In each preset period, the parameters of the phase prediction model are adjusted using the parameters of the phase training target model. For example, the time corresponding to the current sample is time step t, every n time steps is a training round, and the parameters of the phase prediction model are updated every M training rounds.
[0124] For the current training sample among N training samples, the traffic state of the next sample in the current training sample is input into the phase prediction model, and the corresponding sample phase policy and policy value prediction value are output (the phase prediction model selects the sample phase policy with the largest policy value prediction value as the output sample phase policy, which is used here to assist in building training labels); based on the policy value prediction value and the reward value in the current training sample, the training label of the current training sample for the phase training target model corresponding to the phase prediction model is determined.
[0125] The traffic state of each training sample is input into the phase prediction model, and the corresponding policy value prediction value is output. Based on the difference between the policy value prediction value of each training sample traffic state and the training label of each training sample, a third loss function is constructed. The phase training target model is trained based on the third loss function. The parameters in the phase prediction model are updated based on the parameters in the trained phase training target model.
[0126] In the method provided in the above embodiment, the phase prediction model is updated and trained by soft update to obtain a more accurate phase prediction model, thereby improving the accuracy of phase type selection in signal control and the robustness of the traffic light control system.
[0127] In one embodiment, see Figure 5 The process of obtaining training samples includes:
[0128] Step 502: Obtain sample traffic status, input the sample traffic status into the phase prediction model and the time prediction model respectively, and output the corresponding sample phase type and sample control duration;
[0129] Step 504: Perform traffic light simulation control according to the sample traffic state, the corresponding sample phase type and the sample control duration to obtain the next sample traffic state and the corresponding sample reward value after the traffic light simulation control. The sample reward value is used to characterize the degree of traffic condition improvement in the next sample traffic state.
[0130] Step 506: Take the sample traffic state as the previous sample traffic state and the next sample traffic state as the subsequent sample traffic state. The previous traffic state, the corresponding sample phase type, the corresponding sample control duration, the corresponding sample reward value, and the subsequent traffic state constitute the training sample.
[0131] The sample traffic state refers to traffic state data of a randomly acquired target intersection. For example, historical vehicle distribution data and corresponding phase data at each moment collected by roadside devices (such as millimeter-wave radar, lidar, cameras, etc.) passing through the target intersection can be imported into the traffic light simulator to prepare for interactive training between the time prediction model and the phase prediction model and the environment. The traffic light simulator is built based on the parameters of the target intersection, and this embodiment does not specifically limit the simulator's construction process. When training is required, at time step t, the traffic condition data and phase data of the intersection at the current time step t are obtained from the traffic light simulator to generate the current traffic state, which serves as a sample traffic state.
[0132] For the current sample traffic state s (based on the previous sample traffic state), after inputting it into the phase prediction model, the phase allocation action a is obtained. phase After being input into the time prediction model, the time allocation action 'a' is obtained. time Phase allocation action a phase and time allocation action a time Combining actions a = (a phase ,a time After the data is sent to the traffic light simulator for simulation control, the new traffic state of the target intersection will be obtained, which is the next sample traffic state s′ (the traffic state in the later sample). At the same time, the simulator may judge the degree of improvement of the traffic conditions of the target intersection by the joint action a based on the new traffic state of the target intersection, and feed back the reward r, which is the reward r corresponding to the current sample traffic state s.
[0133] The above process yields an interaction experience (s,a,r,s′), which can be stored as a training sample in the training data buffer pool. When it is necessary to train the phase prediction model and the time prediction model, a batch of training samples is obtained from the training data buffer pool, and the phase prediction model and the time prediction model are trained simultaneously.
[0134] In the method provided in the above embodiments, by simultaneously inputting sample traffic states into the time prediction model and the phase prediction model, a joint action is obtained, establishing the correlation between phase allocation and action allocation. The phase allocation action and the time allocation action are used as an interaction experience to constitute a training sample. During training, the same data source is used, and the two agents in the distributed decision region of this method use the same training data in each training iteration, thereby learning a matching optimal policy and achieving a joint optimal policy.
[0135] In one embodiment, see Figure 6 The training samples include the previous traffic conditions and the corresponding sample control duration; based on the training samples and the second-time training target model, the first-time training target model is trained, including:
[0136] Step 602: Based on the previous traffic state, the target model is trained in the second time to evaluate the sample control duration of the previous traffic state and obtain the corresponding score value.
[0137] Step 604: Construct a first loss function based on the score value, and train the parameters in the target model in the first time step using the first loss function.
[0138] For the current training sample (s,,,′), see Figure 4 The previous sample traffic state s will be input into the first-time training target model Eval Actor network to obtain the new time allocation action (sample control duration) a under the current sample traffic state s. new Then, the target model Eval Critic network, trained in a second time step, calculates the current sample traffic state action pair (,a) new ) rating in, π(s;μ) represents the weights of the Eval Critic network, where π(s;μ) is the time allocation strategy for training the target model Eval Actor network in the first time step, and μ is the weight of the target model Eval Actor network in the first time step.
[0139] It should be noted that the evaluation value This represents the cumulative expected return. In one embodiment, the gradient ascent algorithm can be used to maximize the cumulative expected return and update the parameters of the target model Eval Actornetwork trained at the first moment.
[0140] In one embodiment, the first loss function can be:
[0141]
[0142] The parameter μ in the target model trained in the first time step is updated using the first loss function.
[0143] In the method provided in the above embodiments, the parameters of the target model trained in the first time are updated by establishing a first loss function, thereby improving the model training effect and thus improving the control effect of the phase corresponding duration of the traffic lights.
[0144] In one embodiment, the training samples include the preceding traffic state, the corresponding sample control duration, the corresponding sample reward value, and the subsequent traffic state; based on the training samples, the time prediction model, and the time evaluation model, a second time training target model is trained, including:
[0145] The post-traffic state will be input into the time prediction model, and the corresponding sample control duration in the post-traffic state will be output.
[0146] Based on the subsequent traffic conditions, the control duration of the corresponding samples in the subsequent traffic conditions is evaluated using a time evaluation model to obtain the first score value;
[0147] Based on the first score and the corresponding sample reward value in the previous traffic state, the training labels of the training samples are determined for the training target model in the second time.
[0148] Based on the previous traffic conditions, the control duration of the corresponding samples under the previous traffic conditions is evaluated using a time evaluation model to obtain a second score value.
[0149] Based on the difference between the second score and the training label, a second loss function is constructed, and the parameters in the target model are trained using the second loss function at the second time step.
[0150] The subsequent traffic states in the training samples are input into the time prediction model, and the corresponding sample duration strategy of the subsequent traffic states in the training samples is output. Based on the subsequent traffic states in the training samples, the corresponding sample duration strategy of the subsequent traffic states in the training samples is scored by the time evaluation model to obtain the first score value. According to the first score value and the reward value of the training samples, the training label of the training samples for the second time training target model corresponding to the time evaluation model is determined.
[0151] Based on the preceding traffic states in the training samples, a time evaluation model is used to score the corresponding sample duration strategies of the preceding traffic states in the training samples to obtain a second score value; based on the difference between the second score value and the training label, a second loss function is constructed, and the parameters in the second time training objective model are trained based on the second loss function; based on the parameters in the trained second time training objective model and the parameters in the time evaluation model, the parameters in the time evaluation model are updated.
[0152] In one embodiment, see Figure 4 The second loss function is constructed as follows:
[0153]
[0154] Where, a′ time =(s′;μ′) represents the sample control duration corresponding to the subsequent traffic conditions; a time =π(s;μ) represents the sample control duration corresponding to the previous traffic conditions; As the first score, As the second score, The network weights of the Target Criticnetwork are used for time evaluation models.
[0155] In one embodiment, the training samples include the preceding traffic state, the corresponding sample phase type, the corresponding sample reward value, and the following traffic state; based on multiple training samples and the phase prediction model, the phase training target model is trained, including:
[0156] For the current training sample among multiple training samples, the subsequent traffic state of the current training sample is input into the phase prediction model, and the corresponding sample phase type and strategy value prediction value of the subsequent traffic state are output.
[0157] Based on the predicted value of the strategy corresponding to the subsequent traffic state and the sample reward value corresponding to the preceding traffic state, the training label of the current training sample for the phase training target model is determined.
[0158] The preceding traffic conditions are input into the phase prediction model, and the corresponding policy value prediction value is output in the preceding traffic conditions.
[0159] A third loss function is constructed based on the differences between the predicted policy value values corresponding to the previous traffic states and the corresponding training labels in each training sample.
[0160] The parameters in the phase training target model are trained using a third loss function.
[0161] In one embodiment, see Figure 4 The third loss function can be:
[0162] L(θ)=E[(r+γQ′ target (s′,a′ phase ;θ′)-Q eval (s,a phase ;θ)) 2 ]
[0163] Where θ represents the network weights of the phase training target model Eval network; θ′ represents the network weights of the phase prediction model Target network; a′ phase Represents the maximum Q′ target The corresponding phase selection action is then performed. The network weights θ of the Eval network are updated using the third loss function described above.
[0164] In the method provided in the above embodiments, the parameters of the phase training target model are updated by establishing a third loss function, thereby improving the model training effect and thus improving the control effect of the phase corresponding duration of the traffic lights.
[0165] In one embodiment, a traffic light control method based on an intelligent agent is provided; see [link to relevant documentation]. Figure 7 The training process of an intelligent agent includes:
[0166] Get real-time intersection information;
[0167] The information is processed to obtain the current phase information and vehicle distribution. The phase information is one-hot encoded, and the vehicle distribution includes the queue length of each lane and the distance the vehicle travels at the intersection. Then, the state s is constructed and input into the neural network.
[0168] The traffic light timing task is decomposed into two subtasks: a traffic light phase decision subtask and a traffic light time allocation subtask. Each subtask is decided by a different decision region. Then, corresponding discrete phase decision agents and continuous time allocation agents are constructed. This method uses the DQN (Deep Q Network) algorithm and the DDPG (Deep Deterministic Policy Gradient) algorithm to learn the optimal discrete phase selection strategy and the optimal continuous time allocation strategy, respectively. These two algorithms simultaneously complete the traffic light timing task at a given moment.
[0169] During the training data collection phase, the two decision-making agents continuously interact with the intersection. This method employs a distributed execution approach, simultaneously inputting the current time step state s into the phase decision region and the time allocation region, allowing their respective agents to simultaneously decide on the action a. phase and a timeThe decision result is applied to the intersection, which provides the agent with a reward r. Simultaneously, the intersection state changes, transitioning from state s to the next state s′. Then, (s, a, r, s′) (where a = (a phase ,a time The data is stored in the training data cache pool H, and the interaction ends at the current time step.
[0170] During the agent learning phase, the agent for each decision region draws a batch of interaction experiences from the cache pool for training to improve its decision-making ability. The training data collection phase and the agent learning phase alternate continuously until the proposed algorithm converges and the optimal joint policy is obtained. Specifically, DQN updates the Target network every M training rounds using a soft update method; all Target network weights in DDPG are updated at every time step.
[0171] This distributed execution method simultaneously distributes the state containing implicit coupling information at the current time step to different decision regions. The agent can extract the coupling relationship from the state through a neural network to implement a matching joint strategy, thereby achieving joint optimization and improving the traffic efficiency brought about by the agent controlling the traffic lights.
[0172] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0173] Based on the same inventive concept, this application also provides a traffic signal control device for implementing the traffic signal control method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more traffic signal control device embodiments provided below can be found in the limitations of the traffic signal control method described above, and will not be repeated here.
[0174] In one embodiment, such as Figure 8 As shown, a traffic signal light control device is provided, including: a data acquisition module 801, a state generation module 802, an action prediction module 803, and an action execution module 804, wherein:
[0175] The data acquisition module 801 is used to acquire the current phase type of the traffic lights being used and the current traffic conditions of the target intersection for the traffic lights in the target intersection.
[0176] The status generation module 802 is used to generate the current traffic status of the target intersection based on the current phase type and current road condition data.
[0177] The motion prediction module 803 is used to input the current traffic state into the trained phase prediction model and the trained time prediction model respectively, and output the target phase type and target control duration.
[0178] The action execution module 804 is used to control the traffic lights to switch to the target phase corresponding to the target phase type, and to control the duration of the target phase according to the target control duration.
[0179] In one embodiment, the state generation module 802 is further configured to:
[0180] Based on the current phase type and the number of all phase types, generate a first feature vector to characterize the current phase type;
[0181] A second feature vector is generated based on the current number of vehicles in each lane and the total number of lanes. A third feature vector is generated based on the average driving distance in each lane and the total number of lanes.
[0182] The current traffic status of the target intersection is generated based on the first feature vector, the second feature vector, and the third feature vector.
[0183] In one embodiment, the traffic signal control device further includes a first training module for:
[0184] Obtain a first-time training target model with the same structure as the time prediction model, and obtain a time evaluation model and a second-time training target model with the same structure as the time evaluation model;
[0185] Obtain training samples, and train the second time training target model based on the training samples, the time prediction model, and the time evaluation model;
[0186] The first-time training target model is trained based on the training samples and the second-time training target model.
[0187] Based on the parameters in the training target model and the parameters in the time evaluation model after training, the parameters in the time evaluation model are updated;
[0188] The parameters in the time prediction model are updated based on the parameters in the target model and the parameters in the time prediction model at the first moment after training.
[0189] In one embodiment, the traffic signal control device further includes a second training module for:
[0190] Obtain multiple training samples and a phase training target model with the same structure as the phase prediction model;
[0191] The phase training target model is trained based on multiple training samples and a phase prediction model;
[0192] The parameters in the phase prediction model are updated based on the parameters in the trained phase training target model.
[0193] In one embodiment, the traffic signal control device further includes a training sample acquisition module, used for:
[0194] Obtain sample traffic status, input the sample traffic status into the phase prediction model and the time prediction model respectively, and output the corresponding sample phase type and sample control duration;
[0195] Traffic light simulation control is performed based on the sample traffic state, the corresponding sample phase type, and the sample control duration. The next sample traffic state and the corresponding sample reward value are obtained after the traffic light simulation control. The sample reward value is used to characterize the degree of traffic condition improvement in the next sample traffic state.
[0196] The sample traffic state is taken as the previous sample traffic state, and the next sample traffic state is taken as the subsequent sample traffic state. The previous traffic state, the corresponding sample phase type, the corresponding sample control duration, the corresponding sample reward value, and the subsequent traffic state constitute the training sample.
[0197] In one embodiment, the first training module is further configured to:
[0198] Based on the previous traffic conditions, the target model is trained in the second time to evaluate the control duration of the corresponding samples in the previous traffic conditions and obtain the corresponding evaluation value.
[0199] A first loss function is constructed based on the evaluation value, and the parameters in the target model are trained using the first loss function.
[0200] In one embodiment, the first training module is further configured to:
[0201] The post-traffic state will be input into the time prediction model, and the corresponding sample control duration in the post-traffic state will be output.
[0202] Based on the subsequent traffic conditions, the control duration of the corresponding samples in the subsequent traffic conditions is evaluated using a time evaluation model to obtain the first score value;
[0203] Based on the first score and the corresponding sample reward value in the previous traffic state, the training labels of the training samples are determined for the training target model in the second time.
[0204] Based on the previous traffic conditions, the control duration of the corresponding samples under the previous traffic conditions is evaluated using a time evaluation model to obtain a second score value.
[0205] Based on the difference between the second score and the training label, a second loss function is constructed, and the parameters in the target model are trained using the second loss function at the second time step.
[0206] In one embodiment, the second training module is further configured to:
[0207] For the current training sample among multiple training samples, the subsequent traffic state of the current training sample is input into the phase prediction model, and the corresponding sample phase type and strategy value prediction value of the subsequent traffic state are output.
[0208] Based on the predicted value of the strategy corresponding to the subsequent traffic state and the sample reward value corresponding to the preceding traffic state, the training label of the current training sample for the phase training target model is determined.
[0209] The preceding traffic conditions are input into the phase prediction model, and the corresponding policy value prediction value is output in the preceding traffic conditions.
[0210] A third loss function is constructed based on the differences between the predicted policy value values corresponding to the previous traffic states and the corresponding training labels in each training sample.
[0211] The parameters in the phase training target model are trained using a third loss function.
[0212] Each module in the aforementioned traffic signal control device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the computer device's memory as software, so that the processor can call and execute the corresponding operations of each module.
[0213] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores traffic status data. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a traffic light control method.
[0214] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0215] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform all the steps provided in the above embodiments.
[0216] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, performs all the steps provided in the above embodiments.
[0217] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements all the steps provided in the above embodiments.
[0218] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0219] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0220] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A traffic signal light control method, characterized in that, The method includes: For the traffic lights at the target intersection, obtain the current phase type of the traffic lights and the current traffic conditions of the target intersection; Based on the current phase type and the current traffic data, generate the current traffic state of the target intersection; The current traffic state is simultaneously input into the trained phase prediction model and the trained time prediction model, and the target phase type and target control duration are output. The phase prediction model is a discrete phase selection agent, which is used to determine the target phase type corresponding to the next time step of the current time step. The time prediction model is a continuous time allocation agent, which is used to obtain the target control duration corresponding to the target phase type. Control the traffic light to switch to the target phase corresponding to the target phase type, and control the duration of the target phase according to the target control duration.
2. The method according to claim 1, characterized in that, The current traffic data includes the current number of vehicles in each lane at the target intersection and the average travel distance of vehicles in each lane after the current phase type is put into use. The step of generating the current traffic state of the target intersection based on the current phase and the current road condition data includes: Based on the current phase type and the number of all phase types, a first feature vector is generated to characterize the current phase type; A second feature vector is generated based on the current number of vehicles in each lane and the total number of lanes; a third feature vector is generated based on the average driving distance in each lane and the total number of lanes. Based on the first feature vector, the second feature vector, and the third feature vector, the current traffic status of the target intersection is generated.
3. The method according to claim 1, characterized in that, The training process of the time prediction model includes: Obtain a first time training target model with the same structure as the time prediction model, and obtain a time evaluation model and a second time training target model with the same structure as the time evaluation model; Obtain training samples, and train the second time training target model based on the training samples, the time prediction model, and the time evaluation model; Based on the training samples and the second time-training target model, the first time-training target model is trained. Based on the parameters in the second-time training target model after training and the parameters in the time evaluation model, the parameters in the time evaluation model are updated; The parameters in the time prediction model are updated based on the parameters in the target model trained at the first moment after training and the parameters in the time prediction model.
4. The method according to claim 3, characterized in that, The process of obtaining the training samples includes: Obtain sample traffic status, input the sample traffic status into the phase prediction model and the time prediction model respectively, and output the corresponding sample phase type and sample control duration; According to the sample traffic state, the corresponding sample phase type and the sample control duration, traffic light simulation control is performed to obtain the next sample traffic state and the corresponding sample reward value after the traffic light simulation control. The sample reward value is used to characterize the degree of traffic condition improvement in the next sample traffic state. The sample traffic state is taken as the previous sample traffic state, and the next sample traffic state is taken as the subsequent sample traffic state. The previous sample traffic state, the corresponding sample phase type, the corresponding sample control duration, the corresponding sample reward value, and the subsequent sample traffic state constitute the training sample.
5. The method according to claim 3, characterized in that, The training samples include the previous traffic conditions and the corresponding sample control duration; the training of the first time-based training target model based on the training samples and the second time-based training target model includes: Based on the current traffic conditions, the target model is trained using the second time to evaluate the sample control duration corresponding to the current traffic conditions, and obtain the corresponding evaluation value. A first loss function is constructed based on the evaluation value, and the parameters in the target model at the first time are trained using the first loss function.
6. The method according to claim 3, characterized in that, The training samples include the preceding traffic state, the corresponding sample control duration, the corresponding sample reward value, and the subsequent traffic state; the training of the second time training target model based on the training samples, the time prediction model, and the time evaluation model includes: The subsequent traffic state is input into the time prediction model, and the corresponding sample control duration for the subsequent traffic state is output. Based on the subsequent traffic state, the time evaluation model is used to evaluate the corresponding sample control duration of the subsequent traffic state to obtain a first score value; Based on the first score and the sample reward value corresponding to the previous traffic state, the training label of the training sample for the second time training target model is determined; Based on the current traffic conditions, the time evaluation model is used to evaluate the sample control duration corresponding to the current traffic conditions to obtain a second score value. Based on the difference between the second score and the training label, a second loss function is constructed, and the parameters in the second time training target model are trained using the second loss function.
7. A traffic signal light control device, characterized in that, The device includes: The data acquisition module is used to acquire the current phase type of the traffic lights in the target intersection and the current traffic conditions of the target intersection. The status generation module is used to generate the current traffic status of the target intersection based on the current phase type and the current traffic data. The action prediction module is used to simultaneously input the current traffic state into the trained phase prediction model and the trained time prediction model, and output the target phase type and target control duration. The phase prediction model is a discrete phase selection agent, used to determine the target phase type corresponding to the next moment of the current moment. The time prediction model is a continuous time allocation agent, used to obtain the target control duration corresponding to the target phase type. The action execution module is used to control the traffic light to switch to the target phase corresponding to the target phase type, and to control the duration of the target phase according to the target control duration.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.