A training method and device of a trunk traffic signal control model

By using a hierarchical trunk traffic signal control model and leveraging the collaborative efforts of the master control agent and the working agents, the challenges of trunk traffic signal control have been solved, enabling efficient traffic management in complex environments.

CN118171097BActive Publication Date: 2025-11-04HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211584906.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-09
Publication Date
2025-11-04
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

How to effectively control traffic signals on main roads to alleviate traffic congestion and improve traffic efficiency.

Method used

A hierarchical trunk traffic signal control model is adopted. The master control agent performs global traffic data analysis to generate global tasks, and the working agents perform intersection task analysis to generate traffic signal control strategies. The agent parameters are adjusted based on observation data until the training conditions are met.

Benefits of technology

It enables adaptive adjustment in highly dynamic and complex traffic environments, taking into account both global and local objectives, thereby improving the feasibility and efficiency of traffic signal control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118171097B_ABST
    Figure CN118171097B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of arterial traffic signal control model training method and device, comprising: obtaining the sample global traffic data of to be controlled arterial road generates the global traffic feature of to be controlled arterial road;Global task of to be controlled arterial road is obtained by using master control agent analysis;Using each working agent, the intersection task corresponding to the working agent is analyzed to obtain the traffic signal control strategy of the intersection corresponding to the working agent;According to the traffic signal control strategy of each intersection in to be controlled arterial road, the observation global traffic data of to be controlled arterial road is determined;Adjust the parameter of at least one of master control agent and each working agent;Select other sample global traffic data of to be controlled arterial road to continue to train arterial traffic signal control model, until meeting model training end condition, obtain the trained arterial traffic signal control model.The application realizes the training of arterial traffic signal control model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of traffic management and control technology, and in particular to a training method and apparatus for a trunk line traffic signal control model. Background Technology

[0002] With the continuous advancement of urbanization, the issue of arterial road traffic signal control has gradually become one of the most important problems in the transportation system. Reasonable traffic signal control can effectively alleviate traffic congestion and improve the overall efficiency of traffic operations. Therefore, how to control arterial road traffic signals has become a pressing technical problem to be solved. Summary of the Invention

[0003] The purpose of this application is to provide a training method and apparatus for a trunk line traffic signal control model, so as to achieve control of trunk line traffic signals. The specific technical solution is as follows:

[0004] In a first aspect, embodiments of this application provide a training method for a trunk line traffic signal control model, wherein the trunk line traffic signal control model includes a master control agent and multiple working agents, each of the working agents corresponding to an intersection on the trunk line to be controlled; the method includes:

[0005] Obtain sample global traffic data of the arterial road to be controlled; generate global traffic features of the arterial road to be controlled based on the sample global traffic data;

[0006] The global traffic features are analyzed using the master control agent to obtain the global task of the trunk line to be controlled. The global task includes multiple intersection tasks, and each intersection task corresponds to a working agent.

[0007] For each working agent, the task of the intersection corresponding to the working agent is analyzed to obtain the traffic signal control strategy of the intersection corresponding to the working agent.

[0008] Based on the traffic signal control strategies of each intersection in the arterial road to be controlled, determine the observed global traffic data of the arterial road to be controlled;

[0009] Based on the observed global traffic data, the parameters of the master control agent and at least one of the working agents are adjusted; other sample global traffic data of the trunk line to be controlled are selected, and the trunk line traffic signal control model is trained until the model training termination condition is met, thus obtaining the trained trunk line traffic signal control model. The trunk line traffic signal control model is used to output the traffic signal control strategy of each intersection on the trunk line to be controlled when the current global traffic data of the trunk line to be controlled is input.

[0010] Secondly, embodiments of this application provide a training device for a trunk line traffic signal control model, wherein the trunk line traffic signal control model includes a master control agent and multiple working agents, each of the working agents corresponding to an intersection on the trunk line to be controlled; the device includes:

[0011] The first data acquisition module is used to acquire sample global traffic data of the trunk line to be controlled; and to generate global traffic features of the trunk line to be controlled based on the sample global traffic data.

[0012] The global task acquisition module is used to analyze the global traffic features using the main control agent to obtain the global task of the trunk line to be controlled. The global task includes multiple intersection tasks, and each intersection task corresponds to a working agent.

[0013] The first strategy acquisition module is used to analyze the intersection task corresponding to each working agent and obtain the traffic signal control strategy for the intersection corresponding to the working agent.

[0014] The data determination module is used to determine the observed global traffic data of the trunk line to be controlled based on the traffic signal control strategies of each intersection in the trunk line to be controlled.

[0015] The model training module is used to adjust the parameters of the master control agent and at least one of the working agents based on observed global traffic data; select other sample global traffic data of the trunk line to be controlled, and continue to train the trunk line traffic signal control model until the model training termination condition is met, thereby obtaining a trained trunk line traffic signal control model. The trunk line traffic signal control model is used to output the traffic signal control strategy of each intersection on the trunk line to be controlled when the current global traffic data of the trunk line to be controlled is input.

[0016] Thirdly, embodiments of this application provide an electronic device, including:

[0017] Memory, used to store computer programs;

[0018] When the processor executes the program stored in the memory, it implements the training method or the trunk traffic signal control method of any of the trunk traffic signal control models described in the embodiments of this application.

[0019] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the training method or the trunk traffic signal control method of any of the trunk traffic signal control models described in the embodiments of this application.

[0020] Fifthly, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to execute any of the above-described methods for training or controlling trunk traffic signals.

[0021] Beneficial effects of the embodiments in this application:

[0022] The training method for the arterial traffic signal control model provided in this application, after acquiring sample global traffic data of the arterial line to be controlled and generating global traffic features of the arterial line to be controlled, firstly, uses a master control agent to analyze the global traffic features, that is, to perform a holistic global analysis of the arterial line to be controlled, to obtain global tasks including intersection tasks corresponding one-to-one with each working agent in the arterial line to be controlled; then, for each working agent, it uses the agent to perform local analysis of its corresponding intersection task, thereby obtaining the traffic signal control strategy for the intersection corresponding to that working agent. Then, based on the traffic signal control strategy of each intersection, the observed global traffic data of the arterial line to be controlled is determined, and the parameters of the master control agent and at least one agent in each working agent are adjusted accordingly. Afterwards, other sample global traffic data of the arterial line to be controlled are selected again to continue training the arterial traffic signal control model until the model training termination condition is met, resulting in a trained arterial traffic signal control model. The arterial traffic signal control model is used to output the traffic signal control strategy for each intersection in the arterial line to be controlled, given the current global traffic data of the arterial line to be controlled. This application innovatively proposes a hierarchical arterial coordinated signal control scheme. It utilizes a master control agent to determine the overall global coordinated objective of the arterial to be controlled based on global traffic data. Then, each lower-level working agent determines its local objective based on its respective intersection task. Simultaneously, the scheme optimizes and achieves the global coordinated objective. This hierarchical structure allows for better task decomposition and organically combines the global objectives of the arterial to be controlled with the local objectives of each working agent. While achieving the global coordinated objective of the arterial to be controlled, it also considers the local objectives of each working agent. Furthermore, it can adaptively adjust according to the traffic scenario, making the trained model more adaptable to the highly dynamic and complex traffic environment of the arterial to be controlled, thus providing a more feasible solution for traffic signal control.

[0023] Of course, implementing any product or method of this application does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.

[0025] Figure 1a This is a first schematic diagram of a training method for a trunk traffic signal control model according to an embodiment of this application;

[0026] Figure 1b This is a second schematic diagram of a training method for a trunk traffic signal control model according to an embodiment of this application;

[0027] Figure 2a This is a flowchart illustrating the workflow of a working intelligent agent according to an embodiment of this application.

[0028] Figure 2b This is a schematic diagram of a global traffic feature generation method according to an embodiment of this application;

[0029] Figure 2c This is a schematic diagram of a global task generation method according to an embodiment of this application;

[0030] Figure 3 This is a third schematic diagram of a training method for a trunk traffic signal control model according to an embodiment of this application;

[0031] Figure 4 This is a fourth schematic diagram of a training method for a trunk traffic signal control model according to an embodiment of this application;

[0032] Figure 5 This is a first schematic diagram of the reward value calculation method according to an embodiment of this application;

[0033] Figure 6 This is a schematic diagram of a trunk traffic signal control method according to an embodiment of this application;

[0034] Figure 7 This is a schematic diagram of a training device for a trunk traffic signal control model according to an embodiment of this application;

[0035] Figure 8 This is a schematic diagram of a trunk traffic signal control device according to an embodiment of this application;

[0036] Figure 9 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0037] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.

[0038] First, the terminology used in the embodiments of this application will be explained:

[0039] Markov Decision Process (MDP): Used to simulate the stochastic policies and rewards that an agent can achieve in systems with Markov properties.

[0040] Reinforcement Learning (RL): Also known as reinforcement learning, it is used to describe and solve the problem of how an agent learns strategies to maximize rewards or achieve specific goals during its interaction with the environment.

[0041] Multi-agent reinforcement learning: A system that includes multiple agents, which share an environment and influence each other.

[0042] Hierarchical Deep Reinforcement Learning (HRL): A framework for reinforcement learning that typically consists of two layers, where the upper layer is responsible for generating sub-objectives and the lower layer is responsible for completing the given sub-objectives.

[0043] The Actor-Critic Algorithm consists of two parts: an Actor and a Critic. The Actor selects actions based on probabilities, while the Critic evaluates the actions based on the Actor's performance. The Actor then adjusts the probability of selecting actions based on the Critic's rating.

[0044] Deep Deterministic Policy Gradient (DDPG): An online deep reinforcement learning algorithm within the Actor-Critic framework that can solve continuous control problems.

[0045] Environment: This specification includes real traffic environments, simulated traffic environments, and intelligent agent training environments.

[0046] Intelligent Agent: refers to an abstraction of an entity that can learn on its own and interact with its environment. In this specification, a single intersection can be abstracted as an intelligent agent.

[0047] Policy: What action should be taken in a specific state to maximize the cumulative reward.

[0048] Reward: refers to quantifiable information fed back to the agent in relation to the target environment for optimizing the agent, used to judge the quality of the agent's actions.

[0049] Arterial Traffic Signal Control: A traffic signal control method that links the traffic signal controllers of multiple intersections on a main road to implement coordinated control.

[0050] Phase difference (Offset): In coordinated control, the difference between the start time of the phase or cycle of the coordinated intersection and that of the designated reference intersection, or the time difference between the start time of the phase or cycle of the coordinated intersection and the specified reference time.

[0051] Signal Cycle: The complete process of a signal light color changing in a set signal phase sequence for one cycle.

[0052] Signal Phase: The sequence of consecutive signal light colors displayed when one or more traffic flows simultaneously have the right of way within a signal cycle.

[0053] Ring: A set of two or more flows that are given timing in sequence, where only one flow in the set can receive the green light at any given time.

[0054] Cycle-based control: This scheme assigns a certain green light duration to each phase of the intersection in a cycle to achieve the intersection control objective.

[0055] Switch-based control: The switching rules in switch-based traffic light control follow a certain phase sequence pattern. Before the end of each phase, it is decided whether to switch to the next phase.

[0056] Long Time Horizon and Sparse Reward Problem: During the exploration process, the agent may experience difficulties in obtaining rewards for extended periods or the rewards may be delayed, leading to slow learning.

[0057] To achieve control of arterial traffic signals, this application provides a training method and apparatus for an arterial traffic signal control model, which are described in detail below:

[0058] See Figure 1a and Figure 1b This application provides a flowchart illustrating a first method for training a trunk traffic signal control model. The trunk traffic signal control model includes a master control agent and multiple working agents, each working agent corresponding to an intersection on the trunk line to be controlled. The method includes:

[0059] Step S11: Obtain sample global traffic data of the trunk line to be controlled; generate global traffic features of the trunk line to be controlled based on the sample global traffic data.

[0060] A controlled artery has multiple intersections, each corresponding to a worker agent. The entire controlled artery corresponds to a manager agent, which acquires traffic data from each intersection along the entire artery. In one example, a worker agent can observe the traffic data of its corresponding intersection and send it to the manager agent, thus enabling the manager agent to observe traffic data from all intersections. The manager agent acquires sample global traffic data for the controlled artery, specifically sample traffic data from each intersection. This sample global traffic data can include queue lengths, traffic flow, number of arriving vehicles, vehicle waiting times, green light execution times, and other relevant traffic and signal data for each intersection, which can be configured according to actual needs.

[0061] In one example, the global traffic data sample can be data collected in real time based on devices such as counters or cameras, or it can be a historically collected and organized sample dataset. Specifically, the global traffic data sample can be the traffic data of each intersection on the main road to be controlled within a preset time period. The preset time period can be set according to actual needs, such as 3 minutes, 10 minutes, 1 hour, 12 hours, 1 day, or 1 week.

[0062] After obtaining the global traffic data, features are extracted from it to obtain the global traffic features of the arterial road to be controlled. Specifically, feature extraction can be performed using feature extraction neural networks or feature extraction algorithms, etc., which are not limited here. The obtained global traffic features can represent the traffic characteristics of each intersection on the arterial road to be controlled and reflect the overall traffic characteristics of the arterial road to be controlled.

[0063] Step S12: Analyze the global traffic features using the main control agent to obtain the global task of the trunk line to be controlled.

[0064] The global task includes multiple intersection tasks, each corresponding to a worker agent. Specifically, an intersection task can be the traffic signal cycle, green light execution time, signal phase difference, and cycle difference between the corresponding intersection and other intersections. The global task can represent the overall traffic signal control scheme for the entire arterial road to be controlled.

[0065] In one example, the controlling agent analyzes global traffic features using reinforcement learning algorithms. Specifically, Actor-Critic algorithms can be employed, such as A2C (Advantage Actor-Critic), A3C (Asynchronous Advantage Actor-Critic), and DDPG (deep deterministic policy gradient). This involves first generating example global tasks using global traffic features, then evaluating the feasibility of these example global tasks and the potential traffic conditions on the controlled artery after application. During the evaluation process, better example global tasks are continuously generated, ultimately yielding the optimal global task for the controlled artery.

[0066] In one embodiment of this application, the global task of the trunk line to be controlled represents the target phase difference between adjacent intersections in the trunk line to be controlled; for any two adjacent intersections, the target phase difference between the two adjacent intersections represents the difference between the start times of the green light signals between the two adjacent intersections; for any intersection, the intersection task of that intersection represents the target phase difference corresponding to that intersection. In one example, the target phase difference is constant within a preset time period.

[0067] Therefore, the global task of the arterial road to be controlled obtained in this application embodiment is the difference between the start times of the green light signals at adjacent intersections. In other words, the training objective of the arterial road traffic signal control model in this application is periodic signal control, which involves configuring a certain green light signal duration for each phase of each intersection according to the cycle. This application embodiment, while ensuring the signal cycle remains unchanged, achieves coordinated control of the arterial road by changing the phase difference between adjacent intersections. Compared to the switching control strategy in related technologies, this reduces the probability of traffic safety problems that may occur due to frequent signal switching and improves the feasibility of the traffic control scheme determined based on the arterial road traffic signal control model obtained in this application embodiment.

[0068] Step S13: For each working agent, analyze the intersection task corresponding to the working agent to obtain the traffic signal control strategy for the intersection corresponding to the working agent.

[0069] After receiving the global task, the master agent distributes each intersection task within the global task to the corresponding worker agents. Each worker agent then analyzes its assigned intersection task. For example, each worker agent can use reinforcement learning algorithms to analyze the intersection task, or it can use an Actor-Critic Algorithm, such as A2C, A3C, or DDPG. This involves generating example traffic signal control strategies based on the intersection tasks, evaluating the potential traffic conditions at the intersection after applying these strategies, and continuously generating better example traffic signal control strategies during the evaluation process. Finally, the optimal traffic signal control strategy for the intersection is obtained. Specifically, the traffic signal control strategy can include the intersection's traffic signal phases (the continuous signal display sequence of one or more traffic flows simultaneously gaining the right-of-way within a cycle), traffic signal control cycle (the complete process of signals changing in a set signal phase sequence for one cycle), and green light execution time.

[0070] In one example, the working process of a working intelligent agent can be as follows: Figure 2a As shown, the intersection task is embedded into the Worker agent. After analysis by the Worker agent, the traffic signal control strategy for the intersection is output and executed. In addition, the Worker agent can also be used to observe the traffic data of the intersection (intersection sample data).

[0071] Step S14: Determine the observed global traffic data of the trunk line to be controlled based on the traffic signal control strategies of each intersection in the trunk line to be controlled.

[0072] Global traffic data can be obtained through a simulation model of the arterial road to be controlled, or through sensors installed on the arterial road. After each working agent on the arterial road obtains the traffic signal control strategy for the corresponding intersection, the entire arterial road is run based on the traffic signal control strategy for each intersection. Then, the global traffic data of the arterial road within a preset time period is acquired again, that is, the traffic data of each intersection on the arterial road, as the observed global traffic data of the arterial road.

[0073] In one example, the training of the arterial traffic signal control model can be pre-trained before the model is put into formal use on the arterial traffic. In this case, the operation of the above-mentioned traffic signal control strategy based on each intersection on the entire arterial road to be controlled can be a simulated operation, and the obtained global traffic data observed on the arterial road to be controlled is the simulated operation data. Alternatively, the training of the arterial traffic signal control model can be real-time training and updating after the model has been put into formal use on the arterial traffic. In this case, the operation of the above-mentioned traffic signal control strategy based on each intersection on the entire arterial road to be controlled can be a formal operation, and the obtained global traffic data observed on the arterial road to be controlled is the real operation data.

[0074] Step S15: Based on the observed global traffic data, adjust the parameters of the master control agent and at least one of the working agents; select other sample global traffic data of the trunk line to be controlled, and continue to train the trunk line traffic signal control model until the model training termination condition is met, and obtain the trained trunk line traffic signal control model.

[0075] The trunk line traffic signal control model is used to output traffic signal control strategies for each intersection on the trunk line to be controlled, given the current global traffic data of the trunk line to be controlled.

[0076] After obtaining the global traffic data, the parameters of the master control agent and at least one of the working agents are adjusted. This can be done by adjusting the parameters of both the master control agent and each working agent before continuing training, or by adjusting the parameters of only each working agent or only some of the working agents before continuing training.

[0077] In one example, the decision to adjust parameters can be based on observed global traffic data. For instance, if the global traffic data indicates problems with overall traffic conditions on the controlled artery, the parameters of both the master controller and each worker agent can be adjusted. Conversely, if the global traffic data indicates problems with traffic conditions at intersections corresponding to some worker agents on the controlled artery, only the parameters of those worker agents can be adjusted. Alternatively, the decision to adjust the parameters of the master controller and each worker agent can be based on actual training needs during the training process.

[0078] During the continued training of the arterial traffic signal control model, the selected global traffic data of other samples of the arterial line to be controlled can be global traffic data of samples that have not been selected during the training process. In one example, it can be global traffic data of samples of the arterial line to be controlled within a preset time period after a period of time. Based on this, the arterial traffic signal control model is trained until the preset model training end condition is met, and the trained arterial traffic signal control model is obtained.

[0079] In one example, the end condition for model training could be that the number of training iterations has reached a pre-set threshold, such as two thousand or five thousand times; the end condition could also be that the loss of the mainline traffic signal control model converges or the reward value of the working agent converges.

[0080] As mentioned above, the training of the arterial traffic signal control model can be performed in real-time after the model has been put into formal use on arterial traffic. Therefore, after the model training is completed, the model used for arterial traffic signal control is updated to the trained arterial traffic signal control model. For example, if the characteristics of global traffic data no longer meet the pre-set standard characteristics within several consecutive preset time periods—for instance, if a certain proportion of intersection queue lengths exceed the pre-set maximum queue length, the longest vehicle waiting time exceeds the pre-set maximum waiting time, or the traffic flow is not within the pre-set standard flow range—then the arterial traffic signal control model can be trained and updated again in real-time.

[0081] As can be seen from the above, the training method for the arterial traffic signal control model provided in this application, after acquiring sample global traffic data of the arterial line to be controlled and generating global traffic features of the arterial line to be controlled, firstly uses the master control agent to analyze the global traffic features, that is, to perform a holistic global analysis of the arterial line to be controlled, to obtain global tasks including intersection tasks corresponding one-to-one with each working agent in the arterial line to be controlled; then, for each working agent, it uses the working agent to perform local analysis of its corresponding intersection task, thereby obtaining the traffic signal control strategy for the intersection corresponding to the working agent. Then, based on the traffic signal control strategy of each intersection, the observed global traffic data of the arterial line to be controlled is determined, and the parameters of the master control agent and at least one agent in each working agent are adjusted accordingly. Afterwards, other sample global traffic data of the arterial line to be controlled are selected again, and the arterial traffic signal control model is trained again until the model training termination condition is met, resulting in a trained arterial traffic signal control model. The arterial traffic signal control model is used to output the traffic signal control strategy of each intersection in the arterial line to be controlled when the current global traffic data of the arterial line to be controlled is input.

[0082] In the training method of the arterial traffic signal control model in this application embodiment, a hierarchical arterial coordinated signal control scheme is proposed. This scheme utilizes a master control agent to determine the overall global coordinated goal of the arterial to be controlled through global traffic data. Then, each working agent in the lower layer determines its own local goal based on its respective intersection task. At the same time, the optimization and realization of the global coordinated goal are completed. This hierarchical structure can better decompose the task and organically combine the global goal of the arterial to be controlled and the local goals of each working agent. While realizing the global coordinated goal of the arterial to be controlled, it also takes into account the local goals of each working agent. Furthermore, it can adaptively adjust according to the traffic scenario, so that the trained model can better adapt to the highly dynamic and complex traffic environment of the arterial to be controlled, providing a more feasible solution for traffic signal control.

[0083] In one possible implementation, the sample global traffic data includes intersection sample data for each intersection on the arterial road to be controlled; such as... Figure 2b As shown, step S11 above generates the global traffic characteristics of the arterial road to be controlled based on the sample global traffic data, including:

[0084] Step S21: For each intersection, generate the intersection features based on the intersection sample data of that intersection;

[0085] Step S22: Concatenate the intersection features of each intersection to obtain a concatenation matrix.

[0086] Step S23: Perform feature extraction and feature embedding processing on the spliced ​​matrix to obtain the global traffic features of the trunk line to be controlled.

[0087] In this embodiment, the intersection sample data may include queue lengths and traffic flow at each intersection. Therefore, the intersection features can be two-dimensional data. The intersection features of each intersection are concatenated, i.e., their respective intersection dimensions are added, to obtain a three-dimensional concatenated matrix. Then, feature extraction and feature embedding processing are performed on the concatenated matrix to obtain the global traffic features of the arterial road to be controlled. The global traffic features can represent the overall traffic situation of the arterial road to be controlled, or the traffic situation of all intersections within the arterial road to be controlled. Specifically, feature extraction and feature embedding processing can be implemented based on a feature extraction network composed of convolutional neural networks and fully connected neural networks.

[0088] In one example, the intersection features of each intersection in this embodiment can be obtained according to the following formula:

[0089]

[0090] in, Let f be the intersection feature of intersection i, and let f be the f-th traffic feature. h represents the intersection sample data obtained by observing all approach lanes of intersection i within the time period T. f (·) indicates the feature processing method, and T represents the observation time window. w This represents the step size of the feature extraction network. Indicates the number of phases. It represents a real number field with dimensions equal to the number of phases multiplied by the number of steps.

[0091] Finally, the intersection features of all intersections are concatenated to obtain the global traffic feature s. M .

[0092] In one example, the method for generating global tasks can be as follows: Figure 2c As shown. The arterial road to be controlled includes N intersections from intersection 1 to intersection N. Traffic data (intersection sample data) of each intersection is acquired. The intersection features of each intersection are concatenated to obtain a concatenation matrix. Then, feature extraction is performed by a feature extraction module, and feature embedding is performed by a feature embedding module. Finally, the main control agent is used for processing to obtain the global task of the arterial road to be controlled.

[0093] As can be seen from the above, the training method for the arterial traffic signal control model provided in this application concatenates the intersection features of each intersection in the arterial road to be controlled, which are included in the global traffic data of the sample. Based on the concatenated matrix, feature extraction and feature embedding are performed to obtain global traffic features that can reflect both the overall arterial road to be controlled and the local features of each intersection. This allows the subsequent training of the arterial traffic signal control model to simultaneously consider both the global aspects of the arterial road to be controlled and the local aspects of each intersection, thus improving the feasibility of the arterial traffic signal control model.

[0094] In one example, the master agent and each worker agent can be directly trained together; in another example, each worker agent can be pre-trained first, and then the master agent and each worker agent can be trained together.

[0095] In one possible implementation, such as Figure 3 As shown, before step S11 above, which involves obtaining sample global traffic data for the trunk line to be controlled, the method further includes:

[0096] Step S31: Generate a pre-trained global task using the main control agent based on random sampling;

[0097] The pre-trained global task includes multiple pre-trained crossover tasks;

[0098] Step S32: For each working agent, analyze the pre-trained intersection task corresponding to the working agent to obtain the pre-trained traffic signal control strategy for the intersection corresponding to the working agent.

[0099] Step S33: Determine the pre-trained intersection observation data for each intersection based on the pre-trained traffic signal control strategy for each intersection;

[0100] Step S34: For each working agent, adjust the parameters of the working agent based on the pre-trained intersection observation data corresponding to that working agent;

[0101] Step S35: Return to step S31 and use the master agent to generate a pre-trained global task based on random sampling to continue execution until the agent's pre-training termination condition is met.

[0102] Steps S31-S35 above constitute the pre-training process of the working agent. During the pre-training process, only the parameters of the working agent are adjusted, while the parameters of the master agent are not adjusted.

[0103] In this embodiment, before jointly training the master agent and the worker agents, the worker agents in the trunk line traffic signal control model are first pre-trained. During pre-training, only the parameters of each worker agent are adjusted; that is, only each worker agent is trained. Pre-training ends after the agent training termination condition is met. The agent training termination condition can be customized according to actual conditions, such as reaching a preset number of training iterations, or the convergence of the loss or reward value of each worker agent.

[0104] In pre-training, the master agent generates a global task based on random sampling. Each worker agent receives a corresponding pre-trained intersection task from the global task and analyzes it to obtain a pre-trained traffic signal control strategy for each worker agent. Based on each pre-trained traffic signal control strategy, the system simulates the operation for a period of time (e.g., a preset time period) to obtain pre-trained intersection observation data, including traffic data of the intersection during the simulated operation, such as queue length, traffic flow, number of arriving vehicles, vehicle waiting time, and green light execution time. Based on this, the parameters of each worker agent are adjusted, and then the system returns to the master agent, which generates a new global task based on random sampling and distributes it to each worker agent, allowing each worker agent to train again until the preset agent pre-training termination condition is met, indicating that the training of each worker agent is complete and the model pre-training ends.

[0105] After the pre-training is completed, the joint training process of the master control agent and the working agent in the above embodiment begins until the model training end condition is met, and the trained trunk traffic signal control model is obtained.

[0106] As can be seen from the above, the training method for the trunk traffic signal control model provided in this application first performs pre-training before jointly training the master control agent and the working agents. In the pre-training stage, only the ability and efficiency of each working agent to perform intersection tasks are trained. In the subsequent joint training stage, the trunk traffic signal control model is trained as a whole, thereby planning different training objectives for different training stages and improving the efficiency of model training.

[0107] In one possible implementation, the master control agent employs the DDPG algorithm; adjusting the parameters of the master control agent based on observed global traffic data includes: determining a first gradient value to be adjusted for the parameters in the master control agent based on the observed global traffic data and the current values ​​of the parameters in the master control agent, and adjusting the parameters in the master control agent according to the first gradient value to be adjusted.

[0108] And / or,

[0109] The working agent adopts the DDPG algorithm, and the observed global traffic data includes the intersection observation data of each intersection. The step of adjusting the parameters of each working agent according to the observed global traffic data includes: for each working agent to be adjusted, determining a second gradient value of the parameters of the working agent to be adjusted based on the intersection observation data of the intersection of the working agent to be adjusted and the current value of the parameters of the working agent to be adjusted, and adjusting the parameters of the working agent to be adjusted according to the second gradient value.

[0110] In this embodiment, the master control agent and each working agent adopt the DDPG algorithm to determine their respective gradient values ​​to be adjusted based on the current values ​​of their respective parameters, and then adjust their respective parameters according to the gradient values ​​to be adjusted.

[0111] In one example, the first gradient value to be adjusted for the parameters in the controlling agent can be updated according to the following formula:

[0112]

[0113] Among them, s M Represents global traffic characteristics, π M Indicates the policy of the controlling agent, φ M For network parameters, g M For global tasks, Q MThis is an approximate Q-value calculated based on the network.

[0114] As can be seen from the above, the training method for the arterial traffic signal control model provided in this application uses the DDPG algorithm to update the parameters of both the master control agent and each working agent. This allows for the acquisition of better global tasks and traffic signal control strategies for each intersection, improving the effectiveness of the arterial traffic signal control scheme and the applicability of the model. Furthermore, it enables the arterial traffic signal control model to adaptively adjust according to traffic scenarios, allowing the trained model to better adapt to highly dynamic traffic environments without being limited by specific models and rules.

[0115] In one embodiment of this application, such as Figure 4 As shown, a flowchart illustrating the training method for a second arterial traffic signal control model is also provided, wherein the observed global traffic data includes intersection observation data for each of the aforementioned intersections; the method further includes:

[0116] Step S41: Calculate the cumulative throughput and average maximum queue length of the trunk line to be controlled based on the intersection observation data of each intersection;

[0117] Step S42: Calculate the global reward value of the master control agent based on the cumulative throughput and the average maximum queue length;

[0118] Step S43: For each working agent, calculate the intersection reward value of the working agent based on the intersection observation data corresponding to the working agent;

[0119] Step S44: If the global reward value and the reward values ​​of each intersection converge, determine that the model training termination condition is met.

[0120] In this embodiment, the intersection observation data includes the queue length and traffic flow of each intersection within a preset time period. Based on this, the cumulative throughput and average maximum queue length of the entire arterial road to be controlled are calculated. Then, the global reward value of the master control agent is calculated. The global reward value can represent the optimality of the global task obtained by the master control agent, that is, the effectiveness in improving the traffic operation of the arterial road to be controlled. Then, the corresponding intersection reward value is calculated for each working agent. The intersection reward value can represent both the degree of achievement of the global task by the working agent and the optimality of the traffic signal control strategy of the intersection obtained by each working agent, that is, the effectiveness in improving the traffic operation of the intersection.

[0121] When the global reward value and the reward values ​​at each intersection converge, it indicates that the global task obtained by the master agent is the optimal global task. This means that the phase difference between adjacent intersections on the controlled artery is the most optimized solution, maximizing the improvement of the overall traffic flow on the controlled artery. Furthermore, the phase difference between the intersection corresponding to each worker agent and its adjacent intersections has approximated the target phase difference in the global task, and the signal control scheme in the traffic signal control strategy obtained by each worker agent is the most optimized solution, maximizing the improvement of traffic flow at each intersection. At this point, the model training termination condition is met, and model training ends.

[0122] In one embodiment of this application, the calculation of the global reward value of the master control agent based on the cumulative throughput and the average maximum queue length includes:

[0123] The global reward value of the controlling agent is calculated according to the following formula:

[0124]

[0125] in, This represents the global reward value, and vart represents the cumulative throughput. This represents the average maximum queue length, and w is a preset weighting coefficient.

[0126] As can be seen from the above, the training method of the arterial traffic signal control model provided in this application ends when the global reward value and the reward value of each intersection converge, that is, when the observed global traffic data shows that the overall traffic operation of the arterial road to be controlled and the local traffic operation of each intersection can be improved to the greatest extent. The trained model can provide the most optimized arterial traffic signal control scheme and provide a more feasible improvement scheme for arterial traffic operation.

[0127] In one possible implementation, such as Figure 5 As shown, step S43 above calculates the intersection reward value for each working agent based on the intersection observation data corresponding to that agent, including:

[0128] Step S51: For each working agent, based on the intersection observation data corresponding to the working agent, determine the maximum queue length, cumulative delay time, effective green light time, observation phase difference, and target phase difference of the intersection corresponding to the working agent;

[0129] Step S52: Calculate the first reward value of the working agent based on the maximum queue length, cumulative delay time, and effective green light time of the intersection corresponding to the working agent;

[0130] Step S53: Calculate the second reward value of the working agent based on the observation phase difference and target phase difference of the intersection corresponding to the working agent;

[0131] Step S54: Calculate the intersection reward value of the working agent based on the first reward value and the second reward value of the working agent.

[0132] In this embodiment of the application, the intersection observation data includes the queue length and traffic flow of the intersection. Based on this, the maximum queue length, cumulative delay time (total delay time of traffic flow), effective green light time (green light time after a certain traffic flow has passed), observed phase difference (current phase difference of the intersection), and target phase difference in the above-mentioned intersection task can be determined.

[0133] Then, based on the maximum queue length, cumulative delay time, and effective green light time of the intersection corresponding to each working agent, the first reward value of the working agent is calculated. The first reward value can represent the optimality of the current traffic signal control strategy of each working agent.

[0134] Based on the observation phase difference and target phase difference of the intersection corresponding to each working agent, the second reward value of the working agent is calculated. The second reward value can represent the degree of achievement of the global task by each working agent.

[0135] Based on these first and second reward values, the intersection reward value for each working agent is calculated, which can simultaneously reflect the optimality of the current traffic signal control strategy of each working agent and the degree of achievement of the global task.

[0136] In one embodiment of this application, step S52 above calculates the first reward value of the working agent based on the maximum queue length, cumulative delay time, and effective green light time of the intersection corresponding to the working agent, including:

[0137] The first reward value for the working agent is calculated using the following formula:

[0138]

[0139] in, The first reward value for the i-th working agent. Let be the maximum queue length for the i-th working agent. Let be the cumulative delay time of the i-th working agent. The effective green light time is the time for the i-th working agent, where w1, w2, and w3 are preset weight coefficients; the working agent is the i-th working agent.

[0140] Step S53 above calculates the second reward value for the working agent based on the observed phase difference and target phase difference of the intersection corresponding to the working agent, including:

[0141] The second reward value for the working agent is calculated using the following formula:

[0142]

[0143] in, Let f1(·) be the second reward value for the i-th working agent, and let f1(·) represent the L1 regularization function. Let be the observation phase difference of the i-th working agent. Let be the target phase difference for the i-th working agent;

[0144] Step S54 above calculates the intersection reward value of the working agent based on the first reward value and the second reward value, including:

[0145] Calculate the intersection reward value for the working agent using the following formula:

[0146]

[0147] in, Let α1 and α2 be the intersection reward value for the i-th working agent, and let α1 and α2 be the preset weight coefficients.

[0148] The aforementioned preset weighting coefficients are all set according to actual needs.

[0149] As can be seen from the above, the training method for the arterial traffic signal control model provided in this application calculates the intersection reward value together with a first reward value calculated based on the maximum queue length, cumulative delay time, and effective green light time of the intersection corresponding to each working agent, and a second reward value calculated based on the observation phase difference and target phase difference of the intersection corresponding to each working agent. This allows the intersection reward value to simultaneously reflect the optimality of the current traffic signal control strategy of each working agent and the degree of achievement of the global task, further providing a better solution for arterial traffic signal control and improving the feasibility of the arterial traffic signal control model.

[0150] See Figure 6 This application also provides a flowchart of a trunk line traffic signal control method, including:

[0151] Step S61: Obtain global traffic data for the trunk line to be controlled;

[0152] Step S62: Input the global traffic data into the trunk line traffic signal control model, and obtain the traffic signal control strategy for each intersection in the trunk line to be controlled based on the output of the trunk line traffic signal control model. The trunk line traffic signal control model is trained by any of the above-mentioned trunk line traffic signal control model training methods.

[0153] Step S63: Control the traffic signals at each intersection on the trunk line to be controlled according to the traffic signal control strategies described above.

[0154] As can be seen from the above, the trunk line traffic signal control method provided in this application obtains the global traffic data of the trunk line to be controlled, processes the global traffic data based on the trunk line traffic signal control model, obtains the traffic signal control strategy for each intersection in the trunk line to be controlled according to the output result of the trunk line traffic signal control model, and controls the traffic signals of each intersection in the trunk line to be controlled accordingly. This achieves adaptive adjustment of the traffic signal control strategy according to the current traffic scenario, which is more adaptable to the highly dynamic traffic environment and effectively improves traffic operation.

[0155] See Figure 7 This application also provides a schematic diagram of the structure of a training device for a trunk traffic signal control model. The trunk traffic signal control model includes a master control agent and multiple working agents, each of which corresponds to an intersection on the trunk line to be controlled. The device includes:

[0156] The first data acquisition module 701 is used to acquire sample global traffic data of the trunk line to be controlled; and generate global traffic features of the trunk line to be controlled based on the sample global traffic data.

[0157] The global task acquisition module 702 is used to analyze the global traffic features using the main control agent to obtain the global task of the trunk line to be controlled. The global task includes multiple intersection tasks, and each intersection task corresponds to a working agent.

[0158] The first strategy acquisition module 703 is used to analyze the intersection task corresponding to each working agent and obtain the traffic signal control strategy of the intersection corresponding to the working agent.

[0159] The data determination module 704 is used to determine the observed global traffic data of the trunk line to be controlled based on the traffic signal control strategies of each intersection in the trunk line to be controlled.

[0160] The model training module 705 is used to adjust the parameters of the master control agent and at least one of the working agents based on observed global traffic data; select other sample global traffic data of the trunk line to be controlled, and continue to train the trunk line traffic signal control model until the model training termination condition is met, thereby obtaining a trained trunk line traffic signal control model. The trunk line traffic signal control model is used to output the traffic signal control strategy of each intersection on the trunk line to be controlled when the current global traffic data of the trunk line to be controlled is input.

[0161] As can be seen from the above, the training device for the arterial traffic signal control model provided in this application, after acquiring sample global traffic data of the arterial line to be controlled and generating global traffic features of the arterial line to be controlled, first uses the master control agent to analyze the global traffic features, that is, to perform a holistic global analysis of the arterial line to be controlled, to obtain global tasks including intersection tasks corresponding one-to-one with each working agent in the arterial line to be controlled; then, for each working agent, it uses the working agent to perform local analysis of its corresponding intersection task, thereby obtaining the traffic signal control strategy for the intersection corresponding to the working agent. Then, based on the traffic signal control strategy of each intersection, the observed global traffic data of the arterial line to be controlled is determined, and the parameters of the master control agent and at least one agent in each working agent are adjusted accordingly. Afterwards, other sample global traffic data of the arterial line to be controlled are selected again to continue training the arterial traffic signal control model until the model training termination condition is met, and a trained arterial traffic signal control model is obtained. The arterial traffic signal control model is used to output the traffic signal control strategy of each intersection in the arterial line to be controlled when the current global traffic data of the arterial line to be controlled is input.

[0162] The training device for the arterial traffic signal control model in this application proposes an arterial coordinated signal control scheme based on a hierarchical structure. The trained arterial traffic signal control method can use the master control agent to determine the overall global coordinated goal of the arterial to be controlled through global traffic data, and then use each working agent in the lower layer to determine its own local goal according to its respective intersection task. At the same time, the optimization and realization of the global coordinated goal are completed. The upper and lower structure can better decompose the task and can organically combine the global of the arterial to be controlled and the local of each working agent. While realizing the global coordinated goal of the arterial to be controlled, it also takes into account the local goals of each working agent. At the same time, it can also adaptively adjust according to the traffic scenario, so that the trained model can better adapt to the highly dynamic and highly complex traffic environment of the arterial to be controlled, providing a more feasible solution for traffic signal control.

[0163] In one embodiment of this application, the sample global traffic data includes intersection sample data of each intersection in the trunk line to be controlled;

[0164] The data acquisition module is specifically used for: generating intersection features for each intersection based on the intersection sample data; concatenating the intersection features of each intersection to obtain a concatenation matrix; and performing feature extraction and feature embedding processing on the concatenation matrix to obtain the global traffic features of the arterial road to be controlled.

[0165] As can be seen from the above, the training device for the arterial traffic signal control model provided in this application splices the intersection features of each intersection in the arterial road to be controlled, which are included in the global traffic data of the sample, and performs feature extraction and feature embedding processing based on the spliced ​​matrix. This results in global traffic features that can reflect both the overall arterial road to be controlled and the local features of each intersection, so that the subsequent training of the arterial traffic signal control model can simultaneously consider both the global aspects of the arterial road to be controlled and the local aspects of each intersection, thereby improving the feasibility of the arterial traffic signal control model.

[0166] In one embodiment of this application, the global task of the trunk line to be controlled represents the target phase difference between each adjacent intersection in the trunk line to be controlled; for any two adjacent intersections, the target phase difference between the two adjacent intersections represents the difference between the start times of the green light signals between the two adjacent intersections; for any intersection, the intersection task of the intersection represents the target phase difference corresponding to the intersection.

[0167] In one embodiment of this application, the apparatus further includes:

[0168] The pre-training task generation module is used to generate pre-training global tasks using the main control agent based on random sampling, wherein the pre-training global tasks include multiple pre-training intersection tasks.

[0169] The pre-trained strategy acquisition module is used to analyze the pre-trained intersection task corresponding to each working agent and obtain the pre-trained traffic signal control strategy for the intersection corresponding to the working agent.

[0170] The pre-training data determination module is used to determine the pre-training intersection observation data of each intersection based on the pre-training traffic signal control strategy of each intersection.

[0171] The pre-training parameter adjustment module is used to adjust the parameters of each working agent based on the pre-training intersection observation data corresponding to that working agent.

[0172] The step return module is used to return the step: using the main control agent to generate a pre-trained global task based on random sampling and continue execution until the agent pre-training end condition is met;

[0173] As can be seen from the above, the training device for the trunk traffic signal control model provided in this application performs pre-training before jointly training the master control agent and the working agents. In the pre-training stage, only the ability and efficiency of each working agent to perform intersection tasks are trained. In the subsequent joint training stage, the trunk traffic signal control model is trained as a whole, thereby planning different training objectives for different training stages and improving the efficiency of model training.

[0174] In one embodiment of this application, the master control agent adopts the DDPG algorithm; the first model training submodule is specifically used to: determine the first gradient value to be adjusted for the parameters in the master control agent based on the observed global traffic data and the current value of the parameters in the master control agent, and adjust the parameters in the master control agent according to the first gradient value to be adjusted.

[0175] And / or,

[0176] The working agent adopts the DDPG algorithm. The second model training submodule is specifically used for: for each working agent to be adjusted, based on the intersection observation data of the intersection of the working agent to be adjusted and the current value of the parameter in the working agent to be adjusted, determining the second gradient value of the parameter in the working agent to be adjusted, and adjusting the parameter of the working agent to be adjusted according to the second gradient value.

[0177] As can be seen from the above, the training device for the arterial traffic signal control model provided in this application uses the DDPG algorithm to update the parameters of the master control agent and each working agent. This enables the acquisition of better global tasks and traffic signal control strategies for each intersection, improving the effectiveness of the arterial traffic signal control scheme and the applicability of the model. Furthermore, it allows the arterial traffic signal control model to adaptively adjust according to traffic scenarios, making the trained model more adaptable to highly dynamic traffic environments without being limited by specific models and rules.

[0178] In one embodiment of this application, the observed global traffic data includes intersection observation data for each of the intersections; the device further includes:

[0179] The data calculation module is used to calculate the cumulative throughput and average maximum queue length of the trunk line to be controlled based on the intersection observation data of each intersection.

[0180] The global reward value calculation module is used to calculate the global reward value of the main control agent based on the cumulative throughput and the average maximum queue length.

[0181] The intersection reward value calculation module is used to calculate the intersection reward value for each working agent based on the intersection observation data corresponding to that working agent.

[0182] The training judgment module is used to determine whether the model training termination condition is met when the global reward value and the reward values ​​of each intersection converge.

[0183] As can be seen from the above, the training device for the arterial traffic signal control model provided in this application ends the model training when the global reward value and the reward values ​​of each intersection converge, that is, when the observed global traffic data indicates that the overall traffic operation of the arterial road to be controlled and the local traffic operation of each intersection can be improved to the greatest extent. The trained model can provide the most optimized arterial traffic signal control scheme and provide a more feasible improvement scheme for arterial traffic operation.

[0184] In one embodiment of this application, the global reward value calculation module is specifically used for:

[0185] The global reward value of the controlling agent is calculated according to the following formula:

[0186]

[0187] in, Represents the global reward value, v art This indicates the cumulative traffic volume. This represents the average maximum queue length, and w is a preset weighting coefficient.

[0188] As can be seen from the above, the training device for the arterial traffic signal control model provided in this application ends the model training when the global reward value and the reward values ​​of each intersection converge, that is, when the observed global traffic data indicates that the overall traffic operation of the arterial road to be controlled and the local traffic operation of each intersection can be improved to the greatest extent. The trained model can provide the most optimized arterial traffic signal control scheme and provide a more feasible improvement scheme for arterial traffic operation.

[0189] In one embodiment of this application, the intersection reward value calculation module includes:

[0190] The data calculation submodule is used to calculate, for each working agent, the maximum queue length, cumulative delay time, effective green light time, observation phase difference, and target phase difference of the intersection corresponding to that working agent, based on the intersection observation data of that working agent.

[0191] The first reward value calculation submodule is used to calculate the first reward value of the working agent based on the maximum queue length, cumulative delay time, and effective green light time of the intersection corresponding to the working agent.

[0192] The second reward value calculation submodule is used to calculate the second reward value of the working agent based on the observation phase difference and target phase difference of the intersection corresponding to the working agent.

[0193] The intersection reward value calculation submodule is used to calculate the intersection reward value of the working agent based on the first reward value and the second reward value of the working agent;

[0194] The first reward value calculation submodule is specifically used for:

[0195] The first reward value for the working agent is calculated using the following formula:

[0196]

[0197] in, The first reward value for the i-th working agent. Let be the maximum queue length for the i-th working agent. Let be the cumulative delay time of the i-th working agent. The effective green light time for the i-th working agent is defined by w1, w2, and w3, which are all preset weight coefficients.

[0198] The second reward value calculation submodule is specifically used for:

[0199] The second reward value for the working agent is calculated using the following formula:

[0200]

[0201] in, Let f1(·) be the second reward value for the i-th working agent, and let f1(·) represent the L1 regularization function. Let be the observation phase difference of the i-th working agent. Let be the target phase difference for the i-th working agent;

[0202] The intersection reward value calculation submodule is specifically used for:

[0203] Calculate the intersection reward value for the working agent using the following formula:

[0204]

[0205] in, Let α1 and α2 be the intersection reward value for the i-th working agent, and let α1 and α2 be preset weight coefficients.

[0206] As can be seen from the above, the training device for the arterial traffic signal control model provided in this application calculates the intersection reward value together with a first reward value calculated based on the maximum queue length, cumulative delay time, and effective green light time of the intersection corresponding to each working agent, and a second reward value calculated based on the observation phase difference and target phase difference of the intersection corresponding to each working agent. This allows the intersection reward value to simultaneously reflect the optimality of the current traffic signal control strategy of each working agent and the degree of achievement of the global task, further providing a better solution for arterial traffic signal control and improving the feasibility of the arterial traffic signal control model.

[0207] See Figure 8 This application also provides a schematic diagram of the structure of a trunk line traffic signal control device, including:

[0208] The second data acquisition module 801 is used to acquire global traffic data of the trunk line to be controlled.

[0209] The second strategy acquisition module 802 is used to input the global traffic data into the trunk line traffic signal control model, and obtain the traffic signal control strategy for each intersection in the trunk line to be controlled based on the output of the trunk line traffic signal control model. The trunk line traffic signal control model is trained by any of the trunk line traffic signal control model training devices described above.

[0210] Traffic signal control module 803 is used to control the traffic signals at each intersection on the trunk road to be controlled in accordance with the traffic signal control strategies described above.

[0211] As can be seen from the above, the trunk line traffic signal control device provided in this application obtains the global traffic data of the trunk line to be controlled, processes the global traffic data based on the trunk line traffic signal control model, obtains the traffic signal control strategy for each intersection in the trunk line to be controlled according to the output result of the trunk line traffic signal control model, and controls the traffic signals of each intersection in the trunk line to be controlled accordingly. This realizes the adaptive adjustment of the traffic signal control strategy according to the current traffic scenario, which is more adaptable to the highly dynamic traffic environment and effectively improves the traffic operation.

[0212] This application also provides an electronic device, such as... Figure 9 As shown, it includes:

[0213] Memory 901 is used to store computer programs;

[0214] The processor 902, when executing the program stored in the memory 901, implements the training method and the trunk traffic signal control method of any of the above-mentioned trunk traffic signal control models.

[0215] Furthermore, the aforementioned electronic device may also include a communication bus and / or a communication interface, with the processor 902, communication interface, and memory 901 communicating with each other via the communication bus.

[0216] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0217] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0218] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0219] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0220] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the training method of any of the above-described trunk traffic signal control models and the steps of the trunk traffic signal control method.

[0221] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute the training method of any trunk traffic signal control model and the trunk traffic signal control method in the above embodiments.

[0222] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a solid-state drive (SSD), etc.

[0223] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0224] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, the embodiments for apparatus, electronic devices, and storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0225] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. A training method for a trunk line traffic signal control model, characterized in that, The mainline traffic signal control model includes a master control agent and multiple working agents, each working agent corresponding to an intersection on the mainline to be controlled; the method includes: Obtain sample global traffic data of the arterial road to be controlled; generate global traffic features of the arterial road to be controlled based on the sample global traffic data; The global traffic features are analyzed using the master control agent to obtain the global task of the trunk line to be controlled. The global task includes multiple intersection tasks, and each intersection task corresponds to a working agent. For each working agent, the task of the intersection corresponding to the working agent is analyzed to obtain the traffic signal control strategy of the intersection corresponding to the working agent. Based on the traffic signal control strategies of each intersection in the arterial road to be controlled, determine the observed global traffic data of the arterial road to be controlled; Based on the observed global traffic data, the parameters of the master control agent and at least one of the working agents are adjusted; other sample global traffic data of the trunk line to be controlled are selected, and the trunk line traffic signal control model is trained until the model training termination condition is met, thus obtaining the trained trunk line traffic signal control model. The trunk line traffic signal control model is used to output the traffic signal control strategy of each intersection on the trunk line to be controlled when the current global traffic data of the trunk line to be controlled is input.

2. The method according to claim 1, characterized in that, The sample global traffic data includes the intersection sample data of each intersection in the trunk line to be controlled; The step of generating global traffic features for the arterial road to be controlled based on the sample global traffic data includes: For each intersection, generate the intersection features based on the intersection sample data for that intersection; The intersection features of each intersection are concatenated to obtain a concatenation matrix; The spliced ​​matrix is ​​subjected to feature extraction and feature embedding processing to obtain the global traffic features of the trunk line to be controlled.

3. The method according to claim 1, characterized in that, The global task of the trunk line to be controlled represents the target phase difference between each adjacent intersection in the trunk line to be controlled; for any two adjacent intersections, the target phase difference between the two adjacent intersections represents the difference between the start times of the green light signals between the two adjacent intersections; for any intersection, the intersection task of the intersection represents the target phase difference corresponding to the intersection.

4. The method according to claim 1, characterized in that, Prior to the step of acquiring sample global traffic data for the trunk line to be controlled, the method further includes: The master control agent generates a pre-trained global task based on random sampling, wherein the pre-trained global task includes multiple pre-trained intersection tasks; For each working agent, the pre-trained intersection task corresponding to the working agent is analyzed to obtain the pre-trained traffic signal control strategy for the intersection corresponding to the working agent. Based on the pre-trained traffic signal control strategies of each intersection, determine the pre-trained intersection observation data for each intersection; For each working agent, adjust the parameters of the working agent based on the pre-trained intersection observation data corresponding to that working agent; Return step: Use the master control agent to generate a pre-trained global task based on random sampling and continue execution until the agent pre-training termination condition is met.

5. The method according to claim 1, characterized in that, The master control agent adopts the DDPG algorithm; the step of adjusting the parameters of the master control agent according to the observed global traffic data includes: determining the first gradient value to be adjusted for the parameters in the master control agent according to the observed global traffic data and the current value of the parameters in the master control agent, and adjusting the parameters in the master control agent according to the first gradient value to be adjusted. And / or, The working agent adopts the DDPG algorithm, and the observed global traffic data includes the intersection observation data of each intersection. The step of adjusting the parameters of each working agent according to the observed global traffic data includes: for each working agent to be adjusted, determining a second gradient value of the parameters of the working agent to be adjusted based on the intersection observation data of the intersection of the working agent to be adjusted and the current value of the parameters of the working agent to be adjusted, and adjusting the parameters of the working agent to be adjusted according to the second gradient value.

6. The method according to claim 1, characterized in that, The observed global traffic data includes intersection observation data for each of the aforementioned intersections; the method further includes: Based on the intersection observation data of each intersection, calculate the cumulative throughput and average maximum queue length of the trunk line to be controlled; The global reward value of the master control agent is calculated based on the cumulative throughput and the average maximum queue length. For each working agent, calculate the intersection reward value for that working agent based on the intersection observation data corresponding to that working agent; If the global reward value and the reward values ​​at each intersection converge, the model training termination condition is determined to be met.

7. The method according to claim 6, characterized in that, The step of calculating the global reward value of the master control agent based on the cumulative throughput and the average maximum queue length includes: The global reward value of the controlling agent is calculated according to the following formula: in, Represents the global reward value, v art This indicates the cumulative traffic volume. This represents the average maximum queue length, and w is a preset weighting coefficient.

8. The method according to claim 6, characterized in that, For each working agent, the intersection reward value is calculated based on the intersection observation data corresponding to that agent, including: For each working agent, based on the intersection observation data corresponding to that working agent, determine the maximum queue length, cumulative delay time, effective green light time, observation phase difference, and target phase difference of the intersection corresponding to that working agent; Calculate the first reward value for the working agent based on the maximum queue length, cumulative delay time, and effective green light time of the intersection corresponding to the working agent. The second reward value of the working agent is calculated based on the observation phase difference and target phase difference of the intersection corresponding to the working agent. Calculate the intersection reward value of the working agent based on its first reward value and second reward value.

9. The method according to claim 8, characterized in that, The calculation of the first reward value for the working agent based on the maximum queue length, cumulative delay time, and effective green light time at the intersection corresponding to the working agent includes: The first reward value for the working agent is calculated using the following formula: in, The first reward value for the i-th working agent. Let be the maximum queue length for the i-th working agent. Let be the cumulative delay time of the i-th working agent. The effective green light time for the i-th working agent is defined by w1, w2, and w3, which are all preset weight coefficients. The step of calculating the second reward value of the working agent based on the observed phase difference and the target phase difference of the intersection corresponding to the working agent includes: The second reward value for the working agent is calculated using the following formula: in, Let f1(·) be the second reward value for the i-th working agent, and let f1(·) represent the L1 regularization function. Let be the observation phase difference of the i-th working agent. Let be the target phase difference for the i-th working agent; The step of calculating the intersection reward value of the working agent based on the first reward value and the second reward value includes: Calculate the intersection reward value for the working agent using the following formula: in, Let α1 and α2 be the intersection reward value for the i-th working agent, and let α1 and α2 be preset weight coefficients.

10. A method for controlling mainline traffic signals, characterized in that, include: Acquire global traffic data for the trunk line to be controlled; The global traffic data is input into the trunk line traffic signal control model, and the traffic signal control strategy for each intersection in the trunk line to be controlled is obtained based on the output of the trunk line traffic signal control model. The trunk line traffic signal control model is trained by the method described in any one of claims 1-9. The traffic signals at each intersection on the arterial road to be controlled are controlled according to the traffic signal control strategies described above.

11. A training device for a trunk line traffic signal control model, characterized in that, The mainline traffic signal control model includes a master control agent and multiple working agents, each working agent corresponding to an intersection on the mainline to be controlled; the device includes: The first data acquisition module is used to acquire sample global traffic data of the trunk line to be controlled; and to generate global traffic features of the trunk line to be controlled based on the sample global traffic data. The global task acquisition module is used to analyze the global traffic features using the main control agent to obtain the global task of the trunk line to be controlled. The global task includes multiple intersection tasks, and each intersection task corresponds to a working agent. The first strategy acquisition module is used to analyze the intersection task corresponding to each working agent and obtain the traffic signal control strategy for the intersection corresponding to the working agent. The data determination module is used to determine the observed global traffic data of the trunk line to be controlled based on the traffic signal control strategies of each intersection in the trunk line to be controlled. The model training module is used to adjust the parameters of the master control agent and at least one of the working agents based on observed global traffic data; select other sample global traffic data of the trunk line to be controlled, and continue to train the trunk line traffic signal control model until the model training termination condition is met, thereby obtaining a trained trunk line traffic signal control model. The trunk line traffic signal control model is used to output the traffic signal control strategy of each intersection on the trunk line to be controlled when the current global traffic data of the trunk line to be controlled is input.

12. The apparatus according to claim 11, characterized in that, The sample global traffic data includes the intersection sample data of each intersection in the trunk line to be controlled; The data acquisition module is specifically used for: generating intersection features for each intersection based on the intersection sample data; concatenating the intersection features of each intersection to obtain a concatenation matrix; and performing feature extraction and feature embedding processing on the concatenation matrix to obtain the global traffic features of the arterial road to be controlled. The global task of the trunk line to be controlled represents the target phase difference between each adjacent intersection in the trunk line to be controlled; for any two adjacent intersections, the target phase difference between the two adjacent intersections represents the difference between the start times of the green light signal between the two adjacent intersections. For any intersection, the intersection task represents the target phase difference corresponding to that intersection; The device further includes: The pre-training task generation module is used to generate pre-training global tasks using the main control agent based on random sampling, wherein the pre-training global tasks include multiple pre-training intersection tasks. The pre-trained strategy acquisition module is used to analyze the pre-trained intersection task corresponding to each working agent and obtain the pre-trained traffic signal control strategy for the intersection corresponding to the working agent. The pre-training data determination module is used to determine the pre-training intersection observation data of each intersection based on the pre-training traffic signal control strategy of each intersection. The pre-training parameter adjustment module is used to adjust the parameters of each working agent based on the pre-training intersection observation data corresponding to that working agent. The step return module is used to return the step: using the main control agent to generate a pre-trained global task based on random sampling and continue execution until the agent pre-training end condition is met; The master control agent adopts the DDPG algorithm; the first model training submodule is specifically used to: determine the first gradient value to be adjusted for the parameters in the master control agent based on the observed global traffic data and the current value of the parameters in the master control agent, and adjust the parameters in the master control agent according to the first gradient value to be adjusted. And / or, The working agent adopts the DDPG algorithm. The second model training submodule is specifically used for: for each working agent to be adjusted, based on the intersection observation data of the intersection of the working agent to be adjusted and the current value of the parameter in the working agent to be adjusted, determining the second gradient value of the parameter in the working agent to be adjusted, and adjusting the parameter of the working agent to be adjusted according to the second gradient value. The observed global traffic data includes intersection observation data for each of the aforementioned intersections; the device further includes: The data calculation module is used to calculate the cumulative throughput and average maximum queue length of the trunk line to be controlled based on the intersection observation data of each intersection. The global reward value calculation module is used to calculate the global reward value of the main control agent based on the cumulative throughput and the average maximum queue length. The intersection reward value calculation module is used to calculate the intersection reward value for each working agent based on the intersection observation data corresponding to that working agent. The training judgment module is used to determine whether the model training termination condition is met when the global reward value and the reward values ​​of each intersection converge. The global reward value calculation module is specifically used for: The global reward value of the controlling agent is calculated according to the following formula: in, Represents the global reward value, v art This indicates the cumulative traffic volume. This represents the average maximum queue length, where w is a preset weighting coefficient. The intersection reward value calculation module includes: The data calculation submodule is used to determine, for each working agent, the maximum queue length, cumulative delay time, effective green light time, observation phase difference, and target phase difference of the intersection corresponding to that working agent, based on the observation data of the intersection corresponding to that working agent. The first reward value calculation submodule is used to calculate the first reward value of the working agent based on the maximum queue length, cumulative delay time, and effective green light time of the intersection corresponding to the working agent. The second reward value calculation submodule is used to calculate the second reward value of the working agent based on the observation phase difference and target phase difference of the intersection corresponding to the working agent. The intersection reward value calculation submodule is used to calculate the intersection reward value of the working agent based on the first reward value and the second reward value of the working agent; The first reward value calculation submodule is specifically used for: The first reward value for the working agent is calculated using the following formula: in, The first reward value for the i-th working agent. Let be the maximum queue length for the i-th working agent. Let be the cumulative delay time of the i-th working agent. The effective green light time for the i-th working agent is defined by w1, w2, and w3, which are all preset weight coefficients. The second reward value calculation submodule is specifically used for: The second reward value for the working agent is calculated using the following formula: in, Let f1(·) be the second reward value for the i-th working agent, and let f1(·) represent the L1 regularization function. Let be the observation phase difference of the i-th working agent. Let be the target phase difference for the i-th working agent; The intersection reward value calculation submodule is specifically used for: Calculate the intersection reward value for the working agent using the following formula: in, Let α1 and α2 be the intersection reward value for the i-th working agent, and let α1 and α2 be preset weight coefficients.

13. A trunk line traffic signal control device, characterized in that, include: The second data acquisition module is used to acquire global traffic data for the trunk line to be controlled. The second strategy acquisition module is used to input the global traffic data into the trunk line traffic signal control model, and obtain the traffic signal control strategy for each intersection in the trunk line to be controlled based on the output of the trunk line traffic signal control model, wherein the trunk line traffic signal control model is trained by the device described in any one of claims 11-12. A traffic signal control module is used to control the traffic signals at each intersection on the main road to be controlled in accordance with the traffic signal control strategies described above.

14. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method described in any one of claims 1-10.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-10.

Citation Information

Patent Citations

  • Traffic signal lamp control method and system based on reinforcement learning

    CN112863206A

  • Multi-mode traffic artery signal coordination control method and device based on multi-agent cooperation

    CN113299078A