Traffic signal control method and computer readable storage medium

CN122821780APending Publication Date: 2026-09-25ZHEJIANG ELECTROMECHANICAL VOCATIONAL & TECH COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610993815.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]在本实施例中提供了一种交通信号控制方法和计算机可读存储介质,以解决相关技术中存在复杂交通环境时交通信号控制效率较低的问题

Benefits of technology

[0034]与相关技术相比,在本实施例中提供的交通信号控制方法,通过采集目标路口的交通视频数据;对交通视频数据进行各流向的车道状态特征提取,得到各个流向的车道状态特征向量;根据各个流向的车道状态特征向量,得到交通状态特征向量;并确定交通状态特征向量对应的交通模式标签;将交通状态特征向量和交通模式标签输入至训练后的全局策略网络,得到全局策略网络输出的策略引导向量;策略引导向量表征各通行方向的优先级;将交通状态特征向量和策略引导向量输入至训练后的局部策略网络,得到局部策略网络输出的交通信号控制指令。其能够通过全局策略网络和局部策略网络的分级策略决策框架,将交通信号控制过程从单一映射转变为多粒度协同优化,提高了复杂交通环境时交通信号控制的效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821780A_ABST
    Figure CN122821780A_ABST
Patent Text Reader

Abstract

The application relates to a traffic signal control method and a computer readable storage medium, wherein the traffic signal control method comprises the following steps: collecting traffic video data of a target intersection; extracting lane state feature vectors of each flow direction from the traffic video data; obtaining a traffic state feature vector according to the lane state feature vectors of each flow direction; determining a corresponding traffic mode label; inputting the traffic state feature vector and the traffic mode label into a trained global strategy network to obtain a strategy guide vector; and inputting the traffic state feature vector and the strategy guide vector into a trained local strategy network to obtain a traffic signal control instruction output by the local strategy network. The hierarchical strategy decision framework of the global strategy network and the local strategy network can change the traffic signal control process from single mapping to multi-granularity collaborative optimization, thereby improving the efficiency of traffic signal control in a complex traffic environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent transportation and intelligent control technology, and in particular to traffic signal control methods and computer-readable storage media. Background Technology

[0002] With the continuous growth of urban motor vehicle ownership, traffic congestion has become increasingly prominent. As a key node in the urban traffic system, the operational efficiency of signalized intersections directly affects road capacity and overall traffic operation.

[0003] Traditional traffic signal control methods mainly include two categories: fixed timing control and adaptive control based on inductive detection. Fixed timing methods rely on historical statistical data for periodic configuration, making it difficult to adapt to dynamic changes in traffic flow. While inductive control methods can make adjustments to some extent based on local detection information, they rely on geomagnetic coils or ground-based detection equipment, resulting in high deployment costs, complex maintenance, and susceptibility to environmental influences. Therefore, these technologies suffer from low traffic signal control efficiency in complex traffic environments.

[0004] There is currently no effective solution to the problem of low traffic signal control efficiency in complex traffic environments in related technologies. Summary of the Invention

[0005] This embodiment provides a traffic signal control method and a computer-readable storage medium to address the problem of low traffic signal control efficiency in complex traffic environments in related technologies.

[0006] Firstly, this embodiment provides a traffic signal control method, including:

[0007] Collect traffic video data at the target intersection;

[0008] Lane state features for each direction are extracted from the traffic video data to obtain lane state feature vectors for each direction; traffic state feature vectors are obtained based on the lane state feature vectors for each direction; and traffic mode labels corresponding to the traffic state feature vectors are determined.

[0009] The traffic state feature vector and the traffic mode label are input into the trained global policy network to obtain the policy guidance vector output by the global policy network; the policy guidance vector represents the priority of each traffic direction.

[0010] The traffic state feature vector and the policy guidance vector are input into the trained local policy network to obtain the traffic signal control command output by the local policy network.

[0011] In some embodiments, the extraction of lane state features from the traffic video data for each direction of traffic flow to obtain lane state feature vectors for each direction of traffic flow includes:

[0012] Using a pre-trained target detection and tracking model, lane state features for each direction are extracted from the traffic video data to obtain lane state feature vectors for each direction.

[0013] In some embodiments, determining the traffic pattern label corresponding to the traffic state feature vector includes:

[0014] The K-means algorithm is used to perform cluster analysis on the traffic state feature vector to obtain the traffic pattern label corresponding to the traffic state feature vector.

[0015] In some embodiments, the training process of the global policy network includes:

[0016] Collect traffic video data from intersections as training samples;

[0017] Lane state features for each direction are extracted from the traffic video data training samples to obtain lane state feature vector training data for each direction; traffic state feature vector training data is obtained based on the lane state feature vector training data for each direction; and traffic mode label training data corresponding to the traffic state feature vector training data is determined.

[0018] The initial global policy network is trained using the traffic state feature vector training data and the traffic mode label training data to obtain the trained global policy network that outputs policy guidance vector training data.

[0019] In some embodiments, the trained global policy network, which uses the traffic state feature vector training data and the traffic pattern label training data to train an initial global policy network to obtain output policy guidance vector training data, includes:

[0020] Construct the initial global policy network; the global policy network is a multilayer perceptron;

[0021] The traffic state feature vector training data and the traffic mode label training data are input into the initial global policy network to train the initial global policy network with the goal of maximizing the first cumulative reward, resulting in the trained global policy network that outputs the policy guidance vector training data.

[0022] In some embodiments, the training process of the global policy network further includes:

[0023] Based on the traffic mode label training data, a generative adversarial network is used to obtain traffic state feature vector synthesis data.

[0024] Using the traffic state feature vector training data, the traffic state feature vector synthesis data, and the traffic mode label training data, the trained global policy network is optimized and trained to obtain an optimized global policy network with output policy guidance vector synthesis data.

[0025] In some embodiments, the training process of the local policy network includes:

[0026] Obtain the traffic state feature vector training data and the strategy guidance vector training data;

[0027] Using the traffic state feature vector training data and the policy guidance vector training data, the initial local policy network is trained to obtain the trained local policy network that outputs traffic signal control command training data.

[0028] In some embodiments, the trained local policy network, which uses the traffic state feature vector training data and the policy guidance vector training data to train an initial local policy network to obtain training data for outputting traffic signal control commands, includes:

[0029] Construct the initial local policy network; the local policy network is a multilayer perceptron;

[0030] The traffic state feature vector training data and the policy guidance vector training data are input into the initial local policy network to train the initial local policy network with the goal of maximizing the second cumulative reward, resulting in the trained local policy network that outputs the traffic signal control command training data.

[0031] In some embodiments, the training process of the local policy network further includes:

[0032] Using the traffic state feature vector training data, the traffic state feature vector synthesis data, the strategy guidance vector training data, and the strategy guidance vector synthesis data, the trained local strategy network is optimized and trained to obtain an optimized local strategy network that outputs synthesized traffic signal control commands.

[0033] Secondly, this embodiment provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the traffic signal control method described in the first aspect.

[0034] Compared with related technologies, the traffic signal control method provided in this embodiment collects traffic video data at the target intersection; extracts lane state features for each direction from the traffic video data to obtain lane state feature vectors for each direction; obtains traffic state feature vectors based on the lane state feature vectors for each direction; determines the traffic mode label corresponding to the traffic state feature vector; inputs the traffic state feature vector and traffic mode label into a trained global policy network to obtain a policy guidance vector output by the global policy network; the policy guidance vector represents the priority of each traffic direction; and inputs the traffic state feature vector and policy guidance vector into a trained local policy network to obtain traffic signal control instructions output by the local policy network. This method, through a hierarchical policy decision-making framework of global and local policy networks, transforms the traffic signal control process from a single mapping to multi-granularity collaborative optimization, improving the efficiency of traffic signal control in complex traffic environments.

[0035] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description

[0036] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0037] Figure 1 This is a hardware structure block diagram of the terminal of the traffic signal control method in this embodiment;

[0038] Figure 2 This is a flowchart of the traffic signal control method in this embodiment;

[0039] Figure 3 This is a flowchart of the training process of the global policy network in this embodiment;

[0040] Figure 4 This is a flowchart of the training process of the local policy network in this embodiment. Detailed Implementation

[0041] To better understand the purpose, technical solution, and advantages of this application, the application is described and illustrated below in conjunction with the accompanying drawings and embodiments.

[0042] Unless otherwise defined, the technical or scientific terms used in this application shall have the general meaning understood by one of ordinary skill in the art to which this application pertains. Words such as “a,” “an,” “an,” “the,” “the,” and “these” used in this application do not indicate quantitative limitation and may be singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include steps or modules (units) not listed, or may include other steps or modules (units) inherent to these processes, methods, products, or devices. Words such as “connected,” “linked,” and “coupled” used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. Normally, the character " / " indicates that the objects before and after it are in an "or" relationship. The terms "first," "second," "third," etc., used in this application are merely to distinguish similar objects and do not represent a specific order of objects.

[0043] The method embodiments provided in this example can be executed on a terminal, computer, or similar computing device. For example, it can run on a terminal. Figure 1 This is a hardware structure block diagram of the terminal of the traffic signal control method in this embodiment. For example... Figure 1 As shown, a terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 and a memory 104 for storing data are also included. The processor 102 may be, but is not limited to, a microprocessor (MCU) or a programmable logic device (FPGA). The terminal may also include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that… Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the terminal described above. For example, the terminal may also include components that are larger than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown are illustrated.

[0044] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the traffic signal control method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thereby implementing the aforementioned method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0045] The transmission device 106 is used to receive or send data via a network. This network includes a wireless network provided by the terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 can be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0046] This embodiment provides a traffic signal control method. Figure 2 This is a flowchart of the traffic signal control method in this embodiment, as shown below. Figure 2 As shown, the process includes the following steps:

[0047] Step S201: Collect traffic video data at the target intersection.

[0048] Specifically, the target intersection can be an urban intersection. High-definition video cameras are deployed at each of the four entrances to the target intersection, with uniformly configured parameters such as resolution and video frame rate. These four entrances refer to the east, south, west, and north entrances. The high-definition cameras can continuously collect traffic video data from the target intersection. This traffic video data includes vehicle traffic images from each lane in each of the four entrances to the target intersection.

[0049] Step S202: Extract lane state features for each direction from the traffic video data to obtain lane state feature vectors for each direction; obtain traffic state feature vectors based on lane state feature vectors for each direction; and determine the traffic mode label corresponding to the traffic state feature vectors.

[0050] Specifically, the flow direction refers to the combination of the entrance direction and the vehicle travel direction. The entrance directions include east entrance, south entrance, west entrance, and north entrance; the vehicle travel directions include straight ahead and left turn. Therefore, there are a total of eight flow directions: east entrance straight ahead, south entrance straight ahead, west entrance straight ahead, north entrance straight ahead, east entrance left turn, south entrance left turn, west entrance left turn, and north entrance left turn. The lane state feature vector includes average lane flow, average lane queue length, average lane speed, and lane accident label. This traffic state feature vector is composed of lane state feature vectors for the eight flow directions. The traffic mode label can include morning peak mode, off-peak mode, evening peak mode, low flow mode, accident mode, etc. Using a pre-trained target detection and tracking model, lane state features for each flow direction are extracted from traffic video data to obtain lane state feature vectors for each flow direction. The lane state feature vectors for each flow direction are combined to obtain the traffic state feature vector. Using the K-means algorithm, cluster analysis is performed on the traffic state feature vectors to determine the traffic mode labels corresponding to the traffic state feature vectors.

[0051] Step S203: Input the traffic state feature vector and traffic mode label into the trained global policy network to obtain the policy guidance vector output by the global policy network.

[0052] Specifically, the traffic state feature vector consists of lane state feature vectors for eight different traffic directions. These lane state feature vectors include average lane flow rate, average lane queue length, average lane speed, and lane accident label. The traffic mode label can include morning peak mode, off-peak mode, evening peak mode, low flow mode, and accident mode. The global policy network is a multilayer perceptron. The policy guidance vector represents the priority of each traffic direction. This policy guidance vector includes the priority of straight-ahead traffic in the north-south direction, the priority of left turns in the north-south direction, the priority of straight-ahead traffic in the east-west direction, and the priority of left turns in the east-west direction.

[0053] Step S204: Input the traffic state feature vector and the policy guidance vector into the trained local policy network to obtain the traffic signal control command output by the local policy network.

[0054] Specifically, the traffic state feature vector consists of lane state feature vectors for eight flow directions. These lane state feature vectors include average lane flow rate, average lane queue length, average lane speed, and lane accident labels. The policy guidance vector includes the priority of north-south straight travel, north-south left turn, east-west straight travel, and east-west left turn. The local policy network is a multilayer perceptron. The traffic signal control command includes the permitted direction and the duration of the permitted direction. The permitted direction can be north-south straight travel, north-south left turn, east-west straight travel, or east-west left turn.

[0055] In this embodiment, the traffic signal control method acquires traffic video data at the target intersection; extracts lane state features for each direction from the traffic video data to obtain lane state feature vectors for each direction; obtains traffic state feature vectors based on the lane state feature vectors for each direction; determines the traffic mode label corresponding to the traffic state feature vector; inputs the traffic state feature vector and traffic mode label into a trained global policy network to obtain a policy guidance vector output by the global policy network; the policy guidance vector represents the priority of each traffic direction; and inputs the traffic state feature vector and policy guidance vector into a trained local policy network to obtain traffic signal control instructions output by the local policy network. This method, through a hierarchical policy decision-making framework of global and local policy networks, transforms the traffic signal control process from a single mapping to multi-granularity collaborative optimization, improving the efficiency of traffic signal control in complex traffic environments.

[0056] In some of these embodiments, lane state features for each direction are extracted from traffic video data to obtain lane state feature vectors for each direction. This includes: using a pre-trained target detection and tracking model to extract lane state features for each direction from traffic video data to obtain lane state feature vectors for each direction.

[0057] Specifically, the traffic video data includes images of vehicle traffic in each lane from the four entrances to the target intersection. Each frame... Lane-level vehicle detection results are obtained using a pre-trained object detection algorithm (such as YOLO V9). These vehicle detection results are represented by vehicle bounding boxes. ,in The index represents the detected first... a car, and These are the x and y coordinates of the center point of the bounding box, respectively. and These represent the width and height of the bounding box, respectively.

[0058] After obtaining the vehicle bounding box, a pre-trained target tracking algorithm (such as SORT) is used to track the vehicle, resulting in the... vehicle trajectory ,in This is the frame window length. Based on the tracking results, statistics for each lane are calculated. In time interval Number of vehicles passing through And calculate the traffic flow in the lanes. Simultaneously, the average vehicle speed is estimated using the positional changes of adjacent frames. The calculation formula is expressed as:

[0059] ;

[0060] in For the first The car is Location at any given moment For the first The car is Location at any given moment This is a scaling constant for converting pixel distance to actual spatial distance. Calculate the queue length for vehicles in each lane. The calculation formula is expressed as:

[0061] ;

[0062] in For vehicles with speeds below a threshold before the stop line The number of vehicles, This represents the average length of each vehicle. If a vehicle is detected to be stationary in the lane for an extended period, i.e., the stopping time exceeds a predefined hyperparameter... If so, it is determined to be an accident, and it is represented by a binary variable. The tag is the accident label.

[0063] This flow direction refers to the combination of the entrance direction and the vehicle travel direction. The entrance directions include east entrance, south entrance, west entrance, and north entrance; the vehicle travel directions include straight ahead and left turn. Therefore, there are a total of eight flow directions: east entrance straight ahead, south entrance straight ahead, west entrance straight ahead, north entrance straight ahead, east entrance left turn, south entrance left turn, west entrance left turn, and north entrance left turn. The lane state feature vectors for each flow direction include average lane flow rate, average lane queue length, average lane speed, and lane accident label.

[0064] In some of these embodiments, determining the traffic pattern label corresponding to the traffic state feature vector includes: using the K-means algorithm to perform cluster analysis on the traffic state feature vector to obtain the traffic pattern label corresponding to the traffic state feature vector.

[0065] Specifically, the lane state feature vectors of each flow direction are combined to obtain the traffic state feature vector. , is represented as:

[0066] ;

[0067] Calculate the traffic state feature vector Calculate the distance to each cluster center (usually using Euclidean distance), find the cluster center with the smallest distance, and use the cluster center number as the traffic mode label corresponding to the traffic state feature vector. The calculation formula is:

[0068] ;

[0069] in This is a traffic state feature vector. For the first The cluster center of each cluster, This represents the total number of clusters.

[0070] In some of these embodiments, Figure 3 This is a flowchart of the training process of the global policy network in this embodiment, as shown below. Figure 3 As shown, the training process of the global policy network includes:

[0071] Step S301: Collect traffic video data training samples at the intersection;

[0072] Step S302: Extract lane state features for each direction from the traffic video data training samples to obtain lane state feature vector training data for each direction; obtain traffic state feature vector training data based on the lane state feature vector training data for each direction; and determine the traffic mode label training data corresponding to the traffic state feature vector training data.

[0073] Step S303: Using traffic state feature vector training data and traffic pattern label training data, train the initial global policy network to obtain the trained global policy network with output policy guidance vector training data.

[0074] Specifically, this intersection can be an urban intersection. High-definition cameras are deployed at each of the four entrances to the intersection, with uniformly configured parameters such as resolution and video frame rate. These four entrances refer to the east, south, west, and north entrances. The high-definition cameras can continuously collect traffic video data training samples from the intersection. These training samples include vehicle traffic images from each lane in each of the four entrances.

[0075] Each frame of images Lane-level vehicle detection results are obtained using a pre-trained object detection algorithm (such as YOLO V9). These vehicle detection results are represented by vehicle bounding boxes. ,in The index represents the detected first... a car, and These are the x and y coordinates of the center point of the bounding box, respectively. and These represent the width and height of the bounding box, respectively.

[0076] After obtaining the vehicle bounding box, a pre-trained target tracking algorithm (such as SORT) is used to track the vehicle, resulting in the... vehicle trajectory ,in This is the frame window length. Based on the tracking results, statistics for each lane are calculated. In time interval Number of vehicles passing through And calculate the traffic flow in the lanes. Simultaneously, the average vehicle speed is estimated using the positional changes of adjacent frames. The calculation formula is expressed as:

[0077] ;

[0078] in For the first The car is Location at any given moment For the first The car is Location at any given moment This is a scaling constant for converting pixel distance to actual spatial distance. Calculate the queue length for vehicles in each lane. The calculation formula is expressed as:

[0079] ;

[0080] in For vehicles with speeds below a threshold before the stop line The number of vehicles, This represents the average length of each vehicle. If a vehicle is detected to be stationary in the lane for an extended period, i.e., the stopping time exceeds a predefined hyperparameter... If so, it is determined to be an accident, and it is represented by a binary variable. The tag is the accident label.

[0081] This flow direction refers to the combination of the entrance direction and the vehicle travel direction. The entrance directions include east entrance, south entrance, west entrance, and north entrance; the vehicle travel directions include straight ahead and left turn. Therefore, there are a total of eight flow directions: east entrance straight ahead, south entrance straight ahead, west entrance straight ahead, north entrance straight ahead, east entrance left turn, south entrance left turn, west entrance left turn, and north entrance left turn. The lane state feature vector training data for each flow direction includes average lane flow rate, average lane queue length, average lane speed, and lane accident labels.

[0082] The traffic state feature vector training data is obtained by combining the lane state feature vector training data of each direction. , is represented as:

[0083] ;

[0084] At different times Traffic state feature vector training data Then, the K-means unsupervised clustering algorithm is used to perform cluster analysis on the traffic state feature vector training data, with the goal of minimizing the overall clustering error. The calculation formula is expressed as:

[0085] ;

[0086] in For training data of traffic state feature vectors, For the first The cluster center of each cluster, The total number of clusters, The calculation formula for traffic pattern label training data is as follows:

[0087] ;

[0088] in For training data of traffic state feature vectors, For the first The cluster center of each cluster, This represents the total number of clusters.

[0089] Construct an initial global policy network. This global policy network is a multilayer perceptron. Input traffic state feature vector training data and traffic pattern label training data into the initial global policy network, and train the initial global policy network with the objective of maximizing the first cumulative reward, to obtain the trained global policy network with output policy guidance vector training data.

[0090] In some embodiments, an initial global policy network is trained using traffic state feature vector training data and traffic pattern label training data to obtain a trained global policy network that outputs policy guidance vector training data, including:

[0091] Construct the initial global policy network; the global policy network is a multilayer perceptron;

[0092] The traffic state feature vector training data and traffic pattern label training data are input into the initial global policy network. The initial global policy network is trained with the goal of maximizing the first cumulative reward, and the trained global policy network is obtained by outputting the policy guidance vector training data.

[0093] Specifically, this global policy network is a multilayer perceptron. The input layer of this global policy network is 33-dimensional and contains two hidden layers. The first hidden layer uses 128 neurons and employs the ReLU activation function, and the second hidden layer uses 64 neurons and also employs the ReLU activation function. The output of this global policy network is represented as follows:

[0094] ;

[0095] in For a trainable weight matrix, For training data of traffic state feature vectors and traffic mode labels, For a trainable weight matrix, For bias, For bias. This first cumulative reward Represented as:

[0096] ;

[0097] in At the starting time, For the preset time step, As a discount factor, For immediate rewards, the calculation formula is as follows:

[0098] ;

[0099] in For flow direction exist Average queue length at any given time. To maximize the first cumulative reward. The initial global policy network is trained for the target. Policy-guided vector training data. , The meta-policy library is a predefined discrete action space, represented as:

[0100] ;

[0101] in This is the timing target vector, where each dimension represents the priority of a traffic type. Priority for north-south straight travel; Priority for left turns in north-south direction, Priority for east-west straight travel; Priority for left turns in the east-west direction. The sum of the four elements is 1.

[0102] In some embodiments, the training process of the global policy network further includes:

[0103] Based on the traffic mode label training data, a generative adversarial network is used to obtain traffic state feature vector synthesis data;

[0104] By using traffic state feature vector training data, traffic state feature vector synthesis data, and traffic pattern label training data, the trained global policy network is optimized to obtain an optimized global policy network with output policy guidance vector synthesis data.

[0105] Specifically, the generative adversarial network (GAN) comprises a generator network G and a discriminator network D. The generator network G is implemented using a fully connected neural network with two hidden layers, the first of which is... Represented as:

[0106] ;

[0107] Second hidden layer Represented as:

[0108] ;

[0109] Output of generator network G Represented as:

[0110] ;

[0111] in For a trainable weight matrix, For a trainable weight matrix, For a trainable weight matrix, For bias, For bias, As a bias. The random noise vector Traffic pattern label training data Input the generator network to obtain synthesized traffic state feature vector data. .

[0112] The discriminator network D is implemented by a fully connected neural network containing two hidden layers. The first hidden layer... Represented as:

[0113] ;

[0114] Second hidden layer Represented as:

[0115] ;

[0116] The output S of the discriminator network D is expressed as:

[0117] ;

[0118] in For a trainable weight matrix, For a trainable weight matrix, For a trainable weight matrix, For bias, For bias, For bias. The output of the discriminator network D. As S approaches 1, the traffic state feature vector synthesis data... The closer S is to real data, the better the traffic state feature vector synthesized data becomes; when S is closer to 0, the better the traffic state feature vector synthesized data becomes. The data differs significantly from the actual data.

[0119] The loss function of the generator network G is expressed as:

[0120] ;

[0121] in For mathematical expectation, The score output by the discriminator. The generator network G is optimized to minimize the loss function of the generator network. The loss function of the discriminator network D is expressed as:

[0122] ;

[0123] in For mathematical expectation, The distribution of traffic state feature vector training data is defined. The discriminator network D is optimized with the objective of maximizing its loss function. Using the traffic state feature vector training data, the synthesized traffic state feature vector data, and the traffic pattern label training data, the trained global policy network is optimized to obtain the optimized global policy network with synthesized output policy guidance vector data. This optimization training process is the same as the global policy network training process described above, and its details are not elaborated here.

[0124] In some of these embodiments, Figure 4 This is a flowchart of the training process of the local policy network in this embodiment, as shown below. Figure 4 As shown, the training process of the local policy network includes:

[0125] Step S401: Obtain traffic state feature vector training data and strategy guidance vector training data;

[0126] Step S402: Using traffic state feature vector training data and strategy guidance vector training data, train the initial local policy network to obtain the trained local policy network that outputs traffic signal control command training data.

[0127] Specifically, an initial local policy network is constructed. This local policy network is a multilayer perceptron. Traffic state feature vector training data and policy guidance vector training data are acquired. The traffic state feature vector training data and policy guidance vector training data are input into the initial local policy network, and the initial local policy network is trained with the goal of maximizing the second cumulative reward, resulting in a trained local policy network that outputs traffic signal control command training data.

[0128] In some embodiments, an initial local policy network is trained using traffic state feature vector training data and policy guidance vector training data to obtain a trained local policy network that outputs traffic signal control command training data, including:

[0129] Construct the initial local policy network; the local policy network is a multilayer perceptron;

[0130] The traffic state feature vector training data and the policy guidance vector training data are input into the initial local policy network. The initial local policy network is trained with the goal of maximizing the second cumulative reward, resulting in the trained local policy network that outputs traffic signal control command training data.

[0131] Specifically, this local policy network is a multilayer perceptron. The input layer of this local policy network is 36-dimensional and contains two hidden layers. The first hidden layer uses 128 neurons and employs the ReLU activation function, as does the second hidden layer. The output of this local policy network is represented as follows:

[0132] ;

[0133] in For a trainable weight matrix, The training data are for traffic state feature vectors and policy guidance vectors. For a trainable weight matrix, For bias, For bias. This second cumulative reward Represented as:

[0134] ;

[0135] in At the starting time, For the preset time step, As a discount factor, For immediate rewards, the calculation formula is as follows:

[0136] ;

[0137] in For lane At any moment Average queue length The average waiting time for vehicles at the intersection. These are the weighting coefficients. These are weighting coefficients. The goal is to maximize the second cumulative reward. The initial local policy network is trained for the target. Traffic signal control command training data. Represented as:

[0138] ;

[0139] in The permitted direction can be straight north-south, left turn north-south, straight east-west, or left turn east-west. The duration of the permitted direction.

[0140] In some embodiments, the training process of the local policy network further includes:

[0141] By using traffic state feature vector training data, traffic state feature vector synthesis data, strategy guidance vector training data, and strategy guidance vector synthesis data, the trained local policy network is optimized to obtain the optimized local policy network that outputs traffic signal control command synthesis data.

[0142] Specifically, the training data of traffic state feature vectors, the synthesized data of traffic state feature vectors, the training data of policy guidance vectors, and the synthesized data of policy guidance vectors are input into the trained local policy network. The local policy network is optimized and trained with the goal of maximizing the second cumulative reward, resulting in an optimized local policy network that outputs synthesized traffic signal control commands. This optimization training process is the same as the local policy network training process described above, and its specific details will not be elaborated upon.

[0143] Furthermore, in conjunction with the traffic signal control methods provided in the above embodiments, this embodiment can also provide a computer-readable storage medium for implementation. The storage medium stores a computer program; when executed by a processor, the computer program implements any of the traffic signal control methods in the above embodiments.

[0144] It should be understood that the specific embodiments described herein are merely illustrative of the application and not intended to limit it. All other embodiments derived by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.

[0145] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0146] Obviously, the accompanying drawings are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar situations based on these drawings without any creative effort. Furthermore, it is understood that although the work done in this development process may be complex and lengthy, for those skilled in the art, certain design, manufacturing, or production modifications made based on the technical content disclosed in this application are merely conventional technical means and should not be considered as insufficient disclosure of this application.

[0147] The term "embodiment" in this application refers to a specific feature, structure, or characteristic described in connection with an embodiment that may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily imply the same embodiment, nor does it imply that it is mutually exclusive with or independent of other embodiments. It will be clearly or implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0148] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.

Claims

1. A traffic signal control method, characterized in that, include: Collect traffic video data at the target intersection; Lane state features for each direction are extracted from the traffic video data to obtain lane state feature vectors for each direction; traffic state feature vectors are obtained based on the lane state feature vectors for each direction; and traffic mode labels corresponding to the traffic state feature vectors are determined. The traffic state feature vector and the traffic mode label are input into the trained global policy network to obtain the policy guidance vector output by the global policy network; the policy guidance vector represents the priority of each traffic direction. The traffic state feature vector and the policy guidance vector are input into the trained local policy network to obtain the traffic signal control command output by the local policy network.

2. The traffic signal control method according to claim 1, characterized in that, The step of extracting lane state features for each direction from the traffic video data to obtain lane state feature vectors for each direction includes: Using a pre-trained target detection and tracking model, lane state features for each direction are extracted from the traffic video data to obtain lane state feature vectors for each direction.

3. The traffic signal control method according to claim 1, characterized in that, Determining the traffic mode label corresponding to the traffic state feature vector includes: The K-means algorithm is used to perform cluster analysis on the traffic state feature vector to obtain the traffic pattern label corresponding to the traffic state feature vector.

4. The traffic signal control method according to claim 1, characterized in that, The training process of the global policy network includes: Collect traffic video data from intersections as training samples; Lane state features for each direction are extracted from the traffic video data training samples to obtain lane state feature vector training data for each direction; traffic state feature vector training data is obtained based on the lane state feature vector training data for each direction; and traffic mode label training data corresponding to the traffic state feature vector training data is determined. The initial global policy network is trained using the traffic state feature vector training data and the traffic mode label training data to obtain the trained global policy network that outputs policy guidance vector training data.

5. The traffic signal control method according to claim 4, characterized in that, The step of training an initial global policy network using the traffic state feature vector training data and the traffic pattern label training data to obtain the trained global policy network that outputs policy guidance vector training data includes: Construct the initial global policy network; the global policy network is a multilayer perceptron; The traffic state feature vector training data and the traffic mode label training data are input into the initial global policy network to train the initial global policy network with the goal of maximizing the first cumulative reward, resulting in the trained global policy network that outputs the policy guidance vector training data.

6. The traffic signal control method according to claim 4, characterized in that, The training process of the global policy network also includes: Based on the traffic mode label training data, a generative adversarial network is used to obtain traffic state feature vector synthesis data. Using the traffic state feature vector training data, the traffic state feature vector synthesis data, and the traffic mode label training data, the trained global policy network is optimized and trained to obtain an optimized global policy network with output policy guidance vector synthesis data.

7. The traffic signal control method according to claim 6, characterized in that, The training process of the local policy network includes: Obtain the traffic state feature vector training data and the strategy guidance vector training data; The initial local policy network is trained using the traffic state feature vector training data and the policy guidance vector training data to obtain the trained local policy network that outputs traffic signal control command training data.

8. The traffic signal control method according to claim 7, characterized in that, The step of training an initial local policy network using the traffic state feature vector training data and the policy guidance vector training data to obtain the trained local policy network that outputs traffic signal control command training data includes: Construct the initial local policy network; the local policy network is a multilayer perceptron; The traffic state feature vector training data and the policy guidance vector training data are input into the initial local policy network to train the initial local policy network with the goal of maximizing the second cumulative reward, resulting in the trained local policy network that outputs the traffic signal control command training data.

9. The traffic signal control method according to claim 7, characterized in that, The training process of the local policy network also includes: Using the traffic state feature vector training data, the traffic state feature vector synthesis data, the strategy guidance vector training data, and the strategy guidance vector synthesis data, the trained local strategy network is optimized and trained to obtain an optimized local strategy network that outputs synthesized traffic signal control commands.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the traffic signal control method according to any one of claims 1 to 9.