Traffic signal control method, electronic device and system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
- Filing Date
- 2025-05-07
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]本申请的目的在于提供一种交通信号控制方法,从而解决交通信号控制无法响应交通流的波动,容易造成拥堵,对用户的使用体验不佳的问题
[0049] Based on the traffic control system provided in the third aspect of this application, historical traffic data of the target intersection can be acquired, and historical traffic sequence elements and traffic sequence elements to be decided can be concatenated to generate a decision sequence. A sequence-based decision-making method is then used to assist in the current traffic control cycle decision-making based on the historical traffic conditions of the target intersection. This traffic control system can fully utilize historical information, thereby better adapting to traffic flow fluctuations, effectively reducing vehicle congestion, and improving the user experience.
Smart Images

Figure CN120452192B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of signal control technology, and in particular to a traffic signal control method, electronic device and system. Background Technology
[0002] With the increasing number of motor vehicles in cities, higher application requirements are being placed on signal control at single-point intersections to ensure the growing demands for traffic safety and efficiency. Since single-step decision-making is often based on the flow rate of the most recent period (generally calculated on a periodic basis or every 5 to 10 minutes), it cannot respond to fluctuations in traffic flow and is prone to decision-making errors. This often results in signal control cycles that are too large or too small, increasing delays, number of stops, and other indicators, easily causing congestion and a poor user experience. Summary of the Invention
[0003] The purpose of this application is to provide a traffic signal control method to solve the problem that traffic signal control cannot respond to fluctuations in traffic flow, which easily causes congestion and results in a poor user experience.
[0004] Based on the above objectives, in a first aspect, this application provides a traffic signal control method, the control method comprising:
[0005] Acquire historical traffic data for the target intersection. The historical traffic data includes multiple historical traffic sequence elements. Each historical traffic sequence element includes a report, a status, and a signal control cycle. The status represents the critical lane demand flow corresponding to different phases in the signal control cycle, and the report represents the environmental feedback when different signal control cycles are adopted.
[0006] The state of the target decision cycle is acquired using radar-guided equipment, and traffic sequence elements to be decided are generated based on the state.
[0007] Generate a decision sequence based on historical traffic sequence elements and traffic sequence elements to be decided;
[0008] The decision sequence is input into a preset offline reinforcement learning model. The offline reinforcement learning model infers based on the reward, state and signal control period in the historical traffic sequence elements to obtain the signal control period inference value of the traffic sequence elements to be decided.
[0009] Traffic light duration allocation for different phases is based on the signal control cycle inference value.
[0010] Based on this, this application acquires historical traffic data of the target intersection, concatenates historical traffic sequence elements and traffic sequence elements to be decided to generate a decision sequence, and adopts a sequence-based decision-making approach. By leveraging the historical traffic conditions of the target intersection, it assists in decision-making during the current traffic control cycle. This method fully utilizes historical information, thereby better adapting to traffic flow fluctuations, effectively reducing vehicle congestion, and improving the user experience.
[0011] Furthermore, the control method also includes training an offline reinforcement learning model, including:
[0012] The initial element state of the initial traffic sequence element in the training sequence is defined according to the magnitude of the demand flow of the key lane. Based on the initial element state, the reward and signal control period of the corresponding element are obtained through the preset traffic mechanism model.
[0013] The initial element state is multiplied by the change ratio vector using a Hadma multiplication to obtain the derived element state of multiple derived traffic sequence elements in the training sequence. Based on the derived element state, the reward and signal control cycle of the corresponding element are obtained through the traffic mechanism model.
[0014] The training sequence is input into the initialized offline reinforcement learning model. The loss function is calculated based on the output of the offline reinforcement learning model. The parameters of the offline reinforcement learning model are updated with the goal of reducing the loss function.
[0015] Based on this, the initial element states are defined according to the demand flow of the key lanes, making the element states in the training sequence more consistent with the actual environment. A large number of training sequences are then input into the offline reinforcement learning model. According to the training method of the offline reinforcement learning model provided in this application, the offline reinforcement learning model can be trained to ensure that when making signal control cycle decisions on the decision sequence to be decided, the optimal signal control cycle inference value can be output based on the trained offline reinforcement learning model, thereby improving the traffic efficiency of the intersection and reducing congestion.
[0016] Furthermore, through a pre-defined traffic mechanism model, the rewards and signal control cycles corresponding to each traffic sequence element are obtained, including:
[0017] Based on signal delay, stopping time, and lane queue length, a comprehensive performance index is formed. This performance index is used to characterize the environmental feedback reward when different signal control cycles are adopted under different traffic sequence element states.
[0018] The state of traffic sequence elements is input into the traffic mechanism model, and the signal control cycle output by the traffic mechanism model is obtained with the goal of minimizing the performance index.
[0019] Based on this, data augmentation technology is introduced to generate a large number of training sequences using traffic mechanism models for offline reinforcement learning models, effectively solving the problem that the number of samples at single-point intersections is small and it is difficult to provide enough raw sample data for offline reinforcement learning models to train.
[0020] Furthermore, through a pre-defined traffic mechanism model, the rewards and signal control cycles corresponding to each traffic sequence element are obtained, including:
[0021] The state of the traffic sequence element is input into the traffic mechanism model, and the model is traversed within the feasible range of the signal control period to obtain the performance index corresponding to different signal control periods. The signal control period that minimizes the performance index is selected as the optimal signal control period of the traffic sequence element. The performance index corresponding to the optimal signal control period is weighted and inverted to obtain the reward of the traffic sequence element.
[0022] Based on this, this application performs a weighted inverse of the performance metrics to obtain the return of the traffic sequence element (i.e., The smaller the performance index PI, the lower the negative impact of factors such as delays, parking, and queuing on traffic efficiency. Therefore, the minimized performance index, after weighted inversion, can be used to characterize the optimal environmental feedback. Offline reinforcement learning models use historical action sequences (i.e., signal control cycles) and reward sequences (i.e.,...) to represent the optimal environmental feedback. By performing inference, the optimal signal control cycle inference value can be output, thereby improving the traffic efficiency of the target single-point intersection and reducing traffic congestion.
[0023] Furthermore, the control method also includes parameter calibration of the traffic mechanism model, including:
[0024] Obtain a real sample dataset, which includes multiple real sample data. Each real sample data includes the mapping relationship between the state of traffic sequence elements and the optimal signal control cycle.
[0025] Set the range of values for the parameters to be calibrated, and use a genetic algorithm to randomly generate a population of values for the parameters within the range of the calibration parameters;
[0026] By inputting real sample data into the traffic mechanism model, performance indicators under different values of the parameters to be calibrated in the population are obtained.
[0027] The optimal calibration parameters are obtained by aiming to minimize the sum of performance indicators corresponding to each real sample data.
[0028] Based on this, parameter calibration of the traffic mechanism model can ensure the reliability of the signal control cycle and reward generated by data augmentation, thus providing a foundation for the training of subsequent offline reinforcement learning models.
[0029] Furthermore, the traffic mechanism model adjusts the parameters and signal control cycle using the following formula:
[0030]
[0031] In the formula, w1, w2, w3, w4, w5, and w6 represent the parameters of the traffic mechanism model, d(C,f,w4) represents the signal delay function corresponding to the signal control period, s(C,f,w5) represents the parking time function corresponding to the signal control period, q(C,f,w6) represents the queue length function corresponding to the signal control period, C represents the signal control period of the traffic sequence element as a sample, and f represents the state of the traffic sequence element as a sample.
[0032] Based on this, the performance index of the traffic mechanism model is formed by comprehensively considering the intersection delay d, parking time s, and queuing time q. The input of the traffic mechanism model is the element state f = (f1, f2, ..., f p ), p is the number of phases, f p This represents the critical lane demand flow in the p-th phase, with the optimization objective being to minimize the performance index PI, within the feasible space of the signal control cycle [C]. min C max By iterating through the [database], the optimal signal control period C* and the corresponding optimal performance index PI can be obtained. * In this way, a large number of training sequences can be generated for offline reinforcement learning models to train, effectively solving the problem that the number of samples at a single intersection is small, making it difficult to provide enough raw sample data for offline reinforcement learning models to train.
[0033] Furthermore, the signal delay function is calculated using the following formula:
[0034]
[0035] In the formula, This indicates that the traffic flow is adjusted as a percentage (f). pk Calculate the delay per vehicle, where G is the total green light duration for all phases and s is the saturation flow rate.
[0036] f pk This indicates that traffic flow is adjusted according to the percentile k. f represents the state of the traffic sequence elements used as samples, z pk This indicates that based on the value of percentile k, when k is 10, z pk When z takes the value -1.28 and k is 30 pk When z takes the value -0.52 and k is 50, pk When z takes the value 0 and k is 70 pk When z takes the value 0.52 and k is 90, pk The value is 1.28.
[0037] Based on this, a method for calculating the signal delay function in the traffic mechanism model is provided. Parameters w1 and w4 are designed in the signal delay function to make the evaluation of the impact of signal delay in the performance index more accurate, thereby improving the reliability of the traffic mechanism model.
[0038] Furthermore, the parking time function is calculated using the following formula:
[0039]
[0040] In the formula, G represents the average number of stops in the i-th phase; s represents the saturation flow rate. i f represents the effective green light time for the i-th phase. i This represents the critical lane demand flow in the i-th phase.
[0041] Based on this, a method for calculating the parking time function in the traffic mechanism model is provided. Parameters w2 and w5 are designed in the parking time function to make the evaluation of the impact of parking time on performance indicators more accurate, thereby improving the reliability of the traffic mechanism model.
[0042] Furthermore, the queue length function is calculated using the following formula:
[0043]
[0044] In the formula, L represents the headway, f represents the state of the traffic sequence element as a sample, R represents the red light time for each phase, s represents the saturation flow, and F... u This indicates the lane utilization coefficient.
[0045] Based on this, a method for calculating the queue length function in the traffic mechanism model is provided. Parameters w3 and w6 are designed in the queue length function to make the evaluation of the impact of queue length on performance indicators more accurate, thereby improving the reliability of the traffic mechanism model.
[0046] Secondly, this application also provides an electronic device, comprising: at least one memory and at least one processor, wherein the at least one memory stores executable code, and the at least one processor is used to execute the executable code in the at least one memory to implement the above-described control method.
[0047] Based on the electronic device provided in the second aspect of this application, historical traffic data of the target intersection can be acquired, and historical traffic sequence elements and traffic sequence elements to be decided can be concatenated to generate a decision sequence. Using a sequence-based decision-making approach, the historical traffic conditions of the target intersection can assist in decision-making during the current traffic control cycle. This electronic device can fully utilize historical information, thereby better adapting to traffic flow fluctuations, effectively reducing vehicle congestion, and improving the user experience.
[0048] Thirdly, this application also provides a traffic control system, including: traffic lights and the aforementioned electronic equipment.
[0049] Based on the traffic control system provided in the third aspect of this application, historical traffic data of the target intersection can be acquired, and historical traffic sequence elements and traffic sequence elements to be decided can be concatenated to generate a decision sequence. A sequence-based decision-making method is then used to assist in the current traffic control cycle decision-making based on the historical traffic conditions of the target intersection. This traffic control system can fully utilize historical information, thereby better adapting to traffic flow fluctuations, effectively reducing vehicle congestion, and improving the user experience. Attached Figure Description
[0050] Figure 1 This is a schematic diagram of a traffic signal control method provided in one embodiment of this application;
[0051] Figure 2 This is a schematic diagram of the offline reinforcement learning model training process provided in the embodiments of this application;
[0052] Figure 3 This is a schematic diagram of the training sequence generation process provided in an embodiment of this application;
[0053] Figure 4 This is a schematic diagram of the traffic mechanism model calibration process provided in the embodiments of this application;
[0054] Figure 5 This is a schematic diagram of a traffic signal control method provided in another embodiment of this application;
[0055] Figure 6 A schematic diagram of an electronic device provided in an embodiment of this application;
[0056] Figure 7 A schematic diagram of a traffic control system provided in an embodiment of this application. Detailed Implementation
[0057] The present application will be described in detail below with reference to the specific embodiments shown in the accompanying drawings. However, these embodiments do not limit the present application. Any structural, methodological, or functional modifications made by those skilled in the art based on these embodiments are included within the protection scope of the present application.
[0058] like Figure 1 As shown, this application provides a traffic signal control method, the control method including:
[0059] S101: Obtain historical traffic data of the target intersection. The historical traffic data includes multiple historical traffic sequence elements. Each historical traffic sequence element includes a status and a signal control period. The status of the element represents the critical lane demand flow corresponding to different phases in the signal control period, and the status represents the environmental feedback when different signal control periods are adopted.
[0060] S102: Use radar-guided equipment to acquire the status of the target decision cycle and generate traffic sequence elements to be decided based on the status;
[0061] S103: Generate a decision sequence based on historical traffic sequence elements and traffic sequence elements to be decided;
[0062] S104: Input the decision sequence to be decided into the preset offline reinforcement learning model. The offline reinforcement learning model infers based on the reward, state and signal control period in the historical traffic sequence elements to obtain the signal control period inference value of the traffic sequence elements to be decided.
[0063] S105: Allocate signal light duration for different phases based on signal control cycle inference values.
[0064] Specifically, the control method provided in this application belongs to single-point periodic signal control. Single-point periodic signal control is a type of intersection signal control where a single intersection does not interact with other intersections. It only uses the information of the current intersection to solve for the optimal cycle and green light duration for each phase, thereby optimizing the overall performance of the intersection. Based on this, this application obtains historical traffic data of the target intersection, concatenates historical traffic sequence elements and traffic sequence elements to be decided to generate a decision sequence, and adopts a sequence decision-making approach. By using the historical traffic conditions of the target intersection, the current traffic control cycle decision is assisted. In this way, historical information can be fully utilized, thereby better adapting to traffic flow fluctuations, effectively reducing vehicle congestion, and improving the user experience.
[0065] For ease of explanation, the traffic sequence elements in this application embodiment are defined as follows: any traffic sequence element represents a complete signal cycle, including three elements: feedback, state, and signal control cycle. The critical lane demand flow corresponding to different phases in the signal control cycle is taken as the state of the historical traffic sequence element. The feedback is used to describe the environmental feedback brought about by controlling the traffic signal using the set signal control cycle in the state of the traffic sequence element.
[0066] For example, the signal control period of a traffic sequence element is represented as C. i* Assuming there are p phases in the signal control cycle, the state of the traffic sequence element can be represented as f = (f1, f2, ..., f...). p ), p is the number of phases, f p This represents the critical lane demand flow in the p-th phase. The return is represented as... (The calculation method for the reward will be explained in detail later.) Therefore, the traffic sequence element can be represented as:
[0067] i represents the sequence number of the traffic sequence element.
[0068] The generated decision sequence can be represented as:
[0069] T = (T1, T2, ..., T) d ,T d+1 In the formula: T1, T2, ... T d T represents the historical traffic sequence element. d+1 This represents the traffic sequence element to be decided, where the return and signal control cycle of the traffic sequence element to be decided are set to zero.
[0070] Let the decision sequence be T = (T1, T2, ..., T d ,T d+1 The input is fed into the offline reinforcement learning model. After the internal calculation of the offline reinforcement learning model, the reward, state and signal control period in the historical traffic sequence elements are inferred, and the signal control period inference value of the traffic sequence elements to be decided in the sequence can be obtained.
[0071] Furthermore, considering the balance between processing efficiency and accuracy of the offline reinforcement learning model, this embodiment selects fifteen historical traffic sequence elements closest to the current time. These fifteen historical traffic sequence elements are arranged in chronological order, with the first element being the furthest from the current time and the fifteenth element being the closest to the current time. It should be noted that if the data storage module contains fewer than 15 historical traffic elements, all historical data are concatenated into a sequence. By setting the number of traffic sequence elements in the sequence, both the processing efficiency and accuracy of the offline reinforcement learning model can be balanced.
[0072] As an optional implementation, the decision sequence can first be normalized by transforming the original data range of each traffic sequence element in the decision sequence to form a standardized form suitable for the sequence input. After the decision sequence is normalized, the range of the signal control period inference value output by the offline reinforcement learning model is also normalized. Therefore, the output signal control period inference value also needs to be denormalized. Denormalization also requires attention to the range of the signal control period. For example, if the range of the signal control period is 50–220 seconds, then if the denormalized signal control period inference value is less than 50 seconds, then 50 seconds is selected as the final signal control period inference value; if the denormalized signal control period inference value is greater than 220 seconds, then 220 seconds is selected as the final signal control period inference value.
[0073] As an optional implementation method, the allocation of signal light durations for different phases based on signal control cycle inference values includes: allocating the signal control cycle inference values according to the equal saturation principle or other methods to obtain the green light duration for each phase, generating a signal light duration allocation scheme, and then assigning the scheme to the traffic light controllers for operation. The equal saturation principle refers to a method in signal timing where the green light duration allocation is determined by ensuring that the ratio of traffic flow to capacity (i.e., saturation) is equal for each phase. Signal timing according to the equal saturation principle can balance traffic pressure across phases and prevent overload or idle traffic in some phases.
[0074] In this embodiment, the state of traffic sequence elements represents the demand flow of critical lanes corresponding to different phases in the traffic control cycle. As an optional implementation, in this embodiment, the lane with the highest flow in a phase is designated as the critical lane. By using the demand flow of the lane with the highest flow, the traffic state of the current traffic control cycle can be described to the greatest extent, thereby ensuring the rationality of subsequent traffic state-based traffic control cycle decisions.
[0075] For critical lane demand flow of historical traffic sequence elements, historical data can be obtained. For critical lane demand flow of traffic sequence elements to be decided, it can be calculated based on the number of vehicles passing through the critical lane during the green light and the remaining number of vehicles when the light turns red (the number of vehicles passing through and the remaining number of vehicles can be collected and stored by radar-guided equipment in the previous signal control cycle for lane demand flow calculation in the next signal control cycle; the remaining number of vehicles can be the number of vehicles in the critical lane collected by radar-guided equipment immediately after the light turns red, or the number of vehicles in the critical lane collected by radar-guided equipment some time after the light turns red). The lane demand flow can be calculated as follows:
[0076] Lane demand flow = (Number of vehicles passing through the lane during the green light time of the previous cycle + Number of vehicles on the lane segment at the beginning of the red light time of the previous cycle) / Duration of the previous signal control cycle (s) * 3600;
[0077] In this embodiment, the required traffic volume is converted into hourly traffic volume. Depending on the actual situation, it can also be converted into half-hour traffic volume, minute-level traffic volume, etc. This application does not limit it in this way.
[0078] like Figure 2 As shown, as an optional implementation, the control method provided in this application embodiment further includes training an offline reinforcement learning model. Using the trained offline reinforcement learning model, the accuracy of the signal control cycle inference value output for the traffic sequence elements to be decided can be improved, thereby reducing indicators affecting traffic efficiency such as signal delays and the number of stops, and alleviating traffic congestion.
[0079] Training an offline reinforcement learning model includes:
[0080] Step S201: Initialize the offline reinforcement learning model;
[0081] Step S202: Input the training sequence into the initialized offline reinforcement learning model;
[0082] Step S203: Calculate the loss function based on the output of the offline reinforcement learning model;
[0083] Step S204: Update the parameters of the offline reinforcement learning model with the goal of reducing the loss function.
[0084] Specifically, as an optional implementation, the mean squared error (MSE) is chosen as the loss function for the offline reinforcement learning model, and its formula is as follows:
[0085]
[0086] In the formula, N represents the total number of traffic sequence elements in the training sequence, and Y... i This represents the true value of the i-th traffic sequence element (i.e., the signal control period in the input traffic sequence element). This represents the predicted value of the i-th traffic sequence element (i.e., the signal control cycle inference value output by the offline reinforcement model).
[0087] Calculate the average loss function of K training sequences (K is a natural number greater than 1) and update the weights of the offline reinforcement learning model to reduce training error.
[0088] Step S205: Input new training sequence data for training (training data that has already been used will not be used again), and determine whether the preset number of iterations has been reached. If the preset number of iterations has not been reached, return to step S202; otherwise, proceed to step S206.
[0089] Step S206: Determine whether the loss function has decreased. If it has decreased, return to step S202; otherwise, proceed to step S207.
[0090] Step S207: Complete the training of the offline reinforcement learning model. Specifically, if the calculation error of the loss function of the training sequence is lower than the preset value, and the loss function no longer decreases with continued iteration, the network can be considered to have converged, and the training of the offline reinforcement learning model is complete.
[0091] Furthermore, for single-point intersections, the number of samples is usually small, making it difficult to provide enough raw sample data for offline reinforcement learning models to train. In this embodiment, data augmentation technology is used to generate a large number of training sequences using traffic mechanism models for offline reinforcement learning models to train.
[0092] like Figure 3 As shown, as an optional implementation, generating a large number of training sequences using a traffic mechanism model includes:
[0093] Step S301: Define the initial element state of the initial traffic sequence element in the training sequence according to the size of the demand flow of the key lane. Based on the initial element state, obtain the reward and signal control cycle of the corresponding element through the preset traffic mechanism model.
[0094] Step S302: Perform a Hadema multiplication on the initial element state and the change ratio vector to obtain the derived element state of multiple derived traffic sequence elements in the training sequence. Based on the derived element state, obtain the reward and signal control cycle of the corresponding element through the traffic mechanism model.
[0095] Step S303: Concatenate the initial traffic sequence elements with the derived traffic sequence elements to obtain the training sequence.
[0096] For example, in step S301, defining the initial element state of the initial traffic sequence element in the training sequence according to the magnitude of the demand flow of the critical lane can be done in the following way:
[0097] The initial traffic sequence elements are configured to satisfy the following conditions: 25% of the critical lanes have low demand flow, 50% of the critical lanes have medium demand flow, and 25% of the critical lanes have high demand flow.
[0098] Specifically, based on the foregoing, the state of a traffic sequence element can be represented as follows:
[0099] f = (f1, f2, ..., f p In the formula, p is the number of phases, and f p This represents the critical lane demand flow in the p-th phase.
[0100] Assuming the initial traffic sequence element's state is denoted as f1, then 25% of the critical lane demand flow is low, 50% is medium, and 25% is high. This method defines the initial element state according to the magnitude of the critical lane demand flow, making the defined initial element state more closely reflect the actual situation. For example, in this embodiment, the low flow range is defined as 0–75 vehicles / h, the medium flow range as 76–99 vehicles / h, and the high flow range as 100–300 vehicles / h. The magnitude of the critical lane demand flow range can also be set or adjusted via the user's client.
[0101] At this point, the initial element state can be represented as:
[0102] f1=(k1,k2,…,k p );
[0103] The defined initial element states are input into the traffic mechanism model, which is a multi-objective optimization model for finding the optimal cycle. Through the traffic mechanism model, the corresponding signal control cycle can be output. and returns (The specific methods for outputting the signal control cycle and reward from the traffic mechanism model will be explained in detail later.) The state f1 and signal control cycle... and returns By rearranging the traffic flow, we can obtain the initial traffic sequence elements, which can be represented as follows:
[0104] Based on the initial element state f1=(k1,k2,…,k p ), randomly changing a certain proportion r1=(r1,r2,…,r p As an optional implementation, the proportional coefficient r of the randomly varying traffic flow in the critical lane is sampled from a normal distribution, i.e., r ~ N(μ,σ). 2 In this embodiment of the application, the coefficient of the normal distribution is μ = 1.0, σ 2 =1.0.
[0105] In step S302, performing a Hadma multiplication between f1 and the change ratio vector r1 yields the state f2 of the derived traffic sequence element, i.e.:
[0106] f2=f1*r1=(k1r1,k2r2,…,k p r p )
[0107] By repeating the steps above to generate derived element states, multiple different derived element states f3, f4, ... f can be obtained. dWait, call the traffic mechanism model to obtain states f3, f4, ... f d The corresponding optimal signal control period And the optimal objective function value corresponding to the optimal period. By rearranging the state, signal control period, and reward corresponding to each traffic sequence element, any traffic sequence element in the training sequence can be represented as follows:
[0108] In step S303, the initial traffic sequence elements and the derived traffic sequence elements are concatenated to obtain the training sequence. The generated training sequence can be represented as follows: T = (T1, T2, ... T d ).
[0109] Based on the above description, this application employs data augmentation technology to generate a large number of training sequences using a traffic mechanism model for training the offline reinforcement learning model. This effectively addresses the problem of insufficient raw sample data for training the offline reinforcement learning model due to the limited number of samples at single-point intersections. The generated training sequences are then input into the offline reinforcement learning model. Following the training method provided in this application, the offline reinforcement learning model can be trained to ensure that, when making signal control cycle decisions on subsequent decision sequences, the trained offline reinforcement learning model can output the optimal signal control cycle inference value, thereby improving intersection traffic efficiency and reducing congestion.
[0110] As an optional implementation, a pre-defined traffic mechanism model is used to obtain the rewards and signal control cycles corresponding to each traffic sequence element, including:
[0111] Based on signal delay, stopping time, and lane queue length, a comprehensive performance index is formed. This performance index is used to characterize the environmental feedback reward when different signal control cycles are adopted under different traffic sequence element states.
[0112] The state of traffic sequence elements is input into the traffic mechanism model, and the signal control cycle output by the traffic mechanism model is obtained with the goal of minimizing the performance index.
[0113] Optionally, the traffic mechanism model adopts a single-step decision-making, single-point periodic signal control optimization model. The input of the traffic mechanism model is the element state f = (f1, f2, ..., f...). p ), p is the number of phases, f p This represents the critical lane demand flow in the p-th phase. Due to the performance index PI of the traffic mechanism model... i It is a composite of intersection delay d, parking time s, and queueing time q. Therefore, in performance evaluation, the smaller the index, the better. The performance index PI is used as the benchmark. iMinimum is the optimization objective, within the feasible space of the signal control cycle [C] min C max By traversing within [the specified area], the optimal signal control period C can be obtained. * And the optimal performance index PI corresponding to the optimal period * .
[0114] In this embodiment of the application, the performance index is weighted and inverted to obtain the reward of the traffic sequence element (i.e. The smaller the performance index PI, the lower the negative impact of factors such as delays, parking, and queuing on traffic efficiency. Therefore, the minimized performance index, after weighted inversion, can be used to characterize the optimal environmental feedback. Offline reinforcement learning models use historical action sequences (i.e., signal control cycles) and reward sequences (i.e.,...) to represent the optimal environmental feedback. By performing inference, the optimal signal control cycle inference value can be output, thereby improving the traffic efficiency of the target single-point intersection and reducing traffic congestion.
[0115] like Figure 4 As shown, as an optional implementation, the control method provided in this application embodiment further includes: first calibrating the parameters of the traffic mechanism model, and then generating training data based on the calibrated traffic mechanism model.
[0116] Parameter calibration of the traffic mechanism model includes:
[0117] Step S401: Obtain the real sample dataset, which includes multiple real sample data. Each real sample data includes the mapping relationship between the state of traffic sequence elements and the optimal signal control cycle.
[0118] Step S402: Set the range of values for the parameter to be calibrated, and use a genetic algorithm to randomly generate a population of values for the parameter to be calibrated within the range of the calibration parameters;
[0119] Step S403: Input the real sample data into the traffic mechanism model to obtain the performance index under different values of the parameters to be calibrated in the population;
[0120] Step S404: Obtain the optimal calibration parameters by aiming to minimize the sum of performance indicators corresponding to each real sample data.
[0121] Real sample data can be used to select strategies with better on-site control effects from data captured by radar-guided equipment. This can be achieved by using [traffic sequence element state f – optimal signal control cycle C]. * The mapping relationship combination of ].
[0122] The traffic mechanism model adjusts its parameters and signal control cycle using the following formula:
[0123]
[0124] In the formula, w1, w2, w3, w4, w5, and w6 represent the parameters of the traffic mechanism model, d(C,f,w4) represents the signal delay function corresponding to the signal control period, s(C,f,w5) represents the parking time function corresponding to the signal control period, q(C,f,w6) represents the queue length function corresponding to the signal control period, C represents the signal control period of the traffic sequence element as a sample, and f represents the state of the traffic sequence element as a sample.
[0125] As an optional implementation, the signal delay function is calculated using the following formula:
[0126]
[0127] In the formula, This indicates that the traffic flow is adjusted as a percentage (f). pk Calculate the delay per vehicle, where G is the total green light duration for all phases, and s is the saturation flow rate. For example, D p10 This indicates that the traffic flow f is adjusted according to the percentile k = 10 (p10 represents 10%). p10 Calculate the delay for each vehicle.
[0128] f pk This indicates that traffic flow is adjusted according to the percentile k. f represents the state of the traffic sequence elements used as samples, z pk This indicates that based on the value of percentile k, when k is 10 (pk is p10, representing 10%), z pk When z takes the value -1.28 and k is 30 (pk is p30 at this time, representing 10%), pk When z takes a value of -0.52 and k is 5 (pk is p5, representing 5%), pk When z is 0 and k is 70 (pk is p70, representing 70%), pk When z takes a value of 0.52 and k is 90 (pk is p90 at this time, representing 90%), pk The value is 1.28.
[0129] Based on the calculation method of the signal delay function in the traffic mechanism model provided in this application embodiment, parameters w1 and w4 are designed in the signal delay function, which makes the evaluation of the impact of signal delay in the performance index more accurate, thereby improving the reliability of the traffic mechanism model.
[0130] As an optional implementation, the parking time function is calculated using the following formula:
[0131]
[0132] In the formula, G represents the average number of stops in the i-th phase; s represents the saturation flow rate. i f represents the effective green light time for the i-th phase. i This represents the critical lane demand flow in the i-th phase.
[0133] Based on the calculation method of the parking time function in the traffic mechanism model provided in this application embodiment, the impact of parking time on performance indicators can be more accurately evaluated by using the influencing parameters w2 and w5 in the parking time function, thereby improving the reliability of the traffic mechanism model.
[0134] As an optional implementation, the queue length function is calculated using the following formula:
[0135]
[0136] In the formula, L represents the headway, f represents the state of the traffic sequence element as a sample, R represents the red light time for each phase, s represents the saturation flow, and F... u This indicates the lane utilization coefficient.
[0137] Based on the calculation method of the queue length function in the traffic mechanism model provided in this application embodiment, the queue length function can influence parameters w3 and w6, making the evaluation of the influence of queue length on performance indicators more accurate, thereby improving the reliability of the traffic mechanism model.
[0138] For example, w1, w2, w3, w4, w5, and w6 all represent parameters of the traffic mechanism model that simultaneously satisfy the following constraints.
[0139]
[0140] Under the constraints described above, find a set of parameters {w1,w2,w3,...,w5,w6} that minimizes the error between the signal control period output by the traffic mechanism model and the label period (which can be the signal control period in the sample data). The solution for the parameters that minimize the signal control period error can be as follows: Under the constraints of parameters w1 to w6, use a genetic algorithm to randomly generate a population of parameters to be calibrated within the calibration parameter range. Input each real sample data from the real sample dataset into the traffic mechanism model, obtain the performance index under different values of the parameters to be calibrated in the population, and calculate the sum of the total objective function values of the real sample dataset ∑. i PI i (i is the data sample number) The smallest calibration parameter As the optimal calibration parameter.
[0141] Through the above parameter calibration of the traffic mechanism model, after determining the optimal calibration parameters w1 to w6, substituting the optimal calibration parameters w1 to w6 into the traffic mechanism model yields the traffic mechanism model with complete parameter calibration. Inputting the state f of the traffic sequence elements into the traffic mechanism model with complete parameter calibration, within the feasible space range of the signal control period [C]... min C max By traversing within [the specified area], the optimal signal control period C can be obtained. * And the optimal objective function value PI corresponding to the optimal period * Parameter calibration of the traffic mechanism model can ensure the reliability of the signal control cycle and reward generated by data augmentation, thus providing a foundation for the training of subsequent offline reinforcement learning models.
[0142] To further illustrate the traffic signal control method provided in the embodiments of this application, such as Figure 5 As shown, a complete flowchart is provided, including:
[0143] Step S501: Calibrate the parameters of the traffic mechanism model based on a small amount of sample data;
[0144] Step S502: Introduce data augmentation techniques and generate training sequences using the calibrated traffic mechanism model;
[0145] Step S503: Input the training sequence into the initialized offline reinforcement learning model to complete the training of the offline reinforcement learning model;
[0146] Step S504: Generate a decision sequence based on historical traffic sequence elements and decision sequence elements;
[0147] Step S505: Input the decision sequence to be decided into the offline reinforcement learning model that has been trained. The offline reinforcement learning model infers based on the reward, state and signal control period in the historical traffic sequence elements to obtain the signal control period inference value of the traffic sequence elements to be decided.
[0148] Step S506: Based on the signal control cycle inference value, allocate the signal light duration for different phases.
[0149] In summary, the traffic signal control method provided in this application incorporates data augmentation technology, using small sample data to calibrate the traffic mechanism model. The calibrated model can generate a large number of training sequences for offline reinforcement learning model training, and also provide reward information for the training sequences. Furthermore, the control method is built upon offline reinforcement learning technology, employing a sequential decision-making approach to calculate the signal control cycle. This sequential decision-making method fully utilizes historical information to better adapt to traffic flow fluctuations. Compared to the single-step decision-making method of the traffic mechanism model, it reduces decision-making errors and avoids signal control cycles that are too large or too small, thereby reducing delays, the number of stops, and other indicators, improving intersection efficiency, and alleviating congestion.
[0150] like Figure 6 As shown, based on the same inventive concept, this application also provides an electronic device 100, including: at least one memory 11 and at least one processor 12, wherein the at least one memory 11 stores executable code, and the at least one processor 12 is used to execute the executable code in the at least one memory 11 to implement the above-mentioned traffic signal control method.
[0151] At least one of the aforementioned memories can be used to store a computer program, which may include instructions and data to implement the steps of any of the methods described above. Memory 11 may be random access memory, read-only memory, non-volatile, programmable ROM, erasable PROM, electrically erasable, flash memory, optical memory, and registers, etc. Processor 12 may be a general-purpose processor, which can perform specific steps and / or operations by reading and executing the computer program stored in memory 11. The general-purpose processor may use the memory 11 during the execution of the steps and / or operations. The general-purpose processor may be a central processing unit, ASIC, and FPGA, etc. The electronic device 100 may also include a communication interface, which may include input / output interfaces, physical interfaces, and logical interfaces for interconnecting devices within the network device. In implementation, each step of the above method can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The methods disclosed in the embodiments of this application can be directly implemented by a hardware processor or by a combination of hardware and software modules in the processor.
[0152] According to the electronic device 100 provided in the embodiments of this application, data augmentation technology is introduced to calibrate the traffic mechanism model using small sample data. The calibrated traffic mechanism model can generate a large number of training sequences for offline reinforcement learning model training, and can also provide reward information for the training sequences. In addition, the control method is constructed based on offline reinforcement learning technology and uses a sequential decision-making approach to complete the signal control cycle calculation, rather than a single-step decision-making method based on the traffic mechanism model. Through this sequential decision-making approach, historical information can be fully utilized to better adapt to traffic flow fluctuations, reduce decision-making errors, and avoid signal control cycles that are too large or too small. This reduces indicators such as delays and number of stops, improves the traffic efficiency of intersections, and reduces congestion.
[0153] like Figure 7 As shown in the figure, this application embodiment also provides a traffic control system 200, including: a traffic light 21 and the above-mentioned electronic device 100.
[0154] According to the traffic control system provided in this application embodiment, data augmentation technology is introduced to calibrate the traffic mechanism model using small sample data. The calibrated traffic mechanism model can generate a large number of training sequences for offline reinforcement learning model training, and can also provide reward information for the training sequences. Furthermore, the control method is built based on offline reinforcement learning technology and uses a sequential decision-making approach to calculate the signal control cycle, rather than a single-step decision-making method based on the traffic mechanism model. This sequential decision-making approach can fully utilize historical information to better adapt to traffic flow fluctuations, reduce decision-making errors, and avoid signal control cycles that are too large or too small. This reduces indicators such as delays and number of stops, improves intersection efficiency, and reduces congestion.
[0155] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a solid-state drive (SSD), etc.
[0156] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0157] It is understood that the term "exemplary" as used herein means "as an example, illustration, or description." Any embodiment described as "exemplary" is not necessarily preferred or superior to other embodiments and / or does not exclude features in combination with other embodiments. It should be understood that certain features of this application described in the context of a single embodiment for clarity may also be provided in combination in a single embodiment. Conversely, various features of this application described in the context of a single embodiment for clarity may also be provided individually or in any suitable combination or as part of any other described embodiment of this application.
[0158] The above-disclosed embodiments are merely preferred embodiments of this application, but are not intended to limit the scope of this application. Those skilled in the art will understand that any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and scope of this application and the appended claims are equivalent substitutions and still fall within the scope of the invention.
Claims
1. A method for controlling traffic signals, characterized in that, The control method includes: Historical traffic data of the target intersection is obtained. The historical traffic data includes multiple historical traffic sequence elements. Each of the historical traffic sequence elements includes a report, a status, and a signal control cycle. The status represents the critical lane demand flow corresponding to different phases in the signal control cycle, and the report represents the traffic environment feedback when different signal control cycles are adopted. The state of the target decision cycle is acquired using radar-guided equipment, and traffic sequence elements to be decided are generated based on the state of the target decision cycle. A decision-making sequence is generated based on the historical traffic sequence elements and the traffic sequence elements to be decided. The decision sequence is input into a preset offline reinforcement learning model. The offline reinforcement learning model infers based on the reward, state and signal control period in the historical traffic sequence elements to obtain the signal control period inference value of the traffic sequence element to be decided. The signal light duration is allocated for different phases based on the signal control cycle inference value.
2. The traffic signal control method according to claim 1, characterized in that, The control method further includes training the offline reinforcement learning model, including: The initial element state of the initial traffic sequence element in the training sequence is defined according to the magnitude of the demand flow of the critical lane. Based on the initial element state, the reward and signal control period of the corresponding element are obtained through a preset traffic mechanism model. The initial element state is multiplied by the change ratio vector using a Hadma multiplication to obtain the derived element state of multiple derived traffic sequence elements in the training sequence. Based on the derived element state, the reward and signal control period of the corresponding element are obtained through the traffic mechanism model. The training sequence is input into the initialized offline reinforcement learning model. The loss function is calculated based on the output of the offline reinforcement learning model. The parameters of the offline reinforcement learning model are updated with the goal of reducing the loss function.
3. The traffic signal control method according to claim 2, characterized in that, Through a pre-defined traffic mechanism model, the rewards and control cycles corresponding to each traffic sequence element are obtained, including: Based on signal delay, stopping time, and lane queue length, a comprehensive performance index is formed, wherein the performance index is used to characterize the environmental feedback reward when different signal control cycles are adopted under different traffic sequence element states. The state of traffic sequence elements is input into the traffic mechanism model, and the signal control cycle output by the traffic mechanism model is obtained with the goal of minimizing the performance index.
4. The traffic signal control method according to claim 2, characterized in that, Through a pre-defined traffic mechanism model, the rewards and control cycles corresponding to each traffic sequence element are obtained, including: The state of the traffic sequence element is input into the traffic mechanism model, and the model is traversed within the feasible range of the signal control period to obtain the performance index corresponding to different signal control periods. The signal control period that minimizes the performance index is selected as the optimal signal control period of the traffic sequence element. The performance index corresponding to the optimal signal control period is weighted and inverted to obtain the reward of the traffic sequence element.
5. The traffic signal control method according to claim 4, characterized in that, The control method further includes parameter calibration of the traffic mechanism model, including: Obtain a real sample dataset, which includes multiple real sample data, and any real sample data includes the mapping relationship between the state of traffic sequence elements and the optimal signal control cycle; Set the range of values for the parameters to be calibrated, and use a genetic algorithm to randomly generate a population of values for the parameters within the range of the calibration parameters; The real sample data are input into the traffic mechanism model to obtain the performance index of the population under different values of the parameters to be calibrated. The goal is to obtain the optimal calibration parameters by minimizing the sum of performance indicators corresponding to each real sample data.
6. The traffic signal control method according to any one of claims 3 to 5, characterized in that, The traffic mechanism model adjusts its parameters and signal control period using the following formula: In the formula, All of these represent the parameters of the traffic mechanism model. This represents the signal delay function corresponding to the signal control period. This represents the parking time function corresponding to the signal control cycle. The signal control cycle corresponds to the queue length function. C This represents the signal control period of the traffic sequence elements used as samples. f This represents the state of the traffic sequence element that serves as the sample.
7. The traffic signal control method according to claim 6, characterized in that, The signal delay function is calculated using the following formula: ; In the formula, This indicates that traffic flow is adjusted using a percentage. Calculate the delay for each vehicle. The total green light duration for each phase. For saturated flow rate, Indicated by percentile k Adjust traffic flow, , The state of the traffic sequence elements used as samples. Indicated by percentile k of Pick Value, when k It is 10 o'clock The value is -1.
28. k 30 hours The value is -0.
52. k When it is 50 The value is 0. k 70 hours The value is 0.
52. k 90 hours The value is 1.
28.
8. The traffic signal control method according to claim 6, characterized in that, The parking time function is calculated using the following formula: ; In the formula, , indicating the first i Average number of stops per phase; Indicates saturation flow rate. Indicates the first i The effective green time of the phase Indicates the first i The critical lane demand flow of the phase.
9. The traffic signal control method according to claim 6, characterized in that, The queue length function is calculated using the following formula: ; In the formula, L Indicates the distance between the front of the vehicles. f This represents the state of the traffic sequence elements used as samples. R The red light time for each phase is represented by s, where s represents the saturation flow rate. F u This indicates the lane utilization coefficient.
10. An electronic device, characterized in that, include: At least one memory and at least one processor, The at least one memory stores executable code, and the at least one processor executes the executable code in the at least one memory to implement the method as described in any one of claims 1-9.
11. A traffic control system, characterized in that, include: Traffic lights and the electronic device as described in claim 10.
Citation Information
Patent Citations
Multi-intersection traffic light control method and system based on reinforcement learning, and storage medium
CN113223305A
Passenger delay minimization signal control method based on deep reinforcement learning
CN117116064A