Traffic signal control method, electronic equipment and system
By obtaining historical traffic data and offline reinforcement learning models to optimize signal light duration allocation, the problem of traffic signal control being unable to respond to traffic flow fluctuations is solved, and better congestion management and user experience is achieved.
Patent Information
- Application Number
- CN202510589375.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing traffic signal control methods cannot effectively respond to the fluctuations of traffic flow, resulting in a large or small signal control cycle, which can easily cause congestion and affect user experience.
By obtaining the historical traffic data of the target intersection, using the Leishi device to obtain the current status, generating a sequence to be decided, and inputting it into the preset offline reinforcement learning model for inference, combining the traffic mechanism model and data enhancement technology, optimizing the signal light duration allocation.
Effectively reduce vehicle congestion, improve user experience, improve cross-section traffic efficiency, adapt to traffic fluctuations, and reduce delays and parking times.
Smart Images

Figure CN120452192A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of signal control, and in particular to a traffic signal control method, electronic equipment, and system. Background Art
[0002] With the increasing number of motor vehicles in cities, the need for single-point intersection signal control is increasing, driven by the need to ensure ever-increasing traffic safety and efficiency. Single-step decisions are often based on recent traffic flow (typically cycle-by-cycle or 5- to 10-minute intervals), and are unable to respond to traffic flow fluctuations. This can easily lead to incorrect decisions, often resulting in excessively long or short signal control cycles. This increases delays, parking times, and other indicators, leading to congestion and a poor user experience. Summary of the Invention
[0003] The purpose of this application is to provide a traffic signal control method to solve the problem that traffic signal control cannot respond to fluctuations in traffic flow, easily causing congestion and poor user experience.
[0004] Based on the above objectives, in a first aspect, the present application provides a traffic signal control method, the control method comprising:
[0005] Obtain historical traffic data for the target intersection. Historical traffic data includes multiple historical traffic sequence elements. Each historical traffic sequence element includes a report, a state, and a signal control cycle. The state represents the demand flow of key lanes corresponding to different phases in the signal control cycle, and the report represents the environmental feedback when different signal control cycles are adopted.
[0006] Use radar equipment to obtain the status of the target decision cycle and generate traffic sequence elements to be decided based on the status;
[0007] Generate a sequence to be decided based on historical traffic sequence elements and traffic sequence elements to be decided;
[0008] The decision sequence is input into a preset offline reinforcement learning model. The offline reinforcement learning model performs reasoning based on the rewards, states, and signal control cycles of the historical traffic sequence elements to obtain the signal control cycle inference value of the traffic sequence element to be decided.
[0009] The duration of traffic lights in different phases is allocated based on the signal control cycle inference value.
[0010] Based on this, this application obtains historical traffic data for a target intersection, concatenates historical traffic sequence elements with those for a pending decision sequence, and uses a sequence-based decision-making approach to assist in making decisions about the current traffic control cycle based on the historical traffic conditions at the target intersection. This approach fully utilizes historical information, enabling better adaptation to traffic flow fluctuations, effectively reducing congestion, and improving the user experience.
[0011] Furthermore, the control method also includes training an offline reinforcement learning model, including:
[0012] The initial element state of the initial traffic sequence element in the training sequence is defined according to the size of the demand flow in the key lane. Through the preset traffic mechanism model, the reward and signal control cycle of the corresponding element are obtained according to the initial element state;
[0013] Perform Hadamard multiplication on the initial element state and the change ratio vector to obtain the derived element states of multiple derived traffic sequence elements in the training sequence. Based on the derived element states, the reward and signal control period of the corresponding element are obtained through the traffic mechanism model.
[0014] The training sequence is input into the initialized offline reinforcement learning model, a loss function is calculated based on the output result of the offline reinforcement learning model, and the parameters of the offline reinforcement learning model are updated with the goal of reducing the loss function.
[0015] Based on this, the initial element state is defined according to the size of the demand flow in the key lane, so that the element state in the training sequence is more in line with the actual environment. A large number of generated training sequences are input into the offline reinforcement learning model. According to the training method of the offline reinforcement learning model provided in this application, the offline reinforcement learning model can be trained to ensure that when the signal control cycle decision is made for the decision sequence in the future, the optimal signal control cycle inference value can be output based on the offline reinforcement learning model that has completed the training, thereby improving the traffic efficiency of the intersection and reducing congestion.
[0016] Furthermore, through the preset traffic mechanism model, the rewards and signal control cycles corresponding to each traffic sequence element are obtained, including:
[0017] Based on signal delay, parking time and lane queue length, a performance index is formed. The performance index is used to characterize the return of environmental feedback when different signal control cycles are adopted under different traffic sequence elements.
[0018] The status of traffic sequence elements is input into the traffic mechanism model, and the signal control cycle output by the traffic mechanism model is obtained with the goal of minimizing the performance index.
[0019] Based on this, data augmentation technology is introduced to use the traffic mechanism model to generate a large number of training sequences for offline reinforcement learning model training, which effectively solves the problem that the number of single-point intersection samples is small and it is difficult to provide sufficient original sample data for offline reinforcement learning model training.
[0020] Furthermore, through the preset traffic mechanism model, the rewards and signal control cycles corresponding to each traffic sequence element are obtained, including:
[0021] The state of the traffic sequence element is input into the traffic mechanism model, and the model is traversed within the feasible range of the signal control cycle to obtain the performance indicators corresponding to different signal control cycles. The signal control cycle that minimizes the performance indicator is selected as the optimal signal control cycle of the traffic sequence element. The performance indicator corresponding to the optimal signal control cycle is weighted and inverted to obtain the return of the traffic sequence element.
[0022] Based on this, this application performs weighted inversion of the performance index to obtain the return of the traffic sequence element (i.e. ), the smaller the performance index PI is, the lower the negative impact of factors such as delays, parking, and queuing on traffic efficiency is. Therefore, the minimized performance index can be used to characterize the optimal environmental feedback after weighted inversion. The offline reinforcement learning model is based on the historical sequence of action (i.e., signal control cycle)-reward (i.e., ) is used for reasoning, and the optimal signal control cycle reasoning value can be output, which can improve the traffic efficiency of the target single-point intersection and reduce traffic congestion.
[0023] Furthermore, the control method also includes calibrating parameters of the traffic mechanism model, including:
[0024] Acquire a real sample data set, where the real sample data set includes multiple real sample data, and any real sample data includes a mapping relationship between a state of a traffic sequence element and an optimal signal control period;
[0025] Set the value range of the parameter to be calibrated, and use the genetic algorithm to randomly generate a population of parameter values to be calibrated within the calibration parameter range;
[0026] Input each real sample data into the traffic mechanism model to obtain the performance indicators under different values of the parameters to be calibrated in the population;
[0027] The goal is to minimize the sum of the performance indicators corresponding to each real sample data and obtain the optimal calibration parameters.
[0028] Based on this, parameter calibration of the traffic mechanism model can ensure the reliability of the signal control cycle and returns generated by data enhancement, thereby providing a basis for the subsequent training of the offline reinforcement learning model.
[0029] Furthermore, the traffic mechanism model adjusts parameters and signal control cycles through the following formula:
[0030]
[0031] Where w1, w2, w3, w4, w5, and w6 are all parameters of the traffic mechanism model, d(C, f, w4) represents the signal delay function corresponding to the signal control period, s(C, f, w5) represents the parking time function corresponding to the signal control period, q(C, f, w6) represents the queue length function corresponding to the signal control period, C represents the signal control period of the traffic sequence element as a sample, and f represents the state of the traffic sequence element as a sample.
[0032] Based on this, the performance index of the traffic mechanism model is formed by the comprehensive intersection delay d, parking time s, and queue q. The input of the traffic mechanism model is the element state f=(f1,f2,…,f p ), p is the number of phases, f p It represents the critical lane demand flow in the pth phase, with the performance index PI minimized as the optimization goal, in the feasible space range of the signal control period [C min ,C max ], we can obtain the optimal signal control period C* and the optimal performance index PI corresponding to the optimal period. * In this way, a large number of training sequences can be generated for offline reinforcement learning model training, which effectively solves the problem that the number of single-point intersection samples is small and it is difficult to provide sufficient original sample data for offline reinforcement learning model training.
[0033] Furthermore, the signal delay function is calculated by the following formula:
[0034]
[0035] Where, Indicates the percentage adjustment of traffic flow f pk Calculate the delay of each vehicle, G is the total green light duration of each phase, s is the saturation flow rate,
[0036] f pk Indicates adjusting the traffic flow according to percentile k, f is the state of the traffic sequence element as a sample, z pk Indicates that according to the value of percentile k, when k is 10, z pk When the value is -1.28 and k is 30, z pk When the value is -0.52 and k is 50, z pk When the value is 0 and k is 70, z pk When the value is 0.52 and k is 90, z pk The value is 1.28.
[0037] Based on this, a calculation method for the signal delay function in the traffic mechanism model is provided. Parameters w1 and w4 are designed in the signal delay function to make the evaluation of the impact of signal delay on performance indicators more accurate, thereby improving the reliability of the traffic mechanism model.
[0038] Furthermore, the parking time function is calculated by the following formula:
[0039]
[0040] Where, represents the average number of stops in phase i; s represents the saturation flow rate, G i represents the effective green light time of phase i, f i represents the demand flow of the key lane in phase i.
[0041] Based on this, a calculation method for the parking time function in the traffic mechanism model is provided. Parameters w2 and w5 are designed in the parking time function to make the evaluation of the impact of parking time on performance indicators more accurate, thereby improving the reliability of the traffic mechanism model.
[0042] Furthermore, the queue length function is calculated by the following formula:
[0043]
[0044] Where L represents the headway, f represents the state of the traffic sequence element as a sample, R represents the red light time of each phase, s represents the saturated flow, and F u represents the lane utilization coefficient.
[0045] Based on this, a calculation method for the queue length function in the traffic mechanism model is provided. Parameters w3 and w6 are designed in the queue length function to make the evaluation of the impact of queue length on performance indicators more accurate, thereby improving the reliability of the traffic mechanism model.
[0046] In a second aspect, the present application also provides an electronic device comprising: at least one memory and at least one processor, wherein the at least one memory stores executable code, and the at least one processor is used to execute the executable code in the at least one memory to implement the above-mentioned control method.
[0047] The electronic device provided in the second aspect of this application can obtain historical traffic data for a target intersection, concatenate historical traffic sequence elements with elements of a traffic sequence to be decided, and generate a sequence to be decided. Using a sequence decision-making approach, the historical traffic conditions at the target intersection can assist in making decisions during the current traffic control cycle. This electronic device can fully utilize historical information to better adapt to traffic flow fluctuations, effectively reduce vehicle congestion, and enhance the user experience.
[0048] In a third aspect, the present application also provides a traffic control system, comprising: a signal light and the above-mentioned electronic device.
[0049] The traffic control system provided in the third aspect of this application can obtain historical traffic data for a target intersection, concatenate historical traffic sequence elements with elements of the traffic sequence to be decided, and generate a sequence to be decided. Using a sequence decision-making approach, the historical traffic conditions at the target intersection can assist in making decisions about the current traffic control cycle. This traffic control system can fully utilize historical information to better adapt to traffic flow fluctuations, effectively reduce vehicle congestion, and enhance the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 A schematic diagram of a traffic signal control method provided by an embodiment of the present application;
[0051] Figure 2 A schematic diagram of the offline reinforcement learning model training process provided in an embodiment of the present application;
[0052] Figure 3 A schematic diagram of the training sequence generation process provided in an embodiment of the present application;
[0053] Figure 4 A schematic diagram of the traffic mechanism model calibration process provided in the embodiment of this application;
[0054] Figure 5 A schematic diagram of a traffic signal control method provided by another embodiment of the present application;
[0055] Figure 6 A schematic diagram of an electronic device provided in an embodiment of the present application;
[0056] Figure 7 Schematic diagram of the traffic control system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0057] The present application will be described in detail below in conjunction with the specific embodiments shown in the accompanying drawings, but these embodiments do not limit the present application. Structural, methodological, or functional changes made by ordinary technicians in this field based on these embodiments are included in the scope of protection of the present application.
[0058] like Figure 1 As shown, the present application provides a traffic signal control method, the control method comprising:
[0059] S101: Acquire historical traffic data for a target intersection. The historical traffic data includes multiple historical traffic sequence elements. Each historical traffic sequence element includes a state and a signal control period. The state of the element represents the required flow rate of a key lane corresponding to different phases in the signal control period, and the feedback represents the environmental feedback when different signal control periods are adopted.
[0060] S102: using a radar device to obtain the state of the target decision cycle, and generating a traffic sequence element to be decided based on the state;
[0061] S103: Generate a to-be-decided sequence based on historical traffic sequence elements and to-be-decided traffic sequence elements;
[0062] S104: Inputting the sequence to be decided into a preset offline reinforcement learning model. The offline reinforcement learning model performs reasoning based on the rewards, states, and signal control cycles in the historical traffic sequence elements to obtain the signal control cycle inference value of the traffic sequence element to be decided;
[0063] S105: Allocate the duration of traffic lights of different phases based on the signal control cycle inference value.
[0064] Specifically, the control method provided by the present application belongs to single-point periodic signal control, which is a type of intersection signal control method. A single intersection does not interact with other intersections, and only uses the information of this intersection to solve the optimal cycle and the duration of the green light for each phase, so as to optimize the overall indicators of the intersection. Based on this, the present application obtains the historical traffic data of the target intersection, splices the historical traffic sequence elements and the traffic sequence elements to be decided to generate a sequence to be decided, and adopts a sequence decision-making method to assist the current traffic control cycle decision through the historical traffic conditions of the target intersection. In this way, historical information can be fully utilized to better adapt to fluctuations in traffic flow, effectively reduce vehicle congestion, and enhance the user experience.
[0065] For ease of explanation, the traffic sequence elements are defined as follows in the embodiment of the present application: any traffic sequence element represents a complete signal cycle, including three elements: feedback, status, and signal control cycle. Among them, the demand flow of key lanes corresponding to different phases in the signal control cycle is used as the status of the historical traffic sequence element; the feedback is used to describe the environmental feedback brought about when the set signal control cycle is used to control the traffic signal under the state of the traffic sequence element.
[0066] For example, the signal control period of a traffic sequence element is expressed as C i*. Assuming that there are p phases in the signal control cycle, the state of the traffic sequence element can be expressed as f = (f1, f2, ..., f p ), p is the number of phases, f p represents the critical lane demand flow in phase p. The reward is expressed as (The calculation method of the reward will be explained in detail later), then the traffic sequence element can be expressed as:
[0067] i represents the sequence number of the traffic sequence element.
[0068] The generated decision sequence can be expressed as:
[0069] T=(T1,T2,…,T d ,T d+1 ), where: T1, T2, ... T d Represents the historical traffic sequence element, T d+1 Indicates the traffic sequence element to be decided, wherein the return and signal control cycle of the traffic sequence element to be decided are set to zero.
[0070] The decision sequence T=(T1, T2,…, T d ,T d+1 ) is input into the offline reinforcement learning model. After the internal calculation of the offline reinforcement learning model, the rewards, states, and signal control cycles in the historical traffic sequence elements are inferred to obtain the signal control cycle inference value of the traffic sequence element to be decided in the sequence.
[0071] Furthermore, considering the balance between the processing efficiency and accuracy of the offline reinforcement learning model, in the embodiment of the present application, fifteen historical traffic sequence elements closest to the current time are selected. These fifteen historical traffic sequence elements are arranged in chronological order, i.e., the first element is the element farthest from the current time, and the fifteenth element is the element closest to the current time. It should be noted that if the data storage module contains less than 15 historical traffic elements, all historical data is selected and spliced into a sequence. By setting the number of traffic sequence elements in the sequence, the processing efficiency and accuracy of the offline reinforcement learning model can be taken into account.
[0072] As an optional implementation, the decision sequence can be normalized first. This converts the raw data range of each traffic sequence element in the decision sequence into a standardized form suitable for sequence input. After the decision sequence is normalized, the range of the signal control period inference value output by the offline reinforcement learning model is also normalized. Therefore, the output signal control period inference value also needs to be denormalized. Denormalization also requires attention to the value range of the signal control period. For example, if the signal control period ranges from 50 to 220 seconds, if the denormalized signal control period inference value is less than 50 seconds, 50 seconds is selected as the final signal control period inference value. If the denormalized signal control period inference value is greater than 220 seconds, 220 seconds is selected as the final signal control period inference value.
[0073] As an optional implementation, allocating signal durations for different phases based on inferred signal control cycle values includes allocating the inferred signal control cycle values according to the equal saturation principle or other methods to obtain green light durations for each phase, generating a signal duration allocation plan, and applying the plan to traffic light signal operation. The equal saturation principle, in signal timing, determines green light duration allocation by ensuring that the ratio of traffic flow to capacity (i.e., saturation) is equal for each phase. Implementing signal timing according to the equal saturation principle can balance traffic pressure across phases and prevent overloading or unused traffic on some phases.
[0074] In this embodiment of the present application, the state of the traffic sequence element represents the demand flow rate of the critical lane corresponding to different phases in the signal control cycle. As an optional implementation method, in this embodiment of the present application, the lane with the highest flow rate in a phase is designated as the critical lane. The demand flow rate of the lane with the highest flow rate can best describe the traffic state of the current signal control cycle, thereby ensuring the rationality of subsequent signal control cycle decisions based on the traffic state.
[0075] The demand flow of key lanes for historical traffic sequence elements can be obtained by acquiring historical data. For traffic sequence elements to be decided, the demand flow of key lanes can be calculated based on the number of vehicles passing when the key lane has a green light and the number of remaining vehicles when the light switches to red (the number of vehicles passing and the number of remaining vehicles can be collected and stored by the radar device in the previous signal control cycle for the calculation of lane demand flow in the next signal control cycle. The number of remaining vehicles can be the number of vehicles in the key lane collected by the radar device just after the light switches to red, or the number of vehicles in the key lane collected by the radar device a certain time after the light switches to red). The demand flow of the lane can be calculated as follows:
[0076] Lane demand flow = (number of vehicles passing through the lane during the green light time of the previous cycle + number of vehicles on the lane section during the red light time of the previous cycle) / duration of the previous signal control cycle (s) * 3600;
[0077] In the embodiment of the present application, the demand flow is converted into hourly flow. Depending on the actual situation, it can also be converted into half-hourly flow, minute-level flow, etc. This application is not limited to this.
[0078] like Figure 2 As shown, as an optional implementation, the control method provided in the embodiments of the present application also includes training an offline reinforcement learning model. Utilizing the trained offline reinforcement learning model, the accuracy of the signal-control cycle inference value output for decision-making traffic sequence elements can be improved, thereby reducing indicators that affect traffic efficiency, such as signal delays and the number of stops, and alleviating traffic congestion.
[0079] Training an offline reinforcement learning model involves:
[0080] Step S201: Initialize the offline reinforcement learning model;
[0081] Step S202: input the training sequence into the initialized offline reinforcement learning model;
[0082] Step S203: Calculate the loss function based on the output result of the offline reinforcement learning model;
[0083] Step S204: Update the parameters of the offline reinforcement learning model with the goal of reducing the loss function.
[0084] Specifically, as an optional implementation method, the mean square error (MSE) is selected as the loss function of the offline reinforcement learning model, and its formula is:
[0085]
[0086] Where N represents the total number of traffic sequence elements in the training sequence, Y i represents the true value of the i-th traffic sequence element (i.e., the signal control period in the input traffic sequence element), represents the predicted value of the i-th traffic sequence element (i.e., the signal control cycle inference value output by the offline reinforcement model).
[0087] Calculate the average loss function of K training sequences (K is a natural number greater than 1) and update the weights of the offline reinforcement learning model to reduce the training error.
[0088] Step S205: input new training sequence data for training (used training data is no longer used), and determine whether the preset number of iterations has been reached. If not, return to step S202; otherwise, proceed to step S206.
[0089] Step S206: Determine whether the loss function is decreasing. If it is still decreasing, return to step S202; otherwise, proceed to step S207.
[0090] Step S207: Complete the training of the offline reinforcement learning model. Specifically, if the calculated error of the loss function of the training sequence is lower than the preset value and the loss function does not decrease with further iterations, the network is considered to have converged, and the training of the offline reinforcement learning model is completed.
[0091] Furthermore, for single-point intersections, the number of samples is usually small, and it is difficult to provide sufficient original sample data for offline reinforcement learning model training. To this end, data enhancement technology is adopted in the embodiment of the present application, and a large number of training sequences are generated using the traffic mechanism model for offline reinforcement learning model training.
[0092] like Figure 3 As shown in FIG, as an optional implementation method, a large number of training sequences are generated using the traffic mechanism model, including:
[0093] Step S301: defining the initial element state of the initial traffic sequence element in the training sequence according to the size of the critical lane demand flow, and obtaining the reward and signal control period of the corresponding element based on the initial element state through a preset traffic mechanism model;
[0094] Step S302: Perform Hadamard multiplication on the initial element state and the change ratio vector to obtain the derived element states of multiple derived traffic sequence elements in the training sequence. Based on the derived element states, the reward and signal control period of the corresponding element are obtained through the traffic mechanism model.
[0095] Step S303: Concatenate the initial traffic sequence element and the derived traffic sequence element to obtain a training sequence.
[0096] For example, in step S301, the initial element state of the initial traffic sequence element in the training sequence is defined according to the size of the required flow rate of the key lane, which can be done in the following manner:
[0097] The states of the initial traffic sequence elements are such that 25% of the critical lanes have low demand flows, 50% of the critical lanes have medium demand flows, and 25% of the critical lanes have high demand flows.
[0098] Specifically, according to the above, the state of the traffic sequence element can be expressed as:
[0099] f=(f1,f2,…,f p ), where p is the number of phases, f p represents the critical lane demand flow in the p-th phase.
[0100] Assume that the state of the initial traffic sequence element is recorded as f1, then 25% of the critical lane demand flow is small flow, 50% of the critical lane demand flow is medium flow, and 25% of the critical lane demand flow is large flow. In this way, the initial element state is defined according to the size of the critical lane demand flow, so that the defined initial element state is more in line with the actual situation. For example, in an embodiment of the present application, the flow range of small flow is defined as 0 to 75 vehicles / h, the flow range of medium flow is 76 to 99 vehicles / h, and the flow range of large flow is 100 to 300 vehicles / h. The size range of the critical lane demand flow here can also be set or adjusted through the user's client.
[0101] At this point, the generated initial element state can be expressed as:
[0102] f1=(k1,k2,…,k p );
[0103] The defined initial element state is input into the traffic mechanism model. The traffic mechanism model is an optimization model for solving the optimal cycle with multiple objectives. Through the traffic mechanism model, the corresponding signal control cycle can be output. and returns (The specific method of outputting the signal control cycle and the return of the traffic mechanism model will be described in detail later). and returns After rearrangement, the initial traffic sequence elements can be obtained, which can be expressed as follows:
[0104] Based on the initial element state f1=(k1,k2,…,k p ), random changes in a certain proportion r1=(r1,r2,…,r p ), as an optional implementation method, the proportional coefficient r of the random change of the key lane flow is sampled from the normal distribution, that is, r~N(μ,σ 2 ), in the embodiment of the present application, the coefficient of normal distribution μ=1.0, σ 2 =1.0.
[0105] In step S302, f1 is multiplied by the change ratio vector r1 to obtain the state f2 of the derived traffic sequence element, namely:
[0106] f2=f1*r1=(k1r1,k2r2,…,k p r p )
[0107] Repeat the above steps to generate derived element states to obtain multiple different derived element states f3, f4, ....f detc., call the traffic mechanism model to obtain states f3, f4, ...f d Corresponding optimal signal control period And the optimal objective function value corresponding to the optimal period By rearranging the states, signal control cycles, and returns corresponding to the traffic sequence elements, any traffic sequence element in the training sequence can be represented as follows:
[0108] In step S303, the initial traffic sequence element and the derived traffic sequence element are concatenated to obtain a training sequence. The generated training sequence can be expressed as follows: T = (T1, T2, ... T d ).
[0109] According to the above description, the data enhancement technology is adopted in the embodiment of the present application, and a large number of training sequences are generated by using the traffic mechanism model for offline reinforcement learning model training, which effectively solves the problem that the number of single-point intersection samples is small and it is difficult to provide sufficient original sample data for offline reinforcement learning model training. The large number of generated training sequences are input into the offline reinforcement learning model. According to the training method of the offline reinforcement learning model provided in this application, the offline reinforcement learning model can be trained to ensure that when the signal control cycle decision is subsequently made for the decision sequence, the optimal signal control cycle inference value can be output based on the trained offline reinforcement learning model, thereby improving the traffic efficiency of the intersection and reducing congestion.
[0110] As an optional implementation method, the rewards and signal control cycles corresponding to each traffic sequence element are obtained through the preset traffic mechanism model, including:
[0111] Based on signal delay, parking time and lane queue length, a performance index is formed to characterize the return of environmental feedback when different signal control cycles are adopted under different traffic sequence elements;
[0112] The state of the traffic sequence element is input into the traffic mechanism model, and the signal control period output by the traffic mechanism model is obtained with the minimum of the performance index as the goal.
[0113] Optionally, the traffic mechanism model adopts a single-step decision-making single-point periodic signal control optimization model, and the input of the traffic mechanism model is the element state f=(f1,f2,…,f p ), p is the number of phases, f p Indicates the critical lane demand flow in the pth phase. Since the performance index PI of the traffic mechanism model i It is a combination of intersection delay d, parking time s, and queue q. Therefore, when evaluating the return, the smaller the index, the better. The performance index PI iMinimum is the optimization goal, in the feasible space range of signal control period [C min ,C max ] to obtain the optimal signal control period C * And the optimal performance index PI corresponding to the optimal cycle * .
[0114] In the embodiment of the present application, the performance index is weighted and inverted to obtain the reward of the traffic sequence element (i.e. ), the smaller the performance index PI is, the lower the negative impact of factors such as delays, parking, and queuing on traffic efficiency is. Therefore, the minimized performance index can be used to characterize the optimal environmental feedback after weighted inversion. The offline reinforcement learning model is based on the historical sequence of action (i.e., signal control cycle)-reward (i.e., ) is used for reasoning, and the optimal signal control cycle reasoning value can be output, which can improve the traffic efficiency of the target single-point intersection and reduce traffic congestion.
[0115] like Figure 4 As shown, as an optional implementation method, the control method provided in the embodiment of the present application also includes: first calibrating the parameters of the traffic mechanism model, and generating training data based on the traffic mechanism model after parameter calibration.
[0116] Parameter calibration of the traffic mechanism model includes:
[0117] Step S401: obtaining a real sample data set, where the real sample data set includes a plurality of real sample data, and any real sample data includes a mapping relationship between a traffic sequence element state and an optimal signal control period;
[0118] Step S402: setting a value range of the parameter to be calibrated, and using a genetic algorithm to randomly generate a population of values of the parameter to be calibrated within the calibration parameter range;
[0119] Step S403: Input each real sample data into the traffic mechanism model to obtain performance indicators under different values of the parameters to be calibrated in the population;
[0120] Step S404: Optimal calibration parameters are obtained with the goal of minimizing the sum of performance indicators corresponding to the real sample data.
[0121] The real sample data can be selected from the data captured by the radar equipment to select the strategy with better on-site control effect, which can be adopted [traffic sequence element state f——optimal signal control cycle C * ] mapping relationship combination.
[0122] The traffic mechanism model adjusts parameters and signal control cycles through the following formula:
[0123]
[0124] Where w1, w2, w3, w4, w5, and w6 are all parameters of the traffic mechanism model, d(C, f, w4) represents the signal delay function corresponding to the signal control period, s(C, f, w5) represents the parking time function corresponding to the signal control period, q(C, f, w6) represents the queue length function corresponding to the signal control period, C represents the signal control period of the traffic sequence element as a sample, and f represents the state of the traffic sequence element as a sample.
[0125] As an optional implementation, the signal delay function is calculated using the following formula:
[0126]
[0127] Where, Indicates the percentage adjustment of traffic flow f pk Calculate the delay of each vehicle, G is the total green light duration of each phase, s is the saturation flow rate. For example, D p10 Indicates adjusting the traffic flow f according to percentile k = 10 (p10 means 10%) p10 Calculate delays for each vehicle.
[0128] f pk Indicates adjusting the traffic flow according to percentile k, f is the state of the traffic sequence element as a sample, z pk According to the value of percentile k, when k is 10 (pk is p10, representing 10%), z pk When the value is -1.28 and k is 30 (pk is p30, representing 10%), z pk When the value is -0.52 and k is 5 (pk is p5, representing 5%), z pk When the value is 0 and k is 70 (pk is p70, representing 70%), z pk When the value is 0.52 and k is 90 (pk is p90, representing 90%), z pk The value is 1.28.
[0129] Based on the calculation method of the signal delay function in the traffic mechanism model provided in the embodiment of the present application, parameters w1 and w4 are designed in the signal delay function, so that the evaluation of the impact of signal delay on the performance index is more accurate, thereby improving the reliability of the traffic mechanism model.
[0130] As an optional implementation, the parking time function is calculated using the following formula:
[0131]
[0132] Where, represents the average number of stops in phase i; s represents the saturation flow rate, G i represents the effective green light time of phase i, f i represents the demand flow of the key lane in phase i.
[0133] Based on the calculation method of the parking time function in the traffic mechanism model provided in the embodiment of the present application, the influencing parameters w2 and w5 in the parking time function can be used to make the evaluation of the impact of parking time on the performance index more accurate, thereby improving the reliability of the traffic mechanism model.
[0134] As an optional implementation, the queue length function is calculated using the following formula:
[0135]
[0136] Where L represents the headway, f represents the state of the traffic sequence element as a sample, R represents the red light time of each phase, s represents the saturated flow, and F u represents the lane utilization coefficient.
[0137] Based on the calculation method of the queue length function in the traffic mechanism model provided in the embodiment of the present application, the queue length function can be used to influence parameters w3 and w6, making the evaluation of the impact of queue length on the performance index more accurate, thereby improving the reliability of the traffic mechanism model.
[0138] For example, w1, w2, w3, w4, w5, and w6 all represent the parameters of the traffic mechanism model and meet the following constraints:
[0139]
[0140] Under the above constraints, find a set of parameters {w1, w2, w3, .., w5, w6} that minimizes the error between the signal control period output by the traffic mechanism model and the label period (the label period can be the signal control period in the sample data). The solution for the parameters with the smallest signal control period error can be: under the constraints of parameters w1 to w6, use a genetic algorithm to randomly generate a population of parameter values to be calibrated within the calibration parameter range, input each real sample data set into the traffic mechanism model, obtain the performance indicators under different parameter values to be calibrated in the population, and calculate the sum of the total objective function values ∑ i PI i (i is the data sample number) the minimum calibration parameter as the optimal calibration parameter.
[0141] Through the above parameter calibration of the traffic mechanism model, after determining the optimal calibration parameters w1~w6, the optimal calibration parameters w1~w6 are substituted into the traffic mechanism model to obtain the traffic mechanism model with completed parameter calibration. The state f of the traffic sequence element is input into the traffic mechanism model with completed parameter calibration. In the feasible space range of the signal control cycle [C min ,C max ] to obtain the optimal signal control period C * And the optimal objective function value PI corresponding to the optimal period * Parameter calibration of the traffic mechanism model can ensure the reliability of the signal control cycle and returns generated by data enhancement, thus providing a basis for the subsequent training of the offline reinforcement learning model.
[0142] To further illustrate the traffic signal control method provided in the embodiment of the present application, Figure 5 As shown, an overall flow chart is provided, including:
[0143] Step S501: calibrating parameters of the traffic mechanism model based on a small amount of sample data;
[0144] Step S502: Introducing data enhancement technology, using the calibrated traffic mechanism model to generate a training sequence;
[0145] Step S503: inputting the training sequence into the initialized offline reinforcement learning model to complete the training of the offline reinforcement learning model;
[0146] Step S504: generating a sequence to be decided based on the historical traffic sequence elements and the sequence to be decided elements;
[0147] Step S505: Input the sequence to be decided into the trained offline reinforcement learning model. The offline reinforcement learning model performs inference based on the rewards, states, and signal control cycles in the historical traffic sequence elements to obtain the signal control cycle inference value of the traffic sequence element to be decided.
[0148] Step S506: Allocate the duration of traffic lights of different phases based on the inferred value of the signal control cycle.
[0149] In summary, the traffic signal control method provided in the embodiment of the present application introduces data enhancement technology and uses small sample data to calibrate the traffic mechanism model. The calibrated traffic mechanism model can generate a large number of training sequences for offline reinforcement learning technology model training on the one hand, and can also provide reward information for the training sequence on the other hand. In addition, the control method is constructed based on offline reinforcement learning technology and uses a sequential decision-making method to complete the signal control cycle calculation. This sequential decision-making method can fully utilize historical information to better adapt to traffic flow fluctuations. Compared with the method based on single-step decision-making of the traffic mechanism model, it can reduce decision-making errors and avoid causing the signal control cycle to be too large or too small, thereby reducing indicators such as delays and the number of stops, improving the traffic efficiency of the intersection, and reducing congestion.
[0150] like Figure 6 As shown, based on the same inventive concept, an embodiment of the present application further provides an electronic device 100, comprising: at least one memory 11 and at least one processor 12, wherein the at least one memory 11 stores executable code, and the at least one processor 12 is used to execute the executable code in the at least one memory 11 to implement the above-mentioned traffic signal control method.
[0151] The at least one memory can be used to store a computer program, which may include instructions and data, and implement the steps of any of the above methods. Memory 11 can be random access memory, read-only memory, non-volatile memory, programmable ROM, erasable PROM, electrically erasable memory, flash memory, optical memory, registers, etc. Processor 12 can be a general-purpose processor that performs specific steps and / or operations by reading and executing a computer program stored in memory 11. The general-purpose processor may use data stored in memory 11 during the execution of the steps and / or operations. The general-purpose processor can be a central processing unit, an ASIC, an FPGA, etc. The electronic device 100 can also include a communication interface, which can include input / output interfaces, physical interfaces, and logical interfaces for interconnecting devices within the network device. During implementation, the steps of the above method can be implemented by hardware integrated logic circuits in the processor or by software instructions. The methods disclosed in conjunction with the embodiments of the present application can be directly implemented by a hardware processor or implemented by a combination of hardware and software modules in the processor.
[0152] According to the electronic device 100 provided in the embodiment of the present application, data enhancement technology is introduced to calibrate the traffic mechanism model using small sample data. The calibrated traffic mechanism model can generate a large number of training sequences for offline reinforcement learning technology model training on the one hand, and provide reward information for the training sequence on the other hand. In addition, the control method is constructed based on offline reinforcement learning technology and uses a sequential decision-making method to complete the signal control cycle calculation, rather than the traffic mechanism model based on a single-step decision method. Through this sequential decision-making method, historical information can be fully utilized to better adapt to traffic flow fluctuations, reduce decision-making errors, and avoid causing the signal control cycle to be too long or too short, thereby reducing indicators such as delays and the number of stops, improving the traffic efficiency of the intersection, and reducing congestion.
[0153] like Figure 7 As shown, the embodiment of the present application further provides a traffic control system 200 , including: a signal light 21 and the above-mentioned electronic device 100 .
[0154] According to the traffic control system provided in the embodiments of the present application, data augmentation technology is introduced to calibrate the traffic mechanism model using small sample data. The calibrated traffic mechanism model can generate a large number of training sequences for offline reinforcement learning technology model training on the one hand, and provide reward information for the training sequence on the other hand. In addition, the control method is constructed based on offline reinforcement learning technology and uses a sequential decision-making method to complete the signal control cycle calculation, rather than the single-step decision-making method of the traffic mechanism model. Through this sequential decision-making method, historical information can be fully utilized to better adapt to traffic flow fluctuations, reduce decision-making errors, and avoid causing the signal control cycle to be too long or too short, thereby reducing indicators such as delays and the number of stops, improving the traffic efficiency of the intersection, and reducing congestion.
[0155] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a solid-state drive (SSD).
[0156] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0157] It will be understood that the word "exemplary" as used herein means "serving as an example, instance, or illustration." Any embodiment described as "exemplary" is not necessarily preferred or advantageous over other embodiments and / or does not exclude the ability to combine features of other embodiments. It will be understood that certain features of the present application, which are described in the context of separate embodiments for the sake of clarity, may also be provided in combination in a single embodiment. Conversely, various features of the present application, which are described in the context of a single embodiment for the sake of clarity, may also be provided separately or in any suitable combination or as any other described embodiment of the present application.
[0158] The above disclosure is only a preferred embodiment of the present application, but it is not intended to limit the scope of rights of the present application. A person skilled in the art can understand that without departing from the spirit and scope of the present application and the appended claims, changes, modifications, substitutions, combinations, and simplifications should all be equivalent replacement methods and still fall within the scope of the invention.
Claims
1. A traffic signal control method, characterized in that: The control method includes: Acquire historical traffic data for a target intersection, the historical traffic data comprising a plurality of historical traffic sequence elements, each of which comprises a report, a state, and a signal control cycle, wherein the state represents the required traffic flow of a key lane corresponding to different phases in the signal control cycle, and the report represents traffic environment feedback when different signal control cycles are adopted; Using a radar device to obtain a state of a target decision cycle, and generating a traffic sequence element to be decided based on the state of the target decision cycle; generating a sequence to be decided based on the historical traffic sequence elements and the traffic sequence elements to be decided; Inputting the sequence to be decided into a preset offline reinforcement learning model, the offline reinforcement learning model performs reasoning based on the rewards, states, and signal control cycles in the historical traffic sequence elements to obtain the signal control cycle inference value of the traffic sequence element to be decided; The signal light durations of different phases are allocated based on the signal control cycle inference value.
2. The traffic signal control method according to claim 1, characterized in that: The control method further includes training the offline reinforcement learning model, including: The initial element state of the initial traffic sequence element in the training sequence is defined according to the size of the demand flow of the key lane. Through the preset traffic mechanism model, the reward and signal control period of the corresponding element are obtained according to the initial element state; Performing a Hadamard multiplication on the initial element state and the change ratio vector to obtain derived element states of a plurality of derived traffic sequence elements in the training sequence, and obtaining the reward and signal control period of the corresponding element according to the derived element states through the traffic mechanism model; The training sequence is input into the initialized offline reinforcement learning model, a loss function is calculated based on the output result of the offline reinforcement learning model, and the parameters of the offline reinforcement learning model are updated with the goal of reducing the loss function.
3. The traffic signal control method according to claim 2, characterized in that: Through the preset traffic mechanism model, the rewards and signal control cycles corresponding to each traffic sequence element are obtained, including: Based on signal delay, parking time and lane queue length, a performance index is formed, wherein the performance index is used to represent the return of environmental feedback when different signal control cycles are adopted under different traffic sequence elements; The state of the traffic sequence element is input into the traffic mechanism model, and the signal control period output by the traffic mechanism model is obtained with the minimum of the performance index as the goal.
4. The traffic signal control method according to claim 2, characterized in that: Through the preset traffic mechanism model, the rewards and signal control cycles corresponding to each traffic sequence element are obtained, including: The state of the traffic sequence element is input into the traffic mechanism model, and the model is traversed within a feasible range of the signal control cycle to obtain performance indicators corresponding to different signal control cycles. The signal control cycle that minimizes the performance indicator is selected as the optimal signal control cycle of the traffic sequence element. The performance indicator corresponding to the optimal signal control cycle is weighted and inverted to obtain the reward of the traffic sequence element.
5. The traffic signal control method according to claim 4, characterized in that: The control method further includes calibrating parameters of the traffic mechanism model, including: Acquire a real sample data set, wherein the real sample data set includes a plurality of real sample data, and any real sample data includes a mapping relationship between a traffic sequence element state and an optimal signal control period; Set the value range of the parameter to be calibrated, and use the genetic algorithm to randomly generate a population of parameter values to be calibrated within the calibration parameter range; Inputting each of the real sample data into the traffic mechanism model to obtain performance indicators under different values of the parameters to be calibrated in the population; The goal is to minimize the sum of the performance indicators corresponding to each real sample data and obtain the optimal calibration parameters.
6. The traffic signal control method according to any one of claims 3 to 5, characterized in that: The traffic mechanism model adjusts parameters and signal control cycles through the following formula: Where w1, w2, w3, w4, w5, and w6 represent the parameters of the traffic mechanism model, d(C, f, w4) represents the signal delay function corresponding to the signal control period, s(C, f, w5) represents the parking time function corresponding to the signal control period, q(C, f, w6) represents the queue length function corresponding to the signal control period, C represents the signal control period of the traffic sequence element as a sample, and f represents the state of the traffic sequence element as a sample.
7. The traffic signal control method according to any one of claim 6, characterized in that: The signal delay function is calculated by the following formula: The total duration of the green light in the phase, s is the saturation flow rate, f pk Indicates adjusting the traffic flow according to percentile k, f is the state of the traffic sequence element as a sample, z pk Indicates that according to the value of percentile k, when k is 10, z pk When the value is -1.28 and k is 30, z pk When the value is -0.52 and k is 50, z pk When the value is 0 and k is 70, z pk When the value is 0.52 and k is 90, z pk The value is 1.
28.
8. The traffic signal control method according to claim 6, characterized in that: The parking time function is calculated by the following formula: Where, represents the average number of stops in phase i; s represents the saturation flow rate, G i represents the effective green light time of phase i, f i represents the demand flow of the key lane in phase i.
9. The traffic signal control method according to claim 6, characterized in that: The queue length function is calculated by the following formula: Where L represents the headway, f represents the state of the traffic sequence element as a sample, R represents the red light time of each phase, s represents the saturated flow, and F u represents the lane utilization coefficient.
10. An electronic device, characterized in that: include: at least one memory and at least one processor, The at least one memory stores an executable code, and the at least one processor is configured to execute the executable code in the at least one memory to implement the method according to any one of claims 1 to 9.
11. A traffic control system, characterized in that: include: A signal lamp and an electronic device as claimed in claim 10.
Citation Information
Patent Citations
Traffic signal lamp control method and device, electronic equipment and storage medium
CN111564048A
Multi-intersection traffic light control method and system based on reinforcement learning, and storage medium
CN113223305A
Passenger delay minimization signal control method based on deep reinforcement learning
CN117116064A
Intersection adaptive signal control method based on cellular deduction multi-step decision
CN119207124A
Method and system for dynamic traffic control for one or more junctions
WO2022070201A1