Traffic signal online optimization control method and system for low fuel consumption
By building virtual scenes using traffic digital twin technology and utilizing online learning to optimize traffic signal control, the problems of vehicle fuel consumption and carbon emissions in intersection areas in existing technologies have been resolved, and optimized traffic signal control with low fuel consumption and low carbon emissions has been achieved.
Patent Information
- Application Number
- CN202211696609.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2042-12-28
AI Technical Summary
Existing traffic signal control schemes are unable to effectively reduce vehicle fuel consumption and carbon emissions in intersection areas. The lack of an accurate model of the relationship between vehicle fuel consumption and signal control makes it difficult to achieve low-fuel-consumption traffic optimization control.
Traffic digital twin software is used to build virtual scenes, traffic signal control is optimized through online learning, green light timing is optimized using lookup tables and evaluation values, and evaluation values are updated in combination with virtual vehicle fuel consumption to achieve single-phase and multi-phase low fuel consumption control, and finally the real traffic signal is controlled through the main control module.
It significantly reduces the average fuel consumption and carbon emissions of vehicles in urban intersection areas, optimizes signal control schemes to suit real conditions, and avoids losses and greenhouse effects caused by unreasonable signal schemes.
Smart Images

Figure CN115953909B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of ecological transportation technology and transportation digital twin technology, and in particular to an online optimization control method and system for traffic signals with low fuel consumption. Background Art
[0002] Low carbon emissions and energy conservation are the development trends of intelligent transportation. Although the number of new energy vehicles has grown rapidly in recent years, conventional fuel vehicles still account for a large proportion of the market. Carbon emissions from motor vehicle fuel consumption have a significant impact on the environment. Fuel consumption of motor vehicles increases significantly at urban intersections. The signal control schemes at road intersections are closely related to vehicle fuel consumption. Therefore, it is of great significance to study low-energy traffic signal optimization control methods to reduce average fuel consumption per vehicle at intersections, thereby reducing carbon emissions.
[0003] Urban intersections are traffic bottlenecks, and traffic control schemes significantly impact the operating conditions and fuel consumption of vehicles. Existing traffic control schemes aim to reduce average vehicle delays and increase traffic volume. These schemes primarily rely on delay models, such as the Webster delay model and the HCM2000 delay model. While some research has focused on traffic fuel consumption, these studies primarily analyze the estimated fuel consumption of vehicles under specific traffic organization and control conditions. However, research on signal control for low fuel consumption has been lacking.
[0004] The key to optimizing traffic signal control for low fuel consumption lies in establishing a relationship model between traffic control schemes and vehicle fuel consumption. Due to the highly complex, nonlinear coupling relationship between intersection control schemes and vehicle fuel consumption, accurately modeling this relationship is difficult. Therefore, achieving optimized control for low fuel consumption at intersections, and thus reducing carbon emissions, is challenging. Summary of the Invention
[0005] The purpose of the present invention is to provide a low-fuel-consumption online optimization control method and system for traffic signals, which not only effectively reduces the average fuel consumption of vehicles in urban intersection areas, thereby saving energy, but also reduces the carbon emissions of vehicles.
[0006] The purpose of the present invention is achieved through the following technical solutions:
[0007] A traffic signal online optimization control method for low fuel consumption, comprising:
[0008] Use traffic digital twin software to build virtual scenes corresponding to real scenes;
[0009] In each virtual intersection of the virtual scene, each phase consisting of non-conflicting traffic flows is set;
[0010] A reference phase is selected for online learning of a single-phase low-fuel-consumption optimization control scheme. During the current cycle of the learning process, the current state of the road controlled by the reference phase in the virtual scene is determined. A lookup table is searched and the optimal phase green light timing corresponding to the current state is found from the lookup table based on the evaluation value. The optimal phase green light timing is assigned to the reference phase, and the corresponding evaluation value is updated based on the average fuel consumption of the virtual vehicle during the optimal phase green light timing. The learning continues for multiple cycles until the evaluation values converge, an optimized lookup table is obtained, and the online learning of the single-phase low-fuel-consumption optimization control scheme is completed. The state refers to a state related to fuel consumption.
[0011] The optimized lookup table is shared with other phases and transmitted to the main control module by the traffic digital twin software. The main control module uses the optimized lookup table to jointly optimize all phases and control the real traffic lights in the real scene to complete the online optimization control of traffic signals.
[0012] A traffic signal online optimization control system for low fuel consumption, comprising:
[0013] A virtual scene construction unit, used to construct a virtual scene corresponding to the real scene using traffic digital twin software;
[0014] A phase setting unit, configured to set phases consisting of non-conflicting traffic flows in each virtual intersection of the virtual scene;
[0015] The single-phase low-fuel-consumption optimization control scheme online learning unit is used to select a reference phase and conduct online learning of the single-phase low-fuel-consumption optimization control scheme; during the current cycle of the learning process, the current state of the road controlled by the reference phase in the virtual scene is determined, and by searching the lookup table, the optimal phase green light timing corresponding to the current state is found from the lookup table; the optimal phase green light timing is configured in the reference phase, and the corresponding evaluation value is updated based on the average fuel consumption of the virtual vehicle during the optimal phase green light timing; the learning is continued for multiple cycles until the evaluation values converge, an optimized lookup table is obtained, and the online learning of the single-phase low-fuel-consumption optimization control scheme is completed; wherein the state refers to a state related to fuel consumption;
[0016] The multi-phase joint optimization and control unit is used to share the optimized lookup table with other phases. The traffic digital twin software transmits it to the main control module, which then uses the optimized lookup table to jointly optimize all phases and control real traffic signals in real scenarios, completing online traffic signal optimization control.
[0017] A processing device comprising: one or more processors; a memory for storing one or more programs;
[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.
[0019] A readable storage medium stores a computer program, which implements the aforementioned method when the computer program is executed by a processor.
[0020] As can be seen from the technical solutions provided by the present invention, a virtual scene corresponding to the real scene can be constructed; and the optimization of the signal control scheme for low fuel consumption in traffic is achieved based on the traffic digital twin platform, reducing the risks and losses caused by unreasonable signal schemes; because the online optimization of the signal control scheme uses the average fuel consumption of vehicles as an indicator, after multiple learning, the average fuel consumption of vehicles can be significantly reduced, while also reducing carbon emissions. In addition, the system of the present invention interacts with the environment to learn the optimized control scheme for low fuel consumption traffic signals, which conforms to the real situation, effectively solving the energy problem while avoiding the greenhouse effect caused by excessive carbon emissions. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 A flowchart of a method for online optimization and control of traffic signals for low fuel consumption provided by an embodiment of the present invention;
[0023] Figure 2 A schematic diagram of multiple phases consisting of non-conflicting traffic flows provided by an embodiment of the present invention;
[0024] Figure 3 A schematic diagram of a traffic signal online optimization control system for low fuel consumption provided by an embodiment of the present invention;
[0025] Figure 4 A schematic diagram of a processing device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0027] First, the following terms may be used in this article:
[0028] The terms "include," "comprises," "contains," "has," or other similar expressions should be interpreted as non-exclusive. For example, "including certain technical features (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products, or manufactured articles, etc.) should be interpreted as including not only the technical features explicitly listed, but also other technical features known in the art that are not explicitly listed.
[0029] The following describes in detail a fuel-efficient online traffic signal optimization control solution provided by the present invention. Any details not described in detail in the present embodiments are prior art known to those skilled in the art. For any unspecified conditions in the present embodiments, the control procedures will be based on conventional conditions in the art or the manufacturer's recommendations.
[0030] Example 1
[0031] The embodiment of the present invention provides an online optimization control method for traffic signals with low fuel consumption. Through traffic digital twins, the method takes the minimum average fuel consumption per vehicle as the optimization goal, comprehensively considers factors affecting vehicle fuel consumption in the intersection area, including traffic signal control schemes, vehicle driving conditions, and road channelization and geometric shapes of intersections. It better depicts the complex relationship between intersection signal schemes and average vehicle fuel consumption, and is suitable for traffic signal optimization control with low fuel consumption at different types of intersections. Figure 1 As shown, the present invention mainly includes the following steps:
[0032] Step 1: Use traffic digital twin software to build a virtual scene corresponding to the real scene.
[0033] In an embodiment of the present invention, traffic digital twin software can be used to construct a virtual scene including a corresponding virtual intersection and virtual signal according to the geometric shape, road channelization and signal of the intersection in the real scene, set up a virtual vehicle detector at the upstream entrance of the virtual intersection road, and generate a specified number of virtual vehicles of different sizes.
[0034] In the present invention, virtual vehicles of different sizes can be understood as virtual vehicles of different dimensions, for example, large, medium and small virtual vehicles.
[0035] In an embodiment of the present invention, a virtual vehicle detector can be used to detect the length of the virtual vehicle queue in each lane, and the virtual traffic light can execute the relevant traffic light control scheme, which is generally controlled by a main control module. The main control module (not shown in the accompanying drawings) generates a traffic signal control scheme based on the low-fuel-consumption traffic signal online optimization control results or initially uses a preset traffic signal control scheme.
[0036] Step 2: In each virtual intersection of the virtual scene, set each phase consisting of non-conflicting traffic flows.
[0037] like Figure 2 As shown, a phase scheme composed of non-conflicting traffic flows is shown, mainly including: phases composed of opposite straight-moving traffic flows (marked 1, 3) and opposite left-turning traffic flows (2, 4). Of course, the four phases provided here are only examples, and users can set other forms of phases composed of non-conflicting traffic flows according to actual conditions.
[0038] To facilitate understanding, the following industry terminology related to traffic signal control schemes is introduced. Phase refers to the display status of the signal group corresponding to one or more traffic flows that are granted the right of way simultaneously. Phase green light timing is the duration of the green light display during a phase. The signal cycle is the time it takes for the signal light color to change once according to the set signal phase sequence.
[0039] Step 3: Select a reference phase and conduct online learning of the single-phase low fuel consumption optimization control scheme.
[0040] In this embodiment of the present invention, online traffic signal optimization control is defined as a Markov decision process. Since there is no model for the relationship between traffic control schemes and fuel consumption, the average fuel consumption of vehicles at an intersection is minimized. To achieve fuel-efficient traffic signal optimization control, an online learning method is introduced. Through trial and error execution of traffic control schemes, a fuel-efficient traffic signal optimization scheme is found.
[0041] In this embodiment of the present invention, the online optimization control scheme for traffic signals includes a reward indicator determined by the average fuel consumption of each vehicle. The traffic control scheme is represented by a four-tuple: (S, A, P, R), where S is the state set of the roads controlled by each phase, a is the set of phase green light timings corresponding to each phase, P is the state transition probability, that is, the probability that the road state will transition to another state after executing the green light timing corresponding to one phase; R represents the reward, which is calculated using the average fuel consumption of virtual vehicles. Specifically:
[0042] Set S i = i1 ,S i2 ,S i3 ,…,S ik > is a finite set of discrete, joint states, which is a state set of an intersection, where S i1 ~S ik For S i Each sub-state represents a phase-specific state. The phase-specific state can be the length of the vehicle queue and the speed of the leading vehicle at each road entrance. For example, Si1 represents the east-west straight-ahead state (sub-state), Si2 represents the east-west left-turn state (sub-state), and so on. k is the number of phases at the intersection.
[0043] Set A is a discrete, finite set of joint actions, A j = j1 ,A j2 ,A j3 ,…,A jk >, you can set the duration of the green light for the corresponding phase, for example, A j1 For substate S i1 Corresponding phase green light timing set.
[0044] The state transition probability P is the probability that the road changes from one state to another after the signal implements a phase green light timing plan. The state transition probability is calculated using the evaluation value. For an explanation of the evaluation value, please refer to the following text.
[0045] The reward R is related to the average fuel consumption. The lower the average fuel consumption, the higher the reward.
[0046] In the embodiment of the present invention, the online learning of the multi-phase low fuel consumption optimization control scheme is simplified to the online learning of the single-phase low fuel consumption optimization control scheme. By sharing the Q function value (i.e., the evaluation value mentioned later), the computational complexity of the optimization scheme is reduced. Specifically, the computational complexity of the online learning of the multi-phase low fuel consumption optimization control scheme is (|S im |×|A jn |) n , where |S im |、|A jn |represents the set S respectively im 、A jn The number of elements contained in (corresponding to the state of traffic flow in all directions of an intersection and the total number of possible timings for the traffic flow state), n is the number of different traffic flow directions. Such a high computational complexity is difficult to meet the real-time requirements of practical applications. In order to reduce the computational complexity of the optimization process, we select the traffic flow in one direction defined above (such as Figure 2 In 1), for a single phase, we study the relationship between traffic control schemes and fuel consumption. When the Q function converges, we find an optimal control scheme with low fuel consumption, which can be shared with other phases. This means that the optimization timing of a single phase is expanded to the optimization timing of multiple non-conflicting phases. By sharing the Q function value, the computational complexity is reduced to |S im |×|A jn |.
[0047] The main process of online learning of a single-phase low-fuel-consumption optimization control scheme is as follows: During the current cycle of the learning process, determine the current state of the road controlled by the reference phase in the virtual scene (the state related to fuel consumption), search the lookup table, and find the optimal phase green light timing corresponding to the current state from the lookup table based on the evaluation value; configure the optimal phase green light timing to the reference phase, and update the corresponding evaluation value based on the average fuel consumption of the virtual vehicle during the optimal phase green light timing; continue learning for multiple cycles until the evaluation values converge, obtain the optimized lookup table, and complete the online learning of the single-phase low-fuel-consumption optimization control scheme. Specifically:
[0048] In the embodiment of the present invention, the current state includes the length of the virtual vehicle queue and the speed of the leading vehicle (the speed of the first virtual vehicle in front of the intersection stop line). The speed of the leading vehicle is used to distinguish the differences between the same queue lengths during state transition, which has a great impact on the calculation of vehicle fuel consumption. Taking the first phase as an example, the current state of the road it controls is represented by S C1 , where s C1 = <L C1 ,V C1 >, L C1 、V C1 Represent the longer queue length and the head vehicle speed of the two roads controlled by the first phase respectively. Let the current state S C1 =X.
[0049] In the embodiment of the present invention, the first item in the lookup table is the state, the second item is the phase green light timing set, and the third item is the evaluation value set of the combination of the first item and the second item, as shown in Table 1.
[0050] Table 1: Lookup table
[0051]
[0052] In Table 1, S is the state set, which is the Cartesian product of the queue length and the head vehicle speed, and contains two pieces of information: one is the queue length set L, and the other is the head vehicle speed set V, both of which are related to fuel consumption. For example, L∈{0,1,…,30} (unit veh), V∈{0,1,…,12} (unit m / s). A is the phase green light timing set corresponding to the state set (which can be 0-maximum green light time), and each subset in the state set S corresponds to the duration of several phase green light timings; in addition, each subset in the state set and each corresponding phase green light timing have corresponding evaluation values, forming an evaluation value set Q(S, A). Before the self-learning process begins, the evaluation value set can be randomly set. The following is a more detailed introduction to the lookup table:
[0053] In the same state, the phase green light timing set consisting of multiple phase green light timings is recorded as X, and the corresponding phase green light timing set is recorded as in, Indicates the jth phase green light timing corresponding to the current state X, j∈{1,2,…,N max}, N max is the number of phase green light timings. The combination of the current state X and different phase green light timings corresponds to different evaluation values, that is, there are N max evaluation value, the current state X and The evaluation value of For the current state X, the phase green light timing corresponding to the maximum evaluation value is selected from the corresponding phase green light timing set as the optimal phase green light timing, which is recorded as
[0054] In addition, to avoid falling into the local optimum during the evolution process, the Boltzmann method can be used to randomly select the optimal phase green light timing for the current state in the initial stage, and then the optimal phase green light timing can be selected based on the evaluation value using the method described above.
[0055] For the current state X, determine the optimal phase green light timing After that, the master control module can control the virtual signal machine to execute the optimal phase timing. Each virtual vehicle moves according to the microscopic car-following model, and the average fuel consumption of the virtual vehicles at the virtual intersection in the current cycle is calculated, and the evaluation value in the lookup table is updated based on this. Specifically:
[0056] 1) Under the current state X, execute the optimal phase green light timing Then it changes to state Y and uses the best phase green light timing The reward value r corresponding to the average fuel consumption of the virtual vehicle during the period is calculated as follows:
[0057] a) Calculate the specific power (VSP) of each virtual vehicle.
[0058] VSP is a variable that links a vehicle's driving conditions with fuel consumption and emissions, and is closely related to fuel consumption. The calculation formula for the VSP variable (corresponding to small cars) is:
[0059] VSP=v(1.1a+0.132)+0.000302v 2
[0060] Among them, v represents the instantaneous speed of the virtual vehicle, a represents the instantaneous acceleration of the virtual vehicle, and VSP represents the specific power of the virtual vehicle. Generally speaking, the VSP values corresponding to large and medium-sized vehicles are approximately 1.7 and 1.5 times that of small vehicles, respectively.
[0061] b) Calculate the specific power range based on the calculated specific power of the virtual vehicle:
[0062] VSPBin=Int(VSP+0.5)
[0063] Wherein, VSPBin represents the specific power range.
[0064] c) Determine the average fuel consumption rate FR corresponding to VSPBin by looking up the preset specific power and average fuel consumption rate table shown in Table 2 k , and determine the average fuel consumption rate FR0 corresponding to the benchmark specific power range, and calculate the standard fuel consumption corresponding to VSPBin: EFR k =FR k / FR0.
[0065] Table 2: VSP-average fuel consumption rate table
[0066]
[0067] For example, assuming VSPBin=4, the average fuel consumption rate FR k =2.8751; when VSPBin=0 (reference ratio power range), FR0=1.
[0068] d) Calculate the reward value r by combining the standard fuel consumption of all virtual vehicles:
[0069]
[0070] Where Vnum represents the number of virtual vehicles and l represents the calculation interval.
[0071] 2) Calculate the updated evaluation value based on the reward value r, and use the calculated updated evaluation value to update the corresponding evaluation value in the lookup table.
[0072] In the embodiment of the present invention, the formula for calculating the updated evaluation value in combination with the reward value r is:
[0073]
[0074] in, Indicates the optimal phase green light timing corresponding to the current state X in the current cycle. Indicates the current state X in the current cycle and The evaluation value of the arrow on the left Indicates the updated current state X and Evaluation value, Y represents the optimal phase green light timing of the current state X in the current cycle where maxQ(Y,B) represents the highest evaluation value corresponding to state Y, B represents the phase green light timing corresponding to maxQ(Y,B), i.e., the optimal phase green light timing, γ represents the discount factor (e.g., 0.96), and α represents the learning rate, expressed as: α = max((0.45-s / 100000), 0.001). s represents the number of iterations.
[0075] In the embodiment of the present invention, when X=<0,0> (i.e., the queue length and the head vehicle speed are both 0), a cycle of learning ends, and the online learning of the single-phase low fuel consumption optimization control scheme will continue for multiple cycles until the evaluation value converges, so that a more accurate lookup table can be obtained, thereby laying the foundation for the online optimization control of low fuel consumption in traffic.
[0076] Step 4: Share the optimized lookup table with other phases and transmit it to the main control module by the traffic digital twin software. The main control module uses the optimized lookup table to jointly optimize all phases and control the real traffic lights in the real scene to complete the online optimization control of traffic signals.
[0077] In an embodiment of the present invention, the low-fuel-consumption optimization control scheme for a single phase obtained through online learning is expanded to a multi-phase consisting of multiple non-conflicting phases, thereby achieving low-fuel-consumption optimization control for different types of intersections. Specifically: for each phase, the main control module selects the phase green light timing corresponding to the maximum evaluation value from the optimized lookup table based on the current state of the controlled road as the optimal phase green light timing; if the sum of the optimal phase green light timing and the preset yellow light timing time of all phases is greater than the preset maximum signal cycle time, the phase green light timing is adjusted according to the proportion of the road traffic controlled by each phase. If the adjusted phase green light timing is less than the predetermined minimum phase green light timing, no adjustment is made. After the joint optimization is completed, it is transmitted to the real traffic light and called and executed by the real traffic light.
[0078] The above-mentioned scheme provided by the embodiment of the present invention can construct a virtual scene corresponding to the real scene; and the optimization of the signal control scheme for low fuel consumption in traffic is realized based on the traffic digital twin platform, reducing the risks and losses caused by unreasonable signal schemes; because the online optimization of the signal control scheme uses the average fuel consumption of vehicles as an indicator, after multiple learning, the average fuel consumption of vehicles can be significantly reduced, and carbon emissions are also reduced. In addition, the system of the present invention interacts with the environment to learn the optimized control scheme for low fuel consumption traffic signals, which conforms to the real situation, effectively solving the energy problem while avoiding the greenhouse effect caused by excessive carbon emissions.
[0079] Through the description of the above embodiments, those skilled in the art will clearly understand that the above embodiments can be implemented through software or by using software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) and includes a number of instructions for causing a computer device (such as a personal computer, a server, or a network device) to execute the methods described in the various embodiments of the present invention.
[0080] Example 2
[0081] The present invention also provides a traffic signal online optimization control system for low fuel consumption, which is mainly implemented based on the method provided in the above embodiment, such as Figure 3 As shown, the system mainly includes:
[0082] A virtual scene construction unit, used to construct a virtual scene corresponding to the real scene using traffic digital twin software;
[0083] A phase setting unit, configured to set phases consisting of non-conflicting traffic flows in each virtual intersection of the virtual scene;
[0084] The single-phase low-fuel-consumption optimization control scheme online learning unit is used to select a reference phase and conduct online learning of the single-phase low-fuel-consumption optimization control scheme; in the current cycle of the learning process, the current state of the road controlled by the reference phase in the virtual scene is determined, and the optimal phase green light timing corresponding to the current state is found from the lookup table by searching the lookup table based on the evaluation value; the optimal phase green light timing is configured in the reference phase, and the corresponding evaluation value is updated based on the average fuel consumption of the virtual vehicle during the optimal phase green light timing; the learning is continued for multiple cycles until the evaluation values converge, an optimized lookup table is obtained, and the online learning of the single-phase low-fuel-consumption optimization control scheme is completed; wherein the state refers to a state related to fuel consumption;
[0085] The multi-phase joint optimization and control unit is used to share the optimized lookup table with other phases. The traffic digital twin software transmits it to the main control module, which then uses the optimized lookup table to jointly optimize all phases and control real traffic signals in real scenarios, completing online traffic signal optimization control.
[0086] Those skilled in the art will clearly understand that for the convenience and brevity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.
[0087] Example 3
[0088] The present invention also provides a processing device, such as Figure 4 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided by the aforementioned embodiment.
[0089] Furthermore, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.
[0090] In the embodiment of the present invention, the specific types of the memory, input device, and output device are not limited; for example:
[0091] The input device can be a touch screen, image acquisition device, physical button or mouse;
[0092] The output device may be a display terminal;
[0093] The memory may be a random access memory (RAM) or a non-volatile memory, such as a disk memory.
[0094] Example 4
[0095] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the above embodiment when the computer program is executed by a processor.
[0096] In the embodiments of the present invention, the computer-readable storage medium may be provided in the aforementioned processing device, for example, as a memory in the processing device. Alternatively, the computer-readable storage medium may be a USB flash drive, a removable hard drive, a read-only memory (ROM), a magnetic disk, or an optical disk, among other media capable of storing program code.
[0097] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A traffic signal online optimization control method for low fuel consumption, characterized in that: include: Use traffic digital twin software to build virtual scenes corresponding to real scenes; In each virtual intersection of the virtual scene, each phase consisting of non-conflicting traffic flows is set; Select a reference phase and conduct online learning of the single-phase low fuel consumption optimization control scheme; In the current cycle of the learning process, the current state of the road controlled by the reference phase in the virtual scene is determined, and the optimal phase green light timing corresponding to the current state is found from the lookup table based on the evaluation value by searching the lookup table; The optimal phase green light timing is assigned to the reference phase, and the corresponding evaluation value is updated based on the average fuel consumption of the virtual vehicle during the optimal phase green light timing. Learning is continued for multiple cycles until the evaluation values converge, obtaining an optimized lookup table, and completing the online learning of the single-phase low-fuel-consumption optimization control scheme. The state refers to a state related to fuel consumption. The optimized lookup table is shared with other phases and transmitted to the main control module by the traffic digital twin software. The main control module uses the optimized lookup table to jointly optimize all phases and control the real traffic lights in the real scene to complete the online optimization control of traffic signals. The first item in the lookup table is the state, the second item is the phase green light timing set, and the third item is the evaluation value of the combination of the first and second items; The state contains two pieces of information: one is the queue length and the other is the speed of the leading vehicle, both of which are related to fuel consumption; In the same state, the phase green light timing set consisting of multiple phase green light timings is recorded as X, and the corresponding phase green light timing set is recorded as in, Indicates the jth phase green light timing corresponding to the current state X, j∈{1,2,…,N max }, N max is the number of phase green light timings. The combination of the current state X and different phase green light timings corresponds to different evaluation values, that is, there are N max evaluation value, the current state X and The evaluation value of For the current state X, the phase green light timing corresponding to the maximum evaluation value is selected from the corresponding phase green light timing set as the optimal phase green light timing, which is recorded as The reward corresponding to the average fuel consumption calculation of the virtual vehicles during the optimal phase green light timing includes: calculating the specific power of each virtual vehicle; calculating the specific power interval VSPBin based on the calculated specific power of the virtual vehicle; calculating the standard fuel consumption corresponding to VSPBin; determining the average fuel consumption rate FR corresponding to VSPBin by querying the preset specific power and average fuel consumption rate table k , and determine the average fuel consumption rate FR0 corresponding to the benchmark specific power range, and calculate the standard fuel consumption corresponding to VSPBin: EFR k =FR k / FR0; Calculate the reward value r based on the standard fuel consumption of all virtual vehicles: Where Vnum represents the number of virtual vehicles and l represents the calculation interval.
2. The method for online optimization and control of traffic signals for low fuel consumption according to claim 1, characterized in that: The online optimization control of traffic signals is a Markov decision process that minimizes the average fuel consumption of vehicles at the intersection through the traffic control scheme and the fuel consumption relationship model; The traffic control scheme is represented by a four-tuple: (S, A, P, R), where S is the set of road states controlled by each phase, A is the set of green light timings corresponding to each phase, P is the state transition probability, that is, the probability that the road state will change to another state after the green light timing corresponding to one phase is executed; R is the reward, calculated using the average fuel consumption of virtual vehicles.
3. The method for online optimization control of traffic signals for low fuel consumption according to claim 1 or 2, characterized in that: The method of updating the corresponding evaluation value by combining the average fuel consumption of the virtual vehicle during the optimal phase green light timing includes: The optimal phase green light timing under the current state X is recorded as Execute optimal phase green light timing Then it changes to state Y and uses the best phase green light timing The reward value r corresponding to the average fuel consumption of the virtual vehicle during the period is calculated, and the updated evaluation value is calculated in combination with the reward value r. The updated evaluation value is used to update the corresponding evaluation value in the lookup table.
4. The method for online optimization and control of traffic signals for low fuel consumption according to claim 3, characterized in that: The specific power of each virtual vehicle is calculated using the following formula: VSP=v(1.1a+0.132)+0.000302v 2 Where v represents the instantaneous speed of the virtual vehicle, a represents the instantaneous acceleration of the virtual vehicle, and VSP represents the specific power of the virtual vehicle; Calculate the specific power range based on the calculated specific power of the virtual vehicle: VSPBin=Int(VSP+0.5) Wherein, VSPBin represents the specific power range.
5. The method for online optimization and control of traffic signals for low fuel consumption according to claim 3, characterized in that: The formula for calculating the updated evaluation value combined with the reward value r is: in, Indicates the optimal phase green light timing corresponding to the current state X in the current cycle. Indicates the current state X in the current cycle and The evaluation value of the arrow on the left Indicates the updated current state X and Evaluation value, Y represents the optimal phase green light timing of the current state X in the current cycle The state after, maxQ(Y,B) represents the highest evaluation value corresponding to state Y, B is the phase green light timing corresponding to maxQ(Y,B), that is, the optimal phase green light timing, γ is the discount factor, and α is the learning rate.
6. The method for online optimization and control of traffic signals for low fuel consumption according to claim 1, characterized in that: The optimized lookup table is shared with other phases and transmitted to the main control module by the traffic digital twin software. The main control module uses the optimized lookup table to jointly optimize all phases, including: For each phase, a real traffic light in a real scene selects the phase green light timing corresponding to the maximum evaluation value from the optimized lookup table according to the current state of the controlled road as the optimal phase green light timing; If the sum of the optimal phase green light timing and the preset yellow light timing for all phases is greater than the preset maximum signal cycle time, the phase green light timing is adjusted according to the proportion of road traffic controlled by each phase. If the adjusted phase green light timing is less than the preset minimum phase green light timing, no adjustment is made.
7. A traffic signal online optimization control system for low fuel consumption, characterized in that: The method according to any one of claims 1 to 6 is implemented, and the system comprises: A virtual scene construction unit, used to construct a virtual scene corresponding to the real scene using traffic digital twin software; A phase setting unit, configured to set phases consisting of non-conflicting traffic flows in each virtual intersection of the virtual scene; The single-phase low-fuel-consumption optimization control scheme online learning unit is used to select a reference phase and conduct online learning of the single-phase low-fuel-consumption optimization control scheme; during the current cycle of the learning process, the current state of the road controlled by the reference phase in the virtual scene is determined, and by searching the lookup table, the optimal phase green light timing corresponding to the current state is found from the lookup table; the optimal phase green light timing is configured in the reference phase, and the corresponding evaluation value is updated based on the average fuel consumption of the virtual vehicle during the optimal phase green light timing; the learning is continued for multiple cycles until the evaluation values converge, an optimized lookup table is obtained, and the online learning of the single-phase low-fuel-consumption optimization control scheme is completed; wherein the state refers to a state related to fuel consumption; The multi-phase joint optimization and control unit is used to share the optimized lookup table with other phases, which is transmitted to the main control module by the traffic digital twin software. The main control module uses the optimized lookup table to jointly optimize all phases and control the real traffic lights in the real scene to complete the online optimization control of traffic signals.
8. A processing device, characterized in that include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.