Intersection reinforcement learning control method and device in mixed traffic environment and medium

By introducing CAV-specific phases and lanes in a hybrid traffic environment, and using the agent in the deep reinforcement learning framework to dynamically adjust the lane and signal light status, the traffic flow coordination problem between HV and CAV is solved, and traffic efficiency and real-time response capabilities are improved.

CN120236413AActive Publication Date: 2025-07-01CHONGQING UNIV OF POSTS & TELECOMM
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510506250.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-01
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

In a hybrid traffic environment, it is difficult for the prior art to effectively coordinate the traffic flow between artificially driven vehicles (HVs) and networked autonomous vehicles (CAVs), resulting in traffic congestion and inefficiency in traffic efficiency. Especially when the penetration rate and traffic flow of different CAVs change, traditional methods cannot respond in a timely manner, and computing resources are consumed largely and real-time is poor.

Method used

CAV-specific lanes with CAV-specific phase and free lane direction are introduced, and the CAV-specific lane mode and signal light status are dynamically adjusted through three agents (Lane-Pattern Agent, Traffic-Signal Agent and CAV-Coordination Agent) in the deep reinforcement learning framework to achieve coordinated control of the hybrid traffic environment.

Benefits of technology

Dynamically optimize lane mode and signal light status under different CAV penetration rates, reduce the average delay time and average queue length of vehicles, improve traffic efficiency, and maximize overall traffic efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236413A_ABST
    Figure CN120236413A_ABST
Patent Text Reader

Abstract

The invention discloses an intersection reinforcement learning control method and device in a mixed traffic environment and a medium, a CAV special phase and a CAV special lane in a free lane direction are introduced, a CAV special lane mode is adjusted through a first intelligent agent according to CAV permeability, a second intelligent agent executes a corresponding signal lamp phase scheme according to the CAV special lane mode, and the intersection reinforcement learning control method and device in the mixed traffic environment are achieved. The third intelligent agent dynamically adjusts the state of the signal lamp according to the real-time traffic flow, and when the signal lamp is in a CAV special phase, the third intelligent agent dynamically determines the passing right of the lane according to the current CAV traffic state and releases the vehicle based on the passing right; by means of the control method, the lane mode and the signal lamp state can be dynamically optimized under different CAV permeability, meanwhile, vehicle passing is flexibly coordinated under the CAV special phase, and compared with a dynamic signal lamp and a traditional fixed signal lamp, the average delay time and the average queuing length of vehicles are shortened, and passing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent transportation, and particularly to an intersection reinforcement learning control method, device and medium in a mixed traffic environment. Background Art

[0002] The rapid growth of the number of automobiles has led to the lag in the construction of urban road networks, and the problem of traffic congestion has become increasingly serious. The economic losses caused by traffic congestion account for 20% of the disposable income of urban residents, which has become one of the key problems restricting urban development. As the core hub and traffic bottleneck area of urban traffic, the traffic demand at road intersections is higher than that of the connected roads when vehicles turn here. Vehicles in different directions may conflict, reducing the traffic capacity and further exacerbating congestion. Therefore, optimizing intersection traffic management and improving traffic capacity are crucial for alleviating traffic congestion.

[0003] Currently, traffic signal lights at intersections mainly adopt fixed-cycle control. However, the traffic flow changes dynamically, which may lead to a mismatch between the actual flow and the expected one. To solve this problem, intelligent traffic signal control strategies have emerged, which dynamically adjust the signal light states through real-time traffic flow information to improve traffic efficiency.

[0004] In addition, in a pure connected and autonomous vehicle (CAV) environment, the control strategy for signal-free intersections realizes vehicle coordination by means of vehicle-to-everything (V2X) technology, significantly improving traffic efficiency. In recent years, the development of autonomous driving technology has been rapid. For example, Baidu's "Apollo Go" has carried out tests in many places, and multiple intelligent connected vehicle test demonstration areas and pilot cities have been established across the country, providing support for the implementation of the control strategy for signal-free intersections.

[0005] However, for a long time to come, human-driven vehicles (HV) and CAVs will coexist in a mixed traffic environment. Although the control strategy for signal-free intersections can improve efficiency, due to the lack of signal light guidance, the behavior of HVs is difficult to predict and the safety cannot be guaranteed.

[0006] Therefore, most studies still rely on signal lights to manage CAVs and HVs. The conventional method is to let CAVs and HVs share lanes, and improve efficiency by adjusting the CAV trajectories and signal light states. However, this method cannot fully exert the potential of CAVs and is also difficult to overcome the uncertainty risks of HVs. The control strategy based on dedicated lanes for CAVs physically isolates CAVs and HVs, reduces interaction conflicts, enables CAVs to operate more efficiently, and optimizes the traffic efficiency at intersections.

[0007] Existing research on CAV (Connected and Automated Vehicle) dedicated lanes mostly focuses on static lane allocation. However, under different traffic flows and CAV penetration rates, it may lead to waste of lane resources or exacerbation of congestion. Some scholars have started research on dynamic dedicated lane control. For example, some research has proposed that left-turning and straight-going CAVs use independent dedicated lanes during their respective phases. However, this separated design requires more lane infrastructure, and when the proportion of oncoming vehicles is unbalanced, the utilization rate of some lanes is insufficient, resulting in waste of resources. Other research has proposed a dynamic allocation method where left-turning and straight-going CAVs share the dedicated lane. However, when the traffic flow fluctuates greatly, the lane function switches frequently, which may reduce the overall traffic efficiency.

[0008] In addition, with the development of artificial intelligence, machine learning has begun to be applied to intelligent transportation systems. As the learning system closest to the human brain, deep reinforcement learning (DRL) can efficiently solve complex decision-making problems. However, its application in mixed traffic environments is still relatively rare, and most research relies on optimization-based algorithms. These traditional methods have theoretical advantages, but they require a large amount of computing resources to solve the global optimal solution, which may lead to computational delays and inability to respond promptly to changes in traffic conditions. In practical applications, with the increase in traffic flow and the improvement of environmental complexity, it is difficult to ensure real-time performance. Summary of the Invention

[0009] Based on the problems proposed in the above background technology, the purpose of the present invention is to provide an intersection reinforcement learning control method, device, and medium in a mixed traffic environment. By introducing a CAV dedicated phase and a CAV dedicated lane in the free lane direction, the conflict between CAVs and HVs (Human-driven Vehicles) is reduced. The mode of the CAV dedicated lane is dynamically adjusted according to the current CAV penetration rate, and then the state of the traffic signal is dynamically adjusted according to the current traffic flow situation. The control method is updated and trained based on the interaction between the intersection and the deep reinforcement learning framework in the mixed traffic environment, so as to solve the intersection coordination problem in the mixed traffic environment.

[0010] The present invention is realized through the following technical solutions:

[0011] The first aspect of the present invention provides an intersection reinforcement learning control method in a mixed traffic environment, including the following steps:

[0012] Set the intersection in the mixed traffic environment as the environment; wherein, the environment includes lanes, and the lanes are divided into ordinary lanes and CAV dedicated lanes;

[0013] Construct a first intelligent agent, a second intelligent agent, and a third intelligent agent;

[0014] The first intelligent agent perceives the state of the environment and executes the ε-greedy policy to select the CAV dedicated lane mode;

[0015] The second agent senses the state of the environment affected by the CAV dedicated lane mode and executes the ε-greedy policy to select a phase;

[0016] The third agent senses the state of the environment in the phase and executes the ε-greedy policy to allocate the right of way to the vehicles on the CAV dedicated lane.

[0017] In the above technical solution, the intersection in the mixed traffic environment is set as the environment. In the CAV and HV mixed traffic environment, this method divides the lanes into ordinary lanes and CAV dedicated lanes, and isolates the mutual influence between CAV and HV by introducing CAV dedicated lanes; the CAV dedicated lanes are set as free lane directions, that is, right turn, straight and left turn are allowed; the ordinary lanes are set as fixed lane modes, that is, left-turn HV can only pass through the intersection from the fixed left-turn ordinary lane. On this basis, a CAV dedicated phase is provided for the CAV dedicated lanes. In the CAV dedicated phase, this method is used to freely regulate the passage of CAVs through the intersection; in other phases, signal lights are used to separate conflicting vehicles, and the conflicting HV traffic flows are allocated to different phases, and HVs completely follow the traditional method and pass through the intersection according to the signal light instructions.

[0018] The deep reinforcement learning framework includes three agents: the first agent (Lane-Pattern Agent), the second agent (Traffic-Signal Agent), and the third agent (CAV-Coordination Agent). The first agent, the second agent, and the third agent are dynamically coupled with the environment in sequence.

[0019] First, the first agent is used to adjust the CAV dedicated lane mode according to the CAV penetration rate. Among them, the first agent senses the quantity ratio of CAVs and HVs in the environment, determines the action by executing the ε-greedy policy - selects the CAV dedicated lane mode, and executes the action to apply the selected CAV dedicated lane mode to the environment.

[0020] Second, the second agent is used to dynamically adjust the signal light phase and determine the phase according to the CAV dedicated lane mode and the real-time traffic flow. Among them, the second agent senses the state of the environment after applying the selected CAV dedicated lane mode to the environment, determines the action by executing the ε-greedy policy - selects the phase, and executes the action to apply the selected phase to the environment.

[0021] Finally, the third intelligent agent is used to allocate the right of way for the vehicles on the CAV dedicated lane at the selected phase. The third intelligent agent senses the state of the environment after applying the selected phase to the environment, determines the action of allocating the right of way for the vehicles on the CAV dedicated lane by implementing the ε-greedy strategy, and regulates the passing of CAVs through the intersection based on the allocated right of way for the vehicles, so as to achieve the control of the mixed traffic flow.

[0022] In an alternative embodiment, the first intelligent agent performs the following steps:

[0023] Obtain the number of CAVs and the number of HVs in the environment at time t, and form the ratio of the number of CAVs to the number of HVs as the state s at time t t ;

[0024] Take the selection of the CAV lane mode as the action, and determine the action a at the next time t of the state s through the ε-greedy strategy t ; t ;

[0025] Execute the action a t , and obtain the number of vehicles passing through after executing the action a t , and calculate the reward r at time t according to the number of vehicles passing through t .

[0026] In an alternative embodiment, after the first intelligent agent finishes execution, it further includes: integrating the state s t , the action a t , the reward r t and the state s of the environment at time t + 1 t+1 into an experience and putting it into the experience replay pool, and training the first intelligent agent using the experience in the experience replay pool.

[0027] In an alternative embodiment, the second intelligent agent performs the following steps:

[0028] Obtain the queue lengths of the lanes corresponding to each signal light phase in the environment at time t, and take the queue lengths of the lanes corresponding to each signal light phase as the state s at time t t ;

[0029] Take the selection of the phase of the CAV dedicated lane mode as the action, and determine the action a at the next time t of the state s through the ε-greedy strategy t ; t ;

[0030] Execute the action a t , and obtain the number of vehicles passing through after executing the action a t , and calculate the reward r at time t according to the number of vehicles passing throught .

[0031] In an alternative embodiment, after the second agent finishes execution, it further includes: integrating the state s t , the action a t , the reward r t and the state s of the environment at time t + 1 t+1 into an experience and putting it into the experience replay pool, and training the second agent using the experience in the experience replay pool.

[0032] In an alternative embodiment, the third agent performs the following steps:

[0033] Obtain the traffic state in the environment at time t, and use the traffic state as the state s at time t t ; wherein, the traffic state is encoded using discrete traffic state encoding, including a position matrix, a speed matrix, a path matrix, and a release priority matrix;

[0034] Take the allocation of the right of way of the vehicles on the dedicated lane for CAVs as an action, and determine the action a at time t in the next time step for the state s t through an ε-greedy policy; t ;

[0035] Execute the action a t , and obtain the vehicle passing data, queue length, and release priority after executing the action a t , and calculate the reward r at time t according to the number of vehicle passes, the queue length, and the release priority t .

[0036] In an alternative embodiment, calculate the reward r at time t according to the vehicle passing data, the queue length, and the release priority t , and the calculation process is as follows:

[0037]

[0038] In the above formula, n t is the number of vehicle passes at time t, M is the number of boom arms, N is the number of lanes, q t is the queue length at time t, f t is the release priority at time t, and α1, α2, and α3 are weights.

[0039] In an alternative embodiment, after the third agent finishes execution, it further includes: integrating the state s t , the action a t , the reward r t and the state s of the environment at time t + 1 t+1Integrate the experience into the experience replay pool, and use the experience in the experience replay pool to train the third agent.

[0040] The second aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements an intersection reinforcement learning control method in a mixed traffic environment.

[0041] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements an intersection reinforcement learning control method in a mixed traffic environment.

[0042] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0043] The present invention introduces a CAV dedicated phase and a CAV dedicated lane in the free lane direction. The first agent adjusts the CAV dedicated lane mode according to the CAV penetration rate, the second agent executes the corresponding signal light phase scheme according to the CAV dedicated lane mode, and dynamically adjusts the signal light state according to the real-time traffic flow. When the signal light is in the CAV dedicated phase, the third agent dynamically determines the right of way of the lane according to the current CAV traffic state and releases the vehicle based on the right of way. Through this control method, it is possible to dynamically optimize the lane mode and signal light state under different CAV penetration rates, and at the same time flexibly coordinate vehicle passing under the CAV dedicated phase, thereby maximizing the overall traffic efficiency and reducing intersection delays. Compared with dynamic signal lights and traditional fixed signal lights, it reduces the average delay time and average queue length of vehicles and increases the passing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings. In the drawings:

[0045] Figure 1 It is the overall architecture diagram of the intersection reinforcement learning control method in a mixed traffic environment provided by Embodiment 1 of the present invention;

[0046] Figure 2 It is the structural schematic diagram of the intersection in a mixed traffic environment provided by Embodiment 1 of the present invention;

[0047] Figure 3 It is the schematic diagram of the lane direction provided by Embodiment 1 of the present invention;

[0048] Figure 4 Schematic diagram of the agent training architecture provided in Embodiment 1 of the present invention;

[0049] Figure 5 Schematic diagram of the CAV dedicated lane mode provided in Embodiment 1 of the present invention;

[0050] Figure 6 Schematic diagram of the structure of an electronic device provided in Embodiment 2 of the present invention. Detailed implementation manners

[0051] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the embodiments and the accompanying drawings. The illustrative embodiments of the present invention and their descriptions are only used to explain the present invention and are not intended to limit the present invention.

[0052] Embodiment 1

[0053] Figure 1 Overall architecture diagram of the intersection reinforcement learning control method in a mixed traffic environment provided in Embodiment 1 of the present invention, as Figure 1 shown. The intersection reinforcement learning control method in a mixed traffic environment includes the following steps:

[0054] Set the intersection in the mixed traffic environment as the environment; wherein, the environment includes lanes, and the lanes are divided into ordinary lanes and CAV dedicated lanes;

[0055] Construct a first agent, a second agent and a third agent;

[0056] The first agent selects the CAV dedicated lane mode by perceiving the state of the environment and executing the ε-greedy policy;

[0057] The second agent selects the phase by perceiving the state of the environment affected by the CAV dedicated lane mode and executing the ε-greedy policy;

[0058] The third agent distributes the right of way of the vehicles on the CAV dedicated lane by perceiving the state of the environment under the phase and executing the ε-greedy policy.

[0059] It should be noted that the present invention includes two parts: an environment and a deep reinforcement learning framework. By the interaction between the environment and the deep reinforcement learning framework, a control method is generated to solve the intersection coordination problem in a mixed traffic environment.

[0060] Among them, the intersection in the mixed traffic environment is set as the environment, as Figure 2As shown, in the CAV-HV mixed traffic environment, this method divides the lanes into general lanes and CAV-exclusive lanes to isolate the mutual influence between CAVs and HVs by introducing CAV-exclusive lanes; as Figure 3 shown, the CAV-exclusive lanes are set as free lane directions, allowing right turns, straight-throughs, and left turns; the general lanes are set as fixed lane modes, that is, left-turning HVs can only pass through the intersection from the fixed left-turn general lanes. On this basis, a CAV-exclusive phase is provided for the CAV-exclusive lanes. Within the CAV-exclusive phase, this method is used to freely regulate the passage of CAVs through the intersection; in other phases, signal lights are used to separate conflicting vehicles, and the conflicting HV traffic flows are assigned to different phases. HVs completely follow the traditional method and pass through the intersection according to the signal light indications.

[0061] The deep reinforcement learning framework includes three agents: the first agent (Lane-Pattern Agent), the second agent (Traffic-Signal Agent), and the third agent (CAV-Coordination Agent). The first agent, the second agent, and the third agent are dynamically coupled with the environment in sequence.

[0062] First, the first agent is used to adjust the CAV-exclusive lane mode according to the CAV penetration rate. Among them, the first agent senses the quantity ratio of CAVs and HVs in the environment, determines the action by executing the ε-greedy policy - selects the CAV-exclusive lane mode, and executes the action to apply the selected CAV-exclusive lane mode to the environment.

[0063] Second, the second agent is used to dynamically adjust the signal light phase based on the CAV-exclusive lane mode and the real-time traffic flow. Among them, the second agent senses the state of the environment after applying the selected CAV-exclusive lane mode to the environment, determines the action by executing the ε-greedy policy - selects the phase, and executes the action to apply the selected phase to the environment.

[0064] Finally, the third agent is used to allocate the right of way for the vehicles on the CAV-exclusive lane under the selected phase. Among them, the third agent senses the state of the environment after applying the selected phase to the environment, determines the action by executing the ε-greedy policy - allocates the right of way for the vehicles on the CAV-exclusive lane, and regulates the passage of CAVs through the intersection based on the allocated vehicle right of way to achieve the control of the mixed traffic flow.

[0065] By introducing the CAV-exclusive phase and CAV-exclusive lanes, it is possible to dynamically adjust the exclusive lane mode and signal light status according to the real-time traffic conditions and CAV penetration rate, maximize the CAV traffic efficiency while ensuring the HV traffic efficiency, thereby significantly reducing the average delay time and average queue length of vehicles at intersections in the mixed traffic environment.

[0066] Further, the first agent, the second agent, and the third agent are constructed and trained based on the DDQN network structure. The training processes of the first agent, the second agent, and the third agent based on the DDQN network structure are as follows Figure 4 shown. The agent obtains the state from the environment, determines the action by executing the ε-greedy policy, executes the action to obtain the reward, the environment transitions to the next state, and the state, action, reward, and next state are saved in the experience replay pool for training the agent. In this embodiment, the state obtained by the first agent, the second agent, and the third agent from the environment, the action determined based on the state, the reward obtained after executing the action, and the next state obtained after the environment transition are all determined based on their implemented functions.

[0067] In an alternative embodiment, for the first agent, its implemented function is to adjust the CAV dedicated lane mode according to the CAV penetration rate.

[0068] In this embodiment, the CAV penetration rate is composed of the ratio of the number of CAVs to the number of HVs. Therefore, the first agent obtains the number of CAVs and the number of HVs in the environment at time t, calculates the ratio of the number of CAVs to the number of HVs, and takes the ratio as the state s at time t t , where the state s at time t t is expressed as follows:

[0069]

[0070] In the above formula, is the ratio composition of CAVs, is the ratio composition of HVs.

[0071] Taking the selection of the CAV lane mode as the action, the action a at time t under the state s is determined through the ε-greedy policy t , and the action a t is expressed as follows: t

[0072] a t ∈A = {1, 2, 3}

[0073] It should be noted that A is the set of CAV dedicated lane modes, which includes a total of three CAV dedicated lane modes, represented by 1, 2, and 3 respectively. In this embodiment, taking a three-lane scenario as an example, the CAV dedicated lane modes are as follows Figure 5As shown, the first CAV dedicated lane mode does not contain a CAV dedicated lane; the second CAV dedicated lane mode includes a CAV dedicated lane, and the CAV dedicated lane is located in the middle; the third CAV dedicated lane mode includes two CAV dedicated lanes, and the two CAV dedicated lanes are adjacent. .

[0074] The first agent adopts the ε-greedy strategy π(a t |s t ) Determine the state s t The next action a at time t t That is, a selection action of selecting from the above three CAV-only lane modes is constituted based on the ratio of the number of CAV vehicles to the number of HV vehicles at time t.

[0075] Execute action a at time t t That is, the environment is set according to the CAV dedicated lane mode, the number of vehicles passing in the selected CAV dedicated lane mode is obtained, and the reward r at time t is calculated based on the number of vehicles passing in the selected CAV dedicated lane mode. t , reward r t The calculation is as follows:

[0076] r t =n t+T -n t

[0077] In the above formula, n t+T is the number of vehicles passing at time t+T, n t is the number of vehicles passing at time t, and T represents the length of time the lane mode is executed.

[0078] At this point, the environment transitions to the next state s t+1 ∈S; the state s t 、Action a t , Reward t and the state s of the environment at time t+1 t+1 Integrate into experience t ,a t ,r t ,s t+1 ) is put into the experience replay pool, and the experience in the experience replay pool is used to train the first agent.

[0079] In an optional embodiment, the second intelligent agent monitors the environment after the first intelligent agent selects the CAV dedicated lane mode and applies it to the environment, and the function it implements is to dynamically adjust the phase of the traffic light and determine the phase based on the CAV dedicated lane mode and real-time traffic flow.

[0080] In this embodiment, the second agent obtains the queue length of the lane corresponding to each signal light phase in the environment at time t, and uses the queue length of the lane corresponding to each signal light phase as the state s at time t. t , state s t It is expressed as follows:

[0081] s t = {q i,t},i∈{1,2,...,I}

[0082] In the above formula, q i,t is the queue length of the lane corresponding to the i-th signal light phase at time t, and I is the total number of signal light phases.

[0083] The phase of the CAV dedicated lane mode is selected as an action, and the state s is determined by the ε-greedy strategy. t The next action a at time t t , action a t It is expressed as follows:

[0084] a t ∈A={1,2,...,I}

[0085] Among them, the second agent uses the ε-greedy strategy π(a t |s t ) Determine the state s t The next action a at time t t , that is, a phase is selected from I signal light phases based on the queue length of the lane corresponding to the signal light phase.

[0086] Execute action a at time t t , that is, applying the selected phase to the environment,

[0087] Obtain the number of vehicles passing in the selected CAV lane mode, and calculate the reward r at time t based on the number of vehicles passing in the selected CAV lane mode t , reward r t The calculation is as follows:

[0088] r t =n t+τ -n t

[0089] In the above formula, n t+τ is the number of vehicles passing at time t+τ, and τ is the duration of a phase.

[0090] At this point, the environment transitions to the next state s t+1 ∈S; the state s t 、Action a t , Reward tand the state s of the environment at time t+1 t+1 Integrate into the experience (s t , a t , r t , s t+1 ) and put it into the experience replay pool, and use the experience in the experience replay pool to train the second intelligent agent.

[0091] The second intelligent agent executes the corresponding signal light phase plan according to the CAV dedicated lane mode, and dynamically adjusts the signal light state according to the real-time traffic flow.

[0092] In an alternative embodiment, for the third intelligent agent, it allocates the right of way to the vehicles on the CAV dedicated lane under the CAV dedicated phase after the second intelligent agent selects the CAV dedicated phase of the CAV dedicated lane mode.

[0093] In this embodiment, the third intelligent agent obtains the traffic state in the environment at time t, and takes the traffic state as the state s at time t t , the state s t is represented as follows:

[0094] s t = {W t , V t , P t , F t}

[0095] In the above formula, W t is the position matrix at time t, V t is the speed matrix at time t, P t is the path matrix at time t, F t is the release priority matrix at time t.

[0096] It should be noted that the traffic state in this embodiment adopts the discrete traffic state encoding (DTSE), which is composed of four matrices: the position matrix, the speed matrix, the path matrix, and the release priority matrix. Each matrix is an X·X matrix, and the X·X matrix represents the conflict area in the intersection under the mixed traffic environment. Therefore, the traffic state obtained in this embodiment is the traffic state of the conflict area in the intersection under the mixed traffic environment. Each position in the matrix corresponds to the grid position in the intersection. If there is a vehicle at the corresponding position, each matrix inputs the corresponding information. If there is no vehicle, it is 0. The release priority corresponds to the delay time of the vehicle. The greater the delay time, the higher the priority.

[0097] Taking the allocation of the right of way for the vehicles on the CAV dedicated lane as an action, in this embodiment, an M×N binary vector is used to represent the allocation action of the right of way for the vehicles on the CAV dedicated lane, and the action a at time t under the state s t is determined by the ε-greedy strategyt , action a t is expressed as follows:

[0098] a t = {h 1,1,t ,..., h m,n,t ,..., h M,N,t}

[0099] In the above formula, M represents the number of vehicle arms, N represents the number of lanes, and h m,n,t represents the right of way on the nth lane of the mth vehicle arm at time t.

[0100] Execute action a t , and obtain the vehicle passing data, queue length, and release priority after executing action a t . Calculate the reward r at time t by comprehensively considering the three factors of the number of vehicle passages, queue length, and release priority t , the reward r t is expressed as follows:

[0101]

[0102] In the above formula, n t is the number of vehicle passages at time t, M is the number of vehicle arms, N is the number of lanes, q t is the queue length at time t, f t is the release priority at time t, and α1, α2, and α3 are weights.

[0103] At this time, the environment transitions to the next state s t+1 ∈S; Integrate the state s t , action a t , reward r t and the state s of the environment at time t + 1 t+1 into an experience (s t , a t , r t , s t+1 ) and put it into the experience replay pool, and use the experiences in the experience replay pool to train the third intelligent agent.

[0104] In summary, the first intelligent agent adjusts the CAV dedicated lane mode according to the CAV penetration rate, the second intelligent agent executes the corresponding signal phase plan according to the CAV dedicated lane mode, and dynamically adjusts the signal state according to the real-time traffic flow. When the signal is in the CAV dedicated phase, the third intelligent agent dynamically determines the right of way of the lane according to the current CAV traffic state and releases the vehicle based on the right of way; Through this control method, it is possible to dynamically optimize the lane mode and signal state under different CAV penetration rates, and at the same time flexibly coordinate vehicle passages in the CAV dedicated phase, thereby maximizing the overall traffic efficiency and reducing intersection delays.

[0105] Example 2

[0106] Figure 6 As shown in the schematic structural diagram of an electronic device provided in Example 2 of the present invention, Figure 6 the electronic device includes a processor 21, a memory 22, an input device 23, and an output device 24; the number of processors 21 in the computer device can be one or more, Figure 6 taking one processor 21 as an example; the processor 21, the memory 22, the input device 23, and the output device 24 in the electronic device can be connected by a bus or other means, Figure 6 taking connection by bus as an example.

[0107] The memory 22, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules. The processor 21 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 22, that is, implements the intersection reinforcement learning control method in the mixed traffic environment of Example 1.

[0108] The memory 22 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory 22 can include high-speed random access memory, and can also include non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 22 can further include a memory remotely set relative to the processor 21, and these remote memories can be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0109] The input device 23 can be used to receive user input such as an id and a password. The output device 24 is used to output a network configuration page.

[0110] Example 3

[0111] The present invention also provides a computer-readable storage medium in Example 3, and the computer-executable instructions are used to implement the intersection reinforcement learning control method in the mixed traffic environment provided in Example 1 when executed by a computer processor.

[0112] A storage medium containing computer-executable instructions provided in the embodiments of the present invention, the computer-executable instructions are not limited to the method operations provided in Example 1, and can also execute related operations in the intersection reinforcement learning control method in the mixed traffic environment provided in any embodiment of the present invention.

[0113] The specific embodiments described above further elaborate on the object, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A reinforcement learning control method for intersections in a mixed traffic environment, characterized in that: The steps include: An intersection in a mixed traffic environment is set as an environment; wherein the environment includes lanes, and the lanes are divided into ordinary lanes and CAV-dedicated lanes; Constructing a first agent, a second agent, and a third agent; The first agent senses the state of the environment and executes an ε-greedy strategy to select a CAV dedicated lane mode; The second agent senses the state of the environment affected by the CAV dedicated lane mode and selects a phase by executing an ε-greedy strategy; The third agent senses the state of the environment at the phase and executes an ε-greedy strategy to allocate the right of way to vehicles on the CAV dedicated lane.

2. The intersection reinforcement learning control method in a mixed traffic environment according to claim 1 is characterized in that: The first agent performs the following steps: Obtain the number of CAV vehicles and HV vehicles in the environment at time t, and use the ratio of the number of CAV vehicles to the number of HV vehicles as the state s at time t t ; The CAV lane mode is selected as the action and the state s is determined by the ε-greedy strategy. t The next action a at time t t ; Perform the action a t , and obtain the execution of the action a t The reward r at time t is calculated based on the number of vehicles passing through after t .

3. The intersection reinforcement learning control method in a mixed traffic environment according to claim 2 is characterized in that: After the first agent completes execution, the method further includes: t , the action a t , the reward r t and the state s of the environment at time t+1 t+1 The experience is integrated into an experience replay pool, and the first agent is trained using the experience in the experience replay pool.

4. The intersection reinforcement learning control method in a mixed traffic environment according to claim 1 is characterized in that: The second agent performs the following steps: Get the queue length of the lane corresponding to each signal light phase in the environment at time t, and use the queue length of the lane corresponding to each signal light phase as the state s at time t t ; The phase is selected as the action and the state s is determined by the ε-greedy strategy. t The next action a at time t t ; Perform the action a t , and obtain the execution of the action a t The reward r at time t is calculated based on the number of vehicles passing through after t .

5. The intersection reinforcement learning control method in a mixed traffic environment according to claim 4 is characterized in that: After the second agent completes execution, the process further includes: t , the action a t , the reward r t and the state s of the environment at time t+1 t+1 The experience is integrated into an experience replay pool, and the second agent is trained using the experience in the experience replay pool.

6. The intersection reinforcement learning control method in a mixed traffic environment according to claim 1 is characterized in that: The third agent performs the following steps: Obtain the traffic state in the environment at time t, and use the traffic state as the state s at time t t ; Wherein, the traffic state adopts discrete traffic state coding, including position matrix, speed matrix, path matrix and release priority matrix; Assigning the right of way to the vehicle on the CAV dedicated lane is taken as an action, and the state s is determined by the ε-greedy strategy. t The next action a at time t t ; Perform the action a t , and obtain the execution of the action a t The reward r at time t is calculated based on the number of vehicles passing, the queue length and the release priority. t .

7. The intersection reinforcement learning control method in a mixed traffic environment according to claim 6 is characterized in that: The reward r at time t is calculated based on the vehicle traffic data, the queue length and the release priority. t , the calculation process is as follows: In the above formula, n t is the number of vehicles passing at time t, M is the number of arms, N is the number of lanes, q t is the queue length at time t, f t is the release priority at time t, and α1, α2, and α3 are weights.

8. The method for controlling intersections by reinforcement learning in a mixed traffic environment according to claim 6, characterized in that: After the third agent completes execution, it also includes: t , the action a t , the reward r t and the state s of the environment at time t+1 t+1 The experience is integrated into an experience replay pool, and the experience in the experience replay pool is used to train the third agent.

9. An electronic device, characterized in that: It comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the intersection reinforcement learning control method in a mixed traffic environment described in any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the intersection reinforcement learning control method in a mixed traffic environment described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Method for setting special phase of intersection automatic driving vehicle

    CN112071074A

  • Traffic capacity estimation method and device considering intelligent network connection vehicle CAV special lane

    CN113382064A

  • Bus priority traffic signal cooperative control method based on multi-agent deep reinforcement learning

    CN117746654A

  • Signal management and control method based on vehicle infrastructure cooperation, and related apparatus and program product

    WO2023246066A1