A method, device, equipment and medium for traffic control of a merging area of an expressway
By establishing a microscopic simulation environment in the merging zone of expressways and using reinforcement learning algorithms to optimize speed limits and signal control, the problems of traffic congestion and pollution in the mixed traffic environment of CAV and HV were solved, achieving efficient traffic flow management and low emissions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- AUTOMOBILE RES INST OF TSINGHUA UNIV IN SUZHOU XIANGCHENG
- Filing Date
- 2023-10-10
- Publication Date
- 2026-04-10
AI Technical Summary
In the context of mixed traffic flow of CAV and HV, the problems of traffic congestion, energy consumption and pollution in the merging zone of expressways are difficult to solve effectively with existing technologies, especially under different traffic flow conditions and the penetration rate of intelligent connected vehicles, there is a lack of adaptive control strategies.
By establishing a micro-simulation environment, constructing a state space, joint action space, and reward function, and using reinforcement learning algorithms to determine the target cooperative control model, the control strategies for variable speed limit zones and ramps are adjusted, including speed limit values and traffic light cycles, to optimize traffic flow density and intelligent connected vehicle penetration rate.
To reduce traffic pressure at expressway merging zones, improve traffic efficiency, reduce pre-set gas emissions, and optimize traffic flow.
Smart Images

Figure CN117315956B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of traffic management and control, and particularly relates to a method and device for controlling traffic at a merging area of an expressway, equipment and a medium. BACKGROUND
[0002] Under the background of rapid development of urban social economy and motorization, with the increasing number of private cars, the traffic problems caused by the private cars are hindering the development of urban economy. Especially at the merging area of the expressway, the ramp vehicles merge into the main line and compete with the main line vehicles for the traffic space, which frequently interferes with the main line traffic flow and causes traffic oscillation, energy consumption, pollution increase, traffic accidents and frequent congestion.
[0003] At present, the development of CAV (Connected and Autonomous Vehicle, intelligent connected vehicle) technology provides a new idea and broad prospects for alleviating traffic congestion at the merging area. Compared with HV (Human-driving Vehicles, human-driven vehicles), CAV technology can obtain more accurate and comprehensive traffic information through V2X (Vehicle to X, vehicle wireless communication) technology, and can more accurately and timely execute dynamic driving tasks.
[0004] Under the mixed driving environment of CAV and HV, the control strategy for improving the traffic capacity of the merging area by combining intelligent algorithms such as reinforcement learning has been relatively mature, and the relationship between the total travel time and the preset gas emission under different traffic flow states and different intelligent connected vehicle penetration rates, how to provide a technical solution that can adaptively adjust the control strategy according to the vehicle, is a technical problem that needs to be solved by the person skilled in the art. SUMMARY
[0005] The present application provides a method and device for controlling traffic at a merging area of an expressway, equipment and a medium, which reduces the traffic pressure at the merging area of the expressway and improves the traffic efficiency at the merging area of the expressway by controlling the variable speed limit and the ramp.
[0006] According to an aspect of the present application, a method for controlling traffic at a merging area of an expressway is provided, which comprises:
[0007] According to the actual road structure and historical traffic data of the target merging area of the expressway, a micro-simulation environment of the target merging area of the expressway is established; wherein the road structure at least includes a variable speed limit area, an acceleration area, an entrance ramp and a bottleneck area;
[0008] construct a state space, a joint action space and a reward function; wherein the reward function comprises at least one of a first reward function, a second reward function and a penalty function; the first reward function is used to feedback total travel time, and the second reward function is used to feedback a preset gas emission amount;
[0009] Based on the microscopic simulation environment, reinforcement learning is performed on the initial cooperative control model of the target expressway merging area according to the state space, the joint action space and the reward function, to determine a target cooperative control model;
[0010] Based on the target cooperative control model, a target speed limit value of the variable speed limit area and a target passing time length in a signal lamp cycle of the entrance ramp are determined according to the obtained real-time traffic flow density of the target expressway merging area, real-time penetration rate of intelligent networked vehicles and a target weight coefficient; wherein the target weight coefficient is used to determine the attention degree to the first reward function and the second reward function.
[0011] According to another aspect of the present application, a passing control device for an expressway merging area is provided, which comprises:
[0012] A simulation environment establishing module is configured to establish a microscopic simulation environment of the target expressway merging area according to actual road structure and historical traffic data of the target expressway merging area; wherein the road structure at least comprises a variable speed limit area, an acceleration area, an entrance ramp and a bottleneck area;
[0013] A simulation parameter constructing module is configured to construct a state space, a joint action space and a reward function; wherein the reward function comprises at least one of a first reward function, a second reward function and a penalty function; the first reward function is used to feedback total travel time, and the second reward function is used to feedback a preset gas emission amount;
[0014] A control model determining module is configured to, based on the microscopic simulation environment, perform reinforcement learning on an initial cooperative control model of the target expressway merging area according to the state space, the joint action space and the reward function, to determine a target cooperative control model;
[0015] A passing scheme determining module is configured to, based on the target cooperative control model, determine a target speed limit value of the variable speed limit area and a target passing time length in a signal lamp cycle of the entrance ramp according to the obtained real-time traffic flow density of the target expressway merging area, real-time penetration rate of intelligent networked vehicles and a target weight coefficient; wherein the target weight coefficient is used to determine the attention degree to the first reward function and the second reward function.
[0016] According to another aspect of the present application, an electronic device is provided, the device comprising:
[0017] at least one processor; and
[0018] a memory in communication with the at least one processor; wherein
[0019] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the traffic control method of the expressway merging area according to any one of the embodiments of the present application.
[0020] According to another aspect of the present application, a computer readable storage medium is provided, the computer readable storage medium stores computer instructions for enabling a processor to implement the traffic control method of the expressway merging area according to any one of the embodiments of the present application when executed by the processor.
[0021] The technical solution provided by the present application establishes a micro-simulation environment of the target expressway merging area according to the actual road structure and historical traffic data of the target expressway merging area; constructs a state space, a joint action space and a reward function; based on the micro-simulation environment, according to the state space, the joint action space and the reward function, the initial cooperative control model of the target expressway merging area is subjected to reinforcement learning to determine the target cooperative control model; based on the target cooperative control model, according to the real-time traffic flow density of the target expressway merging area, the real-time penetration rate of intelligent networked vehicles and the target weight coefficient, the target speed limit value of the variable speed limit zone and the target passing time in the signal lamp cycle of the entrance ramp are determined. The technical solution reduces the traffic pressure of the expressway merging area and improves the passing efficiency of the expressway merging area by controlling the variable speed limit and the ramp.
[0022] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0024] Figure 1 A flowchart of a traffic control method of an expressway merging area provided by Embodiment One of the present application;
[0025] Figure 2 A flow chart of a traffic control method for a merging area of an expressway is provided for Embodiment Two of the present application.
[0026] Figure 3 A structural schematic diagram of a traffic control device for a merging area of an expressway is provided for Embodiment Three of the present application.
[0027] Figure 4 A structural schematic diagram of a device for implementing a traffic control method for a merging area of an expressway is provided for Embodiment Four of the present application. DETAILED DESCRIPTION
[0028] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of the present application.
[0029] It should be noted that the terms "first", "second", "target", "candidate", "initial", "history", "simulation" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0030] Embodiment One
[0031] Figure 1 A flow chart of a traffic control method for a merging area of an expressway is provided for Embodiment One of the present application. This embodiment can be applicable to the case of controlling the traffic scheme for a merging area of an expressway. The method can be performed by a traffic control device for a merging area of an expressway, which can be realized in the form of hardware and / or software, and can be configured in a device with data processing capability. As shown in the figure, the method comprises: Figure 1
[0032] S110, a micro-simulation environment of the target expressway merging area is established according to actual road structure and historical traffic data of the target expressway merging area; wherein the road structure at least includes a variable speed limit area, an acceleration area, an entrance ramp and a bottleneck area.
[0033] The expressway merging area can be a connection section of the expressway main road and the ramp, including the entrance ramp, the acceleration area, the variable speed limit area and the bottleneck area. The ramp can be an auxiliary road for vehicles to enter or exit the main line lane of the expressway main road. The bottleneck area can be a vehicle convergence area of the expressway main line lane direction and the ramp direction. The variable speed limit area can be a region where upstream vehicles of the expressway main line lane come from. The acceleration area can be a downstream region of the variable speed limit area in the expressway main line lane, which provides additional space between the variable speed limit area and the bottleneck area to prevent congestion from spreading to the variable speed limit area and affecting the speed limit control of the variable speed limit area.
[0034] The actual road structure can include the distribution, number, length and structure of the lanes in the target expressway merging area. For example, the actual road structure of a certain expressway merging area is 3 main line lanes and 1 ramp, wherein the variable speed limit area is 400 meters, the acceleration area is 361 meters, the entrance ramp is 153 meters, and the bottleneck area is 95 meters.
[0035] The historical traffic data can be vehicle data collected at each time node in a preset time period in the target expressway merging area. Specifically, the vehicle traffic volume at each time can be obtained through image collectors respectively installed in the variable speed limit area, the acceleration area, the entrance ramp and the bottleneck area, or the vehicle data at each time can be obtained through coil detectors respectively installed in the variable speed limit area, the acceleration area, the entrance ramp and the bottleneck area.
[0036] In this scheme, based on the SUMO (Simulation of Urban Mobility, urban traffic simulation) platform, the road network file, the traffic flow file and the simulation file can be determined according to the actual road structure and the historical traffic data of the target expressway merging area, so as to establish the micro-simulation environment of the target expressway merging area.
[0037] Optionally, the micro-simulation environment is established, including:
[0038] The actual road structure, the historical traffic flow data and the initial vehicle following and lane changing model of the target expressway merging area are obtained.
[0039] The simulation road structure of the target expressway merging area is determined according to the actual road structure.
[0040] The parameters of the initial vehicle following and lane changing model are calibrated to obtain a calibrated vehicle following and lane changing model.
[0041] According to the historical traffic flow data, simulation traffic demands are generated;
[0042] According to the calibrated vehicle car-following and lane-changing model, the simulation traffic demands and the simulation road structure, a microscopic simulation environment is established.
[0043] The vehicle car-following and lane-changing model comprises a vehicle car-following model and a vehicle lane-changing model. The vehicle car-following model can be used to study the corresponding behavior of a following vehicle caused by the change of the motion state of a leading vehicle by using a dynamic method. The vehicle lane-changing model can be used for lane selection of a vehicle and speed adjustment of the vehicle when changing lanes in a multi-lane road simulation.
[0044] In the scheme, the particle swarm algorithm can be used to calibrate the parameters of the initial vehicle car-following and lane-changing model to obtain the calibrated vehicle car-following and lane-changing model.
[0045] In the scheme, the historical traffic flow data can be processed to generate simulation traffic demands of multiple traffic scenes by considering the degree of traffic flow fluctuation on the basis of the historical traffic flow data.
[0046] In the scheme, the actual road structure can be simplified to generate the simulation road structure.
[0047] Further, the microscopic simulation environment can be established according to the calibrated vehicle car-following and lane-changing model, the simulation traffic demands and the simulation road structure.
[0048] S120, a state space, a joint action space and a reward function are constructed; the reward function comprises at least one of a first reward function, a second reward function and a penalty function; the first reward function is used for feedback on total travel time, and the second reward function is used for feedback on a preset gas emission amount.
[0049] The state space can be used to represent a set of all possible states that can occur in the process of passing through the expressway merging area. For example, the total number or the change number of vehicles in each road area at each time.
[0050] The joint action space can be used to represent a set of all possible passing control actions that can be performed in the process of passing through the expressway merging area. For example, the limited speed of the variable speed limit area, the number of vehicles released from the ramp, the length of time for releasing vehicles from the ramp, and the like.
[0051] The reward function can be used to represent the instant reward or punishment obtained after taking a joint action in a certain state in the process of passing through the expressway merging area. For example, taking the scheme for improving the passing capacity of the expressway merging area as an example, if it is found that the passing capacity of the expressway merging area decreases after taking a joint action in a certain state, the decision is punished, and if it is found that the passing capacity of the expressway merging area increases after taking a joint action in a certain state, the decision can be rewarded.
[0052] In the scheme, the reward function can be represented by the first reward function, or the second reward function, or the punishment function, or the first reward function and the second reward function, or the first reward function and the punishment function, or the second reward function and the punishment function, or the first reward function, the second reward function and the punishment function. The first reward function is used to feedback the total travel time, which can be represented by the total travel time of the passing vehicle between the current state and the next state. The second reward function is used to feedback the preset gas emission, which can be represented by the volume of the preset gas emitted by the passing vehicle between the current state and the next state, wherein the preset gas can be carbon monoxide, carbon dioxide, nitrogen oxide, hydrocarbon, etc.
[0053] In addition, the scheme can also construct parameters such as state transition probability and discount factor. The state transition probability can be the probability of transitioning to the next state after taking a joint action from the current state. The discount factor can be used to control the importance of future rewards. The larger the discount factor, the more important the future rewards.
[0054] Optionally, the state space, the joint action space and the reward function are constructed, comprising:
[0055] S121, determining the state space according to the traffic flow density of each simulation road structure and the intelligent network connection vehicle penetration rate of the variable speed limit area.
[0056] The traffic flow density can be the number of vehicles passing through in a certain time. For example, a coil detector can be arranged at a suitable position on the road section, and the traffic flow density can be obtained according to the coil time occupancy. Within a certain observation time, the time occupancy and the traffic flow density have the following relationship:
[0057] o=(l+d)k;
[0058] In the formula, o is the time occupancy, l is the vehicle length, d is the length of the coil detector, and k is the traffic flow density.
[0059] The intelligent connected vehicle penetration rate can be the proportion of the number of intelligent connected vehicles to the total number of vehicles. Specifically, the intelligent connected vehicle can obtain more accurate and comprehensive traffic information through V2X (Vehicle to X) technology, and can more accurately and timely perform dynamic driving tasks.
[0060] Due to the randomness of vehicle arrival, the intelligent connected vehicle penetration rate in the road will fluctuate within a short observation time, and the intelligent connected vehicle has a shorter headway than the manually driven vehicle, which will also affect the running characteristics of the traffic flow. Therefore, the influence of the intelligent connected vehicle penetration rate is also considered when constructing the state space.
[0061] Specifically, the number of intelligent connected vehicles in the variable speed limit zone can be determined through the intelligent connected vehicle communication request received by the remote server; the total number of vehicles in the variable speed limit zone can be determined through the image collector or coil detector installed in the variable speed limit zone; and the intelligent connected vehicle penetration rate can be determined according to the proportion of the number of intelligent connected vehicles to the total number of vehicles.
[0062] In the present scheme, the state space can be constructed by the traffic flow density of the variable speed limit zone, the acceleration zone, the entrance ramp and the bottleneck zone, and the intelligent connected vehicle penetration rate of the variable speed limit zone. The traffic flow density of the variable speed limit zone can be used to represent the traffic demand from the upstream of the expressway main lane; the traffic flow density of the acceleration zone can be used to represent the degree of congestion spreading upstream; the traffic flow density of the entrance ramp can be used to represent the traffic demand from the ramp; and the traffic flow density of the bottleneck zone can be used to represent the passing capacity of the bottleneck zone. The state space can be represented by the following formula:
[0063] s={k v ,k A ,k R ,k B ,p v};
[0064] In the formula, s is the state space, k v is the traffic flow density of the variable speed limit zone, k A is the traffic flow density of the acceleration zone, k R is the traffic flow density of the entrance ramp, k B is the traffic flow density of the bottleneck zone, and p v is the intelligent connected vehicle penetration rate of the variable speed limit zone. The intelligent connected vehicle penetration rate p v of the variable speed limit zone can be determined by the following formula:
[0065]
[0066] In the formula, L v is the length of the variable speed limit zone, kv L v Total number of vehicles in the variable speed limit zone.
[0067] S122, determine a joint action space according to the speed limit value of the variable speed limit zone and the passing time length in the signal light cycle of the entrance ramp.
[0068] The speed limit value can be the maximum speed value allowed by the variable speed limit zone, which can be set according to the traffic conditions and weather conditions of the variable speed limit zone. The signal light cycle of the entrance ramp can include a passing period, a forbidden period and a waiting period to control the number of vehicles coming from the entrance ramp direction, so as to avoid affecting the normal driving of vehicles on the main line of the expressway. The passing time length can be the length of time when the green light is on in the traffic signal light.
[0069] In this scheme, at least one joint execution action can be obtained by combining each speed limit value and each passing time length to determine the joint action space.
[0070] Optionally, the joint action space is determined according to the speed limit value of the variable speed limit zone and the passing time length in the signal light cycle of the entrance ramp, including: discretizing the speed limit value of the variable speed limit zone to obtain a variable speed limit action space composed of at least two discrete speed limit values; discretizing the passing time length in the signal light cycle of the entrance ramp to obtain a ramp control action space composed of at least two discrete passing time lengths; determining the joint action space according to the variable speed limit action space and the ramp control action space.
[0071] For the control of the speed limit value of the variable speed limit zone, the use of continuous action space is more helpful for smooth speed control, but the signal phase needs discrete phase duration, and the minimum interval is 1 second. The integrated control needs to determine the speed limit and the green light phase duration at the same time. Since the mixed space of continuous and discrete actions is not conducive to the training of the agent, the optional value space of the variable speed limit is discretized by appropriate interval in this scheme. The variable speed limit action space can be represented by the following formula:
[0072]
[0073] In the formula, A VSL is the variable speed limit action space, v min is the minimum speed allowed by the variable speed limit zone, v max is the maximum speed allowed by the variable speed limit zone, and N is the total interval number.
[0074] Discretizing the passing time length in the signal light cycle of the entrance ramp, the ramp control action space can be represented by the following formula:
[0075] ARM = {t min , t min + e, t min + 2e, …, T - 2e, t - e, t};
[0076] where A RM is the ramp control action space, t min is the minimum allowed green time within the ramp signal cycle, and e is the interval time.
[0077] Therefore, the joint actions that can be taken at each decision time are as follows: a = (a VSL , a RM ), a VSL ∈ A VSL , a RM ∈ A RM .
[0078] S123, determining a reward function according to at least one of the signal cycle of the entrance ramp, the vehicle variation quantity of the target freeway merge area, the preset gas emission quantity of the target freeway merge area, and the vehicle variation quantity of the bottleneck area.
[0079] The vehicle variation quantity of the target freeway merge area can be the difference between the number of vehicles entering from the main lane and the number of vehicles exiting from the bottleneck area. The preset gas emission quantity can be the total emission of the preset gas released by all vehicles in the target freeway merge area. The vehicle variation quantity of the bottleneck area can be the difference between the number of vehicles entering the bottleneck area and the number of vehicles exiting the bottleneck area.
[0080] In this scheme, the first reward function, the second reward function, and / or the penalty function can be determined according to at least one of the signal cycle of the entrance ramp, the vehicle variation quantity of the target freeway merge area, the preset gas emission quantity of the target freeway merge area, and the vehicle variation quantity of the bottleneck area.
[0081] Optionally, the reward function is determined according to at least one of the signal cycle of the entrance ramp, the vehicle variation quantity of the target freeway merge area, the preset gas emission quantity of the target freeway merge area, and the vehicle variation quantity of the bottleneck area, including: determining a first reward function according to the signal cycle of the entrance ramp and the vehicle variation quantity of the target freeway merge area; determining a second reward function according to the preset gas emission quantity of the target freeway merge area; determining a penalty function according to the vehicle variation quantity of the bottleneck area; and determining a reward function according to at least one of the first reward function, the second reward function, and the penalty function.
[0082] In the control of the on-ramp area, traffic efficiency is used to describe the effectiveness of the traffic control scheme, and the total travel time is often used as a performance indicator of traffic efficiency. When the total number of vehicles in the on-ramp area is zero, the total travel time can be determined by the following formula:
[0083]
[0084] In the formula, TTS is the total travel time of all vehicles in the on-ramp area, n(k) is the total number of vehicles in the on-ramp area at the end of the kth period, and T is the signal light cycle of the entrance ramp.
[0085] Therefore, the first reward function with the goal of reducing the total travel time can be represented by the following formula:
[0086]
[0087] In the formula, is the first reward function, d(k') is the number of vehicles entering the on-ramp area at the k'th time in the kth period, and q(k') is the number of vehicles leaving the on-ramp area at the k'th time in the kth period.
[0088] Similarly, the preset gas emission amount can be determined by the following formula:
[0089]
[0090] In the formula, TCE is the total emission amount of the preset gas of all vehicles in the on-ramp area, a i is the acceleration of the ith vehicle, v i is the speed of the ith vehicle, sl i is the slope of the road where the ith vehicle is located, C(a i ,v i ,sl i is the emission rate of the preset gas of the ith vehicle.
[0091] Therefore, the second reward function with the goal of reducing the preset gas emission amount can be represented by the following formula:
[0092]
[0093] In the formula, is the second reward function.
[0094] In order to avoid the agent reducing the total number of input vehicles by creating congestion in the bottleneck area, i.e., the total number of vehicles leaving the on-ramp area is less than the total number of vehicles loaded from the traffic demand, the scheme will give a penalty feedback in this case. Specifically, the penalty function can be represented by the following formula:
[0095]
[0096] wherein, is a penalty function, m is a penalty, is the total number of vehicles leaving the merging area of the expressway in the kth period, is the total number of vehicles loaded from the traffic demand in the kth period.
[0097] S130, based on the micro-simulation environment, according to the state space, the joint action space and the reward function, the initial cooperative control model of the target expressway merging area is reinforced learning to determine the target cooperative control model.
[0098] Wherein, reinforcement learning is that the agent learns in a way of constantly "trial and error", and the reward obtained by interacting with the environment guides the behavior, and the goal is to make the agent obtain the maximum reward. The reinforcement learning algorithm can be TD Learning time difference learning algorithm, Q-Learning algorithm, SARS A algorithm, etc., and the present scheme does not limit this.
[0099] In the present scheme, in order to train the agent to adaptively adjust the control strategy according to the target preference, the optimal reward and trajectory under different target preferences are matched by using the envelope MODDQN (Multi Objective Double Deep Q-Learning) algorithm to determine the target cooperative control model.
[0100] Optionally, according to the simulation requirement, the state space, the joint action space and the reward function, the initial cooperative control model of the target expressway merging area is reinforced learning to determine the target cooperative control model, comprising: initializing the evaluation network, the target network and the memory bank of the initial cooperative control model; in each simulation round, the execution action is determined according to the current state according to a preset strategy; inputting the execution action into the micro-simulation environment to determine the transition state, and determining the reward function value corresponding to the current state; storing the current data in the memory bank, wherein the current data at least includes the current state, the execution action, the transition state and the reward function value; updating the parameters of the evaluation network to perform the next simulation round; randomly extracting a preset batch number of experiences from the memory bank, and updating the parameters of the evaluation network to the parameters of the target network according to a preset frequency, until each simulation round ends, and the target cooperative control model is determined.
[0101] In the present scheme, first, the learning rate a, the reward discount rate g and the weight space W are input into the initial cooperative control model, and the TTS value function fitting evaluation network and TCE value function fitting evaluation network Initialization is performed, and a TTS value function fitting target network is initialized and TCE value function fitting target network Initialization is performed, and a memory library E is initialized.
[0102] Secondly, each simulation episode = 1, 2, …, M is sequentially trained. Specifically, the training steps are as follows:
[0103] First, ω ∈ W is randomly sampled from the weight space W, and the micro-simulation environment of the target expressway merging area is initialized;
[0104] Secondly, for each decision step k = 1, 2, …, K in each simulation episode, the current state s k is first obtained from the environment, and the agent selects an action a k according to the current state k. The specific selection strategy is as follows:
[0105] ε = 0.01 + 0.05k, if ε ≤ 0.99 else 0.99,
[0106]
[0107] The action a k selected by the agent according to the above strategy is input into the micro-simulation environment, and the environment is simulated to the next decision step, and the transition state s' k is obtained from the environment. If this simulation episode is over, done k = 1, otherwise done k = 0.
[0108] Thirdly, the first reward function the second reward function and the penalty function are calculated according to the reward function, and is stored as current data in the memory library D.
[0109] Secondly, the parameters of the evaluation network are updated for the next simulation episode. The specific updating process is as follows:
[0110] First, the TTS value function fitting evaluation network and the TCE value function fitting evaluation network
[0111] N Batch pieces of data are randomly extracted from the memory library D.
[0112] Secondly, for each weight ω j, respectively, are calculated:
[0113]
[0114]
[0115]
[0116] with and Adam (alpha) update and
[0117] Finally, the parameters of the target network are replaced at a fixed frequency, and the specific updating process is as follows:
[0118] θ' = θ,
[0119] The beneficial effects of the above technical solutions are that the delay updated target network is added to the value function fitting network on the basis of Envelope MOQ-Learning, so as to increase the stability of the training process; and the decisions under each weight in the weight space are trained in advance, the optimal rewards and trajectories under different target preferences are matched, so as to improve the running efficiency of the target collaborative control model.
[0120] S140, based on the target collaborative control model, according to the obtained real-time traffic flow density of the target expressway merging area, real-time intelligent network connected vehicle penetration rate and target weight coefficient, determine the target speed limit value of the variable speed limit area and the target passing time in the signal lamp cycle of the entrance ramp; wherein, the target weight coefficient is used to determine the attention degree to the first reward function and the second reward function.
[0121] Wherein, the target weight coefficient can be set according to the actual situation.
[0122] In this scheme, the intelligent agent as the core of control decision constantly interacts with the actual traffic environment, perceives the real-time road traffic flow density, and combines the traffic rules and environmental changes, so that the intelligent agent can quickly make decisions to adjust the road speed limit and the entering time of ramp vehicles, so as to improve the passing capacity of the expressway merging area and reduce the emission of the preset gas.
[0123] The embodiment of the application provides a traffic control method for a rapid road merging area, which comprises the following steps: establishing a micro-simulation environment of a target rapid road merging area according to actual road structure and historical traffic data of the target rapid road merging area; constructing a state space, a joint action space and a reward function; based on the micro-simulation environment, performing reinforcement learning on an initial collaborative control model of the target rapid road merging area according to the state space, the joint action space and the reward function, and determining a target collaborative control model; and based on the target collaborative control model, determining a target speed limit value of a variable speed limit area and a target passing time in a signal lamp cycle of an entrance ramp according to real-time traffic flow density, real-time intelligent network connected vehicle penetration rate and a target weight coefficient of the target rapid road merging area. According to the technical scheme, the variable speed limit and the ramp are controlled, the traffic pressure of the rapid road merging area is reduced, and the passing efficiency of the rapid road merging area is improved.
[0124] Embodiment two
[0125] Figure 2 A flowchart of a traffic control method for a rapid road merging area is provided for the embodiment two of the application, and the embodiment is optimized based on the above-mentioned embodiment. As shown in the figure, the method of the embodiment specifically comprises the following steps: Figure 2
[0126] S210, a micro-simulation environment of a target rapid road merging area is established according to actual road structure and historical traffic data of the target rapid road merging area; wherein the road structure at least comprises a variable speed limit area, an acceleration area, an entrance ramp and a bottleneck area.
[0127] S220, a state space, a joint action space and a reward function are constructed; wherein the reward function comprises at least one of a first reward function, a second reward function and a penalty function; the first reward function is used for feedback on total travel time, and the second reward function is used for feedback on a preset gas emission amount.
[0128] S230, based on the micro-simulation environment, reinforcement learning is performed on an initial collaborative control model of the target rapid road merging area according to the state space, the joint action space and the reward function, and a target collaborative control model is determined.
[0129] S240, when the reward function comprises a first reward function and a penalty function which are used for reducing total travel time, the target rapid road merging area is controlled by the target collaborative control model for at least two preset intelligent network connected vehicle penetration rates, and a first travel time and a first preset gas emission amount of the target rapid road merging area are respectively determined.
[0130] In the present scheme, in order to avoid the influence of multiple targets on travel time and preset gas emission, the cases of only considering reducing total travel time and only considering reducing preset gas emission are analyzed respectively to determine the optimal strategy for different targets.
[0131] Exemplarily, the first travel time and the first preset gas emission of the target expressway merge area when the target cooperative control model only optimizes the total travel time can be determined under preset intelligent connected vehicle penetration rates of 10%, 30% and 50% respectively.
[0132] S250, when the reward function includes a second reward function and a penalty function for reducing the preset gas emission, the target expressway merge area is controlled by the target cooperative control model for at least two preset intelligent connected vehicle penetration rates, and the second travel time and the second preset gas emission of the target expressway merge area are determined respectively.
[0133] Exemplarily, the second travel time and the second preset gas emission of the target expressway merge area when the target cooperative control model only optimizes the preset gas emission can be determined under preset intelligent connected vehicle penetration rates of 10%, 30% and 50% respectively.
[0134] S260, without controlling the target expressway merge area by the target cooperative control model, the third travel time and the third preset gas emission of the target expressway merge area are determined respectively for at least two preset intelligent connected vehicle penetration rates.
[0135] Exemplarily, the third travel time and the third preset gas emission of the target expressway merge area can be determined under preset intelligent connected vehicle penetration rates of 10%, 30% and 50% respectively.
[0136] S270, according to the first travel time, the first preset gas emission, the second travel time, the second preset gas emission, the third travel time and the third preset gas emission, the corresponding relationship of the first reward function and the second reward function under different intelligent connected vehicle penetration rates is determined.
[0137] Exemplarily, the simulation calculation results of the examples of the above steps S240 to S260 are compared and analyzed to obtain the control effect table of each experimental scene under each intelligent connected vehicle penetration rate, as shown in Table 1:
[0138] Table 1
[0139]
[0140] As shown in Table 1, when the intelligent connected vehicle penetration rate is 10% or 30%, reducing the travel time helps to reduce the preset gas emission. However, when the intelligent connected vehicle penetration rate increases to 50%, reducing the travel time is accompanied by an increase in the preset gas emission, and vice versa. Table 1 also shows that when only one of the two targets is optimized, the optimal improvement of the travel time and the preset gas emission cannot be achieved simultaneously. For example, when the intelligent connected vehicle penetration rate is 10% and only the total travel time is considered, the improvements of the travel time and the preset gas emission are 9.21% and 3.82%, respectively, relative to the case without control. However, when only the preset gas emission is considered, the improvement of the preset gas emission at this intelligent connected vehicle penetration rate increases from 3.82% to 4.71%, while the improvement of the travel time decreases from 9.21% to 4.75%. This shows that although the travel time and the preset gas emission can be reduced simultaneously, the optimal strategy for different targets is different.
[0141] S280, determining a target weight coefficient according to the correspondence between the first reward function and the second reward function at different intelligent connected vehicle penetration rates.
[0142] Specifically, the target weight coefficient can be determined according to the difference between the travel time and the preset gas emission at different attention targets, as well as the actual situation demand.
[0143] S290, determining a target speed limit value of the variable speed limit zone and a target passing time length in a signal light cycle of the entrance ramp based on the target collaborative control model, the real-time traffic flow density of the target freeway merging area, the real-time intelligent connected vehicle penetration rate, and the target weight coefficient, wherein the target weight coefficient is used to determine the attention degree to the first reward function and the second reward function.
[0144] The embodiment of the application provides a passing control method for a freeway merging area. The method simulates the total travel time and the preset gas emission under different intelligent connected vehicle penetration rates and different control modes, to determine the correspondence between the first reward function and the second reward function at different intelligent connected vehicle penetration rates, provide data support for the determination of the target weight coefficient, and further improve the passing capacity of the freeway merging area.
[0145] Embodiment Three
[0146] Figure 3 A structural schematic diagram of a passing control device for a freeway merging area is provided for the third embodiment of the application. As shown in the figure, the device comprises: Figure 3
[0147] The simulation environment establishing module 310 is configured to establish a microscopic simulation environment of the target expressway merging area according to actual road structures and historical traffic data of the target expressway merging area; wherein the road structures at least include a variable speed limit area, an acceleration area, an entrance ramp and a bottleneck area.
[0148] The simulation parameter constructing module 320 is configured to construct a state space, a joint action space and a reward function; wherein the reward function includes at least one of a first reward function, a second reward function and a penalty function; the first reward function is configured to feed back a total travel time, and the second reward function is configured to feed back a preset gas emission amount.
[0149] The control model determining module 330 is configured to perform reinforcement learning on an initial cooperative control model of the target expressway merging area according to the state space, the joint action space and the reward function based on the microscopic simulation environment, to determine a target cooperative control model.
[0150] The passing scheme determining module 340 is configured to determine a target speed limit value of the variable speed limit area and a target passing time length in a signal lamp cycle of the entrance ramp based on the target cooperative control model according to the real-time traffic flow density of the target expressway merging area, the real-time penetration rate of intelligent networked vehicles and a target weight coefficient; wherein the target weight coefficient is configured to determine a degree of attention to the first reward function and the second reward function.
[0151] The embodiment of the present application provides a passing control device for an expressway merging area, which establishes a microscopic simulation environment of a target expressway merging area according to actual road structures and historical traffic data of the target expressway merging area; constructs a state space, a joint action space and a reward function; performs reinforcement learning on an initial cooperative control model of the target expressway merging area according to the state space, the joint action space and the reward function based on the microscopic simulation environment, to determine a target cooperative control model; and determines a target speed limit value of a variable speed limit area and a target passing time length in a signal lamp cycle of an entrance ramp based on the target cooperative control model according to real-time traffic flow density of the target expressway merging area, real-time penetration rate of intelligent networked vehicles and a target weight coefficient. The technical scheme reduces the traffic pressure of the expressway merging area and improves the passing efficiency of the expressway merging area by controlling the variable speed limit and the ramp.
[0152] Further, the simulation parameter constructing module 310 includes:
[0153] The state space determining unit is configured to determine the state space according to traffic flow densities of the simulation road structures and intelligent networked vehicle penetration rates of the variable speed limit areas.
[0154] The joint action space determination unit is configured to determine a joint action space according to a speed limit value of the variable speed limit zone and a passing time length in a signal light cycle of the entrance ramp.
[0155] The reward function determination unit is configured to determine a reward function according to at least one of a signal light cycle of the entrance ramp, a vehicle change quantity of the target expressway merging zone, a preset gas emission amount of the target expressway merging zone, and a vehicle change quantity of the bottleneck zone.
[0156] Further, the joint action space determination unit comprises:
[0157] The variable speed limit action space determination sub-unit is configured to discretize the speed limit value of the variable speed limit zone to obtain a variable speed limit action space composed of at least two discrete speed limit values.
[0158] The ramp control action space determination sub-unit is configured to discretize the passing time length in the signal light cycle of the entrance ramp to obtain a ramp control action space composed of at least two discrete passing time lengths.
[0159] The joint action space determination sub-unit is configured to determine a joint action space according to the variable speed limit action space and the ramp control action space.
[0160] Further, the reward function determination unit comprises:
[0161] The first reward function determination sub-unit is configured to determine a first reward function according to the signal light cycle of the entrance ramp and the vehicle change quantity of the target expressway merging zone.
[0162] The second reward function determination sub-unit is configured to determine a second reward function according to the preset gas emission amount of the target expressway merging zone.
[0163] The penalty function determination sub-unit is configured to determine a penalty function according to the vehicle change quantity of the bottleneck zone.
[0164] The reward function determination sub-unit is configured to determine a reward function according to at least one of the first reward function, the second reward function, and the penalty function.
[0165] Further, the apparatus further comprises:
[0166] The first parameter determination module is configured to, after determining the target cooperative control model, when the reward function comprises a first reward function and a penalty function aiming to reduce total travel time, control the target expressway merging zone through the target cooperative control model for at least two preset intelligent connected vehicle penetration rates, and respectively determine a first travel time and a first preset gas emission amount of the target expressway merging zone.
[0167] The second parameter determination module is configured to, when the reward function comprises a second reward function and a penalty function aiming to reduce a preset gas emission amount, control the target expressway merging area by using the target cooperative control model for at least two preset intelligent connected vehicle penetration rates, and determine a second travel time and a second preset gas emission amount of the target expressway merging area respectively.
[0168] The third parameter determination module is configured to, without controlling the target expressway merging area by using the target cooperative control model, determine a third travel time and a third preset gas emission amount of the target expressway merging area for at least two preset intelligent connected vehicle penetration rates respectively.
[0169] The corresponding relationship determination module is configured to determine a corresponding relationship between the first reward function and the second reward function under different intelligent connected vehicle penetration rates according to the first travel time, the first preset gas emission amount, the second travel time, the second preset gas emission amount, the third travel time and the third preset gas emission amount.
[0170] Correspondingly, the determination process of the target weight coefficient comprises:
[0171] According to the corresponding relationship between the first reward function and the second reward function under the different intelligent connected vehicle penetration rates, a target weight coefficient is determined.
[0172] Further, the control model determination module 330 comprises:
[0173] The parameter initialization unit is configured to initialize an evaluation network, a target network and a memory bank of an initial cooperative control model.
[0174] The parameter simulation training unit is configured to, in each simulation round, determine an execution action according to a preset strategy according to a current state; input the execution action into the micro-simulation environment to determine a transition state, and determine a reward function value corresponding to the current state; store current data in the memory bank, wherein the current data at least comprises the current state, the execution action, the transition state and the reward function value; update parameters of the evaluation network to perform a next simulation round.
[0175] The control model determination unit is configured to randomly extract a preset batch number of experiences from the memory bank, update the parameters of the evaluation network to the parameters of the target network according to a preset frequency, until each simulation round ends, and determine a target cooperative control model.
[0176] Further, the simulation environment establishment module 310 comprises:
[0177] An initial parameter acquisition unit is configured to acquire actual road structure, historical traffic flow data and an initial car-following lane-changing model of a target expressway merging area;
[0178] A road structure determination unit is configured to determine a simulation road structure of the target expressway merging area according to the actual road structure;
[0179] A road model calibration unit is configured to calibrate parameters of the initial car-following lane-changing model to obtain a calibrated car-following lane-changing model;
[0180] A traffic demand generation unit is configured to generate simulation traffic demand according to the historical traffic flow data;
[0181] A simulation environment establishment unit is configured to establish a microscopic simulation environment according to the calibrated car-following lane-changing model, the simulation traffic demand and the simulation road structure.
[0182] The expressway merging area passing control device provided by the embodiments of the present application can execute the expressway merging area passing control method provided by any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.
[0183] Embodiment Four
[0184] Figure 4 A structural schematic diagram of a device 10 that can be used to implement embodiments of the present application is shown. The device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices (e.g., headgear, eyewear, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.
[0185] As Figure 4As shown, the device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., connected to the at least one processor 11 in communication. The memory stores computer programs executable by the at least one processor 11, and the processor 11 can perform various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 12 or loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0186] Various components in the device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc., an output unit 17, such as various types of displays, speakers, etc., a storage unit 18, such as a magnetic disk, an optical disk, etc., and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0187] The processor 11 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the traffic control method for a freeway merge area.
[0188] In some embodiments, the traffic control method for a freeway merge area can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the traffic control method for a freeway merge area described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the traffic control method for a freeway merge area by any other appropriate means, such as by means of firmware.
[0189] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip systems (SOCs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0190] Computer programs used to implement the processes of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program
[0191] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store computer programs for use by or in connection with an instruction execution system, apparatus, or device. Computer-readable storage media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0192] To provide for interaction with a user, the systems and techniques described here can be implemented on a device having a display (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0193] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0194] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server is generally established using computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0195] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in series, or executed in different orders, as long as the desired results of the present disclosure are achieved, and the present disclosure is not limited herein.
[0196] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and scope of the disclosure. Any further modifications, equivalents, and / or alternatives come within the scope of the present disclosure as described in the following claims.
Claims
1. A method of traffic control at a merge area of an expressway, characterized by, The method comprises: According to the actual road structure and historical traffic data of the target expressway merging area, a micro-simulation environment of the target expressway merging area is established; wherein the road structure at least includes a variable speed limit area, an acceleration area, an entrance ramp and a bottleneck area; Constructing a state space, a joint action space and a reward function; wherein the reward function includes at least one of a first reward function, a second reward function and a penalty function; the first reward function is used to feedback the total travel time, and the second reward function is used to feedback the preset gas emission amount; Based on the micro-simulation environment, according to the state space, the joint action space and the reward function, reinforcement learning is performed on the initial cooperative control model of the target expressway merging area to determine a target cooperative control model; Based on the target cooperative control model, according to the real-time traffic flow density of the target expressway merging area, the real-time intelligent network connected vehicle penetration rate and the target weight coefficient, the target speed limit value of the variable speed limit area and the target passing time in the signal lamp cycle of the entrance ramp are determined; wherein the target weight coefficient is used to determine the attention degree of the first reward function and the second reward function.
2. The method of claim 1, wherein, Constructing a state space, a joint action space and a reward function, comprising: According to the traffic flow density of each simulation road structure and the intelligent network connected vehicle penetration rate of the variable speed limit area, a state space is determined; According to the speed limit value of the variable speed limit area and the passing time in the signal lamp cycle of the entrance ramp, a joint action space is determined; According to at least one of the signal lamp cycle of the entrance ramp, the vehicle change quantity of the target expressway merging area, the preset gas emission amount of the target expressway merging area and the vehicle change quantity of the bottleneck area, a reward function is determined.
3. The method of claim 2, wherein, According to the speed limit value of the variable speed limit area and the passing time in the signal lamp cycle of the entrance ramp, a joint action space is determined, comprising: Discretize the speed limit value of the variable speed limit area to obtain a variable speed limit action space composed of at least two discrete speed limit values; Discretize the passing time in the signal lamp cycle of the entrance ramp to obtain a ramp control action space composed of at least two discrete passing times; According to the variable speed limit action space and the ramp control action space, a joint action space is determined.
4. The method of claim 2, wherein, According to at least one of the signal lamp cycle of the entrance ramp, the vehicle change quantity of the target expressway merging area, the preset gas emission amount of the target expressway merging area and the vehicle change quantity of the bottleneck area, a reward function is determined, comprising: According to the signal lamp cycle of the entrance ramp and the vehicle change quantity of the target expressway merging area, a first reward function is determined; According to the preset gas emission amount of the target expressway merging area, a second reward function is determined; According to the vehicle change quantity of the bottleneck area, a penalty function is determined; According to at least one of the first reward function, the second reward function and the penalty function, a reward function is determined.
5. The method of claim 1, wherein, After determining the target cooperative control model, the method further comprises: When the reward function comprises a first reward function and a penalty function aiming at reducing total travel time, the target expressway merging area is controlled by the target cooperative control model for at least two preset intelligent connected vehicle penetration rates, and a first travel time and a first preset gas emission of the target expressway merging area are determined respectively. When the reward function comprises a second reward function and a penalty function aiming at reducing preset gas emission, the target expressway merging area is controlled by the target cooperative control model for at least two preset intelligent connected vehicle penetration rates, and a second travel time and a second preset gas emission of the target expressway merging area are determined respectively. When the target expressway merging area is not controlled by the target cooperative control model, a third travel time and a third preset gas emission of the target expressway merging area are determined for at least two preset intelligent connected vehicle penetration rates. According to the first travel time, the first preset gas emission, the second travel time, the second preset gas emission, the third travel time and the third preset gas emission, a corresponding relationship between the first reward function and the second reward function under different intelligent connected vehicle penetration rates is determined. Correspondingly, the determination process of the target weight coefficient comprises: According to the corresponding relationship between the first reward function and the second reward function under different intelligent connected vehicle penetration rates, a target weight coefficient is determined.
6. The method of claim 1, wherein, According to simulation requirements, the state space, the joint action space and the reward function, an initial cooperative control model of the target expressway merging area is subjected to reinforcement learning to determine a target cooperative control model, comprising: An evaluation network, a target network and a memory bank of the initial cooperative control model are initialized; In each simulation round, an execution action is determined according to a preset strategy according to a current state; the execution action is input into the microscopic simulation environment to determine a transition state, and a reward function value corresponding to the current state is determined; current data is stored in the memory bank, wherein the current data at least comprises the current state, the execution action, the transition state and the reward function value; parameters of the evaluation network are updated for the next simulation round; A preset batch number of experiences are randomly extracted from the memory bank, and the parameters of the evaluation network are updated as the parameters of the target network at a preset frequency until the end of each simulation round, and a target cooperative control model is determined.
7. The method of claim 1, wherein, A microscopic simulation environment is established, comprising: Actual road structure, historical traffic flow data and an initial vehicle following and lane changing model of a target expressway merging area are acquired; According to the actual road structure, a simulation road structure of the target expressway merging area is determined; Parameters of the initial vehicle following and lane changing model are calibrated to obtain a calibrated vehicle following and lane changing model; According to the historical traffic flow data, simulation traffic demand is generated; According to the calibrated vehicle following and lane changing model, the simulation traffic demand and the simulation road structure, a microscopic simulation environment is established.
8. A traffic control device for a merging area of an expressway, characterized by comprising: The device comprises: The simulation environment establishment module is configured to establish a micro-simulation environment of the target expressway merging area according to actual road structures and historical traffic data of the target expressway merging area, wherein the road structures at least include a variable speed limit area, an acceleration area, an entrance ramp, and a bottleneck area; The simulation parameter construction module is configured to construct a state space, a joint action space, and a reward function, wherein the reward function includes at least one of a first reward function, a second reward function, and a penalty function, the first reward function is configured to feed back a total travel time, and the second reward function is configured to feed back a preset gas emission amount; The control model determination module is configured to perform reinforcement learning on an initial cooperative control model of the target expressway merging area based on the micro-simulation environment and according to the state space, the joint action space, and the reward function, and determine a target cooperative control model. The passing scheme determination module is configured to determine a target speed limit value of the variable speed limit area and a target passing time length in a signal lamp cycle of the entrance ramp based on the target cooperative control model and according to the obtained real-time traffic flow density of the target expressway merging area, the real-time penetration rate of the intelligent networked vehicle, and a target weight coefficient, wherein the target weight coefficient is configured to determine a degree of attention to the first reward function and the second reward function.
9. An electronic device, comprising: The device includes: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the passing control method of the expressway merging area according to any one of claims 1-7.
10. A computer readable storage medium characterized by The computer readable storage medium stores computer instructions for enabling the processor to perform the passing control method of the expressway merging area according to any one of claims 1-7 when executed.
Citation Information
Patent Citations
Mixed multi-ramp cooperative confluence control method based on multi-agent reinforcement learning
CN115909785A
Multi-vehicle collaborative marshalling intersection method and system in expressway confluence area in mixed driving environment
CN116740945A