A method and equipment for inter-lane control and guidance on highways
By using Actor-Critic network intelligent agents and dynamic-static fusion signs between the inner and outer lanes of the highway, the problem of uneven traffic flow caused by static signs and markings has been solved, achieving efficient utilization of road resources and improving travel efficiency.
Patent Information
- Application Number
- CN202411411166.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-10
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-10-10
AI Technical Summary
In existing technologies, static signs and markings lead to an imbalance in traffic flow between the inner and outer lanes of highways, resulting in insufficient utilization of road resources.
By acquiring traffic state parameters, using the Actor network agent to output guidance action information, and combining it with the Critic network for training, dynamic solid and dashed line state information guidance is achieved, and dynamic and static fusion signs guide vehicles to change lanes, so as to achieve a balance of traffic flow in the inner and outer lanes.
It improves the efficiency of road resource utilization, reduces the difficulty of responding to emergencies, enhances travel efficiency, eliminates reliance on human experience, and achieves fully parameterized amplitude control.
Smart Images

Figure CN119274367B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent traffic control technology, specifically relating to a method and equipment for regulating and guiding traffic between inner and outer lanes of a highway. Background Technology
[0002] The purpose of this invention is to address the problem that using static signs and markings to guide vehicles leads to uneven traffic flow between inner and outer lanes and insufficient utilization of road resources. This invention proposes a method and device for regulating and guiding traffic between inner and outer lanes on highways. Through calculation and analysis, dynamic solid and dashed line status information is output to guide vehicles to change lanes, thereby balancing the rational utilization of road resources.
[0003] To achieve the above-mentioned objectives, the present invention provides the following technical solution:
[0004] A method for regulating and inducing traffic flow between inner and outer lanes of a highway includes the following steps:
[0005] S1, Obtain traffic state parameters, including traffic flow, vehicle density, and average vehicle speed of each road segment;
[0006] S2, the traffic state parameters are input into the trained agent with an Actor network, and the agent outputs guidance action information, which includes guidance direction and guidance destination, to realize real-time inter-amplitude guidance control;
[0007] It also includes the step of training the agent with the Actor network:
[0008] S21, the output of the guidance action information causes a change in the traffic state parameters, resulting in new traffic state parameters;
[0009] S21, Calculate the reward value of the agent based on the new traffic state parameters;
[0010] S23, the agent is trained based on the traffic state parameters, guidance action information, agent reward value, and new traffic state parameters.
[0011] As a preferred embodiment, during the training of the Actor network, a Critic network is used as the evaluation network to calculate the evaluation value of the agent's induced action information using traffic state parameters and induced action information; the parameters of the Actor network are then updated based on the induced action information evaluation value.
[0012] As a preferred embodiment, step S22 specifically includes the following steps:
[0013] S220, Calculate the inner and outer lane balance index based on the new traffic state parameters;
[0014] S221, calculate the reward value of the inner and outer lane balance index according to the inner and outer lane balance index, and take the opposite of the sum of the inner and outer lane balance indices of each road segment as the reward value of the agent.
[0015] As a preferred embodiment, the formula for calculating the reward value is: Where, r t This is the reward value, m is the road segment number, M is the maximum value of the road segment number, and γ is the maximum reward value. m It is an indicator of internal and external amplitude balance.
[0016] As a preferred embodiment, the formula for calculating the internal and external amplitude balance index is as follows:
[0017]
[0018] Where m is the road segment number, γ m Q is the ratio of the inner and outer lane saturation of the m-th road segment. 2m-1 Q 2m C represents the traffic flow of the inner and outer lanes of the m-th road segment, respectively. 2m-1 C 2m These represent the traffic capacity of the inner and outer lanes of the m-th road segment, respectively.
[0019] As a preferred embodiment, the intelligent agent includes a real-dashed line intelligent agent and a virtual-real line intelligent agent;
[0020] Solid-dashed line agent: used to determine whether vehicles induced to travel on the outer lane at solid-dashed line points should enter the inner lane; the action variables output by the solid-dashed line agent include the mid-distance destination and the mid- and long-distance induced directions;
[0021] The solid-dotted line agent is used to determine whether a vehicle traveling within the inner lane of the solid-dotted line guidance has left the outer lane; the action variables output by the solid-dotted line agent include the destination and the guidance direction.
[0022] As a preferred embodiment, the guidance action information includes multiple sets of information, each set of information corresponding to a road segment number, and each set of information functions independently. The output guidance action information realizes distributed real-time inter-amplitude guidance control.
[0023] As a preferred embodiment, the location where the guidance action information is published includes the starting point and midpoint of the solid and dashed lines and the solid and dashed line segments where the vehicle performs the width transition.
[0024] Based on the same concept, a dynamic and static integrated variable information publishing device, including a processor and a splicing screen, was also proposed.
[0025] The processor outputs the guidance action information using any of the above-described highway in-situ and out-of-situ control guidance methods.
[0026] The induced action information is displayed through a video wall.
[0027] As a preferred option, the direction and destination displayed on the splicing screen are variable, including variable information on dynamic and static fusion at the solid and dashed lines and amplitude adjustment indicators for dynamic and static fusion at the solid and dashed lines:
[0028] The variable information signs at solid and dashed lines guide drivers to medium and long-distance destinations. By changing the medium-distance destination, the signs adjust the objects entering the inner lane and change the arrows to variable directional arrows, which have two forms: forward and left forward, to control whether medium and long-distance drivers enter the inner lane at the variable information signs at solid and dashed lines. Variable information boards below the signs also explain the road conditions to help drivers understand the variable information signs at solid and dashed lines.
[0029] The dynamic and static control signs at the dashed and solid lines guide vehicles traveling in the inner lane to determine whether to enter the outer lane via the lower-mounted variable information display.
[0030] Compared with the prior art, the beneficial effects of the present invention are as follows: the use of the method and equipment of the present invention increases the flexibility of the inner and outer lane control sections, improves the ability to respond to emergencies, reduces the waste of road resources, improves the travel efficiency of travelers, eliminates the reliance on manual experience in control schemes, and forms a fully parameterized inter-lane control scheme. Attached Figure Description
[0031] Figure 1 This is a schematic diagram illustrating how dashed lines and solid / dashed lines in the background art represent the areas where vehicles can switch between outward and inward movements, respectively.
[0032] Figure 2 This is a schematic diagram of the dynamic-static fusion mark at the solid-dashed line in Example 1;
[0033] Figure 3 This is a flowchart of a method for regulating and inducing traffic flow between inner and outer lanes of a highway, as described in Example 1.
[0034] Figure 4 This is a schematic diagram of model updating and correction in Example 1;
[0035] Figure 5 This is a schematic diagram of parameter quantization in the simulation environment of Example 1. Detailed Implementation
[0036] The present invention will be further described in detail below with reference to experimental examples and specific embodiments. However, this should not be construed as limiting the scope of the above-mentioned subject matter of the present invention to the following embodiments; all technologies implemented based on the content of the present invention fall within the scope of the present invention.
[0037] The main concept of this invention is to obtain guidance action information, including guidance direction and guidance destination, based on traffic state parameters, and to set up dynamic and static fusion signs in the transition area (solid and dashed lines, dashed and solid lines) between inner and outer lanes. The fusion signs display guidance action information to guide traffic flow between inner and outer lanes and achieve traffic flow balance between inner and outer lanes.
[0038] Example 1
[0039] A flowchart of a method for regulating and inducing traffic flow between inner and outer lanes of a highway is shown below. Figure 3 As shown, the process includes the following steps: S1, acquiring traffic state parameters, including traffic flow, vehicle density, and average vehicle speed at each road segment; S2, inputting the traffic state parameters into a trained agent with an Actor network, the agent outputting guidance action information, including guidance direction and guidance destination, to achieve real-time inter-segment guidance control.
[0040] Step S2 is the core step of this invention. It is crucial for the agent with the Actor network to output accurate induced action information. Therefore, it is necessary to train the agent with the Actor network to improve the accuracy of the agent's output information.
[0041] The agent with the Actor network is trained and updated. A schematic diagram of model update and correction is shown below. Figure 4 As shown, the specific steps include:
[0042] S21, the output of the guidance action information causes a change in the traffic state parameters, resulting in new traffic state parameters;
[0043] S22, calculate the reward value of the agent based on the new traffic state parameters;
[0044] S23, the agent is trained based on the traffic state parameters, guidance action information, agent reward value, and new traffic state parameters.
[0045] Furthermore, during the training of the Actor network, a Critic network is used as the evaluation network to calculate the evaluation value of the agent's induced action information using traffic state parameters and induced action information; the parameters of the Actor network are then updated based on the evaluation value of the induced action information.
[0046] Specifically, step S22 includes the following steps:
[0047] S220, Calculate the inner and outer lane balance index based on the new traffic state parameters;
[0048] S221, calculate the reward value of the inner and outer lane balance index according to the inner and outer lane balance index, and take the opposite of the sum of the inner and outer lane balance indices of each road segment as the reward value of the agent.
[0049] During the real-time output of guidance information by the intelligent agent, a reinforcement learning framework is used to learn and observe the current state, select appropriate actions, and adjust its strategy based on the reward value obtained after executing the action. This process is continuous, and the agent learns an optimal strategy through interaction with the environment to maximize the cumulative reward value. Specifically, using the reinforcement learning framework, a guidance environment between inner and outer lanes is constructed. The road traffic flow at different times is used as the state variable, and the information released by various dynamic and static signs is used as the action. In a certain state, the agent needs to make a decision and take a certain action. Travelers choose to drive in the inner or outer lane based on the guidance information, and then enter the next state, and so on. The cycle of states and actions constitutes the main part of reinforcement learning. The reward is the negative number of the deviation of the balance index. The purpose of reinforcement learning is to maximize the total reward, that is, the balance between inner and outer lanes.
[0050] With the goal of balancing traffic flow between inner and outer lanes, an inner and outer lane balance index is constructed. The inner and outer lane balance index is: γ m Q is the ratio of the inner and outer lane saturation of the m-th road segment. 2m-1 Q 2m C represents the traffic flow of the inner and outer lanes of the m-th road segment, respectively. 2m-1 C 2m Let be the traffic capacity of the inner and outer lanes of the m-th road segment, respectively. The objective function is to minimize the deviation of the overall road balance index. Traffic capacity refers to the maximum hourly flow rate that a highway facility is expected to be able to handle under the design service level. The maximum hourly flow rate is calculated by referring to tables based on design indicators such as design speed and road width.
[0051] The input to the intelligent agent is the traffic state parameter, which is denoted as s. t =[s t,1 s t,2 , ..., s t,N ], where N is the number of road segments, s t,i The traffic flow state of the i-th road segment is represented by flow rate, speed, and density, which are state variables that persist throughout the entire method. These three state variables—flow rate, average speed, and density—interact with each other. The internal and external amplitude balance index is the agent's reward function, and the state variables influence the calculation of the reward value.
[0052] As a preferred embodiment of the present invention, two types of intelligent agents are defined, including real-dashed line intelligent agents and real-dashed line intelligent agents;
[0053] The solid-dashed line agent is used to determine whether vehicles traveling on the outer lane should enter the inner lane at the solid-dashed line. The action variables output by the solid-dashed line agent include the mid-distance destination and the mid-to-long-distance guidance direction. The solid-dashed line agent is responsible for guiding vehicles traveling on the outer lane to enter the inner lane at the solid-dashed line. The action variables of the solid-dashed line agent include the mid-distance destination and the mid-to-long-distance guidance direction, i.e., solving for [π, ζ, η], where π represents the corresponding number of the mid-distance destination, and ζ and η represent the mid-distance destination and the mid-to-long-distance guidance direction, respectively, with values ranging from [1, 0]. These represent entering the inner lane and maintaining travel on the outer lane, respectively.
[0054] The solid-dotted line agent is used to determine whether vehicles traveling within the solid-dotted line guide should leave the outer lane. The agent's output action variables include the destination and the guiding direction. The solid-dotted line agent is responsible for determining whether vehicles traveling within the solid-dotted line guide should enter the outer lane. The agent's action variables include the destination and the guiding direction, i.e., solving for [ρ, σ], where ρ represents the destination's corresponding number and σ represents the guiding direction, with values ranging from [1, 0], representing entering the outer lane and remaining within the inner lane, respectively.
[0055] Each agent employs an Actor-Critic architecture, where:
[0056] Actor Network: This serves as the policy generation network for each agent. The input to each agent's Actor Network is the global traffic state s. t =[s t,1 s t,2 , ..., s t,N This includes the traffic status of all road segments. This allows each agent to consider the global traffic flow when making decisions. Based on the global state, the Actor network outputs the agent's speed-limiting action 'a'. t,i And apply it to the environment.
[0057] Critic Network: As a policy evaluation network, the Critic Network is also based on the global state and the actions of all agents. t =[a t,1 a t,2 , ..., a t,K The Q-value of each agent is calculated to evaluate the value of the current action.
[0058] At each time step t, each agent's Actor network observes the global traffic state parameters s. t Select the induction action a t,i This action is then applied to the traffic environment. Under the intervention of the induced action, the traffic environment transforms into a new state. t+1 And calculate the corresponding reward value r. tAccording to traffic state parameter s t Induced action information a t,i The agent's reward value r t and new traffic state parameters s t+1 By using the Critic network to update the parameters of the agent's Actor network, the agent with the Actor network can update in real time and accurately obtain induced action information.
[0059] The execution phase includes the following steps:
[0060] Step 1: Observe the global state
[0061] At each time t, the intelligent agent system observes the state s of the entire transportation network. t =[s t,1 s t,2 , ..., s t,N ];
[0062] Step 2: Action selection for each agent
[0063] Actor network for each agent i Based on global traffic state s t Select the corresponding inter-amplitude induction action
[0064] Real-dashed line agents and virtual-real line agents generate different speed-limiting control actions according to their respective policies, which depend on the policy network of each agent.
[0065] Step 3: Apply Actions and Update the Environment
[0066] Each agent i applies a rate-limiting action a t,i The traffic environment is then updated to a new state s based on the actions of all agents. t+1 .
[0067] Step 4: Calculate the reward
[0068] For each agent i, based on the applied action and the new state s t+1 Calculate the reward value r for all agents. t .
[0069] Step 5: Store Experience
[0070] The current interaction data (s) t a t r t s t+1 It is stored in the global experience pool for use in subsequent training.
[0071] The training phase includes the following steps:
[0072] Step 1: Sampling empirical data
[0073] Randomly sample a batch of experiences (s) from the global experience pool. t a t r t s t+1 This is used to update the Actor and Critic networks of the intelligent agent.
[0074] Step 2: Critic Network Update
[0075] Critic network for each agent i Based on global state s t and the actions of all intelligent agents a t Value estimation of computational strategies
[0076] Update the Critic network parameter φ using TD error. i To minimize the loss function, the formula for calculating the loss function is as follows:
[0077]
[0078] Where γ is the discount factor, s t It represents the global traffic status, s t+1 It is the new state after the action is performed, a t It is the action of the intelligent agent, a t+1 It is the action of the next new intelligent agent. It is the value estimate of the agent's policy, r t It is the reward value of the agent. It is an expectation.
[0079] As a preferred solution, the Actor network for each agent i Alternatively, updates can be performed using the policy gradient method. The main steps include: calculating the policy gradient for each agent based on the Q-value of the Critic network.
[0080]
[0081] in, It is the policy gradient, θ i These are parameters used for Actor network updates, enabling the generation of better rate-limiting strategies. It is an expectation.
[0082] Example 2
[0083] Based on the same concept, a dynamic and static integrated variable information publishing device, including a processor and a splicing screen, was also proposed.
[0084] The processor outputs the guidance action information using any of the highway in-situ and out-of-situ control guidance methods described in Embodiment 1; the guidance action information is displayed through a video wall. The direction and destination displayed on the video wall are variable, including dynamic and static fusion variable information at solid and dashed lines and dynamic and static fusion in-situ control indicators at solid and dashed lines.
[0085] The variable message signs at solid and dashed lines, which integrate static and dynamic information, guide drivers to both medium- and long-distance destinations. By changing the medium-distance destination, the signs adjust the direction of entry into the inner lane, changing the arrows to variable directional arrows with two forms: forward and left-forward. This controls whether medium- and long-distance drivers enter the inner lane at the sign. A variable information board below the sign provides road condition explanations to help drivers understand the signs. A schematic diagram of the static and dynamic information signs at solid and dashed lines is shown below. Figure 2 As shown.
[0086] The dynamic and static integrated variable information display equipment is installed at the starting and midpoints of the solid and dashed lines and the solid and dashed lines where vehicles switch between directions; the dynamic and static integrated variable information display equipment is spliced with LED screens to make the direction and destination variable.
[0087] The dynamic and static integration control sign at the dashed and solid line, via a lower-mounted variable information display board, guides vehicles traveling on the inner lane to determine whether to enter the outer lane. For example... Figure 2 As shown, the bottom line of the video wall displays the message: "High traffic volume on the outer lane." Above this, arrows indicate the direction vehicles need to change lanes due to the high traffic volume. At the top, the destination for each lane is displayed. For example, "Chang'an" indicates the lane is towards Chang'an, and "Shenzhen" indicates the lane is towards Shenzhen. Drivers can use the video wall's prompts to change lanes in time, ensuring balanced traffic flow across lanes and improving traffic efficiency.
[0088] Example 3
[0089] Furthermore, to obtain parameters and verify the feasibility of the method, this embodiment also presents the construction of a simulation environment and the design of the intelligent agent within that environment. The specific steps of the simulation environment construction and intelligent agent design process include:
[0090] Step 1: Calibration of the inter-amplitude transition probability model
[0091] Using historical vehicle trajectory data, features related to inter-horizontal transitions are selected, mainly including: vehicle position. x l y speed v x v y Acceleration a, current span (inner or outer span) B, distance to destination D, distance to the end of the dashed / solid line or solid / dashed line Dtrans The consistency C between the current lane and the information board guidance scheme, and the distance d between the target vehicle and the vehicle in front and behind. front d rear The speed difference Δv between the vehicle before and after the target amplitude front Δv rear The traffic density K in the area.
[0092] Set the target variable y for inter-amplitude conversion, with values of 1 (inner amplitude to outer amplitude), -1 (outer amplitude to inner amplitude), and 0 (no conversion). Train the model based on the loss function. After the inter-amplitude conversion probability model is calibrated, it is used for simulation in step 2.
[0093] Step 2: Vehicle Intelligent Agent Design and Road Intelligent Agent Design
[0094] Multiple vehicle agents are created in the simulation environment, each representing a vehicle. The state of each vehicle agent is initialized, including its position, speed, acceleration, current location, lane, and destination. An inter-vehicle transition probability model is used as the behavioral rule for the vehicle agents.
[0095] The road is divided into M segments based on the road marking type. Combining the inner and outer lanes, the road is abstracted into N road agents, where N = 2 × M. Each road agent has a number N and a road length L. n The area where P is located n (Inner width P) n =0 outer width P n =1), Marker type M n (Double solid line M) n =0, solid and dashed lines M n =1, dashed and solid lines M n Attributes such as -1). Each road agent is responsible for monitoring traffic flow and density within its jurisdiction. A schematic diagram of parameter quantization in the simulation environment is shown below. Figure 5 As shown.
[0096] Within each simulation step, the state of each vehicle agent is updated, and its next action is calculated based on its current state and surrounding environmental information (occupancy status of road agents ahead and behind, speed difference, etc.); the traffic density and flow of each road agent are updated based on the actions of the vehicle agents.
[0097] Model Application: The trained agent is applied to a real traffic environment, with the traffic status (flow, density, speed) of each road segment collected by radar as the state variable and the content displayed by the fusion of dynamic and static signs as the action variable.
[0098] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for regulating and inducing traffic flow between inner and outer lanes of a highway, characterized in that, Includes the following steps: S1, Obtain traffic state parameters, including traffic flow, vehicle density, and average vehicle speed of each road segment; S2, the traffic state parameters are input into the trained agent with an Actor network, and the agent outputs guidance action information, which includes guidance direction and guidance destination, to realize real-time inter-amplitude guidance control; It also includes the step of training the agent with the Actor network: S21, the output of the guidance action information causes a change in the traffic state parameters, resulting in new traffic state parameters; S22, calculate the reward value of the agent based on the new traffic state parameters; S23, the agent is trained based on the traffic state parameters, guidance action information, agent reward value, and new traffic state parameters; Step S22 specifically includes the following steps: S220, Calculate the inner and outer lane balance index based on the new traffic state parameters; S221, Calculate the reward value of the inner and outer lane balance index according to the inner and outer lane balance index, and take the negative number of the sum of the inner and outer lane balance indices of each road segment as the reward value of the agent. The formula for calculating the reward value is as follows: ; where r t This is the reward value, where m is the road segment number, and M is the maximum value for the road segment number. It is an indicator of internal and external amplitude balance; The formula for calculating the internal and external amplitude balance index is as follows: in, Q is the ratio of the inner and outer lane saturation of the m-th road segment. 2m-1 Q 2m C represents the traffic flow of the inner and outer lanes of the m-th road segment, respectively. 2m-1 C 2m These represent the traffic capacity of the inner and outer lanes of the m-th road segment, respectively.
2. The method for regulating and guiding traffic between inner and outer lanes of a highway as described in claim 1, characterized in that, During the training of the Actor network, a Critic network is used as the evaluation network to calculate the evaluation value of the agent's induced action information using traffic state parameters and induced action information; the parameters of the Actor network are then updated based on the induced action information evaluation value.
3. The method for regulating and guiding traffic between inner and outer lanes of a highway as described in claim 1, characterized in that, The intelligent agent includes real-dashed line intelligent agents and real-dashed line intelligent agents; Solid-dashed line agent: used to determine whether vehicles induced to travel on the outer lane at solid-dashed line points should enter the inner lane; the action variables output by the solid-dashed line agent include the mid-distance destination and the mid- and long-distance induced directions; The solid-dotted line agent is used to determine whether a vehicle traveling within the inner lane of the solid-dotted line guidance has left the outer lane; the action variables output by the solid-dotted line agent include the destination and the guidance direction.
4. The method for regulating and guiding traffic between inner and outer lanes of a highway as described in claim 3, characterized in that, The guidance action information includes multiple sets of information, each set of information corresponding to a road segment number. Each set of information functions independently, and the output guidance action information realizes distributed real-time inter-amplitude guidance control.
5. The method for regulating and guiding traffic between inner and outer lanes of a highway according to claim 1, characterized in that, The location where the guidance action information is published includes the starting point and midpoint of the solid and dashed lines and the solid and dashed line segments where the vehicle makes a transition between widths.
6. A dynamic and static integrated variable information publishing device, characterized in that, Including processors and video walls; The processor outputs the guidance action information using a highway in-span and out-of-span regulation and guidance method as described in any one of claims 1-5. The induced action information is displayed through a video wall.
7. The dynamic and static integrated variable information publishing device as described in claim 6, characterized in that, The direction and destination displayed on the splicing screen are variable, including variable information on dynamic and static fusion at the solid and dashed lines and amplitude adjustment indicators for dynamic and static fusion at the solid and dashed lines: The variable information signs at solid and dashed lines guide drivers to medium and long-distance destinations. By changing the medium-distance destination, the signs adjust the objects entering the inner lane and change the arrows to variable directional arrows, which have two forms: forward and left forward, to control whether medium and long-distance drivers enter the inner lane at the variable information signs at solid and dashed lines. Variable information boards below the signs also explain the road conditions to help drivers understand the variable information signs at solid and dashed lines. The dynamic and static control signs at the dashed and solid lines guide vehicles traveling in the inner lane to determine whether to enter the outer lane via the lower-mounted variable information display.
Citation Information
Patent Citations
Traffic control method and device
CN108053661A
Model-free reinforcement learning
US20210086798A1