An intelligent traffic flow optimization method based on Internet of Vehicles

By introducing multi-objective game optimization module, V2V communication network, multi-stage dynamic planner and reinforced learning feedback loop into the Internet of Vehicles technology, the problem of difficulty in dealing with emergencies and global load balancing in the existing technology is solved, and rapid response and efficient optimization of traffic flow are achieved.

CN119863933BActive Publication Date: 2025-05-23QINGDAO TRAFFIC TECH INFORMATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510354083.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-05-23
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The existing Internet of Vehicles technology is difficult to deal with emergencies and global load balancing in traffic flow optimization, resulting in lagging optimization strategies and local optimization problems.

Method used

The multi-objective game optimization module, V2V communication network, multi-stage dynamic planner and reinforced learning feedback loop are adopted to dynamically adjust the weight coefficient and path planning parameters to achieve rapid information exchange and real-time decision-making optimization between vehicles.

Benefits of technology

The system can quickly respond to sudden changes in traffic flow, avoid the limitations of relying on periodic data updates, realize global load balancing and efficient operation of the traffic network, and reduce congestion and energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119863933B_ABST
    Figure CN119863933B_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent traffic flow optimization method based on vehicle networking, which relates to the field of intelligent traffic technology and comprises the following steps: dynamic adjustment of weight coefficients in a multi-objective game optimization module, the multi-objective game optimization module dynamically adjusts the weight coefficients of time, congestion and energy consumption according to different time periods; a V2V communication network realizes rapid exchange of vehicle status information; a multi-stage dynamic planner generates a path sequence based on a state transfer equation, predicts future states, and calculates an optimal path according to an optimization target; a reinforcement learning method is used to continuously optimize the multi-objective game optimization module and the multi-stage dynamic planner, and dynamically adjusts the weight parameters in the multi-objective game optimization module and the path planning parameters in the multi-stage dynamic planner; the invention can capture and respond to sudden changes in traffic flow in a short time, avoids the limitation of relying on periodic data updates, and prevents global congestion problems caused by local optimality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field related to intelligent transportation, and specifically to an intelligent traffic flow optimization method based on Internet of Vehicles. Background Art

[0002] The Internet of Vehicles is a technology that uses wireless communication technology to effectively utilize and manage all vehicle dynamic information on the information network platform through on-board equipment on vehicles. It can provide different functional services during vehicle operation. With the development of time, the Internet of Vehicles technology has gradually been applied to the field of traffic control. Intelligent traffic flow is becoming an important direction for urban traffic management. By adopting advanced technology and data analysis, urban traffic systems can operate and manage more efficiently, making traffic monitoring, traffic light control, route planning and other tasks more intelligent and precise, improving the overall efficiency of the traffic system. At the same time, using the Internet of Vehicles technology to control traffic lights can make traffic signal timing more flexible and can be adjusted according to real-time traffic flow, reducing traffic congestion and the occurrence of traffic accidents.

[0003] However, the existing Internet of Vehicles technology still has the following technical problems when in use: 1. The existing traffic flow optimization model relies on periodic or static data updates (such as sensor data at fixed time intervals), which makes it difficult to cope with sudden changes in traffic flow (such as sudden accidents, tidal congestion, etc.), resulting in delayed optimization strategies; 2. Traditional methods are based on single-objective optimization (such as the shortest travel time), and do not adequately consider the load balancing of the global traffic network. For example, optimizing a route for a vehicle may cause congestion to shift to other areas. Summary of the invention

[0004] In order to solve the defects of the prior art, the present invention provides an intelligent traffic flow optimization method based on the Internet of Vehicles.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0006] The present invention provides an intelligent traffic flow optimization method based on vehicle networking, comprising the following steps:

[0007] S1. Dynamic adjustment of weight coefficients in a multi-objective game optimization module, wherein the multi-objective game optimization module dynamically adjusts the weight coefficients of time, congestion, and energy consumption according to different time periods, so that the traffic flow system achieves a balance between vehicle speed, traffic capacity, and energy consumption;

[0008] S2, V2V communication network realizes rapid exchange of vehicle status information, and vehicles exchange information through V2V communication;

[0009] S3, multi-stage dynamic planner generates a path sequence based on the state transition equation. It generates the path sequence of the vehicle through multi-stage dynamic planning and updates the path planning parameters in real time. It uses the state transition equation to predict the future state according to the current state of the vehicle, control actions, and environmental factors, and calculates the best path according to the optimization goal.

[0010] S4. Use deep reinforcement learning methods to continuously optimize the multi-objective game optimization module and the multi-stage dynamic planner. By continuously feeding back reward signals through the system, the traffic flow system can adapt to different road environments and traffic conditions, and dynamically adjust the weight parameters in the multi-objective game optimization module and the path planning parameters in the multi-stage dynamic planner, thereby further optimizing the vehicle's driving strategy.

[0011] S5. Introduce the multi-agent system into the process of traffic flow optimization. Each vehicle acts as an agent and shares status information through V2V communication to form a global collaborative traffic network. Under the reinforcement learning framework, each agent makes the best decision based on the real-time traffic conditions to achieve a balance between local and global goals.

[0012] As a preferred technical solution of the present invention, the multi-objective game optimization module satisfies the following weight coefficient calculation method:

[0013] ;

[0014] Among them, α(t) represents the weight coefficient of the time target, T avg Expressed as the historical average travel time, T 0 It is expressed as the target travel time, and k is the adjustment coefficient of the function;

[0015] ;

[0016] Among them, β(t) represents the weight coefficient of the congestion target, ρ(t) is the regional vehicle density, which represents the number of vehicles per unit area, The impact of congestion on vehicle speed;

[0017] ;

[0018] Among them, γ(t) represents the weight coefficient of the energy consumption target, and the total weight of multiple targets should be 1.

[0019] As a preferred technical solution of the present invention, the state transfer equation of the multi-stage dynamic planner is as follows:

[0020] ;

[0021] Among them, Δt k→k+1It is represented as the path selection update amount between time k and time k+1, v k is the current speed of the vehicle, p k is the current position of the vehicle, E budget is the energy consumption budget, η is the congestion sensitivity coefficient, C local For local traffic information.

[0022] As a preferred technical solution of the present invention, the parameter update of the reinforcement learning follows the following update rules:

[0023] ;

[0024] θ t+1 Expressed as the update amount of the policy parameters, θ t is the parameter vector at the current moment, Denoted as the learning rate, δ TD Expressed as time difference error, π θ (a t |s t ) is the policy function, expressed as t Next, take action a according to the current parameter θ t The probability distribution of

[0025] ;

[0026] Among them, r t is represented by the current reward value, χ is represented by the discount factor, which is used to adjust the impact of future rewards, V(s t ) means in state s t The expected cumulative rewards.

[0027] As a preferred technical solution of the present invention, the V2V communication network adopts a hybrid communication protocol, including DSRC and C-V2X. The DSRC protocol is used when the vehicle distance is less than x meters, and the C-V2X protocol is used when the vehicle distance is greater than x meters.

[0028] As a preferred technical solution of the present invention, the feedback mechanism in the reinforcement learning includes a double update cycle,

[0029] Fast update cycle: The update cycle of local parameters is 50ms, which is mainly used to adjust the local strategies between vehicles so that the system can respond quickly to traffic changes;

[0030] Slow update cycle: The update cycle of global weight is 500ms, which is mainly used to adjust the weight parameters of the global game to ensure the realization of the global optimization goal.

[0031] As a preferred technical solution of the present invention, the multi-objective game optimization module adopts a game strategy based on time expansion, and the optimal strategy of the game adapts to the changing traffic conditions in different time periods.

[0032] As a preferred technical solution of the present invention, the V2V communication network adopts a layered transmission mechanism to improve the information transmission efficiency in a high-density traffic environment.

[0033] As a preferred technical solution of the present invention, the multi-stage dynamic planner introduces a path prediction model so as to timely adjust the path selection of the vehicle in the case of unstable or sudden traffic flow.

[0034] As a preferred technical solution of the present invention, the reinforcement learning feedback loop handles the learning task of large-scale vehicle networks through a sampling strategy.

[0035] The beneficial effects of the present invention are:

[0036] Through high-speed V2V communication, dynamic weight adjustment, multi-stage dynamic planner and reinforcement learning feedback loop, the system can capture and respond to sudden changes in traffic flow in a short time, avoiding the limitation of relying on periodic data updates;

[0037] By utilizing multi-objective game strategies, strategies based on time expansion, and global path prediction models, the system takes into account multiple objectives such as time, congestion, and energy consumption. It achieves effective local and global coordination through dynamically updated weight coefficients and dual update cycles, thus preventing global congestion problems caused by local optimality. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0039] Figure 1 Schematic diagram of the workflow of the intelligent traffic flow optimization method of the present invention. DETAILED DESCRIPTION

[0040] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0041] like Figure 1 As shown, an intelligent traffic flow optimization method based on vehicle networking includes the following steps:

[0042] S1. Dynamic adjustment of weight coefficients in the multi-objective game optimization module. The multi-objective game optimization module dynamically adjusts the weight coefficients of time, congestion, and energy consumption according to different time periods, so that the traffic flow system achieves a balance between vehicle speed, traffic capacity, and energy consumption; in view of the interaction between different vehicles, each vehicle adjusts its driving strategy according to the current traffic status, road condition information, and its own operating goals, so as to achieve collaborative optimization of the entire fleet; this module realizes dynamic weight adjustment of optimization goals through the Nash equilibrium mechanism in game theory. The weight coefficients of multi-objective optimization (time, congestion, and energy consumption) are dynamically changing. These coefficients are adjusted based on real-time data. Specifically, these parameters are adjusted by weighted sum method through factors such as historical travel time, real-time road flow, vehicle density, and energy consumption, so as to ensure that the system reduces congestion and energy consumption while ensuring efficiency;

[0043] The Nash equilibrium mechanism is as follows: each vehicle, as a decision-making subject, selects the optimal driving strategy based on its own operating goals (such as the shortest travel time or the lowest energy consumption) based on the current traffic status and road information, and constructs a utility function. For example, the following utility function can be defined for the i-th vehicle:

[0044] ;

[0045] Among them, T i Indicates the travel time, C i represents the congestion cost, E i Represents energy consumption. Each vehicle competes with each other in the decision-making process and eventually reaches the Nash equilibrium state, that is, the overall system achieves dynamic balance when no single vehicle can unilaterally change its strategy to obtain higher utility.

[0046] S2. V2V communication network realizes rapid exchange of vehicle status information. Vehicles exchange information through V2V communication, and the communication delay is controlled within 100ms (preferably within 50ms); this is to ensure that vehicles can quickly and in real time share information such as driving status, location, speed, etc., and make dynamic decisions and adjustments based on this. The network adopts a hybrid communication protocol. This design ensures the efficiency and coverage of the network, especially in high-speed driving or complex urban environments, which can ensure the real-time and accuracy of data transmission;

[0047] S3, multi-stage dynamic planner generates path sequence based on state transition equation, generates vehicle path sequence through multi-stage dynamic planning, and updates path planning parameters in real time. It uses state transition equation to predict future state according to vehicle current state, control action, environmental factors (such as traffic flow, road conditions), etc., and calculates the best path according to optimization goal. The dynamic planner can handle complex time-extended state transition so as to gradually optimize path planning in multiple state transitions. During vehicle driving, it updates the optimal path in real time to avoid congestion, high energy consumption, etc.

[0048] S4. Use deep reinforcement learning methods to continuously optimize the multi-objective game optimization module and the multi-stage dynamic planner. Through the system's continuous feedback of reward signals (such as traffic flow, energy consumption, travel time, etc.), the traffic flow system adapts to different road environments and traffic conditions, and dynamically adjusts the weight parameters in the multi-objective game optimization module and the path planning parameters in the multi-stage dynamic planner, thereby further optimizing the vehicle's driving strategy. This module uses reinforcement learning algorithms to provide real-time feedback and updates to the game optimization model, reducing congestion, reducing energy consumption, and improving traffic flow;

[0049] S5. Introduce the multi-agent system into the process of traffic flow optimization. Each vehicle acts as an agent and shares state information through V2V communication to form a global collaborative traffic network. Under the reinforcement learning framework, each agent makes the best decision based on the real-time traffic conditions to achieve a balance between local and global goals.

[0050] Specifically, in order to handle complex environments in multi-agent systems, deep reinforcement learning algorithms (such as deep Q networks DQN, A3C, or PPO) can be used to optimize the decision-making process of each vehicle. When making decisions, each vehicle inputs the following state variables:

[0051] State space: The current state of each vehicle includes position, speed, acceleration, traffic conditions ahead (such as traffic density, congestion level, signal light status, etc.), energy consumption, target time, etc. The state of the vehicle will change continuously, and these changes will affect the vehicle's action decision;

[0052] Action space: The possible actions of each vehicle include acceleration, deceleration, lane change, and maintaining speed. The vehicle selects the best action at each moment based on the current state to achieve the goal of optimizing traffic flow;

[0053] Reward function: In order to combine with the multi-objective game optimization module, the reward function needs to consider multiple objectives (time, congestion, energy consumption) at the same time and design the following reward mechanism:

[0054] ;

[0055] Among them, TP(t) stands for Time Penalty(t), which is the time penalty at time t. Usually, the time penalty is related to the time delay required for the vehicle, which may be related to the difference in target time or traffic delay. CP(t) stands for CongestionPenalty(t), which is the congestion penalty at time t. This penalty term is related to the position of the vehicle in the traffic flow and the degree of congestion, which may be related to traffic density or traffic volume. EE(t) stands for Energy Efficiency(t), which is the energy efficiency performance at time t. This indicator is usually related to the energy consumption of the vehicle, and the goal is to optimize energy consumption as much as possible.

[0056] In summary, the intelligent traffic flow optimization method achieves the following goals through the synergy of three core modules: multi-objective game optimization module, multi-stage dynamic planner and reinforcement learning feedback loop:

[0057] Break through the bottleneck of traditional data update periodicity or static update: use real-time data and online learning mechanism to achieve rapid response and dynamic decision-making to sudden traffic situations;

[0058] Achieve global load balancing optimization: Through multi-objective collaborative optimization and dynamic weight adjustment, effectively avoid local optimal problems caused by single-objective optimization, and ensure the balance and efficient operation of the entire transportation network.

[0059] Furthermore, the multi-objective game optimization module satisfies the following weight coefficient calculation method:

[0060] ;

[0061] Among them, α(t) represents the weight coefficient of the time target, T avg It is expressed as the historical average travel time, which is used to measure the current traffic congestion level of the road. 0 It is expressed as the target travel time, which represents the ideal travel time. k is the adjustment coefficient of the sigmoid function, which controls the response sensitivity to the traffic time target. It is dynamically adjusted through a sigmoid function, reflecting the relationship between the real-time traffic conditions of the road and the expected travel time. It uses an exponential function to make the weight change nonlinearly when the actual travel time deviates from the target, which plays a "threshold" effect.

[0062] When the historical average travel time T avg Above target time T 0 When the exponential term -k (T avg -T 0 ) is a negative value, the entire denominator tends to be less than 2, which increases α(t). At this time, the system will emphasize reducing travel time to alleviate congestion;

[0063] When the historical average travel time T avg Below target time T 0 When the exponential term -k (T avg -T 0 ) is a positive value, which makes α(t) decrease. At this time, the system can appropriately relax the pursuit of time optimization and pay more attention to other goals (such as reducing energy consumption);

[0064] ;

[0065] Among them, β(t) represents the weight coefficient of the congestion target, ρ(t) is the regional vehicle density, which represents the number of vehicles per unit area, directly affects the road congestion, and is an important parameter for measuring the degree of road congestion. λ is a constant parameter used to balance the impact between congestion and energy consumption. The impact of congestion on vehicle speed reflects the relationship between vehicle speed and traffic flow. It is calculated through the relationship with vehicle density and speed to reflect the road congestion. For example, when a small change in vehicle speed will significantly affect fuel consumption or emissions, the value of this item will be larger, thereby increasing β(t), prompting the system to give priority to smooth driving and energy saving and emission reduction;

[0066] ;

[0067] Among them, γ(t) is expressed as the weight coefficient of the energy consumption target. γ(t) is usually associated with energy consumption or other unclear indicators to ensure that the sum of the three weight coefficients is 1, so that the multi-objective optimization has a normalized characteristic. When the system increases α(t) and β(t) due to high travel time and congestion pressure, γ(t) decreases accordingly, indicating that the priority of energy consumption optimization is relatively low at this time; conversely, when the traffic conditions are good, γ(t) will increase to further optimize the energy consumption performance.

[0068] By collecting historical travel time, vehicle density, road condition information and energy consumption data in real time, the weight coefficients α(t), β(t) and γ(t) can be dynamically adjusted in different time periods, so that the optimization strategy can respond quickly to emergencies (such as accidents, tidal congestion, etc.);

[0069] By adopting the Nash equilibrium mechanism of game theory, each vehicle can consider the load balance of the entire system while considering its own operating goals, avoiding the situation where the optimal path of a single vehicle causes congestion in other areas, and contributing to the balanced scheduling of the overall transportation network.

[0070] Furthermore, the state transition equation of the multi-stage dynamic planner is as follows:

[0071] ;

[0072] The equation generates the path in real time through the optimization of multiple parameters and ensures that energy consumption is minimized;

[0073] Among them, Δt k→k+1 It is represented as the path selection update amount between time k and time k+1, v k is the current speed of the vehicle, which determines the priority of the vehicle in the current path selection. A higher speed means a lower congestion impact. k is the current position of the vehicle, which determines the starting point of the next state. This parameter determines the path sequence of the vehicle in dynamic planning. budget is the energy budget, which indicates the energy consumption limit of the vehicle to complete the task. The goal is to complete the task with the minimum energy consumption. η is the congestion sensitivity coefficient, which is used to dynamically adjust the congestion weight during path selection. C local It is local traffic information that affects the vehicle’s perception and decision-making of the current status, such as traffic signals and road conditions.

[0074] The state transition equation is designed to dynamically adjust the path selection of each vehicle, taking into account the following objectives:

[0075] Optimize traffic flow: Reduce traffic congestion through route selection, especially during peak hours or in complex environments, and ensure smooth traffic flow;

[0076] Minimize energy consumption: By choosing a path with lower energy consumption, each vehicle is ensured to complete the task in the shortest time while minimizing energy consumption. budget ) is the key control parameter, and the route selection will be adjusted according to the balance between energy consumption and traffic efficiency;

[0077] Adaptive path planning: Based on real-time traffic conditions and local information (C local ), dynamically update the path selection, the congestion sensitivity coefficient (η) is combined with local traffic information to optimize the flexibility of path selection and ensure effective operation in complex environments;

[0078] The multi-stage dynamic planner can update the path in real time at each stage according to the current status and future predictions, avoiding path congestion or excessive energy consumption caused by fixed path planning. By introducing energy budget and congestion sensitivity coefficient, path planning not only pursues the shortest path, but also pays more attention to global energy consumption and local congestion, so that the vehicle can complete the task in the shortest time while reducing overall energy consumption and congestion risks. The introduction of local traffic information in the state transfer equation can dynamically fine-tune the vehicle path according to the actual traffic conditions, thereby quickly responding to sudden traffic changes.

[0079] Here are some scenarios where the multi-stage dynamic planner and its state transition equations are particularly suitable:

[0080] Mixed scenarios of highways and urban roads: In this environment, vehicles may switch between different road networks, including highways and complex urban road networks. Dynamic routing can adjust the vehicle route in real time according to different road conditions (such as traffic flow, vehicle speed, energy consumption, etc.);

[0081] Congestion prediction and avoidance: During busy traffic hours, vehicles can predict and avoid traffic jams based on the congestion sensitivity coefficient (η) and local traffic information (C local ) Avoid congested areas, optimize routes, and reduce waiting time in congested sections, thereby improving traffic efficiency;

[0082] Long-haul transport and electric vehicle optimization: Especially in electric vehicles or energy-sensitive vehicle applications, the energy budget (E budget ) is crucial for route planning. The system can adjust the route based on factors such as current power, road conditions and traffic flow to ensure that energy consumption is minimized;

[0083] Urban Intelligent Transportation System: Through real-time communication between vehicles (V2V communication), the algorithm can achieve more accurate traffic flow optimization, especially in cooperative driving between multiple vehicles, avoiding congestion and improving overall traffic efficiency.

[0084] Furthermore, the parameter update of the reinforcement learning follows the following update rules:

[0085] ;

[0086] This feedback mechanism continuously optimizes the strategy by learning the relationship between the system's state and actions, ensuring the vehicle's optimal behavior in different traffic environments;

[0087] θ t+1 Expressed as the update amount of the policy parameter, θ t is the parameter vector at the current moment. Reinforcement learning improves the decision-making process by continuously optimizing strategy parameters. Represented as the learning rate, it is used to control the pace of policy updates. It is expressed as the feedback signal that optimizes the policy parameters through gradient descent. The policy parameters are optimized based on the difference between the behavior and reward of the current policy. TD It is expressed as a time difference error, which reflects the difference between the current strategy and the next state and is used to adjust the goal in the learning process, π θ (a t |s t ) is the policy function, expressed as t Next, take action a according to the current parameter θ t This method enables the system to continuously adapt to various traffic conditions and gradually improve decision-making accuracy.

[0088] If an action a t In status t The immediate effect produced under the TD >0), the update will increase the probability of the action in the strategy; conversely, if δ TD <0, the probability of the action will be reduced;

[0089] ;

[0090] Among them, r t It is represented as the current reward value, which means the reward obtained by the vehicle in the current state. χ is represented as the discount factor, which is used to adjust the impact of future rewards. V(s t ) means in state s t The expected cumulative reward r t Represented as the current reward value;

[0091] Through the discount factor χ, the update process not only considers the current instant reward r t , and also taking into account possible future rewards V (s t+1 ), thereby achieving a balance between vehicle speed, traffic capacity and energy consumption.

[0092] Reinforcement learning dynamically adjusts the strategy parameters θ by continuously obtaining feedback rewards, so that the vehicle's decision-making process can adapt to sudden traffic conditions (such as accidents and sudden congestion). Unlike traditional methods that rely on fixed-period data updates, reinforcement learning can perform online learning based on the latest state information at each time step, greatly shortening the response delay. The reinforcement learning feedback loop comprehensively considers multiple indicators such as traffic smoothness, energy consumption, and travel time during the update process, avoiding the local optimal problem caused by single-objective optimization, and helping to achieve balanced scheduling of global network loads.

[0093] In addition, deep Q learning (DQN) can be used as the decision-making mechanism of the intelligent agent. The network structure includes a deep neural network to fit the Q value function. Each vehicle learns a Q value based on the current state and feasible actions. Through continuous trial and feedback, the vehicle can gradually optimize its path selection, reduce congestion time, reduce energy consumption, and improve traffic efficiency. The formula is as follows:

[0094] ;

[0095] Among them, R(s t , a t ) is the reward function.

[0096] Further, the V2V communication network adopts a hybrid communication protocol, including DSRC and C-V2X, using the DSRC protocol when the distance between vehicles is less than x meters (preferably 20 meters), and using the C-V2X protocol when the distance between vehicles is greater than x meters;

[0097] The V2V communication network adopts a layered transmission mechanism to improve the efficiency of information transmission in a high-density traffic environment; specifically, when the distance between vehicles is small, the data transmission rate is improved through a short-distance low-latency transmission method; and when the distance between vehicles is large, the stability of information transmission is improved through a longer-distance communication protocol.

[0098] Among them, the setting of the hybrid communication protocol can ensure the communication reliability under different vehicle distance scenarios and avoid the problem of insufficient coverage of a single protocol. Its communication delay is ≤50ms, which is much higher than that of traditional sensors (usually ≥500ms). When an emergency accident signal occurs, it can be broadcast omnidirectionally within 50ms, and vehicles around the accident point can achieve real-time path replanning through fast protocol switching (DSRC priority).

[0099] When the distance between vehicles is small (such as on highways or urban roads), communication between vehicles needs to ensure low latency and high-speed data transmission, which can reduce the time of information propagation and respond to traffic changes in a timely manner;

[0100] When the distance between vehicles is large, a longer-distance communication protocol is used to ensure the stability and reliability of information transmission.

[0101] For example, on highways, when vehicles are close to each other, low-latency V2V communication protocols allow vehicles to quickly share information about braking or lane changes. In suburban or long-distance transportation, vehicles use remote communications to ensure that information can be transmitted to vehicles farther apart, maintaining smooth traffic flow.

[0102] Furthermore, the feedback mechanism in the reinforcement learning includes a double update cycle to improve the response speed and optimization effect of the system.

[0103] Fast update cycle: The update cycle of local parameters is 50ms, which is mainly used to adjust the local strategies between vehicles so that the system can respond quickly to traffic changes, especially in the real-time information exchange between vehicles. This means that when traffic conditions change rapidly, the system can adjust the decision-making strategy of each vehicle in time. For example, when traffic congestion occurs in a specific area, the vehicle will quickly change its driving strategy to avoid unnecessary stagnation. This update cycle ensures that vehicles can quickly adapt to changes in the surrounding environment through real-time information exchange (for example, through V2V communication), improving fluency in a short period of time;

[0104] Slow update cycle: The update cycle of global weights is 500ms, which is mainly used to adjust the weight parameters of the global game to ensure the realization of global optimization goals and ensure system stability, and will not be affected by sudden changes in a short period of time. When considering global optimization goals (such as shortest path, low emissions, etc.), the slow update cycle helps to adjust the priority of these goals to ensure the stable operation of the system and avoid large-scale fluctuations in the system due to local changes. When local traffic changes rapidly, such as sudden congestion on certain sections of the road, although this information is reflected and local strategies are adjusted in the fast update cycle, the slow update cycle ensures that the global optimization goals will not be disturbed by such instantaneous changes, thereby maintaining the long-term stability and efficiency of the system.

[0105] Furthermore, the multi-objective game optimization module adopts a game strategy based on time expansion. In different time periods, the optimal strategy of the game adapts to the changing traffic conditions and avoids strategy lags caused by over-reliance on historical data. Through the introduction of this strategy, the system can respond to sudden traffic incidents in the short term and achieve optimal traffic flow management through long-term game adjustments. During peak traffic periods, the system can dynamically adjust the game objectives between multiple vehicles in a short period of time by expanding the game strategy, thereby achieving optimal traffic management. In the long run, the adjustment of the game strategy can improve the optimization effect of the entire traffic flow through multi-stage learning and avoid the lag caused by over-reliance on historical data.

[0106] Furthermore, the multi-stage dynamic planner introduces a path prediction model so that the vehicle's path selection can be adjusted in a timely manner when traffic flow is unstable or sudden. The model can analyze traffic data over a period of time in the past, predict traffic conditions within a certain period of time in the future, and generate the vehicle's optimal path based on the prediction results.

[0107] By analyzing traffic data from the past and predicting traffic conditions in the future, combined with reinforcement learning strategy adjustments, vehicles can select the optimal path in the real-time traffic network. The model not only relies on historical data, but also takes into account multi-stage dynamic changes, allowing vehicles to make decisions based on real-time feedback. In the short term, the system can quickly respond to sudden traffic events (such as accidents, road closures, etc.), and in the long term, it continuously adjusts the path planning strategy through reinforcement learning, thereby achieving global traffic flow optimization.

[0108] Assuming an accident occurs on a certain road, the system can immediately adjust the vehicle's route through the path prediction model, and evaluate the optimal alternative route based on the multi-stage planner. In the long run, through dynamic planning, the system will gradually adjust the traffic strategy and optimize the overall traffic flow.

[0109] Furthermore, the reinforcement learning feedback loop handles the learning task of large-scale vehicle networks through a sampling strategy. Through this sampling strategy, the system can effectively learn the interactions between vehicles and their impact on traffic flow in a shorter time, thereby improving the learning efficiency and adaptability of the system.

[0110] By introducing appropriate sampling strategies in reinforcement learning, the system can effectively cope with complex decision-making problems in large-scale traffic networks. Through continuous sampling and updating, the system can obtain interaction information between vehicles in real time, and optimize it in combination with the update cycle to improve learning efficiency and adaptability.

[0111] When the system faces a surge in traffic flow, through sampling strategies, it can timely capture the vehicle behavior in this state and make adjustments within the fast update cycle. At the same time, it adjusts global parameters through slow update cycles to ensure that the system's global optimization goals are not disturbed by emergencies.

[0112] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An intelligent traffic flow optimization method based on Internet of Vehicles, characterized in that: The following steps are involved: S1. Dynamic adjustment of weight coefficients in a multi-objective game optimization module, wherein the multi-objective game optimization module dynamically adjusts the weight coefficients of time, congestion, and energy consumption according to different time periods, so that the traffic flow system achieves a balance between vehicle speed, traffic capacity, and energy consumption; S2, V2V communication network realizes rapid exchange of vehicle status information, and vehicles exchange information through V2V communication; S3, multi-stage dynamic planner generates a path sequence based on the state transition equation. It generates the path sequence of the vehicle through multi-stage dynamic planning and updates the path planning parameters in real time. It uses the state transition equation to predict the future state according to the current state of the vehicle, control actions, and environmental factors, and calculates the best path according to the optimization goal. S4. Use deep reinforcement learning methods to continuously optimize the multi-objective game optimization module and the multi-stage dynamic planner. By continuously feeding back reward signals through the system, the traffic flow system can adapt to different road environments and traffic conditions, and dynamically adjust the weight parameters in the multi-objective game optimization module and the path planning parameters in the multi-stage dynamic planner, thereby further optimizing the vehicle's driving strategy. S5. Introduce the multi-agent system into the process of traffic flow optimization. Each vehicle acts as an agent and shares status information through V2V communication to form a global collaborative traffic network. Under the reinforcement learning framework, each agent makes the best decision based on the real-time traffic conditions to achieve a balance between local and global goals.

2. According to claim 1, the intelligent traffic flow optimization method based on the Internet of Vehicles is characterized in that: The multi-objective game optimization module satisfies the following weight coefficient calculation method: ; Among them, α(t) represents the weight coefficient of the time target, T avg It is represented as the historical average travel time, T0 is represented as the target travel time, and k is the adjustment coefficient of the function; ; Among them, β(t) represents the weight coefficient of the congestion target, ρ(t) is the regional vehicle density, which represents the number of vehicles per unit area, The impact of congestion on vehicle speed; ; Among them, γ(t) represents the weight coefficient of the energy consumption target, and the total weight of multiple targets should be 1.

3. The intelligent traffic flow optimization method based on the Internet of Vehicles according to claim 1 is characterized in that: The state transition equation of the multi-stage dynamic planner is as follows: ; Among them, Δt k→k+1 It is represented as the path selection update amount between time k and time k+1, v k is the current speed of the vehicle, p k is the current position of the vehicle, E budget is the energy consumption budget, η is the congestion sensitivity coefficient, C local For local traffic information.

4. The intelligent traffic flow optimization method based on the Internet of Vehicles according to claim 1 is characterized in that: The parameter update of the reinforcement learning follows the following update rules: ; θ t+1 Expressed as the update amount of the policy parameter, θ t is the parameter vector at the current moment, Denoted as the learning rate, δ TD Expressed as time difference error, π θ (a t |s t ) is the policy function, expressed as t Next, take action a according to the current parameter θ t The probability distribution of ; Among them, r t is represented by the current reward value, χ is represented by the discount factor, which is used to adjust the impact of future rewards, V(s t ) means in state s t The expected cumulative rewards.

5. The intelligent traffic flow optimization method based on the Internet of Vehicles according to claim 1 is characterized in that: The V2V communication network adopts a hybrid communication protocol, including DSRC and C-V2X. The DSRC protocol is used when the vehicle distance is less than x meters, and the C-V2X protocol is used when the vehicle distance is greater than x meters.

6. The intelligent traffic flow optimization method based on the Internet of Vehicles according to claim 1 is characterized in that: The feedback mechanism in the reinforcement learning described above includes a double update cycle, Fast update cycle: The update cycle of local parameters is 50ms, which is mainly used to adjust the local strategies between vehicles so that the system can respond quickly to traffic changes; Slow update cycle: The update cycle of global weight is 500ms, which is mainly used to adjust the weight parameters of the global game to ensure the realization of the global optimization goal.

7. The intelligent traffic flow optimization method based on the Internet of Vehicles according to claim 1 is characterized in that: The multi-objective game optimization module adopts a game strategy based on time expansion. In different time periods, the optimal strategy of the game adapts to the changing traffic conditions.

8. The intelligent traffic flow optimization method based on the Internet of Vehicles according to claim 1 is characterized in that: The V2V communication network adopts a layered transmission mechanism to improve information transmission efficiency in a high-density traffic environment.

9. The intelligent traffic flow optimization method based on the Internet of Vehicles according to claim 1 is characterized in that: The multi-stage dynamic planner introduces a path prediction model so as to timely adjust the path selection of the vehicle in case of unstable or sudden traffic flow.

10. The intelligent traffic flow optimization method based on the Internet of Vehicles according to claim 1, characterized in that: The reinforcement learning feedback loop handles the learning task of large-scale vehicle networks by sampling strategies.

Citation Information

Patent Citations

  • Urban road network path planning method for non-global information

    CN108847037A

  • Vehicle-road collaborative driving system and method combining vehicle correlation degree and game theory

    CN113920740A