Multi-lane game decision-making method for vehicle groups based on collision avoidance risk assessment under zero-trust architecture
By combining the zero-trust architecture and the Nash Q-learning algorithm with the collision risk assessment method of potential field theory, the problems of information security threats and decision-making lags in the intelligent Internet of Vehicles are resolved, safe and efficient decision-making for multi-lane lane changes in vehicle groups is achieved, and the overall safety and efficiency of traffic flow is improved.
Patent Information
- Application Number
- CN202510478903.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-04-16
AI Technical Summary
Existing intelligent vehicle networking decision-making methods face problems such as information security threats, inaccurate collision risk assessment, delayed decision-making, and lack of dynamic adjustment capabilities in complex and changing traffic environments, making it difficult to achieve safe and efficient multi-lane lane changes for vehicle groups.
A collision risk assessment method under a zero-trust architecture is adopted, combined with potential field theory and Nash Q-learning algorithm to obtain vehicle status and road information in real time. Through trust evaluation and game decision-making models, the vehicle lane change strategy is dynamically adjusted, taking into account the dynamic relationship between vehicles and real-time traffic conditions.
It achieves accurate simulation of the dynamic relationship between vehicles in complex traffic environments, improves the accuracy of collision risk assessment and the adaptability of decision-making, and enhances the safety and efficiency of multi-lane lane changes.
Smart Images

Figure CN120220413B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent connected vehicles, and in particular relates to a multi-lane game decision-making method for a vehicle group based on collision avoidance risk assessment under a zero-trust architecture. Background Art
[0002] With the continuous development of intelligent connected vehicle technologies, IoV applications are becoming increasingly widespread in complex traffic environments. This is particularly true in the research of swarm collaboration and adaptive lane change decision-making. Inter-vehicle coordination and decision-making mechanisms are crucial for achieving safe lane changes and efficient traffic flow. However, current IoV swarm decision-making methods still have shortcomings in complex and changing traffic environments.
[0003] Existing intelligent vehicle networking decision-making methods typically assume that inter-vehicle information is authentic and reliable. However, in an open vehicle networking environment, this information may be subject to security threats such as tampering, forgery, or delay. Secondly, existing collision risk assessment models are mostly designed for static traffic environments, making them difficult to adapt to complex multi-lane or dynamically changing scenarios, and their assessment accuracy is insufficient. These models are usually designed for single-lane or simple traffic scenarios and cannot effectively handle vehicle interactions and conflicts in complex multi-lane environments. This results in a lack of reliable basis for lane change decisions in multi-lane traffic environments. Furthermore, existing lane change decision-making mechanisms mostly rely on preset rules or fixed strategies, lack dynamic adjustment capabilities, and fail to fully utilize real-time traffic status and collision risk information, resulting in delayed or ineffective decisions.
[0004] Zero Trust Architecture is a new cybersecurity model whose core concept is "never trust, always verify." In intelligent connected vehicles (IoVs) under a Zero Trust architecture, the overall security of the cyber-physical system is ensured from the source by dynamically assessing the trustworthiness of network nodes and limiting their information processing permissions based on the assessment results. However, even within this framework, issues remain, such as communication security issues affecting the credibility of IoV decisions, the difficulty of dynamically assessing collision risk in multi-lane scenarios, and the lack of dynamic adaptability in game-based decision-making strategies. Accurately simulating the dynamic relationships between vehicles and the complex interactions between multiple vehicles during lane changes, and adapting to the influence of multiple environmental factors such as real-time traffic flow and other vehicle behavior to achieve safe and efficient multi-lane lane changes within a swarm, are key to IoV's ability to respond to emergencies in complex traffic environments, reduce the risk of accidents, and improve overall road safety and efficiency. Summary of the Invention
[0005] In response to the problems existing in the above-mentioned prior art, the present application adopts a collision risk assessment method based on potential field theory under a zero-trust architecture, which comprehensively considers key parameters such as vehicle mass, speed, distance, relative motion, etc., and combines the attenuation characteristics of potential field intensity with distance and the trust between vehicles to achieve accurate simulation of the dynamic relationship between vehicles; by simulating the game interaction between multiple vehicles, the relative position of each vehicle, the collision risk and the trust between vehicles are taken into account, and the cost and benefit of the vehicle in the lane changing process are effectively calculated; Nash Q-learning is used to learn the game strategy, and the parameters in the payment function are dynamically adjusted based on feedback information. Through continuous learning and feedback, the vehicle can adjust its decision according to changes in real-time traffic flow, adjacent vehicle behavior, road conditions, etc. The present invention aims to solve the problem of safe and efficient multi-lane lane changing of vehicle groups in complex traffic environments.
[0006] In order to achieve the above technical objectives, the present invention provides the following technical solutions:
[0007] The multi-lane game decision-making method for vehicle groups based on collision avoidance risk assessment under a zero-trust architecture specifically includes the following steps:
[0008] S1. Build an intelligent vehicle network to obtain real-time vehicle status and road information. Vehicle status information includes the speed, acceleration, position, mass, and size of the vehicle and adjacent vehicles, the distance between the vehicle and adjacent vehicles, historical interaction records of adjacent vehicles, and vehicle trustworthiness. Road information includes lane width, lane line position, and number of lanes.
[0009] S2. Real-time assessment of vehicle collision risk through zero-trust architecture and potential field theory;
[0010] S3. Design a multi-lane lane-changing game decision model based on vehicle collision risk, decomposing the multi-lane lane-changing problem into multiple two-vehicle game models;
[0011] S4. Considering both safety and lane-changing efficiency, the vehicle's lane-changing decisions are learned through the Nash Q-learning algorithm. The weight parameters of the game payoff function are adaptively and dynamically adjusted. This allows the vehicle to adaptively adjust its lane-changing decisions according to different traffic environments and complete multi-lane lane changes in a group of vehicles.
[0012] Furthermore, step S2 specifically includes:
[0013] S21. Under the zero-trust architecture, dynamically and continuously evaluate the trust of vehicle i in vehicle j in the intelligent vehicle network, including evaluating the direct trust of vehicle i in vehicle j. and recommendation trust
[0014] S22. Weight the direct trust and the recommended trust to obtain the comprehensive trust T of vehicle node i to vehicle node j. i,j , the formula is:
[0015]
[0016] Among them, ω is the weight coefficient that balances the influence of direct trust and recommendation trust in the overall trust. When there are more direct interactions between vehicle nodes i and j, they are less affected by external factors and are given a higher direct trust weight. When the number of interactions between vehicle nodes i and j is insufficient, the direct trust data is less and of lower quality, and a higher indirect trust weight is given.
[0017] S23. Establish a collision risk field that dynamically reflects the collision risk between vehicles. The collision risk field is a potential field that repel each other between vehicles. The direction of the field strength is from adjacent vehicles toward the own vehicle. The field strength is determined by the equivalent mass, motion state, spatial position, and overall trust in the own vehicle.
[0018] S24. Obtain the field strength E of vehicle j to vehicle i from the collision risk field ij Based on the potential field theory, we can get the repulsive force F of vehicle j on its own vehicle: ij ; Then calculate the potential energy Ec of vehicle j on vehicle i through repulsive force ij , as the collision risk of vehicle j to vehicle i; sum the potential energy of all vehicles in the scene to the ego vehicle, and obtain the collision risk Ec faced by the ego vehicle i ; The formula is:
[0019] F ij =E ij ·M i ;
[0020]
[0021]
[0022] Among them, M i is the equivalent mass of the vehicle, r ij is the distance between vehicle j and ego vehicle i.
[0023] More specifically, the specific evaluation method of the direct trust and recommended trust of vehicle i to vehicle j in step S21 is:
[0024] Direct trust is based on the historical behavior of the vehicle, evaluating its reliability in the network and introducing a penalty factor to constrain illegal behavior. The calculation formula of the direct trust of vehicle i to vehicle j is expressed as:
[0025]
[0026] in, They represent the direct trust of vehicle i in vehicle j at the kth and k-1th time steps in the trust evaluation process respectively; ρ is the time decay factor, which is a preset fixed parameter;
[0027] η k represents the penalty factor, which quantifies the degree of violation of vehicle j in the current lane change process; the penalty factor η k Based on whether the vehicle spacing and acceleration of vehicle j during the lane change game exceed the safety threshold, and based on the duration and degree of violation, the calculation formula is:
[0028]
[0029] Among them, d k 、a k Indicates the current distance and acceleration, d sa and a sa They represent the preset distance safety threshold and acceleration safety threshold respectively; τ is a parameter used to adjust the sensitivity of the model, and T represents the maximum violation duration during the vehicle game lane change process, which is calculated from the moment the acceleration exceeds the safety threshold.
[0030] Recommended Trust It is equal to the weighted sum of the comprehensive trust of other nodes in the network to vehicle j, where the weight is the comprehensive trust of vehicle i to other nodes; other nodes refer to other vehicles except vehicle i and vehicle j; the recommendation trust at the kth time step is The calculation formula is:
[0031]
[0032] More specifically, the calculation method of the midfield strength in step S23 is as follows:
[0033] The calculation method of the electric field strength of vehicle j to vehicle i is different according to the different motion states of vehicle j. If j = a and j = b are used to represent vehicle j in a stationary state and a moving state respectively, then:
[0034]
[0035] in, Respectively represent the magnitude of the field strength generated by the stationary vehicle and the moving vehicle on the ego vehicle i; M a 、M b are the equivalent masses of the stationary vehicle and the moving vehicle respectively; r ia 、r ib is the position vector between the stationary vehicle, the moving vehicle and the ego vehicle i, and its modulus is the distance between the vehicles; θb is the moving direction and position vector r of the moving vehicle bi The angle between them, v b is the actual speed of the moving vehicle; k1 and k2 are coefficients describing the attenuation of the field strength with distance;
[0036] The formula for calculating the equivalent mass of a stationary vehicle is:
[0037] M a =m a ·g(v a );
[0038]
[0039] Among them, m a is the actual weight of the stationary vehicle; v a is the actual speed of the stationary vehicle;
[0040] The equivalent mass of the moving vehicle and the ego vehicle is calculated in the same way as for a stationary vehicle;
[0041] k1 and k2 are determined by the comprehensive trust of vehicle i to vehicle j, specifically:
[0042] k1=T i,a +1;
[0043] k2=T i,b +1;
[0044] Among them, T i,a and T i,b are the comprehensive trust of vehicle i to the stationary vehicle a and the moving vehicle b respectively.
[0045] Furthermore, step S3 specifically includes:
[0046] S31. Design the game participants and strategy space. The game participants are the ego vehicle (Ego) and the neighboring vehicle (Side_lag). The Ego vehicle's action options are lane change (CL) and lane retention (KCL). The Side_lag vehicle's action options are acceleration (ACC) and deceleration (DEC). Acceleration (ACC) indicates a rejection of the cooperative strategy, while deceleration (DEC) indicates a cooperation strategy.
[0047] S32, game payment design; remember Q mn and P mn They represent the game payoffs of the Ego vehicle and the Side_lag vehicle, respectively, where m and n represent the actions chosen by the Ego vehicle and the Side_lag vehicle, respectively;
[0048] When m=1 and n=1, it indicates that the Ego vehicle chooses to change lanes and the Side_lag vehicle chooses to accelerate;
[0049] When m=1 and n=2, it means that the Ego vehicle chooses to change lanes and the Side_lag vehicle chooses to decelerate;
[0050] When m=2 and n=1, it means that the Ego vehicle chooses to maintain the current lane and the Side_lag vehicle chooses to accelerate;
[0051] When m=2 and n=2, it means that the Ego vehicle chooses to maintain the current lane and the Side_lag vehicle chooses to decelerate;
[0052] S33. The Ego vehicle and the Side_lag vehicle engage in a game. If the game result indicates that the Ego vehicle can cooperate with the Side_lag vehicle, the lane of the Side_lag vehicle is considered as the potential target lane of the Ego vehicle. If the Ego vehicle has multiple potential target lanes, the lane with the highest comprehensive benefit is selected as the final target lane. The comprehensive benefit is the sum of the two game payments. For situations where multiple lanes need to be crossed, the Ego vehicle will repeat the above process in sequence after each game is completed until the final lane change goal is achieved.
[0053] Furthermore, step S4 specifically includes:
[0054] S41. Define the state space and action space. The state space stores the relationship and dynamic information between the ego vehicle (Ego) and the neighboring vehicle (Side_lag), including the relative position, velocity, and acceleration of the ego vehicle and the neighboring vehicle (Side_lag). The action space stores the possible actions that the ego vehicle and the side_lag vehicle can take in each state. For the ego vehicle, the action space stores two options: change lanes (CL) or stay in the current lane (KCL). For the side_lag vehicle, the action space stores two options: accelerate (ACC) or decelerate (DEC).
[0055] S42. Design rewards by considering both safety and lane change efficiency; efficiency rewards r s The formula is expressed as:
[0056]
[0057] Among them, v represents the current speed of the vehicle, v des Indicates the target speed desired by the vehicle;
[0058] Security Reward e The formula is expressed as:
[0059]
[0060] Among them, d min Indicates the minimum distance between the vehicle and the nearest vehicle, d safeIndicates the target distance expected by the vehicle;
[0061] The total immediate reward is:
[0062] R=μr s +(1-μ)r e ;
[0063] Among them, μ represents the average level of trust among game participants, that is, the average value of the comprehensive trust of the ego vehicle in neighboring vehicles, which is used to control the weight balance between safety rewards and efficiency rewards. When the average level of trust is high, efficiency is prioritized, and when the average level of trust is low, safety is prioritized. The formula for μ is expressed as:
[0064]
[0065] Among them, T i,j is the comprehensive trust of vehicle i in other vehicles j; J is the total number of other vehicles participating in the game;
[0066] S43, Q-value update and dynamic parameter adjustment: Whenever a vehicle selects an action, the Nash Q-learning algorithm learns the Nash equilibrium solution based on the total reward it receives, and dynamically adjusts the parameters of the payoff function based on the feedback;
[0067] The Q value update formula is:
[0068] Q(s,a1,a2)←(1-α t )Q(s,a1,a2)+α t [R+γNash Q(s′,a′1,a′2)];
[0069] Among them, s represents the current state, a1 is the action currently taken by the Ego vehicle, a2 is the action currently taken by the Side_lag vehicle, R is the immediate total reward; γ is the discount factor, which represents the impact of future rewards; α t is the learning rate; Nash Q(s',a'1,a'2) is the expected cumulative return of the current vehicle when all vehicles participating in the game act according to the Nash equilibrium strategy in the next state;
[0070] Each time a vehicle selects an action, the Nash Q-learning algorithm learns a Nash equilibrium solution based on the total reward it receives, while also dynamically adjusting the parameters of the payoff function based on feedback. Specifically, if a collision occurs during a lane change and safety is insufficient, Nash Q-learning increases the weighting of safety parameters, emphasizing collision avoidance. If the lane change improves driving efficiency, Nash Q-learning increases the weighting of efficiency parameters, prioritizing decisions that improve traffic flow.
[0071] S44. Through multiple interactions between the vehicle and neighboring vehicles, the Q value and the weight parameters of the payment function are continuously adjusted, and the Q value is converged through experiments and feedback, thereby achieving a balance between safety and efficiency, and ultimately enabling the vehicle to adaptively adjust lane change decisions according to different traffic environments.
[0072] Based on the above technical solution, the present invention has at least the following beneficial effects:
[0073] 1. A collision risk assessment model based on a zero-trust framework comprehensively considers key vehicle parameters such as mass, speed, distance, and relative motion. This model, combined with the attenuation of potential field strength over distance, simulates the dynamic relationships between vehicles and potential collision risks, making collision risk assessment more accurate and dynamic. By combining the attenuation of potential field strength over distance with inter-vehicle trust, this model accurately simulates the dynamic relationships between vehicles, improving the accuracy and real-time nature of risk assessment.
[0074] 2. Cooperative Game Theory Enhances the Synergy of Multi-Lane Decision-Making: Based on game theory, multi-lane lane change decision-making simulates the interactive game between multiple vehicles, taking into account each vehicle's relative position, collision risk, and node trust. This effectively calculates the costs and benefits of a vehicle's lane change. This cooperative game model enables vehicles to dynamically adjust their lane change strategies to avoid conflicts with other vehicles, improving overall traffic flow efficiency and safety.
[0075] 3. Adaptive Game Payoff Adjustment Mechanism: This system uses the Nash Q-learning algorithm to learn Nash equilibrium solutions and adaptively adjust the parameters in the game model. Through continuous learning and feedback, vehicles can adjust their decisions based on real-time traffic flow, neighboring vehicle behavior, and road conditions, improving both adaptability and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0077] Figure 1 This is the overall framework diagram of the method proposed in the present invention;
[0078] Figure 2 Flowchart for real-time assessment of vehicle collision risk;
[0079] Figure 3 This is a flowchart of a multi-lane lane-changing game decision model based on vehicle collision risk;
[0080] Figure 4 Schematic diagram of the process of adaptive optimization of game model parameters. DETAILED DESCRIPTION
[0081] In order to make the purpose, technical solutions and advantages of the present invention more clear, the following Figure 1-4 It should be understood that the specific embodiments described herein are only used to illustrate the present invention and are not intended to limit the present invention.
[0082] Although the steps in the present invention are arranged with numbers, they are not intended to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a step requires other steps as a basis, the relative order of the steps can be adjusted. It is understood that the term "and / or" used herein refers to and covers any and all possible combinations of one or more of the associated listed items.
[0083] like Figure 1 As shown, the present invention shows a multi-lane game decision-making method for a vehicle group based on collision avoidance risk assessment under a zero-trust architecture, which specifically includes the following steps:
[0084] S1. Build an intelligent vehicle network to obtain real-time vehicle status information and road information; vehicle status information includes the speed, acceleration, position, mass, size of the vehicle and adjacent vehicles, the distance between the vehicle and adjacent vehicles, the historical interaction records of adjacent vehicles, and vehicle trust; road information includes lane width, lane line position, and number of lanes; this application performs real-time collision risk assessment based on this vehicle status information, specifically, Figure 1 As shown, the severity of the collision is reflected by the equivalent mass of the vehicle (i.e., taking into account the speed and weight of the vehicle). The higher the equivalent mass, the higher the severity of the vehicle collision; the possibility of collision is evaluated by the vehicle spacing and trust; the collision risk assessment combines the severity of the collision and the possibility of collision, and quantifies them using potential field theory; at the same time, road information will also affect the cooperative game decision between vehicles, so this application also takes it into consideration; the specific risk assessment and game decision will be described in steps S2-S4.
[0085] S2. Real-time assessment of vehicle collision risk through zero-trust architecture and potential field theory;
[0086] As a preferred embodiment, Figure 2 As shown, step S2 specifically includes:
[0087] S21. Under the zero-trust architecture, dynamically and continuously evaluate the trust of vehicle i in vehicle j in the intelligent vehicle network, including evaluating the direct trust of vehicle i in vehicle j. and recommendation trust
[0088] More specifically, the specific evaluation method of the direct trust and recommended trust of vehicle i to vehicle j in step S21 is:
[0089] Direct trust is based on the historical behavior of the vehicle, evaluating its reliability in the network and introducing a penalty factor to constrain illegal behavior. The calculation formula of the direct trust of vehicle i to vehicle j is expressed as:
[0090]
[0091] in, represents the direct trust of vehicle i in vehicle j at the kth and k-1th time steps in the trust evaluation process, respectively; ρ is the time decay factor, which is a preset fixed parameter; in this application, the direct trust and recommended trust are values that are dynamically updated in real time with each time step, that is, the causality with the previous time step is taken into account when calculating the trust;
[0092] η k represents the penalty factor, which quantifies the degree of violation of vehicle j in the current lane change process; the penalty factor η k Based on whether the vehicle spacing and acceleration of vehicle j during the lane change game exceed the safety threshold, and based on the duration and degree of violation, the calculation formula is:
[0093]
[0094] Among them, d k 、a k Indicates the current distance and acceleration, d sa and a sa They represent the preset distance safety threshold and acceleration safety threshold respectively; τ is a parameter used to adjust the sensitivity of the model, and T represents the maximum violation duration during the vehicle game lane change process, which is calculated from the moment the acceleration exceeds the safety threshold.
[0095] Recommended Trust It is equal to the weighted sum of the comprehensive trust of other nodes in the network to vehicle j, where the weight is the comprehensive trust of vehicle i to other nodes; other nodes refer to other vehicles except vehicle i and vehicle j; the recommendation trust at the kth time step is The calculation formula is:
[0096]
[0097] S22. Weight the direct trust and the recommended trust to obtain the comprehensive trust T of vehicle node i to vehicle node j. i,j , the formula is:
[0098]
[0099] Among them, ω is the weight coefficient that balances the influence of direct trust and recommendation trust in the overall trust. When there are more direct interactions between vehicle nodes i and j, they are less affected by external factors and are given a higher direct trust weight. When the number of interactions between vehicle nodes i and j is insufficient, the direct trust data is less and of lower quality, and a higher indirect trust weight is given.
[0100] S23. Establish a collision risk field that dynamically reflects the collision risk between vehicles. The collision risk field is a potential field that repel each other between vehicles. The direction of the field strength is from adjacent vehicles toward the own vehicle. The field strength is determined by the equivalent mass, motion state, spatial position, and overall trust in the own vehicle.
[0101] More specifically, the calculation method of the midfield strength in step S23 is as follows:
[0102] The calculation method of the electric field strength of vehicle j to vehicle i is different according to the different motion states of vehicle j. If j = a and j = b are used to represent vehicle j in a stationary state and a moving state respectively, then:
[0103]
[0104] in, Respectively represent the magnitude of the field strength generated by the stationary vehicle and the moving vehicle on the ego vehicle i; M a 、M b are the equivalent masses of the stationary vehicle and the moving vehicle respectively; r ia 、r ib is the position vector between the stationary vehicle, the moving vehicle and the ego vehicle i, and its modulus is the distance between the vehicles; θ b is the moving direction and position vector r of the moving vehicle bi The angle between them, v b is the actual speed of the moving vehicle; k1 and k2 are coefficients describing the attenuation of the field strength with distance; the present invention introduces trust coefficients k1 and k2 based on trust into the collision assessment model. When other conditions are the same (distance, mass, speed), it can ensure that vehicles with higher trust are considered to have higher driving stability and predictable behavior, thus posing a lower collision risk to the ego vehicle;
[0105] The formula for calculating the equivalent mass of a stationary vehicle is:
[0106] M a =m a ·g(v a );
[0107]
[0108] Among them, m a is the actual weight of the stationary vehicle; v a is the actual speed of the stationary vehicle;
[0109] The equivalent mass of the moving vehicle and the ego vehicle is calculated in the same way as for a stationary vehicle;
[0110] k1 and k2 are determined by the comprehensive trust of vehicle i to vehicle j, specifically:
[0111] k1=T i,a +1;
[0112] k2=T i,b +1;
[0113] Among them, T i,a and T i,b are the comprehensive trust of vehicle i towards stationary vehicle a and moving vehicle b respectively;
[0114] S24. Obtain the field strength F of vehicle j to vehicle i from the collision risk field ij Based on the potential field theory, we can get the repulsive force F of vehicle j on its own vehicle: ij ; Then calculate the potential energy Ec of vehicle j on vehicle i through repulsive force ij , as the collision risk of vehicle j to vehicle i; sum the potential energy of all vehicles in the scene to the ego vehicle, and obtain the collision risk Ec faced by the ego vehicle i ; The formula is:
[0115] F ij =E ij ·M i ;
[0116]
[0117]
[0118] Among them, M i is the equivalent mass of the vehicle, r ij is the distance between vehicle j and ego vehicle i.
[0119] S3. Design a multi-lane lane-changing game decision model based on vehicle collision risk, decomposing the multi-lane lane-changing problem into multiple two-vehicle game models;
[0120] As a preferred embodiment, Figure 3 As shown, step S3 specifically includes:
[0121] S31. Design the game participants and strategy space. The game participants are the ego vehicle (Ego) and the neighboring vehicle (Side_lag). The Ego vehicle's action options are lane change (CL) and lane retention (KCL). The Side_lag vehicle's action options are acceleration (ACC) and deceleration (DEC). Accelerating (ACC) indicates a decision to reject the cooperative strategy (i.e., accelerating appropriately to allow the ego vehicle to choose another possible lane to change to), while decelerating (DEC) indicates a decision to cooperate.
[0122] S32, game payment design; remember Q mn and P mn They represent the game payoffs of the Ego vehicle and the Side_lag vehicle, respectively, where m and n represent the actions chosen by the Ego vehicle and the Side_lag vehicle, respectively. The game payoff matrix can be shown in Table 1 below:
[0123] Table 1 Game payoff matrix
[0124]
[0125] When m=1 and n=1, it indicates that the Ego vehicle chooses to change lanes and the Side_lag vehicle chooses to accelerate;
[0126] When m=1 and n=2, it means that the Ego vehicle chooses to change lanes and the Side_lag vehicle chooses to decelerate;
[0127] When m=2 and n=1, it means that the Ego vehicle chooses to maintain the current lane and the Side_lag vehicle chooses to accelerate;
[0128] When m=2 and n=2, it means that the Ego vehicle chooses to maintain the current lane and the Side_lag vehicle chooses to decelerate;
[0129] More specifically, the game payouts for the Ego vehicle and the Side_lag vehicle in step S32 are:
[0130] Ego vehicle game payment Q mn Considering driving risk-benefit, lane-changing efficiency benefit, and historical cooperation benefit, the driving risk-benefit is determined by the vehicle collision risk evaluated in step S2; the formula is expressed as:
[0131]
[0132] in, is the weight coefficient, which is used to adjust the importance of each part of the income; mn is the error term, which is used to capture unobserved decision variables; R s is the driving risk benefit, R c is the historical cooperation benefit, ΔV is the difference in the upper speed limit of the Ego vehicle before and after the lane change, reflecting the lane change efficiency benefit;
[0133] R s The calculation formula is:
[0134]
[0135] R c Considering the historical interaction information of vehicles, compensation is made based on trust for vehicles that choose to cooperate multiple times. The formula is expressed as:
[0136]
[0137] Where T represents the comprehensive trust of the other partner during the lane change process of the vehicle, f(t) represents the cooperation results of the vehicle in the previous game, t is the number of previous games, f(t) = 1 represents the selection of cooperation strategy, and f(t) = 0 represents the selection of non-cooperation strategy; Acc c represents the collision avoidance acceleration of the Ego vehicle, E c The collision risk is calculated as follows: The collision avoidance acceleration is the minimum acceleration required for the Ego vehicle to avoid a collision with the vehicle in the lane ahead, Side_lead. (Side_lag is to the side and rear of the Ego vehicle, while Side_lead is to the side and front of the Ego vehicle. To ensure that the Ego vehicle does not collide with Side_lead during lane changes, the Ego's upper acceleration limit can be determined based on kinematic theory.)
[0138] λ is a parameter used to adjust the weight of the Ego vehicle's collision avoidance acceleration and vehicle collision risk; v is the vehicle's current speed. When the vehicle is at risk of collision, if the vehicle is traveling at a higher speed, the reaction time to collision is short, and the vehicle needs to take quick action to avoid a collision. Therefore, the collision avoidance acceleration should play a dominant role (1-λ is relatively large). At lower speeds, the vehicle's reaction time is longer, and more attention to long-term collision risk (such as vehicle distance and driving trajectory) helps ensure overall safety, so λ is relatively large.
[0139] Game payoff P of the side_lag vehicle mn Considering driving risk benefits, polite behavior benefits and historical cooperation benefits; the formula is expressed as:
[0140]
[0141] in, is the weight coefficient, which is used to adjust the importance of each part of the income; mn is the error term, which is used to capture the unobserved decision variables;
[0142] R dThe courtesy benefit is that if the Side_lag vehicle chooses to slow down, it will use the maximum deceleration to create more space for the Ego vehicle to change lanes. The formula is:
[0143]
[0144] Among them, Acc d Side_lag represents the minimum deceleration that the vehicle can take, Indicates the anti-collision deceleration of the Side_lag vehicle, that is, the upper limit of acceleration that the Side_lag vehicle needs to comply with to avoid collision with the Ego vehicle in front and side (it may be a negative value, indicating deceleration);
[0145] It should be noted that this application designs two sets of weight parameters; is a constant term representing the participant's basic payment in the absence of other factors. This basic payment reflects the vehicle's driving style (vehicles with different driving styles also have different risks on the road), helping the model better fit the actual data and reduce systematic errors. It will not be dynamically adjusted later, while the remaining weight parameters will be dynamically adjusted according to the real-time lane change situation; This is a safety parameter. If a collision occurs or safety is insufficient during a vehicle lane change, the corresponding vehicle's safety weight increases. If the vehicle lane change is completed safely, the weight increases appropriately. As an efficiency parameter, if the upper speed limit of the vehicle before and after the lane change increases, the efficiency parameter is increased to encourage the vehicle to pass the road section as quickly as possible. If the vehicle's speed is lower or even zero after the lane change, the efficiency parameter is reduced. is the cooperation parameter. If false cooperation is detected (the actual number of times adjacent vehicles give way is less than half of the number of lane changes), the weight is reduced; if active cooperation is continued (the actual number of times adjacent vehicles give way is more than one-to-one times the number of lane changes), the weight is increased. The present application uses weight parameters to flexibly adjust the game payoff (i.e., game cost and benefit) from multiple perspectives in real time, ensuring that the vehicle's lane change strategy not only considers its own interests, but can also be dynamically adjusted to avoid conflicts with other vehicles, thereby improving the efficiency and safety of the overall traffic flow.
[0146] In addition, an error term ε is added mm , δ mnTo capture the unobserved decision variables, decision variables are those that have a direct or significant impact on the final decision result. However, since it is impossible to know all the variables that affect the decision, or these variables cannot be observed and quantified, this application chooses to add an error term to express the overall effect of these unmodeled influencing factors on the decision; when the effects of many independent random variables are superimposed, the final error term often obeys a normal distribution. Therefore, in this embodiment, it is assumed that the error term ε mm , δ mn Obeys normal distribution.
[0147] S33. The Ego vehicle and the Side_lag vehicle engage in a game. If the game result indicates that the Ego vehicle can cooperate with the Side_lag vehicle, the lane of the Side_lag vehicle is considered as the potential target lane of the Ego vehicle. If the Ego vehicle has multiple potential target lanes, the lane with the highest comprehensive benefit is selected as the final target lane. The comprehensive benefit is the sum of the two game payments. For situations where multiple lanes need to be crossed, the Ego vehicle will repeat the above process in sequence after each game is completed until the final lane change goal is achieved.
[0148] S4. Considering both safety and lane-changing efficiency, the Nash Q-learning algorithm is used to learn the ego vehicle's lane-changing decisions. The weight parameters of the game payoff function are dynamically and adaptively adjusted. This allows the vehicle to adaptively adjust its lane-changing decisions based on different traffic conditions and complete multi-lane lane changes within a group of vehicles.
[0149] As a preferred embodiment, Figure 4 As shown, step S4 specifically includes:
[0150] S41. Define the state space and action space. The state space stores the relationship and dynamic information between the ego vehicle (Ego) and the neighboring vehicle (Side_lag), including the relative position, velocity, and acceleration of the ego vehicle and the neighboring vehicle (Side_lag). The action space stores the possible actions that the ego vehicle and the side_lag vehicle can take in each state. For the ego vehicle, the action space stores two options: change lanes (CL) or stay in the current lane (KCL). For the side_lag vehicle, the action space stores two options: accelerate (ACC) or decelerate (DEC).
[0151] S42. Design rewards by considering both safety and lane change efficiency; efficiency rewards r s The formula is expressed as:
[0152]
[0153] Among them, v represents the current speed of the vehicle, v des Indicates the target speed desired by the vehicle;
[0154] Security Reward e The formula is expressed as:
[0155]
[0156] Among them, d min Indicates the minimum distance between the vehicle and the nearest vehicle, d safe Indicates the target distance expected by the vehicle;
[0157] The total immediate reward is:
[0158] R=μr s +(1-μ)r e ;
[0159] Among them, μ represents the average level of trust among game participants, that is, the average value of the comprehensive trust of the ego vehicle in neighboring vehicles, which is used to control the weight balance between safety rewards and efficiency rewards. When the average level of trust is high, efficiency is prioritized, and when the average level of trust is low, safety is prioritized. The formula for μ is expressed as:
[0160]
[0161] Among them, T i,j is the comprehensive trust of vehicle i in other vehicles j; J is the total number of other vehicles participating in the game;
[0162] S43, Q-value update and dynamic parameter adjustment: Whenever a vehicle selects an action, the Nash Q-learning algorithm learns the Nash equilibrium solution based on the total reward it receives, and dynamically adjusts the parameters of the payoff function based on the feedback;
[0163] The Q value update formula is:
[0164] Q(s,a1,a2)←(1-ɑ t )Q(s,a1,a2)+α t [R+γNash Q(s′,a′1,a′2)];
[0165] Among them, s represents the current state, a1 is the action currently taken by the Ego vehicle, a2 is the action currently taken by the Side_lag vehicle, R is the immediate total reward; γ is the discount factor, which represents the impact of future rewards; α t is the learning rate; NashQ(s',a'1,a'2) is the expected cumulative return of the current vehicle when all vehicles participating in the game act according to the Nash equilibrium strategy in the next state;
[0166] Each time a vehicle selects an action, the Nash Q-learning algorithm learns a Nash equilibrium solution based on the total reward it receives, while also dynamically adjusting the parameters of the payoff function based on feedback. Specifically, if a collision occurs during the lane change or safety is insufficient (i.e., the participating vehicles do not maintain a safe distance), Nash Q-learning increases the weighting of safety parameters, emphasizing collision avoidance. If the lane change improves driving efficiency, the weighting of efficiency parameters is increased (increasing the vehicle's speed limit before and after the lane change), prioritizing decisions that improve traffic flow.
[0167] S44. Through multiple interactions between the vehicle and neighboring vehicles, the Q value and the weight parameters of the payment function are continuously adjusted, and the Q value is converged through experiments and feedback, thereby achieving a balance between safety and efficiency, and ultimately enabling the vehicle to adaptively adjust lane change decisions according to different traffic environments.
[0168] Thus, the method proposed in the present invention improves the decision-making adaptability, driving safety and traffic efficiency of vehicle lane changes through real-time collision risk assessment, game lane change decision model and adaptive adjustment of weight parameters under a zero-trust architecture; thus, the method proposed in the present invention enables vehicle groups to effectively respond to emergencies in complex traffic environments, reduce the risk of accidents, and improve the overall safety and efficiency of multi-lane lane changes.
[0169] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.
[0170] The logic and / or steps represented in the flowchart or otherwise described herein may be considered, for example, as an ordered list of executable instructions for implementing logical functions, and may be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device).
[0171] The above embodiments provide a detailed introduction to the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.
Claims
1. A multi-lane game decision-making method for vehicle groups based on collision avoidance risk assessment under a zero-trust architecture, characterized by: The specific steps include: S1. Build an intelligent vehicle network to obtain real-time vehicle status and road information. Vehicle status information includes the speed, acceleration, position, mass, and size of the vehicle and adjacent vehicles, the distance between the vehicle and adjacent vehicles, historical interaction records of adjacent vehicles, and vehicle trustworthiness. Road information includes lane width, lane line position, and number of lanes. S2. Real-time assessment of vehicle collision risk through zero-trust architecture and potential field theory; Step S2 specifically includes: S21. Under the zero-trust architecture, dynamically and continuously evaluate the trust of vehicle i in vehicle j in the intelligent vehicle network, including evaluating the direct trust of vehicle i in vehicle j. and recommendation trust S22. Weight the direct trust and the recommended trust to obtain the comprehensive trust T of vehicle node i to vehicle node j. i,j , the formula is: Where ω is the weight coefficient that balances the influence of direct trust and recommended trust in the overall trust. When there are more direct interactions between vehicle nodes i and j, they are less affected by external factors and are given a higher direct trust weight. When there are fewer interactions between vehicle nodes i and j, there is less direct trust data and the quality is lower, so a higher recommended trust weight is given. S23. Establish a collision risk field that dynamically reflects the collision risk between vehicles. The collision risk field is a potential field that repel each other between vehicles. The direction of the field strength is from adjacent vehicles toward the own vehicle. The field strength is determined by the equivalent mass, motion state, spatial position, and overall trust in the own vehicle. S24. Obtain the field strength E of vehicle j to vehicle i from the collision risk field ij Based on the potential field theory, we can get the repulsive force E of vehicle j on its own vehicle: ij ; Then calculate the potential energy Ec of vehicle j on vehicle i through repulsive force ij , as the collision risk of vehicle j to vehicle i; sum the potential energy of all vehicles in the scene to the ego vehicle, and obtain the collision risk Ec faced by the ego vehicle i ; The formula is: F ij =E ij ·M i ; Among them, M i is the equivalent mass of the vehicle, r ij is the distance between vehicle j and ego vehicle i; S3. Design a multi-lane lane-changing game decision model based on vehicle collision risk, decomposing the multi-lane lane-changing problem into multiple two-vehicle game models; S4. Considering both safety and lane-changing efficiency, the vehicle's lane-changing decisions are learned through the Nash Q-learning algorithm. The weight parameters of the game payoff function in the game decision model are adaptively and dynamically adjusted. This allows the vehicle to adaptively adjust its lane-changing decisions according to different traffic environments and complete multi-lane lane changes in a group of vehicles.
2. The multi-lane game decision-making method for vehicle groups based on collision avoidance risk assessment under the zero-trust architecture according to claim 1 is characterized in that: The specific evaluation method of the direct trust and recommended trust of vehicle i to vehicle j in step S21 is: Direct trust is based on the historical behavior of the vehicle, evaluating its reliability in the network and introducing a penalty factor to constrain illegal behavior. The calculation formula of the direct trust of vehicle i to vehicle j is expressed as: in, They represent the direct trust of vehicle i in vehicle j at the kth and k-1th time steps in the trust evaluation process respectively; ρ is the time decay factor, which is a preset fixed parameter; η k represents the penalty factor, which quantifies the degree of violation of vehicle j in the current lane change process; the penalty factor η k Based on whether the vehicle spacing and acceleration of vehicle j during the lane change game exceed the safety threshold, and based on the duration and degree of violation, the calculation formula is: Among them, d k 、a k Indicates the current distance and acceleration, d sa and a sa They represent the preset distance safety threshold and acceleration safety threshold respectively; τ is a parameter used to adjust the sensitivity of the model, and T represents the maximum violation duration during the vehicle game lane change process, which is calculated from the moment the acceleration exceeds the safety threshold. Recommended Trust It is equal to the weighted sum of the comprehensive trust of other nodes in the network to vehicle j, where the weight is the comprehensive trust of vehicle i to other nodes; other nodes refer to other vehicles except vehicle i and vehicle j; the recommendation trust at the kth time step is The calculation formula is:
3. The multi-lane game decision-making method for vehicle groups based on collision avoidance risk assessment under the zero-trust architecture according to claim 1 is characterized in that: The calculation method of the midfield strength in step S23 is specifically as follows: The calculation method of the electric field strength of vehicle j to vehicle i is different according to the different motion states of vehicle j. If j = a and j = b are used to represent vehicle j in a stationary state and a moving state respectively, then: in, Respectively represent the magnitude of the field strength generated by the stationary vehicle and the moving vehicle on the ego vehicle i; M a 、M b are the equivalent masses of the stationary vehicle and the moving vehicle respectively; r ia 、r ib is the position vector between the stationary vehicle, the moving vehicle and the ego vehicle i, and its modulus is the distance between the vehicles; θ b is the moving direction and position vector r of the moving vehicle bi The angle between them, v b is the actual speed of the moving vehicle; k1 and k2 are coefficients describing the attenuation of the field strength with distance; The formula for calculating the equivalent mass of a stationary vehicle is: M a =m a ·g(v a ); Among them, m a is the actual weight of the stationary vehicle; v a is the actual speed of the stationary vehicle; The equivalent mass of the moving vehicle and the ego vehicle is calculated in the same way as for a stationary vehicle; k1 and k2 are determined by the comprehensive trust of vehicle i to vehicle j, specifically: k1=T i,a +1; k2=T i,b +1; Among them, T i,a and T i,b are the comprehensive trust of vehicle i to the stationary vehicle a and the moving vehicle b respectively.
4. The multi-lane game decision-making method for vehicle groups based on collision avoidance risk assessment under a zero-trust architecture according to claim 1 is characterized in that: Step S3 specifically includes: S31. Design the game participants and strategy space. The game participants are the ego vehicle (Ego) and the neighboring vehicle (Side_lag). The Ego vehicle's action options are lane change (CL) and lane retention (KCL). The Side_lag vehicle's action options are acceleration (ACC) and deceleration (DEC). Acceleration (ACC) indicates a rejection of the cooperative strategy, while deceleration (DEC) indicates a cooperation strategy. S32, game payment design; remember Q mn and P mn They represent the game payoffs of the Ego vehicle and the Side_lag vehicle, respectively, where m and n represent the actions chosen by the Ego vehicle and the Side_lag vehicle, respectively; When m=1 and n=1, it indicates that the Ego vehicle chooses to change lanes and the Side_lag vehicle chooses to accelerate; When m=1 and n=2, it means that the Ego vehicle chooses to change lanes and the Side_lag vehicle chooses to decelerate; When m=2 and n=1, it means that the Ego vehicle chooses to maintain the current lane and the Side_lag vehicle chooses to accelerate; When m=2 and n=2, it means that the Ego vehicle chooses to maintain the current lane and the Side_lag vehicle chooses to decelerate; S33. The Ego vehicle and the Side_lag vehicle engage in a game. If the game result indicates that the Ego vehicle can cooperate with the Side_lag vehicle, the lane of the Side_lag vehicle is considered as the potential target lane of the Ego vehicle. If the Ego vehicle has multiple potential target lanes, the lane with the highest comprehensive benefit is selected as the final target lane. The comprehensive benefit is the sum of the two game payments. For situations where multiple lanes need to be crossed, the Ego vehicle will repeat the above process in sequence after each game is completed until the final lane change goal is achieved.
5. The multi-lane game decision-making method for vehicle groups based on collision avoidance risk assessment under the zero-trust architecture according to claim 4 is characterized in that: The game payouts of the Ego vehicle and the Side_lag vehicle in step S32 are specifically: Ego vehicle game payment Q mn Considering driving risk-benefit, lane-changing efficiency benefit, and historical cooperation benefit, the driving risk-benefit is determined by the vehicle collision risk evaluated in step S2; the formula is expressed as: in, is the weight coefficient, which is used to adjust the importance of each part of the income; mn is the error term, which is used to capture unobserved decision variables; R s is the driving risk benefit, R c is the historical cooperation benefit, ΔV is the difference in the upper speed limit of the Ego vehicle before and after the lane change, reflecting the lane change efficiency benefit; R s The calculation formula is: R c Considering the historical interaction information of vehicles, compensation is made based on trust for vehicles that choose to cooperate multiple times. The formula is expressed as: Where T represents the comprehensive trust of the other partner during the lane change process of the vehicle, f(t) represents the cooperation results of the vehicle in the previous game, t is the number of previous games, f(t) = 1 represents the selection of cooperation strategy, and f(t) = 0 represents the selection of non-cooperation strategy; Acc c represents the collision avoidance acceleration of the Ego vehicle, E c is the vehicle collision risk; λ is a parameter used to adjust the weight of the Ego vehicle's anti-collision acceleration and the vehicle collision risk; v is the current speed of the vehicle; Game payoff P of the side_lag vehicle mn Considering driving risk benefits, polite behavior benefits and historical cooperation benefits; the formula is expressed as: in, is the weight coefficient, which is used to adjust the importance of each part of the income; mn is the error term, which is used to capture the unobserved decision variables; R d The courtesy benefit is that if the Side_lag vehicle chooses to slow down, it will use the maximum deceleration to create more space for the Ego vehicle to change lanes. The formula is: Among them, Acc d Side_lag represents the minimum deceleration that the vehicle can take, Side_lag represents the anti-collision deceleration of the vehicle.
6. The multi-lane game decision-making method for vehicle groups based on collision avoidance risk assessment under the zero-trust architecture according to claim 1 is characterized in that: Step S4 specifically includes: S41. Define the state space and action space. The state space stores the relationship and dynamic information between the ego vehicle (Ego) and the neighboring vehicle (Side_lag), including the relative position, velocity, and acceleration of the ego vehicle and the neighboring vehicle (Side_lag). The action space stores the possible actions that the ego vehicle and the side_lag vehicle can take in each state. For the ego vehicle, the action space stores two options: change lanes (CL) or stay in the current lane (KCL). For the side_lag vehicle, the action space stores two options: accelerate (ACC) or decelerate (DEC). S42. Design rewards by considering both safety and lane change efficiency; efficiency rewards r s The formula is expressed as: Among them, v represents the current speed of the vehicle, v des Indicates the target speed desired by the vehicle; Security Reward e The formula is expressed as: Among them, d min Indicates the minimum distance between the vehicle and the nearest vehicle, d safe Indicates the target distance expected by the vehicle; The total immediate reward is: R=μr s +(1-μ)r e ; Among them, μ represents the average level of trust among game participants, that is, the average value of the comprehensive trust of the ego vehicle in neighboring vehicles, which is used to control the weight balance between safety rewards and efficiency rewards. When the average level of trust is high, efficiency is prioritized, and when the average level of trust is low, safety is prioritized. The formula for μ is expressed as: Among them, T i,j is the comprehensive trust of vehicle i in other vehicles j; J is the total number of other vehicles participating in the game; S43, Q-value update and dynamic parameter adjustment: Whenever a vehicle selects an action, the Nash Q-learning algorithm learns the Nash equilibrium solution based on the total reward it receives and dynamically adjusts the parameters of the payoff function based on the feedback; The Q value update formula is: Q(s,a1,a2)←(1-α t )Q(s,a1,a2)+α t [R+γNash Q(s′,a′1,a′2)]; Among them, s represents the current state, a1 is the action currently taken by the Ego vehicle, a2 is the action currently taken by the Side_lag vehicle, R is the immediate total reward; γ is the discount factor, which represents the impact of future rewards; α t is the learning rate; NashQ(s',a'1,a'2) is the expected cumulative return of the current vehicle when all vehicles participating in the game act according to the Nash equilibrium strategy in the next state; Each time a vehicle selects an action, the Nash Q-learning algorithm learns a Nash equilibrium solution based on the total reward it receives, while also dynamically adjusting the parameters of the payoff function based on feedback. Specifically, if a collision occurs during a lane change and safety is insufficient, Nash Q-learning increases the weighting of safety parameters, emphasizing collision avoidance. If the lane change improves driving efficiency, Nash Q-learning increases the weighting of efficiency parameters, prioritizing decisions that improve traffic flow. S44. Through multiple interactions between the vehicle and neighboring vehicles, the Q value and the weight parameters of the payment function are continuously adjusted, and the Q value is converged through experiments and feedback, thereby achieving a balance between safety and efficiency, and ultimately enabling the vehicle to adaptively adjust lane change decisions according to different traffic environments.
Citation Information
Patent Citations
Method for solving automatic driving vehicle conflict based on differential game decision modeling
CN118298653A
Aiops guided, quantum-safe zero trust data transfer methods in-motion with segmented, data transfer across an overlay network
IN202141058712A