Left turn decision-making method and system under perceptual shielding

By introducing a PPO decision model with virtual vehicles and safety constraints into the autonomous driving system, the uncertainty of left-turn decisions under perception occlusion is solved, and safe and efficient passage in occluded environments is achieved.

CN121043884AActive Publication Date: 2025-12-02ADVANCED TECH RES INST OF BEIJING UNIV OF TECH +3
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511590507.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2025-12-02
Estimated Expiration
2045-11-03

AI Technical Summary

Technical Problem

Existing autonomous driving systems struggle to accurately assess the collision risk of vehicles within obscured areas when perception is obstructed, leading to decision uncertainty, especially in scenarios involving left turns and merging, which can easily cause traffic accidents.

Method used

A left-turn decision-making method under perception occlusion is adopted. The observation space is constructed by introducing a virtual vehicle, and the value distribution network is used for real-time risk assessment. The PPO decision model is optimized by using a near-end strategy with safety constraints to generate decision strategies. The Lagrange multiplier method is used to introduce collision risk as a constraint to generate control actions.

Benefits of technology

It improves the safety and efficiency of autonomous vehicles in obstructed environments, and reduces the risk of traffic accidents through efficient risk assessment and strategy selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121043884A_ABST
    Figure CN121043884A_ABST
Patent Text Reader

Abstract

The invention provides a left turn decision-making method and system under perceptual occlusion, and belongs to the technical field of vehicle automatic driving control, and the method comprises the steps: obtaining vehicle state information and environment vehicle state information, and introducing a virtual vehicle into a perceptual occlusion region corresponding to a vehicle; constructing an observation space based on the vehicle state information, the environment vehicle state information and the shelter boundary point information, and calculating a dynamic baseline to perform real-time risk assessment; based on the observation space and the risk assessment result, optimizing a PPO decision model by adopting a near-end strategy with security constraint to generate a decision strategy; according to the safety constraint, the collision risk serves as a constraint condition to be introduced into a strategy optimization target through a Lagrange multiplier method; and outputting a control action based on the decision strategy to control the vehicle to finish the left turning process. Based on the method, the invention further provides a left turn decision making system under perceptual occlusion. According to the method, the traffic safety and efficiency of the automatic driving vehicle under the uncertainty caused by shielding are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of vehicle autonomous driving control technology, and specifically relates to a method and system for making left-turn decisions under perception obstruction. Background Technology

[0002] At urban intersections, perception occlusion poses a significant challenge to autonomous driving decision-making and planning. Particularly in left-turn merging scenarios, construction barriers, buildings, and other static obstacles often limit the sensor's field of view, preventing the autonomous driving system from promptly identifying the movement of vehicles approaching laterally. Therefore, when a left-turning vehicle attempts to cross an intersection, the collision risk within the occluded area is often inaccurately assessed, leading to considerable uncertainty in the decision-making model. Human drivers can rely on experience to assess the potential risks in occluded areas and make timely decisions to slow down or wait, but most existing autonomous driving decision-making systems ignore this factor, easily resulting in traffic accidents. The uncertainty caused by perception occlusion exacerbates the safety risks at intersections. Especially at intersections without traffic lights, conflicts between left-turning vehicles and vehicles going straight or turning right become a major problem. In these scenarios, obstructions such as tall buildings and construction fences severely limit the detection capabilities of sensors like LiDAR and millimeter-wave radar, causing the system to only detect approaching conflicting vehicles when approaching the intersection. In this case, the lack of prediction of moving vehicles within the occluded area may lead to overly cautious stopping or incorrect acceleration decisions, reducing traffic efficiency and even causing safety accidents. Therefore, how to make efficient and safe decisions when perception is obstructed is a major challenge in autonomous driving technology.

[0003] Existing technologies typically employ traditional reinforcement learning methods, such as DQN and PPO, which can handle some uncertainty but lack modeling of value distribution, making it difficult to accurately assess the value of actions in high-risk scenarios. Alternatively, probabilistic prediction models, such as using Gaussian processes or Bayesian networks to predict vehicle behavior in occluded areas, are computationally complex, have poor real-time performance, and are difficult to integrate end-to-end with decision-making strategies. Safety reinforcement learning methods, such as using constrained MDP (CMDP) or the Lagrange multiplier method to introduce safety constraints, still lack explicit modeling of the uncertainty in value distribution, leaving safety risks in extreme occlusion scenarios. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention proposes a left-turn decision-making method and system under perceived occlusion. This ensures the safety and efficiency of autonomous vehicles' passage under uncertainties caused by occlusion.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: A method for making left-turn decisions under perceptual occlusion includes the following steps: Acquire vehicle status information and environmental vehicle status information, and introduce virtual vehicles within the perception occlusion area corresponding to the vehicle. An observation space containing continuous temporal information is constructed based on the vehicle's status information, the environmental vehicle status information, and the boundary point information of obstructions. The dynamic baseline is calculated using the action value distribution output by the value distribution network for real-time risk assessment. Based on the observation space and risk assessment results, a proximal strategy with safety constraints is adopted to optimize the PPO decision model to generate decision strategies. The PPO decision model uses the dynamic baseline to block action selection in the later stage of training and uses quantile numerical network and quantile proposal network to fit the distribution of action value. Among them, the safety constraint is to introduce collision risk as a constraint condition into the strategy optimization objective by using the Lagrange multiplier method; Based on the generated decision strategy, output control actions to control the vehicle to complete the left turn.

[0006] This invention also proposes a left-turn decision-making system under sensor occlusion, comprising: The state perception module is used to acquire the state information of the vehicle itself and the state information of the surrounding vehicles, and to introduce virtual vehicles within the perception occlusion area corresponding to the vehicle itself. The risk assessment module is used to construct an observation space containing continuous temporal information based on the vehicle's status information, the environmental vehicle status information, and the boundary point information of obstructions, and to calculate a dynamic baseline using the action value distribution output by the value distribution network for real-time risk assessment. The strategy generation module is used to generate decision strategies based on the observation space and risk assessment results, and to optimize the PPO decision model with safety constraints. The PPO decision model uses the dynamic baseline to block the action selection in the later stage of training, and uses a quantile numerical network and a quantile proposal network to fit the distribution of action value. The safety constraint is to introduce collision risk as a constraint condition into the strategy optimization objective through the Lagrange multiplier method. The vehicle control module is used to output control actions based on the generated decision strategy to control the vehicle to complete the left turn process.

[0007] The effects described in the invention are merely those of the embodiments, and not all the effects of the invention. One of the above technical solutions has the following advantages or beneficial effects: This invention proposes a left-turn decision-making method and system under perceptual occlusion, belonging to the field of vehicle autonomous driving control technology. The method includes the following steps: acquiring the vehicle's state information and the environment's vehicle state information, and introducing a virtual vehicle within the perceptual occlusion area corresponding to the vehicle; constructing an observation space containing continuous temporal information based on the vehicle's state information, the environment's vehicle state information, and the boundary point information of the occlusion object, and calculating a dynamic baseline using the action value distribution output by a value distribution network for real-time risk assessment; based on the observation space and the risk assessment results, using a proximal strategy optimization model with safety constraints to generate a decision strategy; the PPO decision model uses the dynamic baseline to perform risk blocking on action selection in the later stages of training, and uses a quantile value network and a quantile proposal network to fit the distribution of action values; wherein, the safety constraint is introduced into the strategy optimization objective by using the Lagrange multiplier method to incorporate collision risk as a constraint condition; and outputting a control action based on the generated decision strategy to control the vehicle to complete the left-turn process. Based on this left-turn decision-making method under perceptual occlusion, a left-turn decision-making system under perceptual occlusion is also proposed. This invention addresses the perception occlusion problem in left-turn merging scenarios. By fusing value distribution networks and safety reinforcement learning methods, it proposes and employs a value distribution PPO decision model with safety constraints. Combining the designed environmental modeling with the perception occlusion representation in the state space, it achieves efficient risk assessment and strategy selection capabilities, effectively deepening the understanding of complex traffic environments in autonomous driving decision-making, thereby ensuring the safety and efficiency of autonomous vehicles under the uncertainty caused by occlusion. Attached Figure Description

[0008] Figure 1 This is a flowchart of a left-turn decision-making method under sensor occlusion proposed in Embodiment 1 of the present invention; Figure 2 This is an architecture diagram of a left-turn decision-making method under perceived occlusion proposed in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the training simulation environment proposed in Embodiment 1 of the present invention; Figure 4 This is a schematic diagram of a left-turn decision system under perception occlusion proposed in Embodiment 2 of the present invention. Detailed Implementation

[0009] To clearly illustrate the technical features of this solution, the invention will be described in detail below through specific embodiments and in conjunction with the accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure of the invention, components and arrangements of specific examples are described below. Furthermore, reference numerals and / or letters may be repeated in different examples. This repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed. It should be noted that the components illustrated in the drawings are not necessarily drawn to scale. Descriptions of well-known components, processing techniques, and processes are omitted in this invention to avoid unnecessarily limiting the invention.

[0010] Example 1 Embodiment 1 of this invention proposes a left-turn decision-making method under perception occlusion, which is used to solve the technical problem of how to achieve efficient and safe decision-making under perception occlusion in the prior art.

[0011] Figure 1 This is a flowchart of a left-turn decision-making method under sensor occlusion proposed in Embodiment 1 of the present invention; In step S1, the vehicle status information and the environmental vehicle status information are obtained, and a virtual vehicle is introduced into the perception occlusion area corresponding to the vehicle.

[0012] In this application, the vehicle's state information includes: the polar coordinate distance and angle of the vehicle relative to the center point of the intersection, the vehicle's speed, and the yaw angle; the environmental vehicle's state information includes: the polar coordinate distance and angle of the vehicle relative to the vehicle, the environmental vehicle's speed, and the yaw angle. Furthermore, the vehicle's state information is represented by state observations at several consecutive time points; the environmental vehicle's state information is represented by state observations at several consecutive time points corresponding to several environmental vehicles.

[0013] To avoid algorithm redundancy and considering lightweight computation, extracting key boundary points for vector encoding is relatively lightweight and directly reflects the geometric positional relationship of occluders. Therefore, the boundary point coordinates are directly vectorized for encoding. This is different from encoding using absolute coordinate systems, which makes it difficult to intuitively reflect the relative relationship between the sensor and the occluders.

[0014] Figure 2 This is an architecture diagram of a left-turn decision-making method under perceived occlusion proposed in Embodiment 1 of the present invention; the present invention uses a relative coordinate system to update the position vector information in real time. Due to the relationship between the laser and the ray of the blind spot occlusion, the present invention chooses polar coordinates for encoding.

[0015] Specifically: ; ; in, Represents an occlusion information vector; This indicates the distance from the vehicle to the nearest point of the construction fence; Indicates the direction from the vehicle to the nearest end of the construction site fence; This indicates the distance from the vehicle to the farthest point of the construction fence; Indicates the direction from the vehicle to the farthest end of the construction site; Represents the complete state vector; Represents the vehicle's state vector; Represents the environmental vehicle state vector; This indicates the current state of the vehicle in the environment. Indicates environmental vehicles The state at any given moment; Indicates environmental vehicles The state at any given moment; This represents the information vector of the occlusion.

[0016] In this application, the action space It is two-dimensional, including the acceleration of the intelligent vehicle and the front wheel steering angle. This indicates the acceleration or deceleration of the vehicle itself, while This represents the front wheel steering angle. When exploring a two-dimensional continuous action space, the infinite action space leads to an excessively large exploration range, resulting in instability during training and uncertainty in convergence. Since this paper does not favor exploring search strategies, to improve the algorithm's stability, we choose to represent environmental information using a rich state space, while simplifying it by discretizing the action space. By appropriately designing the state and action spaces, we accelerate the agent's training speed and improve its performance. Acceleration The set of values ​​is 1.5, 1, 0.5, 0, 0.5, 1, 1.5, units are Steering angle The set of values ​​is 0.5, 0.25, 0, 0.25, 0.5, in rad.

[0017] The scope of protection of this application is not limited to the values ​​listed in Example 1, and those skilled in the art can make reasonable selections based on the actual situation.

[0018] In step S2, an observation space containing continuous temporal information is constructed based on the vehicle's state information, the environmental vehicle state information, and the boundary point information of the obstruction. The dynamic baseline is calculated using the action value distribution output by the value distribution network to perform real-time risk assessment.

[0019] In this application, the specific process of constructing an observation space containing continuous temporal information based on the vehicle's state information, the surrounding vehicle's state information, and the boundary point information of the obstruction is as follows: a polar coordinate system is established with the vehicle as the center; the vehicle's state, the surrounding vehicle's state, and the boundary point information of the obstruction are encoded in the polar coordinate system to form a state vector; the state vectors of multiple consecutive time steps are stacked to construct an observation space containing temporal dynamic relationships.

[0020] This application obtains a frame of the observation space from the environment. This includes the status of your own vehicle and the status of surrounding vehicles: ; ; in, and These represent the polar coordinate distance and angle of the vehicle relative to the center point of the intersection. It is the speed of the vehicle itself. It is the yaw angle; This represents the total number of vehicles that pose a driving risk due to obstruction; each row of the matrix represents information about one vehicle. Represents relative polar coordinate distance; Indicates relative polar angle; Represents relative velocity; Indicates the relative yaw angle; and .

[0021] This application utilizes the action value distribution output by a value distribution network to calculate a dynamic baseline; specifically: ; in, For dynamic baseline; For expectation operators; For policy functions; To sum over quantile indices; The number of quantiles; The quantiles proposed by the quantile proposal network output; The width of the quantile interval; The output of the quantile numerical network in the state Select Action The corresponding number quantile values; This is the Dirac function.

[0022] In step S3, based on the observation space and risk assessment results, a near-end strategy with safety constraints is used to optimize the PPO decision model to generate a decision strategy. The PPO decision model uses the dynamic baseline to perform risk blocking on action selection in the later stage of training, and uses a quantile numerical network and a quantile proposal network to fit the distribution of action value. The safety constraint is to introduce collision risk as a constraint condition into the strategy optimization objective through the Lagrange multiplier method.

[0023] The value distribution-based and safety-oriented reinforcement learning model employs the PPO (Proximal Policy Optimization) model, which utilizes state value distribution information to design a dynamic baseline for the advantage function in the PPO policy gradient. Furthermore, risk blocking is implemented for the agent's action selection in the later stages of training, thereby improving the safety and effectiveness of decision-making. Adjustable quantile scores obtained through FPN (quantile proposal network) are combined with corresponding adjustable quantile values ​​fitted by QVN (quantile value network) to support the final action selection of the vehicle's acceleration and front wheel steering angle. Simultaneously, the reward function design integrates expert trajectory rewards, conflict zone crossing rewards, collision penalties, and target lane arrival rewards to further optimize the overall decision-making performance.

[0024] Based on the index of the discrete action, the quantile position corresponding to the current action is obtained, thus yielding the value distribution of the specific action pair; specifically: ; in, For the distribution of action value, Define the symbol; The symbol indicates that the probability distributions on both the left and right sides are equal.

[0025] The network parameters are updated using TD learning. The TD update formula is as follows: ; in, For the first The error between the predicted quantile and the target quantile; For at any time Quantile numerical networks for current state action pairs The predicted first The quantile value, i.e., the predicted value; At any moment Execute action The instant reward obtained afterward; Discount factor; For the target network to the next state The first one evaluated quantile values; For distributed TD targets.

[0026] In the Bellman equation, the action-value distribution of the current state Distribution of rewards based on expected reward plus discount at the next moment They are equal, the latter also known as the TD objective. However, in the actual process of continuously fitting the distribution and converging, there will inevitably be an error between the two, also known as the TD error, which is the main component of quantile regression. The formula for calculating the overall loss of a quantile numerical network is: ; in, For the threshold is Huber's losses; This refers to timing difference error; The number of quantiles; For the first quantiles.

[0027] Quantile proposal networks employ gradient descent by minimizing the 1-Wasserstein distance between the approximate quantile distribution and the actual quantile distribution. The objective function of the quantile proposal network is: ; in, The gradient of the Wasserstein distance; The distance is 1-Wasserstein. The learnable quantile; Estimation of quantile functions; The target quantile; The value of adjacent target quantiles.

[0028] The objective function to be optimized in the PPO decision model is: ; in, Let be the objective function to be optimized in the PPO decision model; To optimize the strategy parameters; For the optimized strategy parameters; For expectation operators; Importance weight; To optimize the strategy in the state Select action The probability density; To optimize the policy in the state Select action The probability density; This is the dominant function.

[0029] This application aims to minimize the 1-Wasserstein distance between the approximate quantile distribution and the actual quantile distribution for gradient descent, thereby leveraging information from the distribution to help optimize the policy network.

[0030] The reward function used during training of the PPO decision model is: ; in, For the reward function; Rewards for expert trajectory Rewards for crossing conflict zones, As a penalty for collision, A reward is given for reaching the target lane.

[0031] In one or more embodiments, the vehicle obtains state information and reward information by interacting with the environment. Specifically, to meet driving safety requirements, the scheme described in this embodiment uses the real left-turn trajectory as a guiding item in the reward to inspire the agent to learn the left-turn strategy.

[0032] A positive reward is given when the agent follows a left-turn trajectory in the real world (expert data). .

[0033] ; in, ; in, This represents the Euclidean distance between the center point of the vehicle and the center point of the target vehicle at the same moment. This represents the distance between the yaw angle of the vehicle and the yaw angle of the target vehicle. As the first hyperparameter, This is the second hyperparameter. This is the third hyperparameter. This is the fourth hyperparameter. This is the fifth hyperparameter. This is the sixth hyperparameter; the reward is respectively for Item and Information on rewards and punishments will be provided.

[0034] A positive reward is given when the agent successfully traverses a conflict-prone area. In order to enable the agent to explore the environment after imitating the same decision convex space, optimize the trajectory, and improve the passage efficiency, an additional reward for passing through the left-turn conflict zone is adopted as a reward mechanism for efficient passage.

[0035] When an agent violates road boundary constraints or collides with other vehicles, it is given a negative reward; that is: .

[0036] A positive reward is given when the agent reaches the target lane at the final moment. .

[0037] The termination conditions for each round include: the vehicle going beyond the road boundary, the vehicle colliding with an obstacle, reaching the target lane, or exceeding the time limit.

[0038] Extract all tuples and set the state. Feature vectors are obtained from the batch input state coding network, and the baseline is calculated using the upper segment distribution of the value distribution and the formula. Through rewards and cost calculate and Through formula ; Update PPO network K round.

[0039] Through formula Update the Lagrange multiplier λK wheel.

[0040] Discard the future For each time step state, the GRU in the state-coded network discards a corresponding layer.

[0041] After the training process stabilizes, a converged policy network and value network are obtained. The policy network outputs corresponding actions under different states, and the value network predicts the value distribution of the actions. In this stage, the lower quantile interval of the value distribution is used instead of the mean to evaluate the action value for risk control. Compared to taking the minimum value, using the lower half interval makes full use of distribution information. Based on the above principles, the agent adopts the following risk prediction mode in the action selection stage: ; That is, when the agent's actions are distributed The value of the lower half distribution Smaller than the current state distribution When the mean of the lower half of the distribution is reached, the agent reselects an action. This filtering operation during action selection optimizes the experience in the buffer, allowing for more focused increases in the probability of high-quality actions during updates, enabling policy fine-tuning and further avoiding dangerous actions.

[0042] By repeating the above process a certain number of times until the network converges, a trained model is obtained.

[0043] In step S4, a control action is output based on the generated decision strategy to control the vehicle to complete the left turn. Finally, based on the obtained global state features, the trained model is used to obtain the vehicle's output action.

[0044] Figure 3This is a schematic diagram of the training simulation environment proposed in Embodiment 1 of the present invention; the output control actions include: outputting discrete acceleration values ​​and discrete front wheel steering angle values ​​to control the longitudinal and lateral movements of the vehicle.

[0045] This invention proposes a left-turn decision-making method under perceptual occlusion. Based on considerations of perceptual occlusion in left-turn merging scenarios, it proposes and employs a value distribution PPO decision model with safety constraints by fusing value distribution networks and safety reinforcement learning methods. By combining designed environmental modeling with perceptual occlusion representation in the state space, it achieves efficient risk assessment and strategy selection capabilities, effectively deepening the understanding of complex traffic environments in autonomous driving decision-making, thereby ensuring the safety and efficiency of autonomous vehicles under uncertainties caused by occlusion.

[0046] Example 2 Based on the left-turn decision-making method under sensory occlusion proposed in Embodiment 1 of the present invention, Embodiment 2 of the present invention also proposes a left-turn decision-making system under sensory occlusion. Figure 4 This is a schematic diagram of a left-turn decision-making system under sensor occlusion as proposed in Embodiment 2 of the present invention. The system includes: The state perception module executes step S1 of the left turn decision method under perception occlusion proposed in Embodiment 1 of the present invention; it is used to obtain the vehicle state information and the environmental vehicle state information, and to introduce a virtual vehicle in the perception occlusion area corresponding to the vehicle. The risk assessment module executes step S2 of the left-turn decision-making method under perceived occlusion proposed in Embodiment 1 of the present invention; it is used to construct an observation space containing continuous temporal information based on the vehicle's state information, the environment's vehicle state information, and the boundary point information of the occlusion, and to calculate a dynamic baseline using the action value distribution output by the value distribution network for real-time risk assessment. The strategy generation module executes step S3 of the left-turn decision-making method under perceived occlusion proposed in Embodiment 1 of this invention; it is used to generate a decision strategy based on the observation space and risk assessment results, using a near-end strategy optimization PPO model with safety constraints; the PPO decision model uses the dynamic baseline to perform risk blocking on action selection in the later stage of training, and uses a quantile numerical network and a quantile proposal network to fit the distribution of action value; wherein, the safety constraint is to introduce collision risk as a constraint condition into the strategy optimization objective through the Lagrange multiplier method; The vehicle control module executes step S4 of the left-turn decision-making method under perceived occlusion proposed in Embodiment 1 of the present invention; it outputs control actions based on the generated decision strategy to control the vehicle to complete the left-turn process.

[0047] Embodiment 2 of this invention proposes a left-turn decision-making system under perceptual occlusion. Considering the perceptual occlusion problem in left-turn merging scenarios, it proposes and adopts a value distribution PPO decision model with safety constraints by fusing value distribution networks and safety reinforcement learning methods. Combining the designed environmental modeling with the perceptual occlusion representation in the state space, it achieves efficient risk assessment and strategy selection capabilities, effectively deepening the understanding of complex traffic environments in autonomous driving decision-making, thereby ensuring the safety and efficiency of autonomous vehicles under the uncertainty caused by occlusion.

[0048] For a description of the relevant part of the left-turn decision system under perception occlusion provided in Embodiment 2 of this application, please refer to the detailed description of the corresponding part of the left-turn decision system under perception occlusion provided in Embodiment 1 of this application, and will not be repeated here.

[0049] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that the elements inherent in a process, method, article, or apparatus that includes a list of elements are included. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Additionally, portions of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of corresponding technical solutions in the prior art have not been described in detail to avoid excessive elaboration.

[0050] While specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art can make other modifications or variations based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A left-turn decision-making method under perceptual occlusion, characterized in that, Includes the following steps: Acquire vehicle status information and environmental vehicle status information, and introduce virtual vehicles within the perception occlusion area corresponding to the vehicle. An observation space containing continuous temporal information is constructed based on the vehicle's status information, the environmental vehicle status information, and the boundary point information of obstructions. The dynamic baseline is calculated using the action value distribution output by the value distribution network for real-time risk assessment. Based on the observation space and risk assessment results, a near-end strategy with safety constraints is adopted to optimize the PPO decision model to generate decision strategies. The PPO decision model uses the dynamic baseline to mitigate the risk of action selection in the later stages of training, and uses a quantile numerical network and a quantile proposal network to fit the distribution of action value; wherein, the safety constraint is to introduce collision risk as a constraint condition into the policy optimization objective through the Lagrange multiplier method. Based on the generated decision strategy, output control actions to control the vehicle to complete the left turn.

2. The method according to claim 1, characterized in that, The vehicle status information includes: the polar coordinate distance and angle of the vehicle relative to the center point of the intersection, the vehicle speed, and the yaw angle; the environmental vehicle status information includes: the polar coordinate distance and angle of the vehicle relative to the vehicle, the environmental vehicle speed, and the yaw angle.

3. The method according to claim 1, characterized in that, An observation space containing continuous temporal information is constructed based on vehicle state information, environmental vehicle state information, and boundary point information of obstructions; specifically: Establish a polar coordinate system centered on the vehicle; The vehicle's status, the status of surrounding vehicles, and the boundary point information of obstructions are encoded in the polar coordinate system to form a state vector. By stacking the state vectors of multiple consecutive time steps, an observation space containing temporal dynamic relationships is constructed.

4. The method according to claim 1, characterized in that, The dynamic baseline is calculated using the action value distribution output by the value distribution network; specifically: ; in, For dynamic baseline; For expectation operators; For policy functions; To sum over quantile indices; The number of quantiles; The quantiles proposed by the quantile proposal network output; The width of the quantile interval. The output of the quantile numerical network in the state Select Action The corresponding number quantile values; This is the Dirac function.

5. The method according to claim 4, characterized in that, The PPO decision model utilizes the dynamic baseline to perform risk blocking on action selection in the later stages of training, and uses a quantile numerical network and a quantile proposal network to fit the distribution of action values; specifically: ; in, The distribution of action value; This is a symbol for definition.

6. The method according to claim 5, characterized in that, The method also includes: updating network parameters using TD learning, where the TD update formula is: ; in, For the first The error between the predicted quantile and the target quantile; For at any time Quantile numerical networks for current state action pairs The predicted first The quantile value, i.e., the predicted value; At any moment Execute action The instant reward obtained afterward; Discount factor; For the target network to the next state The first one evaluated quantile values; For distributed TD targets; Formula for calculating the overall loss of quantile numerical networks: ; in, For the threshold is Huber's losses; This refers to timing difference error; The number of quantiles; For the first quantiles; Quantile proposal networks employ gradient descent by minimizing the 1-Wasserstein distance between the approximate quantile distribution and the actual quantile distribution. The objective function of the quantile proposal network is: ; in, The gradient of the Wasserstein distance; The distance is 1-Wasserstein. The learnable quantile; Estimation of quantile functions; The target quantile; The value of adjacent target quantiles.

7. The method according to claim 1, characterized in that, The objective function to be optimized in the PPO decision model is: ; in, Let be the objective function to be optimized in the PPO decision model; To optimize the strategy parameters; For the optimized strategy parameters; For expectation operators; Importance weight; To optimize the strategy in the state Select action The probability density; To optimize the policy in the state Select action The probability density; This is the dominant function.

8. The method according to claim 1, characterized in that, The reward function used during training of the PPO decision model is: ; in, For the reward function; Rewards for expert trajectory Rewards for crossing conflict zones, As a penalty for collision, A reward is given for reaching the target lane.

9. The method according to any one of claims 1 to 8, characterized in that, The output control actions include: outputting discrete acceleration values ​​and discrete front wheel steering angle values ​​to control the longitudinal and lateral movements of the vehicle.

10. A left-turn decision-making system under sensory occlusion, characterized in that, include: The state perception module is used to acquire the state information of the vehicle itself and the state information of the surrounding vehicles, and to introduce virtual vehicles within the perception occlusion area corresponding to the vehicle itself. The risk assessment module is used to construct an observation space containing continuous temporal information based on the vehicle's status information, environmental vehicle status information, and boundary point information of obstructions, and to calculate a dynamic baseline using the action value distribution output by the value distribution network for real-time risk assessment. The strategy generation module is used to generate decision strategies by optimizing the PPO decision model based on the observation space and risk assessment results, using a near-end strategy with safety constraints. The PPO decision model uses the dynamic baseline to mitigate the risk of action selection in the later stages of training, and uses a quantile numerical network and a quantile proposal network to fit the distribution of action value; wherein, the safety constraint is to introduce collision risk as a constraint condition into the policy optimization objective through the Lagrange multiplier method. The vehicle control module is used to output control actions based on the generated decision strategy to control the vehicle to complete the left turn process.

Citation Information

Patent Citations

  • Automatic driving automobile decision planning method based on value distribution reinforcement learning

    CN114707359A

  • Vehicle control method and equipment, storage medium and electronic device

    CN114829226A

  • Personified left-turn intersection decision-making method

    CN116714589A

  • Automatic driving decision-making method and system considering shielding uncertainty

    CN118323163A

  • Automatic driving method, device and equipment based on reinforcement learning and storage medium

    CN119356310A