Autonomous Driving Behavior Decision-making Method for Vision Occlusion

By integrating rule mapping and deep reinforcement learning decision-making methods in autonomous driving vehicles, using on-board sensors to obtain direct and indirect observation information, combined with finite state machines and SAC models, the decision-making problems in the field of view occlusion scenario are solved, and the reliability and efficiency of autonomous driving are improved.

CN117284325BActive Publication Date: 2025-07-18XIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311277390.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-28
Publication Date
2025-07-18
Estimated Expiration
2043-09-28

AI Technical Summary

Technical Problem

When facing vision occlusion scenarios, existing autonomous driving technology has too conservative decision strategies, resulting in low traffic efficiency, and the methods that rely on network communication have real-time and security issues, and the existing methods have "black box" problems and require a strong knowledge base to be established.

Method used

The fusion decision-making method based on rule mapping and deep reinforcement learning is adopted, and direct observation information and indirect observation information are obtained through on-board sensors for data fusion, and combined with a finite state machine and a pre-trained SAC model, the optimal driving action is selected.

Benefits of technology

It improves the decision reliability and efficiency of autonomous driving vehicles in the field of vision occlusion scenario, reduces the conservatism of decisions, and improves the accuracy of decisions and the convergence of model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117284325B_ABST
    Figure CN117284325B_ABST
Patent Text Reader

Abstract

The present invention discloses an autonomous driving behavior decision-making method for vision occlusion, which specifically comprises the following steps: Step 1, obtain scene information; Step 2, infer potential risks of the vision occlusion scene by obtaining time, position information and scene element information; Step 3, fuse the data obtained in Step 1 and Step 2; Step 4, generate a decision-making scheme based on rule mapping; Step 5, generate a decision-making scheme and driving behavior based on a deep reinforcement learning model; Step 6, select the optimal driving decision. By mining non-direct observation information in the vision occlusion scene and using numerical encoding of it and direct observation information as the input of the decision-making method, the present invention has better convergence and decision-making accuracy compared with the deep reinforcement learning decision-making model that takes only direct observation values as input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving behavior decision-making methods, and particularly relates to an autonomous driving behavior decision-making method for vision occlusion. Background Art

[0002] Research methods for autonomous driving vision occlusion scenarios mainly start from perspectives such as vehicle networking and vehicle-road cooperation. Although vehicle networking and vehicle-road cooperation methods can indirectly enable autonomous driving vehicles to comprehensively perceive the surrounding environment through network communication, in decision-making research, the first thing to solve is problems such as the real-time performance and security of information transmission. In addition, vehicle-road cooperation relies on the environmental information obtained by roadbed sensors to communicate with surrounding vehicles, but this method has problems such as a large amount of infrastructure construction and maintenance and the randomness of a large number of vision occlusion scenarios.

[0003] On the other hand, in autonomous driving behavior decision-making research, the method based on rule mapping is the most classic and easy to deploy. However, since the rules are preset by researchers, overly conservative driving strategies often appear in previous applications, which instead greatly reduces the traffic efficiency. Secondly, knowledge reasoning and value-based decision-making models are also widely used in the field of autonomous driving, but both of these methods have the "black box" problem, and the former must pre-establish a powerful knowledge base. Summary of the Invention

[0004] The object of the present invention is to provide an autonomous driving behavior decision-making method for vision occlusion, which can improve the reliability and efficiency of autonomous driving vehicle decision-making in the face of potential dangerous scenarios of vision occlusion.

[0005] The technical solution adopted by the present invention is an autonomous driving behavior decision-making method for vision occlusion, which specifically comprises the following steps:

[0006] Step 1: Obtain scene information;

[0007] Step 2: Infer the potential risk of the vision occlusion scene by obtaining time, position information and scene element information;

[0008] Step 3: Fuse the data obtained in Step 1 and Step 2;

[0009] Step 4: Generate a decision-making scheme based on rule mapping;

[0010] Step 5: Generate a decision-making scheme and driving behavior based on a deep reinforcement learning model;

[0011] Step 6: Selection of the optimal driving decision.

[0012] The technical features of the present invention also lie in that

[0013] Step 1 is specifically as follows: obtaining the distance, speed and acceleration information between the vehicle and the obstacle that blocks the view of the vehicle in real time through on-board sensors and calculations.

[0014] In step 3, the direct observations in step 1 and the virtual observations after numerical coding are integrated.

[0015] Step 4 is as follows: by establishing a finite state machine decision model, the driving decision is divided into multiple states, the state trigger conditions and state transfer functions are established, and according to the fusion information of step 3, the driving state is quickly triggered and mapped, and the driving action is generated by the state transfer function.

[0016] Step 5 is specifically as follows: generating a driving strategy through the pre-trained SAC model and the information of step 3, and outputting specific driving action behaviors.

[0017] Step 6 is as follows: select Q(s,a) based on the inferred action value function Q(s,a) R (s,a R ) and Q r (s,a r ), where s represents the state of the environment at a certain moment, and a represents the action taken by the agent corresponding to the state of the environment at this moment; a R represents the action taken by the deep reinforcement learning decision model; a r Indicates actions taken based on the FSM decision model.

[0018] The beneficial effects of the present invention are:

[0019] 1) From the perspective of single-vehicle intelligence, this invention solves the problem of autonomous vehicles’ behavioral decision-making in potentially dangerous traffic scenarios where the city’s traffic vision is blocked. Compared with the problems of Internet of Vehicles and vehicle-road collaboration solutions that rely on network stability and roadbed equipment, this invention has better generalization and practicality;

[0020] 2) Based on the value-based deep reinforcement learning decision model, the present invention combines FSM to construct a decision-making method for autonomous driving vision occlusion potential dangerous traffic scenarios based on deep reinforcement learning and rule fusion. Compared with a single decision model, this method not only solves the problem of single decision and conservative strategy, but also improves the reliability of deep reinforcement learning decision-making;

[0021] 3) The present invention mines the indirect observation information in the field of view occlusion scene, and numerically encodes it together with the direct observation information as the input of the decision-making method. Compared with the deep reinforcement learning decision model that only uses direct observation values as input, it has better convergence and decision accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1It is the process block diagram of the decision-making for the potential dangerous scene of vision occlusion in the present invention;

[0023] Figure 2a It is the top view of the scene of Example 3 of vision occlusion adopted in the present invention;

[0024] Figure 2b It is the rear view of the scene of Example 3 of vision occlusion adopted in the present invention;

[0025] Figure 2c It is the perspective view of the autonomous vehicle of the scene of Example 3 of vision occlusion adopted in the present invention;

[0026] Figure 3 It is the simplified scene of Example 3 in the present invention;

[0027] Figure 4 It is the schematic diagram of selecting the optimal action by generating an action based on the decision-making of the simplified scene diagram of Example 3 in the present invention;

[0028] Figure 5 It is the framework diagram of the training process and the tracked rule-based state value function in the present invention;

[0029] Figure 6 It is the decision-making flow chart of the potential dangerous scene of vision occlusion in the present invention. Specific embodiments

[0030] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0031] The autonomous driving behavior decision-making method for vision occlusion in the present invention is specifically implemented according to the following steps:

[0032] Step 1: Obtain scene information (direct observation value);

[0033] Specifically, step 1 is to obtain the distance, speed, and acceleration information of the vehicle itself and the obstacles in the occluded vision of the vehicle itself in real time through vehicle-mounted sensors (radar, lidar, etc.) and calculations;

[0034] Step 2: Obtain the potential risks (non-direct observation values) of the vision occlusion scene by inferring time, position information, and scene element information;

[0035] Step 3: Fuse the data of the direct observation value and the non-direct observation value;

[0036] In step 3, the direct observation value in step 1 and the virtual observation value after numerical encoding are integrated;

[0037] Step 4: Generate a decision-making scheme based on rule mapping;

[0038] Step 4 is as follows: by establishing a finite state machine decision model (FSM), the driving decision is divided into (initial state, keep driving, acceleration, deceleration, left lane change, right lane change, etc.), the state trigger condition Envents and the state transfer function are established, and according to the fusion information of step 3, the driving state is quickly triggered and mapped, and the driving action is generated by the state transfer function;

[0039] Step 5: Generate decision-making solutions and driving behaviors based on deep reinforcement learning models;

[0040] Step 5 is as follows: generate a driving strategy through the pre-trained deep reinforcement learning algorithm SAC model (soft actorcritic) and the information of step 3, and output specific driving action behaviors;

[0041] Step 6: Selection of optimal driving decision;

[0042] Step 6 is as follows: select Q(s,a) based on the inferred action value function Q(s,a) R (s,a R ) and Q r (s,a r ), since the two methods correspond to the same environmental model and reward function, the action with the largest Q(s,a) is directly selected.

[0043] Example 1

[0044] like Figure 1 As shown, the principle of the present invention is:

[0045] The present invention is a decision-making method for potential traffic hazard scenarios caused by obstruction of vision of an autonomous driving vehicle. This is a decision-making framework based on the fusion of rule mapping and deep reinforcement learning decision-making methods, which can improve the decision-making reliability of the autonomous driving vehicle in this scenario, reduce the conservatism of the decision, and improve traffic efficiency.

[0046] Example 2

[0047] First, the decision-making method adds indirect observation information (virtual observation values) based on the previous decision-making model that only relies on direct observation information as input. The direct observation information includes the speed, distance, lane centerline distance, and angle with the lane centerline of the vehicle and the vehicle that causes the field of view obstruction; the indirect observation information is the non-potential risk value of the environment in the field of view obstruction scene. Then, the model input is combined with the direct observation value through numerical coding. On the one hand, this information fusion method can mine the information of the potential dangerous scene of field of view obstruction and make correct decisions; on the other hand, during the training process, the indirect observation information will also make the decision model converge quickly.

[0048] Secondly, a decision-making framework for deep reinforcement learning that integrates rule mapping is adopted. Since the input contains non-direct observation information, it is actually a constraint on the model during the training process, which is called a virtual constraint and can accelerate the convergence of the model. However, to prevent the model from falling into a local optimum, the SAC algorithm is selected as the decision-making model for deep reinforcement learning. This deep reinforcement learning algorithm incorporates the concept of information entropy into the reward function, which can encourage the agent to expand the search of the action space. On the other hand, the decision-making method based on rule mapping uses a classic finite state machine (FSM) model to represent the driving decisions of an autonomous vehicle as a finite number of states. In this method, these states are the initial state, maintaining driving, accelerating, decelerating, passing slowly, etc. Decisions are made by defining the trigger states and the transitions between states. During the model training stage, refer to Figure 5 Save the decisions generated by the rules as environment-action pairs in the memory pool. The reward function established by the deep reinforcement learning algorithm tracks the environment-action pairs in the memory pool to generate a decision sequence trajectory containing rewards for preservation, thereby inferring the action value function Q r (s,a).

[0049] To train the decision-making model, an environment model for reinforcement learning needs to be established in the simulation software Carla. After the decision-making model is trained, after inputting the non-direct observation values and direct observations of the environment, on the one hand, decisions are generated by the rules and actions are output. On the other hand, decisions and actions are generated by the deep reinforcement learning model. By comparing the action value functions of the actions of the two, the optimal decision action is selected as the output of the autonomous vehicle.

[0050] Embodiment 3

[0051] An autonomous driving behavior decision-making method for vision occlusion, as Figure 6 shown, is specifically implemented according to the following steps:

[0052] Step 1, the input is divided into two parts, as Figure 1 shown in part a) of : One is the direct observation information, which is the speed of the host vehicle, the speed of the obstacle vehicle, the distance, the angle with the center line of the lane, and the distance from the center line of the lane obtained by the on-vehicle sensors and calculations of the autonomous vehicle. The other is the non-direct observation information, which is the potential risk value information of the vision occlusion scenario in this environment inferred based on the position information of the vehicle in the world coordinate system, the surrounding environment information, and the time obtained by the on-vehicle sensors;

[0053] Step 2, use the information encoding method to integrate the non-direct observation information and direct observation information in Step 1, and represent high, medium, and low potential risks as 0, 1, and 2 in numerical representation. The integrated input information obtained is (ego_V y ,o_V y, ego_ds, o_ds......(0 / 1 / 2)), as shown in part b) of Figure 1 ; ego_V y represents the speed of the autonomous vehicle in the lane direction; o_V y represents the speed of the vision-obstructed vehicle in the lane direction; ego_ds represents the angle between the autonomous vehicle and the center line of the lane; o_ds represents the angle between the vision-obstructed vehicle and the center line of the lane; (0 / 1 / 2) indicates that the macro-scene risk of the scenario is low (0), medium (1), or high (2).

[0054] Step 3: Input the integrated information into a rule-based decision-making model, such as Figure 1 shown in part c) of Figure 2a , 2b, 2c, and Figure 3 shown. Since the input information satisfies the judgment that the obstacle vehicle in the right front is parked by the roadside (as shown in 2 ), and since the input information determines that the risk is high at this time, a state transition is triggered, and the autonomous driving vehicle enters the "deceleration state". At this time, the state transition function outputs the action: the vehicle decelerates from 25 km / h to 10 km / h at a rate of 2 m / s Figure 4 ; when decelerating to 10 km / h, a state transition is triggered to slowly pass through, and the output action is 10 km / h. The above rule-based decision is as shown in the rule-based part of r : {(s1, a1), (s2, a2)......};

[0055] Step 4: As shown in part c) of Figure 1 : Input the integrated information into a pre-trained deep reinforcement learning model to obtain the action sequence A R : {(s1, a1), (s2, a2)......};

[0056] Step 5: As shown in d) of Figure 1 and Figure 5 shown: During training, store the scenario-action pairs through the decision-making action storage pool. The action is output to the environment, and then the reward function designed by the deep reinforcement learning model in the environment gives the reward value for each action, and stores the decision-making trajectory O: {s1, a1, r1, s2, a2, r2......} of one training in the decision-making trajectory pool, and obtains the action value function Q(s, a) through speculation. Thus, according to the state and action in Step 3, the specific value Q r (s, a r ) of a certain action is obtained;

[0057] Step 6: Compare the action value Q r(s, a r ) and the action value Q of the action generated by reinforcement learning R (s, a R ), and select the decision-making action with the maximum action value and output it to the autonomous vehicle.

[0058] The autonomous driving behavior decision-making method for vision occlusion in the present invention sets the step process based on the following framework: environmental information (direct observation information acquisition, scene danger state information inference (non-direct observation information), data fusion of direct observation values and non-direct observation values, decision-making model based on rule mapping, decision-making model based on deep reinforcement learning, and selection of the optimal driving decision based on the optimal action value function.

Claims

1. An autonomous driving behavior decision-making method for vision occlusion, characterized in that, The implementation is specifically carried out according to the following steps: Step 1: Obtain scene information, which is direct observation information; Step 2: Obtain potential risks of the field of view occlusion scene inferred from time, position information and scene element information, which is non-direct observation information; Step 3: Integrate the non-direct observation information and the direct observation information using an information coding method; Step 4: Through the established finite state machine decision model, divide the driving decisions into multiple states, establish state trigger conditions and state transition functions, and according to the integrated information in Step 3, quickly trigger and map to driving states, and generate driving actions by the state transition functions; Step 5: Generate a driving strategy through the pre-trained SAC model and the integrated information in Step 3, and output specific driving actions; Step 6: Selection of the optimal driving decision, through the already inferred action value function Select the action value of the action generated by reinforcement learning And the action value of the action generated by the rule The action with the optimal median value, where s represents the environmental state at a certain moment, represents the action taken by the agent corresponding to the environmental state at this moment; represents the action taken by the deep reinforcement learning decision model; represents the action taken based on the FSM decision model.

2. The autonomous driving behavior decision-making method for vision occlusion according to claim 1, wherein The specific content of Step 1 is: Real-time obtain the distance, speed and acceleration information of the vehicle itself and the obstacles in the occluded field of view of the vehicle itself through in-vehicle sensors and calculations.

Citation Information

Patent Citations

  • Method for realizing safety decision control of autonomous vehicle

    CN114644017A

  • Automatic driving lane changing decision control method based on rule fusion reinforcement learning

    CN115257745A