Computer implementation method for determining an action plan to drive an autonomous agent in traffic conditions
The method addresses the limitations of existing occlusion-aware algorithms by constructing a scenario tree with decision deferral and optimizing action plans for autonomous agents, improving safety and adaptability in diverse traffic conditions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2026-03-31
AI Technical Summary
Existing occlusion-aware behavioral planning algorithms for autonomous vehicles and robots are either overly conservative, leading to degraded performance, or lack generality, being applicable only to specific traffic scenarios, and fail to integrate information gathering with collision avoidance.
A computer-implemented method that constructs a scenario tree based on current and predicted occlusion maps, considering both worst-case and best-case scenarios, with a decision deferral time to resolve occlusions, optimizing action plans using a cost function that maximizes information gain and minimizes risk.
Enhances safety and adaptability of autonomous agents by reducing conservative behavior and enabling effective decision-making in various traffic scenarios, mimicking human-like information gathering and risk recognition.
Smart Images

Figure 2026055784000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a computer implementation method for determining an action plan for driving an autonomous agent in traffic conditions. The present invention further relates to a data processing device, a computer program, a computer-readable data carrier, and an autonomous agent. [Background technology]
[0002] Planning the operation of autonomous vehicles and mobile robots in environments with significantly obstructed areas is a critical challenge in the industry. These occlusions are primarily caused by the physical limitations of sensor elements and can reduce the level of safety and performance during operation. There are many real-world examples where occlusion can significantly impact traffic conditions. For example, pedestrians may be obscured behind buildings or cars parked near crosswalks. Another example is when a car or motorcycle is obscured from its own viewpoint while attempting to overtake a truck in front of it.
[0003] One of the most common approaches to address this problem is the development of "occlusion-aware" behavioral planning algorithms. Conventional algorithms in this area employ different approaches, such as: • Reachability analysis - Predict all possible worst-case scenarios, infer the space occupied in time ahead based on the predictions, and control the vehicle to avoid those areas.[1] • Game-theoretic approach – Formulate occlusion-aware planning as a game between the vehicle and other potentially obstructed and conflicting traffic participants. The optimal solution is formulated as a dynamic game problem and then used to plan the vehicle's actions.[2] However, it may lack generality as it is designed for only one traffic scenario. • Potential field method - The environment is represented by a scalar field such as a risk field[3] or an attractive / repulsive field[4], and the action planning process is carried out based on the gradient of that field. • Information-theoretic approach – Introduce an information-theoretic goal related to the level of environmental occlusion. During action planning, this goal is optimized in conjunction with other goals (e.g., performance goals and safety goals) so that the vehicle can proactively collect new information about the occupied area [5, 6].
[0004] Another important segment of prior art relates to techniques for 3D occupancy prediction, which are therefore used as modules for occlusion-aware behavioral planning algorithms. Prior art solutions for 3D occupancy prediction leverage modern AI methods that provide occupancy and semantic predictions in 3D space based on sensor inputs such as cameras [7, 8, 9].
[0005] There are several potential problems related to the technologies described in the prior art section: • The handling of occlusion is too conservative - This problem primarily exists in techniques that consider worst-case scenarios, such as reachability analysis, which can lead to overly conservative vehicle behavior and potentially degraded performance. - They do not combine information gathering capabilities with collision avoidance that recognizes the risks of possible occlusions. On the one hand, [1-4] they may ignore information gathering capabilities but recognize the risks of possible occlusions. On the other hand, [5, 6] they may not recognize the risks but attempt to incorporate information-theoretic objectives into their action planning. • Lack of generality - Some of the references mentioned above are designed for specific traffic scenarios only; that is, they may not be applicable to many other traffic scenarios.
[0006] This specification refers to the following references. [Prior art documents] [Non-patent literature]
[0007] [Non-Patent Document 1] R. Firoozi, A. Mir, GSCamps, and M. Schwager, “Occlusion-Aware MPC for Guaranteed Safe Robot Navigation with Unseen Dynamic Obstacles.” arXiv,Nov.16,2022.Accessed:Mar.16,2024.[Online].Available:http: / / arxiv.org / abs / 2211.09156
Outdoor Tool2
Outdoor Tools3
Outdoor Tools 4
Problems to be Solved by the Invention
[0008] The object of the present invention is to provide an improved method for determining an operation plan for driving an autonomous agent in a traffic situation.
Means for Solving the Problems
[0009] To achieve this object, the present invention provides a computer-implemented method according to claim 1. A data processing device, a computer program, a computer-readable data carrier and an autonomous agent are the subject matter of the parallel claims. Advantageous embodiments of the present invention are the subject matter of the dependent claims.
[0010] In one aspect, the present invention provides a computer-implemented method for determining an operation plan for driving an autonomous agent in a traffic situation, the agent including sensors for capturing sensor data within a sensor field of view, the sensor data indicating the traffic situation, the method comprising a) evaluating the traffic situation at the current time t0 by determining a current occlusion map based on the current sensor data, the current occlusion map indicating one or more occluded regions within the sensor field of view where relevant sensor data is not currently available; b) constructing a scenario tree including a plurality of scenarios regarding the traffic situation during a planning period based on the current occlusion map, and for each scenario, determining an appropriate operation plan for driving the agent, the scenario tree being based on a decision delay that constrains all appropriate operation plans to a common operation plan for a decision delay time T p up to, the decision delay time T p being calculated based on the predicted occlusion map for each scenario and indicating the time at which all but one scenario can be excluded; constructing and determining; c) re-evaluating the traffic situation at a future time t1>t0 until a single scenario remains, and the decision delay time T pSelecting a corresponding action plan to drive the agent in traffic conditions beyond that point. Includes.
[0011] An advantage of this method is that occlusion can be considered during the action planning stage. The agent may be an autonomous vehicle, robot, drone, etc. Occlusion is determined based on the predicted occlusion for each scenario, with a delay time T. p This can be considered by calculating the decision deferral time T. p Until all relevant occlusions are resolved, the system initially follows a common operational plan. This can improve the level of safety during autonomous operation. Decision deferral can be considered passive information gathering.
[0012] Preferably, step b) is b1) Determining an appropriate operation plan by minimizing based on a cost function, wherein the cost function includes a predicted occlusion map for each scenario, and preferably includes a term for maximizing the information gain for one or more currently occluded regions. It also includes.
[0013] An advantage of this method is that a common action plan can be determined based on a cost function that includes a predicted occlusion map. Therefore, the agent can attempt to resolve occlusion while following the common action plan. The common action plan can therefore be determined in a way that optimally resolves the relevant occlusions. This can be considered active information gathering. This can mimic human behavior. For example, in a traffic situation where an agent is following a traffic participant and attempting to overtake it, the agent might first explore the area in front of the traffic participant by abruptly turning towards the center of the road, but still not initiate the overtaking maneuver. This can make the action plan more effective.
[0014] Preferably, step c) is c1) Update the predicted occlusion map based on sensor data at future time t1 > t0, and determine the deferral time T based on the updated predicted occlusion map. p To make it conform, and / or c2) Eliminate one or more scenarios from the scenario tree based on sensor data at future time t1 > t0. It further includes one or both of the above.
[0015] The advantage of this method is that it reduces the decision delay time T while following a common action plan. p This may be applicable. Therefore, if occlusion is resolved faster than expected, the agent can immediately make a decision on the action plan. This can make the action plan more effective.
[0016] Preferably, step b) is b2) Constructing a scenario tree by including one or more worst-case scenarios, each worst-case scenario assuming one or more obstacles in at least one of the currently shielded areas, and / or b3) Constructing a scenario tree by including one or more best-case scenarios, each of which assumes that there are no obstacles in all currently shielded areas. It further includes one or both of the above.
[0017] The advantage of this method may be that it allows for the consideration of both worst-case and best-case scenarios. Therefore, the action plan can recognize risks, but it does not need to be overly conservative.
[0018] Preferably, step b) is b4) To calculate the reachable area of the assumed obstacle in at least one of the currently shielded areas, estimate the maximum possible velocity and / or acceleration of the assumed obstacle. It also includes.
[0019] An advantage of this method is that, in the worst-case scenario, it may be able to consider the maximum possible speed and / or acceleration. Therefore, the action plan can recognize risks, but it does not need to be overly conservative.
[0020] Preferably, step b) is b5) For each action plan, predict the occlusion map for each scenario by mapping the agent's position and / or orientation to one or more currently occluded areas. It also includes.
[0021] Preferably, step a) is a1) Based on the current sensor data, determine the current occupancy map showing one or more regions within the sensor's field of view that are occupied, a2) Determine the current occlusion map based on the current occupancy map. It also includes.
[0022] Preferably, step a2) is a2a) Determining the current occlusion map by extending a straight line from the sensor across one or more occupied areas to the sensor's maximum sensor detection distance. It also includes.
[0023] Preferably, step a) is a3) Filter and / or sort the current occlusion map by ranking one or more currently occluded areas according to a relevance score for evaluating traffic conditions. It also includes.
[0024] An advantage of this method is that occlusions may be filtered by relevance. For example, occlusions far from the agent or road may be less relevant. This can lead to more effective action planning.
[0025] Preferably, step a3) is a3a) Determining a relevance score based on the navigation map, the distance from the agent to one or more currently occupying areas, the reachable area of one or more assumed obstacles, and / or the nature of one or more occupied areas within the sensor's field of view. It also includes.
[0026] An advantage of this method may be that it can take into account the nature of the occupied area. For example, if an area is not occupied by vehicles, pedestrians, or bicycles, the occupied area may be less relevant than other areas. This can lead to more effective action planning.
[0027] Preferably, this method d) Generating control signals to drive agents based on a common and / or selected operation plan. It also includes.
[0028] In another aspect, the present invention provides a data processing device that includes means for performing any of the methods of the prior embodiments.
[0029] Any features, aspects, and / or advantages described herein with respect to embodiments of data processing devices may be optionally applied to embodiments of methods, and vice versa.
[0030] In another aspect, the present invention provides a computer program which, when the program is executed by a computer, includes instructions causing the computer to perform any of the methods of the prior embodiments.
[0031] Any features, aspects, and / or advantages described herein with respect to embodiments of computer programs may be optionally applied to embodiments of methods and / or data processing devices, and vice versa.
[0032] In another aspect, the present invention provides a computer-readable data carrier having a computer program stored thereon.
[0033] Any features, aspects, and / or advantages described herein with respect to embodiments of computer-readable data carriers may be optionally applied to embodiments of methods, data processing devices, and / or computer programs, and vice versa.
[0034] In another embodiment, the present invention provides an autonomous agent including a sensor for capturing sensor data within a sensor field of view, the sensor data indicating traffic conditions, and the agent further includes a data processing device according to any of the prior embodiments.
[0035] Any features, aspects, and / or advantages described herein with respect to embodiments of autonomous agents may be optionally applied to embodiments of methods, data processing devices, computer programs, and / or computer-readable data carriers, and vice versa.
[0036] Preferred embodiments of the present invention can be summarized as follows.
[0037] Based on the map of occluded spaces searched, potentially occluded objects, their relationships, and their future trajectories are evaluated. Using information on occluded spaces and potentially occluded objects, the planning algorithm selects the optimal next action for the vehicle to minimize risk from potentially occluded objects and maximize information gain, i.e., minimize future relevant occluded spaces. One embodiment of this solution is shown in Figure 3 and consists of the following steps.
[0038] 1. Building the occupied grid map: The inventors utilize existing approaches such as [7], [8] or [9] to obtain a 3D occupancy grid map (3D-OGM) or a 2D occupancy grid map (2D-OGM, so-called "bird's-eye view") from current sensor measurements and optionally from measurement history, respectively. Given the resolution of the 3D or 2D space, the 3D-OGM and 2D-OGM contain information on which cells in the space are occupied. Using additional information such as from segmentation and classification, occupied cells can be annotated with a class (e.g., vehicle, pedestrian, road, sidewalk, etc.).
[0039] 2. Add occlusion information: Based on the vehicle's sensor configuration, a sensor model is defined for each sensor (e.g., lidar, camera, radar), and occlusion information as shown in [3] is obtained. Based on each sensor model, the resolution (discretized), sensor detection distance, and field of view (FOV) with the sensor center are obtained. For each cell of the discretized FOV, a 3D ray of equal length to the sensor detection distance, starting from the sensor center, is aligned with the cell (for example, in the case of a pinhole camera model, the center of the camera's coordinate system, the camera's focal length, and each pixel of the camera's image define how the ray should be aligned). After projecting the ray onto a 3D-OGM or a 2D plane, the cells where the OGM intersects can be obtained for each ray, along with the 2D-OGM.
[0040] Starting from the center of the sensor, each cell is marked as visible until it reaches the first occupied cell or the end of the ray. After evaluating each sensor as described above, the result of this step is a 3D-OGM or 2D-OGM containing occlusion information, where any unoccupied cells not marked as visible are occluded.
[0041] 3. Assessment of the situation (scenarios, relevance of shielding areas): Based on 3D-OGM or 2D-OGM, which includes occlusion information (providing information about lanes, intersections, etc.) and infrastructure information from the SD map, occluded areas are filtered by relevance. In this step, all occluded cells are filtered by relevance by projecting each occluded cell onto the SD map, and are retained only if they are projected onto areas relevant to the planning phase (roads, sidewalks, intersections, etc.). Furthermore, the optimal trajectory previously planned for the vehicle is considered to filter the remaining occluded cells by distance metric; that is, occluded cells that are too far away are discarded. In addition, information such as annotations on the 3D-OGM or 2D-OGM map or classification and segmentation of sensor data can be used to detect traffic conditions and areas (scenarios) with higher risk. To take this into account, a relevance score is added to each occluded cell, with each occluded cell assigned to a high-risk scenario having a higher relevance. For example, the relevance score is increased for occluded cells within or near an area where high activity is detected (e.g., many pedestrians detected on a sidewalk) or cells occluded by specific objects such as double-parked cars blocking the opposite lane.
[0042] 4. Calculation of the worst set of obstructed obstacles that can be reached forward: Next, the output 3D map is used to calculate the possible future positions of potentially occluded objects. For this purpose, a worst-case scenario can be used by calculating the forward reachable set (simplifying physics using a point mass model), and the possible maximum velocity and acceleration of the occluded object are limited by the most likely object type extracted from the situation assessment (pedestrian or vehicle). This set can be added to the scenario tree of the branched model predictive control (MPC) as a contingency scenario. The branched MPC (see point 7) can then be planned using all the scenarios added to the scenario tree, i.e., one contingency scenario and one scenario in which potentially occluded objects are ignored.
[0043] 5. Adaptive decision postponement: To calculate an appropriate decision deferral time, the predicted occlusion state calculated by the branch MPC in the last time step (see point 7 for details) is used. Thus, the inventors check all time steps within the planning period to determine whether there is sufficient information available at any given time step (i.e., the occlusion is sufficiently small) and whether or not an object is present. In other words, this block outputs the time step at which the autonomous vehicle is expected to be able to make an informed decision.
[0044] 6. Information gathering cost function: The cost function can be defined as shown in [5] and [6]. The goal is to maximize progress on the reference path, minimize deviation from the path, maximize distance to obstacles, and minimize lateral and longitudinal motion and acceleration for comfortable driving, including minimizing the occluded space adjacent to the cost function term. If there are multiple related occluded spaces, occlusion weights are applied based on the relationships extracted from point 3. To minimize the occluded space, terms are added to the cost function to maximize the information gain between two time steps (the difference in occlusion between two time steps).
[0045] 7. Predictive control of branching models: The branched MPC is constructed using the decision deferral time from point 5 and the information acquisition cost function from point 6. The model used for the branched MPC is simply the kinematic bicycle model, but can be extended with occlusion states. The occlusion states are calculated using occlusion data acquired at point 2 and the associated occlusions identified from point 3. A function is established that maps the current position and orientation of the autonomous system to the occlusion region of each selected occlusion. Essentially, this involves formulating a geometric equation to calculate the field of view (FOV) and returning the occluded space for the input pose of the autonomous system. This occlusion state is also used to formulate the information acquisition cost function.
[0046] The branched MPC plans to minimize the information gathering cost function while simultaneously satisfying the constraints, using all scenarios added to the scenario tree at point 4. This is achieved by the branched MPC by allowing different strategies (i.e., control inputs) that are constrained to be identical in the first time step, i.e., until the calculated time for decision deferral is reached, i.e., until it becomes clear which strategy to take.
[0047] Embodiments of the present invention preferably have the following advantages and effects.
[0048] - Reduce the conservative tendencies of autonomous systems by considering risks in the following ways: • Gather information while still considering the risk of collision with shielded objects. By setting a decision postponement based on the future state of occlusion, the planner will be aware of when they have enough information to decide what action to take. • Allows planners to design different strategies that autonomous vehicles can adapt to depending on which scenario occurs.
[0049] - Enhance the human-like behavior of autonomous systems that are automatically adapted (not hardcoded) as follows: • Nudging for information gain (maximizing field of view) • Delay approaching competitive zones (such as occupying intersections) to wait for further information to come.
[0050] - This approach may be generally applicable to a variety of traffic scenarios involving occlusion. It is not designed for a specific traffic scenario (not just for intersections or obstructed views by large trucks, for example), meaning it does not need to be reconfigured for various traffic scenarios such as [2, 5].
[0051] - This architecture can be easily extended to handle other uncertain scenarios in motion planning (not just for occluded objects). For example, additional branches can be added to handle multimodal predictions regarding the motion of detected traffic participants.
[0052] - Embodiments of the present invention can be further applied to motion planning in robots, drones, and the like.
[0053] Herein, embodiments of the present invention will be described in more detail with reference to the accompanying drawings. [Brief explanation of the drawing]
[0054] [Figure 1] This demonstrates an autonomous agent in a typical traffic scenario. [Figure 2] This document presents a first embodiment of a computer implementation method for determining an action plan to drive an agent based on traffic conditions. [Figure 3] A second embodiment of the computer implementation method is shown. [Modes for carrying out the invention]
[0055] Figure 1 shows an autonomous agent 10 in an exemplary traffic situation 12 at time t0.
[0056] An exemplary traffic situation 12 includes three traffic participants 14: a first traffic participant 14a having the characteristics of a moving pedestrian, a second traffic participant 14b having the characteristics of a moving vehicle, and a third traffic participant 14c having the characteristics of another moving vehicle. Agent 10 moves in the direction of agent travel 16. The third traffic participant 14c moves in directions 18, 18c opposite to the direction of agent travel 16. The third traffic participant 14c may be considered an obstacle 44.
[0057] The exemplary traffic situation 12 further includes a static object 20 as an obstacle 44 in front of agent 10 and in the agent's direction of travel 16. The static object 20 is located between agent 10 and a third traffic participant 14c.
[0058] For autonomous operation, driving, or piloting, agent 10 includes sensors 22 for capturing sensor data 24 within a sensor field of view 26. For example, sensors 22 may be radar devices, lidar devices, and / or camera devices. Sensor data 24 may include data on position, distance, velocity, acceleration, trajectory, and / or the properties of objects 20 and / or traffic participants 14 within the sensor field of view 26. To identify the properties of objects 20 and / or traffic participants 14, machine learning algorithms, such as neural networks, trained for classifying and / or segmenting traffic situation scenes may be used.
[0059] The static object 20 obstructs an area 28 within the sensor field of view 26, preventing the sensor 22 from capturing the associated sensor data 24. The third traffic participant 14c is located within the obstructed area 28 but outside the sensor field of view 26. In other words, due to the static object 20, the agent 10 cannot capture the third traffic participant 14c by the sensor 22, and the sensor data 24 associated with the third traffic participant 14c is not currently available.
[0060] The idea of the embodiment of the present invention is to determine an operation plan 30 for driving an agent 10 in a traffic situation 12, and the operation plan 30 is based on an occlusion area 28.
[0061] FIG. 2 shows a first embodiment of a computer-implemented method for determining an operation plan 30 for driving an agent 10 in a traffic situation 12.
[0062] In step S11, the first embodiment is - evaluating the traffic situation 12 at the current time t0 by determining a current occlusion map 32 based on current sensor data 24, wherein the current occlusion map 32 indicates one or more occlusion areas 28 within the sensor field of view 26 for which the associated sensor data 24 is not currently available, evaluating including.
[0063] In step S12, the first embodiment is - constructing a scenario tree 34 including a plurality of scenarios 36 regarding the traffic situation 12 during a planning period based on the current occlusion map 32, and for each scenario 36, determining an appropriate operation plan 30 for driving the agent 10, wherein the scenario tree 34 constrains all appropriate operation plans 30 to a common operation plan up to a decision delay time T p The decision delay time T p is calculated based on the predicted occlusion map 32 for each scenario 36 and indicates the time at which all but one of the scenarios 36 can be excluded, constructing and determining including.
[0064] In step S13, the first embodiment is - re-evaluating the traffic situation 12 at a future time t1>t0 until a single scenario 36 remains, and selecting the corresponding operation plan 30 for driving the agent 10 in the traffic situation 12 beyond the decision delay time T p including. including.
[0065] Refer to Figure 1 again.
[0066] In step S11, the method can determine all areas 28 within the sensor field of view 26 that are occupying and for which the associated sensor data 24 is not available at time t0. The method can first determine a current occupancy map 38 showing the occupied areas 40 within the sensor field of view 26. Then, the current occlusion map 32 can be determined by extending a straight line from the sensor 22 beyond the occupied areas 40 to the maximum sensor detection distance 42 of the sensor 22. Areas 28 beyond the occupied areas 40 can then be marked as occupying. Thus, the occlusion map 32 can be determined based on the occupancy map 38.
[0067] In step S11, the method may further filter and / or sort the occupying regions 28 according to their relevance scores. For example, an occupying region 28 in the agent's direction of travel 16 may be more relevant than an occupying region 28 in another direction. The method may further compare the occupying regions 28 with a navigation map 58 that may indicate whether or not the occupying regions 28 are located on a road. The relevance scores may further be based on the distance from the agent 10 to the occupying region 28 and / or the nature of the occupied area 40 within the sensor field of view 26.
[0068] In step S12, the method constructs a scenario tree 34 containing one or more scenarios 36 based on the current occlusion map 32. In this step, the method may include one or more worst-case scenarios 48, each worst-case scenario 48 assuming one or more obstacles 44 within the occlusion area 28. For example, one worst-case scenario 48 may assume a third traffic participant 14c as an obstacle 44 within the occlusion area 28. Another worst-case scenario 48 may assume an obstacle 46 that does not exist within the occlusion area 28. Yet another worst-case scenario 48 may assume both a third traffic participant 14c and an obstacle 46 that does not exist within the occlusion area 28.
[0069] In step S12, the method may further estimate the maximum possible speed and / or acceleration and / or trajectory for each assumed obstacle 44 in order to calculate the reachable area of the assumed obstacle 44 within the occlusion area 28. For example, if the method assumes the third traffic participant 14c is a vehicle, the method may assume a maximum possible speed of 70 km / h. If the method assumes the non-existent obstacle 46 is a pedestrian, the method may assume a maximum possible speed of 10 km / h. The reachable area of the assumed obstacle 44 may also be considered in the relevance score.
[0070] In step S12, the method may further include one or more best-case scenarios 50, each best-case scenario 50 assuming there are no obstacles 44 within the shielding area 28. In step S12, the method further determines an appropriate action plan 30 for each scenario 36. In this step, the scenario tree 34 determines all appropriate action plans 30 with a delay time T p This is based on decision deferral, which is constrained to a common action plan. In other words, an appropriate action plan 30 is based on decision deferral time T. p They agree up to this point.
[0071] Decision postponement time T p This is calculated based on the predicted occlusion map 52 for each scenario 36. For example, in one scenario 36, the method may assume that a third traffic participant 14c is an obstacle 44 within the occlusion area 28, which means, for example, a vehicle traveling at 25 km / h. Thus, the method can predict the time when the third traffic participant 14c will emerge from the occlusion area 28. Once the third traffic participant 14c emerges from the occlusion area 28, the best-case scenario 50, in which it is assumed that there are no obstacles 44 within the occlusion area 28, can be excluded.
[0072] For each action plan 30, the predicted occlusion map 52 may further be based on the mapping of the position and / or orientation of agent 10 to the currently occluded area 28. In other words, the decision delay time T p When calculating this, the operation of agent 10 according to each operation plan 30 may be taken into consideration.
[0073] In step S12, a suitable operation plan 30 can be determined by minimizing a cost function 54 that includes the predicted occlusion maps 52 for each scenario 36. For example, the cost function 54 may be a sum over single-scenario cost functions, each of which is associated with a single scenario 36. The cost function 54 may include a term that maximizes the information gain regarding the currently occluded region 28.
[0074] In step S13, the method includes a reassessment of traffic conditions 12 at future time t1 > t0. In step S13, the method may update the occlusion map 52 predicted at time t0 based on sensor data 24 at future time t1 > t0. The updated predicted occlusion map 52 is then used for decision deferral time T p This can be used to adapt. In step S13, one or more scenarios 36 may be further excluded from the scenario tree 34 based on sensor data 24 at future time t1>t0.
[0075] More scenarios 36 may be eliminated from the scenario tree 34 until a single scenario 36 remains, by re-evaluating at a future time t1>t0, or by performing any further re-evaluations at times t2>t1, t3>t2, etc. The corresponding action plan 30 for this remaining scenario 36 is ultimately determined by the (adapted) decision deferral time T p It is selected to drive agent 10 in traffic conditions 12 beyond the specified limit.
[0076] Figure 3 shows a second embodiment of the computer implementation method.
[0077] In step S21, the second embodiment is described as follows: - Construct the current occupancy map 38 based on the current sensor data 24. Includes.
[0078] In step S21, the occupancy map 38 is preferably three-dimensional, but the present invention is not limited thereto. In the occupancy map 38, the sensor field of view 26 can be partitioned into a plurality of cells. Then, the occupancy map 38 can include information regarding whether each cell is occupied or not. In addition, the occupancy map 38 can include an input 56 regarding the nature of the occupancy, i.e., how the cell is occupied, for example, whether the cell is occupied by a building, a vehicle, a pedestrian, a road, and / or a sidewalk. The input 56 can be determined based on a machine learning algorithm such as a neural network that can be trained for annotation of the traffic situation scene, for example, classification and / or segmentation.
[0079] The occupancy map 38 is constructed based on the current sensor data 24 at time t0. Optionally, the occupancy map 38 can be further constructed based on sensor data 24 from one or more previous times t < t0.
[0080] In step S22, the second embodiment - determining the current occlusion map 32 based on the current occupancy map 38 is included.
[0081] In step S22, a sensor model 59 including the maximum sensor detection distance 42 can be constructed. For each cell of the occupancy map 38, three-dimensional rays having a length equal to the maximum sensor detection distance 42 can be aligned starting from the center of the sensor 22. Starting from the center of the sensor 22, each cell can be marked as visible until it reaches the first occupied cell or the end of the ray. After evaluating each sensor 22 as described above, the result of this step S22 can be a three-dimensional occlusion map 32 including occlusion information that any unoccupied cell not marked as visible is occluded. Alternatively, the occlusion map 32 can be two-dimensional.
[0082] In step S23, the second embodiment - Filter and / or sort the current occlusion map 32 by relevance. Includes.
[0083] In step S23, the current occlusion map 32 may be filtered and / or sorted by ranking one or more currently occluded areas 28 by relevance score. The relevance score may be based on a navigation map 58, which may be, for example, a standard-definition (SD) map. All occluded areas 28 may be filtered for relevance by projecting each occluded cell onto the SD map, and cells may be retained only if they are projected onto areas that may be relevant to the planning phase (e.g., roads, sidewalks, intersections). In addition, information such as input 56 regarding annotations (segmentation and / or classification) in the occupancy map 38 and / or occlusion map 32 may be used to detect scenarios 36 and occluded areas 28 with higher risk. The output of step S23 may be an evaluated occlusion map 60 containing one or more occluded areas 28 filtered and / or sorted by relevance.
[0084] In step S24, the second embodiment is described as follows: - One or more worst-case scenarios 48 were included in the scenario tree 34. Includes.
[0085] In step S24, the evaluated occlusion map 60 may be used to calculate the possible future positions of assumed or potential occluded obstacles 44 or objects 20. The worst-case scenario 48 may be included by calculating the set of forward reachable objects where the possible maximum velocity and / or acceleration of assumed or potential occluded obstacles 44 can be limited by the most likely object type (properties) extracted from the situation assessment (pedestrian or vehicle). This set may be added to the scenario tree 34 as a contingency scenario.
[0086] In step S25, the second embodiment includes - including one or more best scenarios 50 in the scenario tree 34 including.
[0087] In step S26, the second embodiment includes - calculating the decision delay time T p including calculating. including.
[0088] In step S26, to calculate the appropriate decision delay time T p the predicted occlusion map 52 of each scenario 36 can be used. The predicted occlusion map 52 can be determined by the output of the prediction module 62 from one or more previous times t < t0.
[0089] The second embodiment further includes steps S27, S28, and S29 performed by the prediction module 62 at time t0. The evaluated occlusion map 60, one or more worst scenarios 48, one or more best scenarios 50, and the decision delay time T p are supplied to the prediction module 62.
[0090] In step S27, the method can determine the occlusion state 64 by expanding the kinematic bicycle model using the evaluated occlusion map 60. In this step, a function that maps the current position and / or orientation of the agent 10 to the occlusion area 28 can be derived. Essentially, this can include formulating a geometric equation that can calculate the sensor field of view 26 and return the occluded space with respect to the position and / or orientation of the agent 10. For this purpose, the kinematic bicycle model can use the sensor model 59.
[0091] In step S28, using the occlusion state 64, the cost function 54 for information collection can be formulated. The prediction module 62 can plan, for example, in steps S24 and S25, using all the scenarios 36, 48, 50 added to the scenario tree 34. The prediction module 62 can minimize the cost function 54 for information collection by satisfying the decision delay time T p and the constraints on the common operation plan. This can be done by the prediction module 62 by allowing different strategies (i.e., control inputs) that are subject to the constraint of being the same until the first time step, i.e., until reaching the calculated decision delay time T p .
[0092] In step S29, the prediction module 62 iterates to predict the occlusion map 52 for each scenario 36. The output of step S29 can be an appropriate operation plan 30 for each scenario 36 including the current optimal trajectory 66 according to the common operation plan. Further, the predicted occlusion map 52 of each scenario 36 from the previous time t < t0 can be updated based on the current sensor data 24, and the decision delay time T p from the previous time t < t0 can be adapted based on the updated predicted occlusion map 52.
[0093] The current optimal trajectory 66 can be fed back to step S23 as the previously planned optimal trajectory for the agent 10 to filter the remaining occluded cells by a distance metric, i.e., occluded cells that are too far can be discarded. The adapted decision delay time T p can be fed back to step S26.
[0094] Then, the method can be repeated at a future time t1 > t0 until the (adapted) decision delay time T p , and a decision can be made over an appropriate operation plan 30 to drive the agent 10 beyond the (adapted) decision delay time T p , i.e., a single scenario 36 remains.
[0095] Embodiments of the present invention may further include generating control signals to drive agent 10 based on a selected action plan 30 (including a common action plan). Embodiments of the present invention further include means adapted to perform embodiments of the described method, or data processing devices or control devices (not shown) adapted to perform embodiments of the described method. Embodiments of the present invention may be implemented in a computer program (not shown) that, when the program is executed by the computer or control means, includes instructions for the computer to perform embodiments of the described method. The computer program may be stored in a computer-readable data carrier (not shown). Embodiments of the present invention further include agent 10, which includes data processing devices or control devices.
[0096] In summary, embodiments of the present invention relate to an action plan based on an occlusion-aware scenario that combines risk recognition and information gathering. [Explanation of Symbols]
[0097] 10 Autonomous Agents 12. Traffic conditions 14 Transportation participants 14a First traffic participant 14b Second traffic participant 14c Third traffic participant 16 Agent's direction of movement 18 Direction of travel for traffic participants 18a Direction of travel of the first traffic participant 18b Direction of travel of the second traffic participant 18c Direction of travel of the third traffic participant 20 static objects 22 sensors 24 Sensor data 26 Sensor field of view 28 Shield area 30 Action Plan 32 Occlusion Maps 34 Scenario Tree 36 Scenarios 38 Occupation Maps 40 Occupied area 42. Maximum sensor detection distance 44 Obstacles 46. Non-existent obstacles (pedestrians) 47 Direction of movement of non-existent obstacles 48 Worst-Case Scenario 50 Best Scenarios 52 Predictive Occlusion Map 54. Cost Function 56 inputs 58 Navigation Map 59 Sensor Models 60 rated occlusion maps 62 Prediction Modules 64 Occlusion State 66 Optimal Trajectory t0 Current time t1 Future time T p Decision postponement time
Claims
1. A computer implementation method for determining an action plan (30) for driving an autonomous agent (10) in traffic conditions (12), wherein the agent (10) includes a sensor (22) for capturing sensor data (24) within a sensor field of view (26), the sensor data (24) represents the traffic conditions (12), and the method is: a) Based on the current sensor data (24), the current occlusion map (32) is determined, and the current time t 0 The evaluation of the traffic conditions (12) in the present, wherein the current occlusion map (32) shows one or more occlusion areas (28) within the sensor field of view (26) for which relevant sensor data (24) is not currently available, b) Based on the current occlusion map (32), construct a scenario tree (34) including multiple scenarios (36) relating to the traffic conditions (12) during the planning period, and determine an appropriate action plan (30) for each scenario (36) to drive the agent (10), wherein the scenario tree (34) determines all appropriate action plans (30) with a delay time T p Based on the decision postponement that is constrained to a common action plan until the decision postponement time T p This is calculated based on the predicted occlusion map (52) for each scenario (36), and indicates the time at which all but one of the aforementioned scenarios (36) can be excluded, and is constructed and determined. c) Until a single scenario (36) remains, future time t 1 >t 0 The traffic conditions (12) are re-evaluated, and the decision postponement time T is determined. p Beyond the traffic conditions (12), select the corresponding action plan (30) for driving the agent (10) and A computer implementation method, including
2. Step b) is b1) Determining the appropriate operation plan (30) by minimizing based on a cost function (54), wherein the cost function (54) includes the predicted occlusion map (52) for each scenario (36), and the cost function (54) preferably includes a term for maximizing the information gain for the one or more currently occluded regions (28). The method according to claim 1, further comprising:
3. Step c) is c1) updating the predicted occlusion map (52) based on the sensor data (24) at the future time t 1 > t 0 and adapting the determination delay time T based on the updated predicted occlusion map (52), and / or p c2) The aforementioned future time t 1 >t 0 Based on the sensor data (24) in the above, one or more scenarios (36) are excluded from the scenario tree (34). The method according to claim 1 or 2, further comprising one or both of the above.
4. Step b) is b2) Constructing the scenario tree (34) by including one or more worst-case scenarios (48), each worst-case scenario (48) assuming one or more obstacles (44) in at least one of the currently shielded areas (28), and / or b3) Constructing the scenario tree (34) by including one or more best-case scenarios (50), wherein each best-case scenario (50) is constructed assuming that there are no obstacles (44) in all of the currently shielded areas (28). The method according to any one of claims 1 to 3, further comprising one or both of the above.
5. Step b) is b4) To calculate the reachable area of the assumed obstacle (44) in at least one of the currently shielded areas (28), estimate the maximum possible speed and / or acceleration of the assumed obstacle (44). The method according to any one of claims 1 to 4, further comprising:
6. Step b) is b5) For each action plan (30), predict the occlusion map (52) for each scenario (36) by mapping the position and / or orientation of the agent (10) to the one or more currently occluded regions (28). The method according to any one of claims 1 to 5, further comprising:
7. Step a) is a1) Based on the current sensor data (24), determine a current occupancy map (38) that shows one or more occupied regions (40) within the sensor field of view (26), a2) Determining the current occlusion map (32) based on the current occupancy map (38) The method according to any one of claims 1 to 6, further comprising:
8. Step a2) is, a2a) Determining the current occlusion map (32) by extending a straight line from the sensor (22) beyond the one or more occupied areas (40) to the maximum sensor detection distance (42) of the sensor (22). The method according to claim 7, further comprising:
9. Step a) is a3) Filtering and / or sorting the current occlusion map (32) by ranking the one or more currently occluded areas (28) according to a relevance score for evaluating the traffic conditions (12). The method according to any one of claims 1 to 8, further comprising:
10. Step a3) is, a3a) Determining the relevance score based on the navigation map (59), the distance from the agent (10) to the one or more currently occupying areas (28), the reachable area of one or more assumed obstacles (44), and / or the nature of one or more occupied areas (40) within the sensor field of view (26). The method according to claim 9, further comprising:
11. d) Generating control signals to drive the agent (10) based on the common operation plan and / or the selected operation plan (30). The method according to any one of claims 1 to 10, further comprising:
12. A data processing device comprising means for performing the method according to any one of claims 1 to 11.
13. A computer program, which, when executed by a computer, includes an instruction causing the computer to perform the method described in any one of claims 1 to 11.
14. A computer-readable data carrier having the computer program described in claim 13 stored thereon.
15. An autonomous agent (10) comprising a sensor (22) for capturing sensor data (24) within a sensor field of view (26), wherein the sensor data (24) indicates traffic conditions (12), and the agent (10) further comprises the data processing device according to claim 12.