Method and device for predicting the movement of objects in the surroundings of an at least semi-autonomously driving vehicle
By using environmental data fused with map data and transmitting it to a Large Language Model for enhanced interpretation, the method addresses the challenge of predicting object movement in complex traffic scenarios, achieving improved accuracy and safety for autonomous vehicles.
Patent Information
- Application Number
- PCT/EP2024/076421
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-23
- Filing Date
- 2024-09-20
- Publication Date
- 2025-05-30
AI Technical Summary
Existing methods for predicting the movement of objects in the environment of an at least partially automated vehicle struggle to accurately interpret complex traffic scenarios, especially when interactions between multiple road users are involved.
The method involves identifying a scene from images of the vehicle's environment, predicting the behavior of objects within that scene, and using environmental data from sensors fused with map data. This data is then transmitted to a Large Language Model (LLM) for enhanced image interpretation and behavior prediction, allowing for more accurate planning of vehicle trajectories.
The solution enables a reliable analysis of complex traffic situations by improving the accuracy of object behavior prediction, thereby enhancing the safety and efficiency of autonomous vehicle operations.
Smart Images

Figure EP2024076421_30052025_PF_FP_ABST
Abstract
Description
[0001] Method and device for predicting the movement of objects in the environment of an at least partially automated vehicle
[0002] The invention relates to a method for predicting the movement of objects in the environment of an at least partially automated vehicle, in which, in order to plan the behavior of the objects, a scene is identified from at least one image of the environment of the vehicle and a behavior of the objects occurring in the scene is predicted, as well as to a device for carrying out the method.
[0003] For autonomous vehicles, it is particularly important to interpret diverse traffic scenarios, as interactions between multiple road users can be highly complex. In addition to the various specifics, country-specific conditions must be considered.
[0004] DE 102021 202 993 A1 discloses a method in which a neural network is assigned to an object to be tracked. Based on the images fed to the neural network, it provides an output containing the positions of the assigned object and / or information about the behavior or a prediction of the behavior of the assigned object. Images from a sequence and / or sections of these images are fed to the neural network for processing.
[0005] The object of the invention is to provide a method and a device for predicting the movement of objects in the environment of an at least partially automated vehicle, which allow a reliable analysis of highly complex traffic situations.
[0006] The invention is based on the features of the independent claims. Advantageous developments and refinements are the subject of the dependent claims. Further features, possible applications, and advantages of the invention will become apparent from the following description and the explanation of exemplary embodiments of the invention illustrated in the figures.
[0007] The problem is solved by the subject matter of patent claim 1 or claim 9.
[0008] In the method explained above for predicting the movement of objects in the environment of an at least partially automated vehicle, in which a scene is identified from at least one image of the vehicle's environment and the behavior of the objects appearing in the scene is predicted to plan the behavior of the objects, environmental data from sensors is determined, preferably as an image, and then fused. The fused environmental data is transmitted together with map data to a scene interpretation module, which, if necessary, generates queries about the scene recognized by the fused environmental data and transmits them to a Large Language Model (LLM).The scene interpretation module refines the scene interpretation based on a response generated by the LLM. The determined data and information characterizing the scene are fed into a behavior planning module for planning a vehicle trajectory and controlling vehicle actuators depending on the planned vehicle trajectory. Using the LLM improves the level of image interpretation using machine learning, as an LLM is trained with cross-domain knowledge and thus enables far more detailed interpretations than a neural network trained solely with data from the autonomous driving domain. For behavior prediction, the LLM is capable of predicting complex intentions of road users. This allows for a more accurate prediction of the movement of individual objects.The environmental data can be determined, for example, from a camera image, from map data, from the scene, or from a grid. No domain-specific training data or annotations are necessary, as large data sets for LLMs with cross-domain training data already exist and can be adopted. Preferably, the environmental data determined by the sensors, the fused environmental data, and / or map data of the environment are made available to the LLM before, i.e., parallel to, the transmission to the scene interpretation module, or alternatively, with the query from the scene interpretation module. The sensors for collecting the environmental data can be designed as cameras for capturing images, as lidar, radar, and / or as ultrasonic sensors.Advantageously, queries from the scene interpretation module to the LLM are sent as soon as the scene interpretation module identifies objects of a certain criticality. For example, when pedestrians are detected at a defined distance from the vehicle's planned trajectory. By enriching the object and scene information with the LLM, the behavior of road users can be determined more precisely, and dangerous situations can be avoided.
[0009] In one embodiment, the queries are sent from the scene interpretation module to the LLM as soon as the scene interpretation module determines a confidence value below a threshold value for a predicted behavior of an object.
[0010] In other words, queries are sent to the LLM as soon as situations arise in which the scene interpretation model alone cannot make a prediction for objects in the environment with sufficient certainty.
[0011] In one variant, the scene interpretation module can query the LLM to annotate training data, i.e., categorize and / or label data sets for processing by a machine or neural network. This training data can then be used to train domain-specific neural networks, such as the scene interpretation module itself.
[0012] In one embodiment, the scene interpretation module marks one or more objects in the image of the environment to which the query directed to the LLM refers. Adjustments and queries are performed without changes to the LLM.
[0013] To obtain the LLM's response in a specific, machine-processable form that is easily interpreted by the scene interpretation module, the LLM is first preconditioned. This can precede the actual query or be integrated into the query itself. For example, an output format can be defined textually, predefining placeholders (e.g., for probabilities of certain behaviors) that are then populated in the LLM's response.
[0014] In a further embodiment, the same query is sent multiple times in different forms from the scene interpretation module to the LLM, whereby the LLM's response is only considered for processing in the scene interpretation module if the responses match in content. This can prevent so-called "hallucinations" of the LLM or refine the confidence by averaging the confidences across multiple queries. The multiple queries can also include different sensor measurements that are transmitted to the LLM. The multiple queries can also be supplemented with map data.
[0015] In another variant, the LLM outputs a probability of occurrence along with the response issued to a query. Starting at a given probability value, behavior planning is then based on the LLM's response, and / or the response issued by the LLM is at least partially output directly to a user. The output to the user can also include a description field populated by the LLM, in which the LLM provides a justification for choosing this probability. This increases the interpretability of the decision-making process for both the developer and the customer.
[0016] A further aspect of the invention relates to a device for predicting the movement of objects in the environment of an at least partially automated vehicle, comprising a behavior planning unit. In a device that allows for reliable analysis of complex traffic situations, the behavior planning unit is configured to execute the method according to at least one feature described in this patent application. The device offers the possibility of providing complex scene interpretations with context from the environmental perception for behavior planning.
[0017] Further advantages, features, and details will become apparent from the following description, in which at least one exemplary embodiment is described in detail—possibly with reference to the drawings. Described and / or illustrated features may form the subject matter of the invention alone or in any meaningful combination, possibly independently of the claims, and may, in particular, also be the subject of one or more separate applications. Identical, similar, and / or functionally equivalent parts are provided with the same reference numerals.
[0018] They show:
[0019] Fig. 1 shows an embodiment of a vehicle with a device according to the invention and Fig. 2 shows an embodiment of an image of a scene to be interpreted.
[0020] Fig. 1 shows an embodiment of a vehicle with a device according to the invention, wherein the vehicle is designed to drive autonomously. In order to create a driving trajectory for this autonomously driving vehicle 1, the vehicle 1 comprises a behavior planning unit 3, which interprets the intentions of the road users moving in the vicinity of the vehicle 1. The behavior planning unit 3 has a perception unit 5, which maps a scene context of the traffic situation in the vicinity of the vehicle 1 using a camera image, a map, or a scene grid and provides it to a Large Language Model 7 (LLM) for processing. The LLM 7 is connected to a map unit 9, in which highly accurate map data is stored.
[0021] In parallel with the connection to the LLM 7, the perception unit 3 is connected to a sensor fusion unit 11, in which the environmental data recorded by the perception unit 3 are fused and transmitted to a scene interpretation module 13, which is also coupled to the map unit 9. The scene interpretation module 13 determines the intentions and possible trajectories of the other road users 15 in a scene such as that shown in Fig. 2 and forwards them to a behavior planning module 17, which evaluates them and determines a travel trajectory of the vehicle 1. This travel trajectory is output to the actuator system 19 of the vehicle 1 in order to move the vehicle 1 safely in the scene recorded in image 21 according to Fig. 2.
[0022] For scene interpretation, the scene interpretation module 13 uses classic algorithms, vehicle models, or even neural networks in the form of concrete trajectories or even so-called intentions, such as lane following, changing lanes left / right, parked, stopped, pedestrian crossing the street, etc. If the behavior of road users 15 does not match the intentions modeled in the scene interpretation module 13, the estimated probability of the intentions is too low, or the scene is poorly assessed, the scene interpretation module 13 queries the LLM 7.
[0023] With the help of the enormous cross-domain knowledge from the training data of the LLM 7 in combination with the information from the perception data, which is no longer present in the scene in the scene interpretation module 13, the LLM 7 can answer the query of the scene interpretation module 13. The interaction of the scene interpretation module 13 and the LLM 7 will be explained in more detail with the help of Fig. 2. In Fig. 2, approximately 20 pedestrians 25, road users 15, are standing very close to the driving corridor 23 of the autonomously driving vehicle 1. With the information from the fused environment model, which is available to the scene interpretation module 13, it cannot be decided with certainty whether the pedestrians 15 want to cross the street or not.
[0024] The scene interpretation module 13 therefore issues a query to the LLM 7. The LLM 7 is conditioned to provide a response in a predefined, interpretable format, such as a table.
[0025] Furthermore, in the image 21, which was recorded by one or more cameras of the perception unit 3, one or more pedestrians 25 are visually marked, for example by a frame 27, to which the textual part of the query of the scene interpretation module 13 refers. For visual marking, the detections of the components of the perception unit 3 of the respective objects 25 can be adopted, e.g. all bounding boxes of all objects 15, 25 to be queried, which are designed as a rectangular coordinate system, are combined to form a rectangle, as in image 21, which is illustrated by the frame 27. However, queries are also conceivable which refer to the entire scene, for which no marking is necessary.
[0026] The query from the scene interpretation module 13 to the LLM 7 may, for example, include the following text:
[0027] "In the image, a group of pedestrians is marked with a box. Fill in the following CSV table with probabilities between 0.0 (not at all likely) and 1.0 (very likely) for each option, with a brief explanation of the estimated probability.
[0028] The table includes the columns Option, Probability, Explanation, where the options: 1. one or more pedestrians in the group will cross the street or 2. the entire group will stand at the pedestrian crossing - are provided, for which a probability and an explanation must be added."
[0029] The query result is used by the scene interpretation module 13 to adjust the probabilities for certain pedestrian intentions 25. If the probability for the option "the group of pedestrians will be standing at the pedestrian crossing" is very high, the probabilities for the pedestrian intentions 25 "crossing the street" can be adjusted downward. The explanation provided by the LLM 7 in the response, such as "the group of pedestrians is waiting in line to enter a restaurant," can be helpful for the analysis of the scene interpretation module 13 or even displayed to the end user in the vehicle 1.
[0030] If the scene interpretation module 13 is certain that the pedestrians 25 will not cross the street, a collision-free vehicle trajectory alongside the group of pedestrians 25 is planned in the behavior planning module 17, which is forwarded to the vehicle actuator system 19 and executed by it. If there is a low probability for the option "one or more of the pedestrians in the group will cross the street," a safe, slow crossing of the street by the pedestrians 25 can be planned.
[0031] In addition to being represented as image 21 from the camera, the scene context can also be transferred to the LLM 7 in other forms. This can be done using various perception data, such as occupancy cells, map data, and / or object data in raster or vector format.
[0032] To prevent further processing of "hallucinations" by the LLM 7, the queries to the LLM 7 are repeated, with the scene interpretation module 13 executing the queries with different seeds (pseudo-random numbers), different perspectives of the scene, such as different camera images, or adapted textual formulations. If the LLM 7's responses to the same queries from the scene interpretation module 13 differ, the scene adaptation may be reduced or suspended entirely.
[0033] The queries can be executed generally or in the case of a situation identified by the scene interpretation module 13 as critical or unclear, e.g., if it is unclear whether a pedestrian 25 or a cyclist is near the lane 23 or a vehicle 15 is parked in a no-stopping zone.
[0034] Depending on the availability of resources in a cloud and in Vehicle 1, the time budget, the type of query, and the quality of the connection to the cloud, the LLM 7 can be inferred directly on Vehicle 1 or the cloud. Statistical queries, such as location-dependent queries, can optionally be cached to avoid repeating previously executed queries. The LLM 7 can also be used to annotate training data. These annotations can be used, for example, to train more performant domain-specific neural networks, such as the scene interpretation module 13.
Claims
Patent claims 1. A method for predicting the movement of objects in the environment of an at least partially automated vehicle, in which a scene is identified from data (21) of the environment of the vehicle (1) determined by at least one sensor and a behavior of the objects (15, 25) occurring in the scene is predicted, characterized in that the environmental data from sensors (5) are determined and fused, wherein the fused environmental data are transmitted together with map data to a scene interpretation module (13), which generates queries based on the scene recognized by the fused environmental data and transmits them to a large language model (7) and carries out further interpretations of the scene described by the fused environmental data on the basis of a response generated by the large language model (7),wherein the determined data and information characterizing the scene are fed to a behavior planning module (17) for planning a vehicle trajectory and for controlling a vehicle actuator (19) depending on the planned vehicle trajectory.
2. Method according to claim 1, characterized in that the queries are sent from the scene interpretation module (13) to the large language module (7) as soon as the scene interpretation module (13) detects further objects (15) in or at a defined distance from the planned vehicle trajectory of the vehicle (1).
3. Method according to claim 1 or 2, characterized in that the queries from the scene interpretation module (13) to the large language model (7) are made as soon as the scene interpretation module (13) determines a confidence value below a threshold value for a predicted behavior of an object (15, 25).
4. The method according to claim 1, 2 or 3, characterized in that the scene interpretation module (13) issues queries to the large language model (7) for annotating training data.
5. Method according to at least one of the preceding claims, characterized in that the scene interpretation module (13) marks one or more objects (25) in the image (21) of the environment to which the query directed to the large language model (7) refers.
6. The method according to claim 5, characterized in that the requests are transmitted in a machine-interpretable format, the response being generated by the Large Language Model (7) by filling the placeholders.
7. Method according to at least one of the preceding claims, characterized in that the same query is sent multiple times from the scene interpretation module (13) to the large language model (7), the answers of the large language model (7) being taken into account for processing in the scene interpretation module (13) only if the answers match.
8. Method according to at least one of the preceding claims, characterized in that the large language model (7) outputs a probability of occurrence with the response output to a query.
9. Device for predicting the movement of objects in the environment of an at least partially automated vehicle, comprising a Behavior planning unit (3), characterized in that the behavior planning unit (3) is designed to carry out the method according to at least one of the preceding claims.
Citation Information
Patent Citations
Method for determining the calibration quality of an optical sensor
DE102021202993A1