Method and system for controlling an at least partially automated vehicle
Patent Information
- Application Number
- PCT/EP2026/054255
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-10
- Filing Date
- 2026-02-17
- Publication Date
- 2026-09-17
Smart Images

Figure EP2026054255_17092026_PF_FP_ABST
Abstract
Description
[0001] R. 417829
[0002] - 1 -
[0003] Description
[0004] title
[0005] Method and system for controlling at least a partially automated vehicle
[0006] The present invention relates to a method, a system and a computer program for controlling an at least partially automated vehicle, particularly with regard to its behavior in unusual driving situations.
[0007] State of the art
[0008] In road traffic, it frequently happens that a vehicle finds itself in an unfamiliar situation that was not anticipated beforehand, especially during the design and development of automated or driver assistance systems. Human drivers, based on their experience both on and off the road, often develop an intuition that leads them to drive with particular caution in such situations. Automated driving systems, however, lack this mechanism.
[0009] German patent application DE 102022214147 A1 discloses a method for vehicle behavior planning, which includes receiving data from an environmental perception system and a planned route of the vehicle. A traffic scenario is determined based on the environmental perception data and the planned route, and at least one geometric behavior option is determined based on the traffic scenario.
[0010] Disclosure of the invention
[0011] It can therefore be considered an object of the invention to provide a reliable and safe method or system for controlling at least a partially operational R. 417829
[0012] - 2 -
[0013] to specify a tomato-equipped vehicle. The invention addresses the aforementioned problem by describing a mechanism by which automated vehicles can react anticipatorily to unfamiliar situations.
[0014] According to a first aspect of the invention, a method for controlling an at least partially automated vehicle is provided. The method comprises receiving sensor data representing a current driving scene and classifying this driving scene into at least one predefined discrete situation class. This is done using a pre-trained vision-language model that acts as a classifier.
[0015] A vision-language model can be understood as an artificial neural network trained to process and relate both visual information (e.g., from camera images) and textual information (e.g., descriptions of scenes). Examples include CLIP (Contrastive Language-Image Pre-training) and ALIGN.
[0016] The vision-language model outputs classification probabilities for each predefined situation class. It then determines a residual probability indicating the likelihood that the current driving scenario does not belong to, or cannot be associated with, any of the predefined situation classes. This residual probability reflects the degree of unusualness of the current driving scenario and allows the system to act cautiously in unfamiliar situations, just as a human driver would. Based on this residual probability and the classification probabilities, the system adjusts the vehicle's driving strategy.
[0017] In a preferred embodiment, the classification result is further processed using a rule-based procedure. According to this procedure, static information such as map data and traffic regulations is received, and a plausibility check of the classification result is performed based on this information. This is also referred to as rule-based "output mapping." The adapted driving strategy is generated based on the classification result, the static information, and the plausibility check. This increases the safety and reliability of the system. (R. 417829)
[0018] - 3 -
[0019] This is because the decisions of the Vision-Language model can be validated by the rules and static information.
[0020] An advantageous embodiment of the invention provides for retraining the vision-language model with raw data from the sensors used. Preferably, raw data from an identical or functionally comparable sensor set can be used for this retraining. The weighting factors of the model preferably remain unchanged after retraining ("freezing the weights"), which supports the stability and safety of the system.
[0021] In a particularly preferred embodiment, the classification is performed based on a database of prompts. Based on these prompts, the vision language model generates regions in a so-called latent feature space. The latent feature space is, in particular, an abstract space in which the features of the sensor data are represented by the vision language model. Similar driving scenes are represented in this space by points located close to one another. Within these regions, a current driving scene is assigned to a predefined situation class. These regions can be approximated or modeled, for example, by a Gaussian mixture model or a Gaussian process.
[0022] In particular, the distance of a feature vector derived from sensor data to the regions in the latent feature space can be calculated. This distance can serve as a measure of the probability of belonging to a specific situation class.
[0023] To optimize computation time, the situation classes can be predefined. During operation, the feature vector is then simply compared with areas of the latent feature space generated in this way and stored, particularly in a database, that correspond to the situation classes. For example, a distance can be approximated as a comparison, especially using a Gaussian mixture model or a Gaussian process. R. 417829
[0024] - 4 -
[0025] In a preferred implementation, the situation classes are predefined. During operation of the at least partially automated vehicle, only a feature vector determined from the sensor data representing the current driving scene needs to be compared with areas of the latent feature space, stored in a database, that correspond to the situation classes. This saves computing power.
[0026] Estimating class membership and model uncertainty can be implemented using subjective logic, with the classifier outputting a multinomial subjective logic opinion. A particular advantage of using subjective logic is that, in addition to the class membership probability, it also provides a measure of reliability indicating the degree of trustworthiness of the probabilities. The subjective logic opinion can be further processed using classical, rule-based consistency and plausibility tests, such as those employing a parallel path for comparison, or calibrated using a validation dataset, thus ensuring the output is based on statistical evidence.
[0027] According to a second aspect of the invention, a system for carrying out the method according to the invention is provided. The system comprises a sensor unit configured to acquire sensor data representing a current driving scene. Furthermore, the system comprises a processing unit configured to classify a current driving scene using a vision-language model, wherein the classification includes assigning the current driving scene to at least one predefined discrete situation class. The processing unit is also configured to determine the degree of unusualness of the current driving scene. This is done based on a residual probability that the current driving scene does not belong to any predefined situation class.Furthermore, the processing unit is configured to initiate at least one risk mitigation strategy when the residual probability exceeds a threshold, wherein the at least one risk mitigation strategy includes reducing the vehicle speed and / or initiating a takeover request to a driver. R. 417829.
[0028] - 5 -
[0029] In a preferred embodiment, the system includes a perception unit. This unit creates an environmental model from the sensor data. The perception unit is a component of the system that processes the sensor data using a neural network and uses it to create a representation of the vehicle's surroundings. The environmental model can, for example, contain information about the position and speed of other road users, lane markings, and traffic signs. The neural network of the perception unit and the vision language model preferably use the same sensor data and the same latent feature space. This is also referred to as the "perception backbone."
[0030] The system according to the invention can, for example, be comprised of an at least partially automated vehicle, in particular a vehicle.
[0031] A third aspect involves proposing a computer program. This computer program comprises instructions that, when executed by a computer, cause it to perform a procedure according to the first aspect. The computer may be comprised of the processing unit of a system according to the second aspect.
[0032] A fourth aspect proposes a machine-readable storage medium on which the computer program is stored according to the third aspect.
[0033] The term "at least partially automated" includes one or more of the following: assisted driving, semi-automated driving, highly automated driving, fully automated driving, driverless control or driving of a vehicle.
[0034] Assisted driving means that the driver of the vehicle is permanently responsible for either the lateral or longitudinal control of the vehicle. The other driving task (i.e., controlling the longitudinal or lateral movement of the vehicle) is performed automatically. This means that with assisted driving, either the lateral or longitudinal control is automatic. R. 417829
[0035] - 6 -
[0036] Semi-automated driving means that in a specific situation (for example: driving on a highway, driving within a parking lot, overtaking an object, driving within a lane defined by lane markings) and / or for a certain period of time, the longitudinal and lateral control of the vehicle is automated. The driver does not need to manually control the vehicle's longitudinal and lateral steering. However, the driver must continuously monitor the automated control of the longitudinal and lateral steering in order to be able to intervene manually if necessary. The driver must be ready to take over full control of the vehicle at any time.
[0037] Highly automated driving means that for a certain period of time in a specific situation (for example: driving on a highway, driving within a parking lot, overtaking an object, driving within a lane defined by lane markings), the longitudinal and lateral control of the vehicle is automated. The driver does not need to manually control the vehicle's longitudinal and lateral steering. The driver does not need to constantly monitor the automated control of longitudinal and lateral steering in order to intervene manually if necessary. If required, a takeover request is automatically issued to the driver to assume control of longitudinal and lateral steering, with a sufficient time buffer. Therefore, the driver must be potentially capable of taking over control of longitudinal and lateral steering. Limits of the automated control of longitudinal and lateral steering are automatically detected.With highly automated control systems, it is not possible to automatically create a low-risk state in every initial situation.
[0038] Fully automated driving means that in a specific situation (for example: driving on a highway, driving within a parking lot, overtaking an object, driving within a lane defined by lane markings), the longitudinal and lateral control of the vehicle is automated. The driver does not need to manually control the vehicle's longitudinal and lateral steering. The driver does not need to monitor the automated control of longitudinal and lateral steering to intervene if necessary. (See also: 417829)
[0039] - 7 -
[0040] to be able to intervene automatically. Before the automatic control of lateral and longitudinal guidance is terminated, the driver is automatically prompted to take over the driving task (controlling the vehicle's lateral and longitudinal guidance), particularly with a sufficient time buffer. If the driver does not take over the driving task, the system automatically returns to a low-risk state. Limits of the automatic control of lateral and longitudinal guidance are automatically detected. In all situations, it is possible to automatically return to a low-risk system state.
[0041] Driverless control means that, regardless of the specific use case (for example: driving on a highway, driving within a parking lot, overtaking an object, driving within a lane defined by lane markings), the vehicle's longitudinal and lateral guidance are automatically controlled. The driver does not need to manually control the vehicle's longitudinal and lateral guidance. The driver does not need to monitor the automatic control of longitudinal and lateral guidance in order to intervene manually if necessary. Thus, the vehicle's longitudinal and lateral guidance are automatically controlled for all road types, speed ranges, and environmental conditions. The driver's entire driving task is therefore automatically taken over. The driver is no longer required. The vehicle can therefore travel from any starting position to any destination position without a driver.Potential problems are solved automatically, without the driver's assistance.
[0042] For the behavior planning of an at least partially automated vehicle, a method known from DE 102022214147 A1 can be used, for example. The driving strategy thus generated can then be adapted according to a method according to the present invention.
[0043] Brief description of the characters
[0044] With reference to the accompanying figures, embodiments of the invention are described in detail. R. 417829
[0045] - 8 -
[0046] Fig. 1 shows the data flow of a method according to a first embodiment of the invention.
[0047] Fig. 2 illustrates the use of a vision-language model in a second embodiment of a method according to the invention.
[0048] Fig. 3 shows an example of an automated vehicle in an urban traffic scene.
[0049] Fig. 4 shows a flowchart of an embodiment of a method according to the invention.
[0050] Fig. 5 schematically shows a storage medium containing a computer program product.
[0051] Preferred embodiments of the invention
[0052] In the following description of exemplary embodiments of the invention, identical elements are designated by the same reference numerals, and a repeated description of these elements may be omitted. The figures represent the subject matter of the invention only schematically.
[0053] Fig. 1 schematically shows the data flow of an exemplary method according to the invention for controlling an at least partially automated vehicle. The method uses a vision language model 103 as a classifier to analyze a current driving scene and adapt the driving strategy of the at least partially automated vehicle.
[0054] The input data for the procedure consists of sensor data 101, originating from various sensors in the vehicle and representing the current driving scene. This sensor data can include, for example, images from cameras, measurement data from lidar or radar systems, or other relevant environmental information. R. 417829
[0055] - 9 -
[0056] In training step 110 (optional, shown as a dashed line), the vision language model 103 is fine-tuned using a training dataset 102. This training dataset 102 contains raw data from recordings of driving scenes. The fine-tuning serves to adapt the pre-trained vision language model 103 to the specific characteristics of the sensor data and the relevant driving situations. The raw data used for training can, for example, originate from the same or an identical sensor set, or from a comparable sensor set, and be stored in a training database.
[0057] In live operation, the vision language model 103 receives the sensor data 101 and classifies the current driving scene into one or more predefined discrete situation classes. The result 120 of the classification, for example, the classification probabilities for each situation class, is forwarded to a rule-based output mapping 104.
[0058] The rule-based output mapping 104 processes the classification result and, if necessary, combines it with static information such as map data and traffic regulations. Additionally, the output mapping performs a plausibility check of the classification result. Based on this information, the output mapping generates a customized driving strategy 105. This driving strategy determines the vehicle's behavior, e.g., speed, steering angle, etc. The driving strategy can be implemented as a set of boundary conditions within which the vehicle control system must operate.
[0059] The adapted driving strategy 105 is then passed to the vehicle control system, which sends the corresponding control commands to the vehicle's actuators.
[0060] Fig. 2 illustrates in more detail the use of a vision-language model 203 in another possible embodiment of a method according to the invention. The vision-language model 203 serves as a classifier to analyze a current driving scene and derive a suitable driving strategy. R. 417829
[0061] - 10 -
[0062] Additionally, Fig. 2 shows a prompt database 208 for predictive driving 208. This database 208 contains prompts that describe the various predefined situation classes. The prompts serve to increase the explainability of the entire system or procedure and to facilitate the definition of the situation classes. The vision language model has the property that it assigns a feature vector in the latent feature space to each prompt and each set of sensor data (input mapping), which describes the individual situation classes. If the sensor data match the prompt, the respective feature vectors are located close together in the latent space.
[0063] Similar to Fig. 1, sensor data 201, representing the current driving scene, are received. This sensor data is transformed into a feature space by the input mapping for sensor data 202. The use of the same input mapping (shared input mapping) as in the Vision Language Model for processing the prompts ensures that the mapping to the latent feature space is consistent within the system, thus enabling a comparison of the feature vectors.
[0064] Since the prompt database 208 does not change at system runtime, the assigned feature vectors in the form of parametric feature distributions can be stored together with further information in a pattern database 204, for example to save computing power.
[0065] The pattern database 204, for example, contains tuples of parametric feature distributions, discrete actions, and explanations, such as in text or speech form. Each tuple can thus represent a learned association between a specific pattern in the sensor data, a corresponding driving action, and an explanation for that action. For example, a pattern representing a ball on the road could be associated with the action "Reduce speed" and the explanation "Increased probability of a child following the ball" (see Fig. 3).
[0066] In step 205, a probability is calculated that the current driving scene matches the stored patterns in the pattern database (Association Likelihood Evaluation). This calculation is based on the distance of the marker. 417829
[0067] - 11 -
[0068] The evaluation compares the probability vectors of the current driving scene to the feature distributions stored in the database. It yields a probability distribution across the different situation classes.
[0069] This Situation Probability Evaluation 205 represents the process of determining how well the current driving scene, in this example represented by its feature vector extracted via Shared Input Mapping, fits the predefined patterns in the pattern database 204. Essentially, it quantifies the probability that the current scene belongs to each of the predefined situation classes. The Situation Probability Evaluation receives the feature vector of the current driving scene from Shared Input Mapping 202. This vector represents the scene in the latent feature space. The core function of the Situation Probability Evaluation is to compare the input feature vector with the feature distributions stored in the pattern database.Each entry in the database represents a specific situation class and contains a description of the typical features observed in that situation, possibly represented as a probability distribution. For each situation class in the sample database, the situation probability score calculates a probability or similarity score. This score indicates how well the feature vector of the current scene matches the expected feature distribution for that class. The specific method for calculating the probability depends on how the feature distributions are represented in the database. It may include, for example, calculating distances in the feature space, comparing probability densities, or using other suitable measures of similarity. The output of the situation probability score is a set of probability values, one for each predefined situation class.This can be represented as a probability distribution over the situation classes, indicating the model's confidence in assigning the current scene to each class. Furthermore, this method determines a residual probability that the current driving scene does not belong to any of the predefined situation classes.
[0070] The probability values for each predefined situation class, as well as the residual probability, are output in 210. This output 210 is then used by the downstream components 206 (Output Mapping, Verhal-R. 417829).
[0071] - 12 -
[0072] The system uses probability generation to determine a suitable driving strategy for the at least partially automated vehicle. In simplified terms, the situation probability assessment (205) acts like a "pattern recognizer." It takes the current scene and compares it to a library of known patterns (the pattern database 204). It then outputs how "well" the scene matches each of these patterns and / or how "unusual" the scene is, thus providing a basis for deciding how the vehicle should react. The higher the probability for a particular situation class, the more certain the system is that the current scene represents that class. This confidence, or conversely the uncertainty (represented by a low probability for all known classes or a high residual probability), is then used to trigger appropriate risk mitigation strategies.
[0073] The result of the Association Likelihood Evaluation 205, i.e., the probability distribution across the situation classes, is forwarded to an Output Mapping 206. This Output Mapping 206 performs a plausibility check, as described in Fig. 1, and generates an adapted driving strategy (Behavior Generation). The driving strategy 207 represents the final decision regarding the vehicle's driving behavior, based on the classification of the driving scene and the tests performed.
[0074] Fig. 3 shows an exemplary and schematic urban traffic scene 300, representing a current driving scene of an automated vehicle 310. The vehicle 310 is moving on a road 320 at the edges of which various static objects 330, such as parked vehicles, buildings, or trees, restrict the field of view 322 of the sensor set 325 of the vehicle 310. Thus, a child 360 behind the parked vehicles on the left edge of the road may not be visible.
[0075] A ball 340 rolls onto the street 320. According to the invention, the Vision Language model 203 has implicitly learned a correlation between ball 340, the possible appearance of a child 360, and a potentially critical situation, such that the feature vector belonging to the situation can be assigned with a high probability to that discrete situation class which corresponds to the prompt "The provided-R. 417829
[0076] - 13 -
[0077] The situation could develop critically, making a reduction in speed necessary.” This information is passed on to the rule-based output mapping 206.
[0078] To perform a rule-based plausibility check, it is advantageous if the prompts in database 208 describe expected situations in more detail. In the example, a prompt such as: "A ball has rolled into the road; a child could run into the road from an obstruction at the roadside; the speed must be reduced" would allow for a more detailed plausibility check. The output mapping 206 detects an occupied cell or a small object 440 in the area in front of the vehicle in the occupancy grid map 400 of the environment model, which could be the ball 340. There are many obscured cells 430 in the occupancy grid map at the roadside where a child could be. The corresponding assignments can be made, for example, by a set of rules, which are primarily "can" relations: An object with a diameter of approximately 50 cm could be a ball.A child can play behind a concealed area, etc.
[0079] Since the classification is often only a possibility and not an established fact, it is a matter of plausibility. The classification thus appears plausible in this situation. Based solely on an object list and occupancy map, it would not be possible to anticipate the situation as critical and adjust the speed accordingly, as such images are not uncommon in the occupancy grid map 400 even for non-critical situations. Nevertheless, the assessment of the Vision Language Model appears plausible in the present situation. Therefore, a reaction corresponding to the prompt and thus the scene classification can be executed. For example, the speed of vehicle 310 is reduced to be able to stop in front of a suddenly appearing child 360 in the event of emergency braking.If a child actually jumps 360 after the ball 340 onto the road 320, the automated vehicle 310 can still stop because the speed was reduced in time.
[0080] Fig. 4 shows a flowchart 500 of a method for an at least partially automated vehicle according to an embodiment of the invention. R. 417829
[0081] - 14 -
[0082] In a first step, sensor data representing a current driving scene is received (510). In a second step, (520), the
[0083] The current driving scene is classified into at least one predefined discrete situation class. For this purpose, a pre-trained vision-language model acting as a classifier is used, with the vision-language model outputting classification probabilities for each predefined situation class. In a third step (530), based on the classification probabilities, a residual probability is calculated that the current driving scene does not belong to any predefined situation class. This residual probability reflects, in particular, the degree of unusualness of the current driving scene with respect to the predefined situation classes. In a fourth step (540), a driving strategy of the at least partially automated vehicle is adapted based on the probability of the situation classes or the residual probability that the current driving scene does not belong to any predefined situation class.
[0084] Fig. 5 schematically shows a storage medium 501 with a computer program product 503. The computer program product 503 comprises instructions which, when the computer program 503 is executed by a computer, cause it to execute a method according to Fig. 4.
Claims
R. 417829 - 15 - Claims 1. Method for controlling at least a partially automated vehicle (310), the method comprising: - Receiving sensor data (101, 201) representing a current driving scene; - Classifying the current driving scene into at least one predefined discrete situation class using a pretrained vision-language model (103, 203) acting as a classifier, wherein the vision-language model (103, 203) outputs classification probabilities for each predefined situation class; - Determining a residual probability that the current driving scene does not belong to any predefined situation class, based on the classification probabilities, wherein the residual probability in particular reflects a degree of unusualness of the current driving scene with respect to the predefined situation classes; and - Adapting a driving strategy (207) of the at least partially automated vehicle (310) based on the classification probabilities and the residual probability that the current driving scene cannot be associated with any predefined situation class 2. The method of claim 1, wherein the method comprises: - Processing the classification result using a rule-based procedure (104, 206), wherein the rule-based procedure (104, 206) includes: Receiving static information, including in particular map data and / or traffic regulations; and / or performing a plausibility check of the classification result; and R. 417829 - 16 - - Generating the adapted driving strategy (207) based on the classification result, the static information and the plausibility check.
3. Method according to one of claims 1 or 2, wherein the vision language model is retrained (110) with raw data (201), in particular raw data from the sensors from which the sensor data (101, 201) are received, or from a sensor set of identical construction or a functionally comparable type.
4. Method according to claim 3, wherein weighting factors of the vision-language model (103, 203) remain unchanged after retraining.
5. Method according to one of the preceding claims, wherein the classification is performed based on a database of prompts (208), wherein the Vision Language Model (203) generates areas in a latent feature space of the sensor data based on the prompts, within which a current driving scene is assigned to a predefined situation class.
6. Method according to claim 5, wherein a distance of a latent feature vector obtained from the sensor data is determined with the regions in the latent feature space, wherein a measure of the probability of belonging to a certain predefined situation class is obtained from the distance, and wherein the distance is approximated in particular by means of a Gaussian Mixture Model or a Gaussian process.
7. A method according to one of the preceding claims, wherein the situation classes are predetermined and, during the operation of the at least partially automated vehicle (310), only a feature vector determined from the sensor data (101, 201) representing the current driving scene is compared with areas of the latent feature space, which correspond to the situation classes and are stored, in particular in a pattern database (204). R. 417829 - 17 - 8. Method according to one of the preceding claims, wherein an estimation of class membership and model uncertainty is implemented based on Subjective Logic, wherein the classifier output is a multinomial Subjective Logic Opinion.
9. System for controlling an at least partially automated vehicle (310) according to a method according to any one of claims 1 to 8, wherein the system comprises: - a sensor unit (325) configured to acquire sensor data (101, 201) representing a current driving scene (300); - a processing unit configured to: o Classifying the current driving scene using a vision-language model (103, 203), wherein the classification includes assigning the current driving scene (300) to at least one predefined discrete situation class; o Determining a degree of unusualness of the current driving scene (300) based on a residual probability that the current driving scene does not belong to any predefined situation class; and o Initiating at least one risk reduction strategy if the residual probability exceeds a threshold, wherein the at least one risk reduction strategy includes reducing the vehicle speed and / or initiating a takeover request to a driver.
10. System according to claim 9, wherein the system comprises a perception unit configured to create an environment model (400) from the sensor data (101, 201), wherein the control of the at least partially automated vehicle is additionally based on the environment model (400), and wherein the environment model is created by means of a neural network, and the neural network and the vision-language model are defined in R. 417829 - 18 - use the same sensor data and use the same latent feature space of the sensor data.
11. Computer program (503) comprising instructions which, when the computer program (503) is executed by a computer, cause it to execute a method according to any one of claims 1 to 8.
12. Machine-readable storage medium (501) on which the computer program (503) according to claim 11 is stored.