Procedures for driving a motor vehicle, driver assistance system and motor vehicle
A physically informed MLLM trained on the Kamm friction circle concept addresses the lack of physical understanding in MLLMs, enabling safe and reliable driving interventions by integrating vehicle dynamics constraints, thus improving autonomous driving systems.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-01-22
- Publication Date
- 2026-03-26
AI Technical Summary
Existing multimodal large language models (MLLMs) in autonomous driving lack a consistent internal representation of physical principles such as inertial forces, friction coefficient limits, and tire force limits, leading to potential faulty interventions by driver assistance systems.
A physically informed MLLM is trained to learn and utilize the Kamm friction circle concept, integrating multimodal time-series data to generate physically valid driving maneuvers, using a physically based module to analyze vehicle dynamics and ensure compliance with physical constraints.
The MLLM can generate physically valid driving actions, ensuring safe and reliable interventions by the driver assistance system, enhancing generalizability and maintaining transparency and trust in autonomous or assisted driving scenarios.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method by which a motor vehicle, in particular autonomously, can be driven, as well as a driver assistance system and a motor vehicle by which the method can be carried out.
[0002] EP 4 145 242 A1 describes a method for autonomous driving of a motor vehicle in which virtual physical forces are assigned to objects detected in a perception field using a trained neural network, the interaction of which with other determined virtual physical forces is investigated in a physical model in order to derive driving actions for the motor vehicle.
[0003] There is a constant need to avoid faulty interventions by a driver assistance system, especially in autonomous driving of a motor vehicle.
[0004] The purpose of the invention is to demonstrate measures that enable the correct intervention of a driver assistance system.
[0005] The problem is solved according to the invention by a method with the features of claim 1, a driver assistance system with the features of claim 9, and a motor vehicle with the features of claim 10. Preferred embodiments of the invention are specified in the dependent claims and the following description, each of which can individually or in combination represent an aspect of the invention, the scope of protection being determined by the claims.
[0006] One aspect of the invention relates to a method for, in particular, autonomous and / or assisted driving of a motor vehicle, in which sensor data, for example from vehicle dynamics control and / or environmental perception, are supplied to a driver assistance system, a possible driving intervention is determined on the basis of the sensor data, wherein several measurement series of vehicle data are taken into account for the determination of the driving intervention with the aid of an MLLM ("multimodal large language model"), wherein for training of the MLLM such measurement series of vehicle data were linked with a physical model for driving stability, the MLLM plausibly verifies the driving intervention determined on the basis of the considered measurement series of vehicle data with the aid of the physical model used in the training and, in the event that the plausibility verification with the aid of the physical model does not determine sufficient driving stability, the driving intervention is corrected untiluntil the plausibility check using the physical model confirms sufficient driving stability.
[0007] This creates a method and system for controlling a vehicle based on a multimodal large language model (MLLM). The MLLM can be trained to learn and utilize its own internal, dynamic representation of the physical driving limits, based in particular on the concept of the Kamm friction circle. This can be achieved through specialized, physics-informed training, in which the MLLM is provided with training input not only from classical sensor data but also from explicit visual representations of the friction circle, such as plots, or an implicit representation through learning a latent representation of the Kamm friction circle, as well as an adapted loss function that penalizes physical inconsistencies. This allows the MLLM to generate physically valid driving maneuvers end-to-end without relying on external controllers or validation modules.Learning the physical boundary conditions that must be strictly adhered to during the use of MLLM enables correct intervention by a driver assistance system.
[0008] It has been recognized that Multimodal Large Language Models (MLLMs) have established themselves as powerful tools for a wide range of tasks in the field of autonomous driving. Their strength lies in the processing and integration of multimodal data sources—such as text, images, and sensor data. While MLLMs can increasingly and reliably represent semantic and anticipatory aspects of driving decisions, they exhibit significant shortcomings in their understanding of physical principles. In particular, they lack consistent internal representations of key vehicle dynamics constraints such as inertial forces, friction coefficient limits, or the limits of motion defined by tire forces. It is possible to address these aspects using external controllers or traditional planning modules, which limits end-to-end capability and generalizability. Consequently, the physical validity of actions is not an integral part of the model inference of LLMs / VLMs / MLLMs.
[0009] According to the invention, a physically informed MLLM is provided that can derive a dynamic, context-dependent representation, in particular of the Kamm friction circle, of the current driving situation from multimodal time-series data (e.g., camera, IMU, CAN bus). This representation can be explicitly presented as a plot or implicitly as a learned latent representation. This representation is intended to describe the current driving dynamics limits and can be directly integrated into decision-making or planning processes. This creates the possibility of generating physically valid actions within the model and using them directly for implementing a driving intervention (e.g., trajectory generation) – without subsequent mechanistic validation – or integrating them as a high-level command / action recommendation for subordinate control systems.The model's generalizability can be used to react appropriately to unprecedented events. This provides a concept that can be used for a driver assistance system to either issue warnings (indirect driving intervention) or actively intervene in the driving task (active driving intervention). This applies whether a human driver is behind the wheel or an automated driving system is controlling the vehicle.
[0010] A physically based module within the MLLM (or as a preceding rule-based module) can be used to analyze time-series data from vehicle communication buses (e.g., lateral and longitudinal acceleration, steering angle rate, friction coefficient estimation, speed, rain intensity, road gradient, ESP, vehicle load, weight distribution in the vehicle, ambient temperature, etc.). This data can be used to reconstruct a normalized friction circle, based on the Kamm friction circle method, which describes the current physical operating range. Environmental factors such as wetness, slipperiness, traffic density, etc., can be context-sensitively considered through supplementary sensor modalities (e.g., camera data). This allows, for example, the generation of four different representations of the driving state.In a standard representation, the current acceleration time series data (and other relevant time series) from the vehicle communication bus can be displayed as a series of numerical values. A first visual representation can show a plot of the current acceleration vector within the friction circle, which can be generated directly from the current time series data using rule-based algorithms. A second visual representation can show a plot of the historical acceleration profile within the friction circle, which can also be generated directly from current and past time series data using rule-based algorithms. This representation can be transferred to a latent space in the same way through training. A further latent vector space representation can include an embedded space that encodes the permissible maneuvers in terms of vehicle dynamics, thus enabling the model to be restricted to physically plausible actions.The latent space can be generated by encoder networks (e.g., variational autoencoders or encoder-decoder architectures) that transform the time-series data and / or plots into a continuous, low-dimensional representation. This space enables the model to detect similarities between different driving states, thus generating a physically interpretable structure. For example, points near the edge of the latent space can represent scenarios close to the friction circle limit. Targeted prompt-based validation allows for checking the physical consistency of the model. For instance, the maximum permissible lateral acceleration for a given coefficient of friction can be queried, or the feasibility of a force vector or trajectory can be verified. This gives the model a physical understanding of the vehicle dynamics, enabling it to derive realistic and safe maneuvers.In particular, the current state (plot / latent representation) and the physically limiting latent structure can be used as two main representations.
[0011] To develop a physical understanding, a ground-truth-based dataset can be used, providing a reference projection of the friction circuit for each driving situation. This projection can be combined with the actually measured force utilization (e.g., percentage utilization of the friction circuit). The data can be processed multimodally by using visual plots of the current and historical states within the friction circuit, time series data from the vehicle communication buses, formula representations to explicitly explain the relationships, and / or verbal descriptions of the situations. This allows for flexible use in various learning settings (e.g., supervised learning, reinforcement learning). Preferably, however, the time series data and plots are contextualized. For this purpose, the model can be conditioned via text prompts to understand what the plot represents (current state, previous utilization, etc.).This approach lays the foundation for a comprehensive physical understanding without limiting generalization capabilities. Specifically, physically consistent training only occurs in a subsequent step, where the target function combines language losses (e.g., cross-entropy) with physical consistency losses (e.g., cosine similarity between the real and predicted force vectors), constraint penalties (e.g., exceeding the friction circle limit), and unit-aware error functions to ensure dimensional correctness. This allows the model to learn to derive physically correct relationships instead of adhering solely to statistical language patterns. Generalization capabilities are thus preserved.
[0012] To further strengthen the physical understanding, image, text, formula data, and numerical scenarios can be explicitly aligned. This is achieved primarily through contrastive learning or distance-based matching. Scenarios near the friction circle limit, in particular, provide valuable learning signals for model calibration. An example of such alignment could involve tasking the model with determining the resulting frictional force based on a given combination of lateral and longitudinal acceleration (e.g., Ax = 0.5g, Ay = 0.7g). This multimodal integration allows the model to make physically sound inferences. The knowledge representations learned in this way are then transferred to new driving conditions. This enables the model to answer questions such as "Is this steering command feasible with a coefficient of friction µ = 0.4?" or "Can safe braking still be performed in this condition?" in a physically consistent manner.The model can transfer physical knowledge to new situations, instead of simply reproducing memorized examples.
[0013] Furthermore, it is possible to integrate explicit vehicle models (e.g., a single-track model) as constraints or differentiable layers, to use real or simulated driving scenarios (e.g., CARLA) to generate physically relevant learning examples, and / or to provide voice-based communication of physical states to users, e.g., "Your driving behavior is currently exceeding the lateral force limit by 4% if you brake now." Physical knowledge is not only used internally within the model but also communicated externally, thereby promoting transparency and trust. This can support situations in which a person is driving the vehicle and increase trust in automated driving functions—instead of relying on a computer program to control the vehicle.
[0014] Resource-efficient fine-tuning can be based on LoRA or QLoRA. Here, only selected weight matrices are supplemented with low-rank adapters. This enables targeted physical conditioning of the model, particularly with regard to friction circuit relationships between longitudinal and lateral accelerations, as well as the coefficient of friction dependence. LoRA ("Low-Rank Adaptation") is a method for efficiently fine-tuning large models, in which not all model weights are adjusted, but only small low-rank matrices are inserted into specific weight matrices. This makes training significantly more memory- and computationally efficient, as the majority of the model remains frozen. QLoRA ("Quantized Low-Rank Adaptation") extends LoRA by adding further quantization of the model weights (e.g., to 4 bits). This allows the base model itself to remain in a quantized state, which further reduces memory requirements.The low-rank adaptations remain at a higher precision (usually 16-bit), so that the model can still be fine-tuned despite quantization.
[0015] In the context of input signals for Large Language Models (LLMs), contextualization refers to the process of enriching raw input data—such as sensor data, images, point clouds, or speech—with additional meaning and situational context before it is passed to the model. Contextualizing input signals describes the transformation and enrichment of raw data into a representation that makes the situational context and meaning of the data understandable to the LLM. LLMs process speech or structured text. However, raw sensor data does not possess direct semantic content for a language model. Therefore, it is advantageous to transfer this data into a linguistically or semantically interpreted context. Only then can the LLM correctly interpret the data, draw conclusions, or derive recommendations for action.
[0016] In particular, the physical model takes into account longitudinal and lateral forces acting on a wheel of the motor vehicle, and insufficient driving stability is determined if a skidding of the motor vehicle is likely.
[0017] Preferably, the physical model uses a vehicle dynamics model based on the Kamm friction circle and / or the Krempel friction ellipse.
[0018] Particularly preferred is the training and prompting of the MLLM with the physical model in the form of a graphical representation of parameters relevant to driving stability, in particular the longitudinal forces and lateral forces acting on the wheel of the motor vehicle.
[0019] In particular, a time series of graphical representations of the physical model generated from measurement data of the motor vehicle is taken into account.
[0020] Preferably, if sufficient driving stability for the driving intervention has been determined, the driving intervention is corrected until it lies within a predefined tolerance window to the historical data of the time series of the physical model, provided that sufficient driving stability is possible within the predefined tolerance window.
[0021] A loss function was particularly preferred for training the MLLM, in which a detected driving intervention that violates the physical model is punished disproportionately.
[0022] In particular, the vehicle data measurement series include time series data on lateral acceleration, longitudinal acceleration, steering angle, steering angle rate, vehicle speed, estimated coefficient of friction of a road surface, weather data influencing the estimated coefficient of friction, road cleaning data influencing the estimated coefficient of friction and / or measurement data from an ESP system.
[0023] Another aspect concerns a driver assistance system for autonomous and / or assisted driving of a motor vehicle, whereby the driver assistance system is configured to carry out the procedure, which can be trained and further developed as described above. By learning the mandatory physical boundary conditions of the MLLM used, correct intervention by a driver assistance system is enabled.
[0024] Another aspect concerns a motor vehicle with a driver assistance system, which can be configured and further developed as described above. The motor vehicle preferably has an environment perception system connected to the driver assistance system for monitoring the vehicle's surroundings, particularly to avoid a collision with an object by intervening after evaluating the sensor data. Correct intervention by the driver assistance system is enabled by learning the physical boundary conditions that must be strictly adhered to during the use of the MLLM (Modular Learning Module).
[0025] The invention is now explained by way of example with reference to the accompanying drawings and preferred embodiments, wherein the features shown below can represent an aspect of the invention, either individually or in combination, and the scope of protection is defined by the claims. The drawings show: Fig. 1: a schematic representation of the physical mapping in a Kamm friction circle, Fig. 2: a schematic representation of a Kamm friction circle created from historical data and Fig. 3 a schematic representation of the principle of the method according to the invention.
[0026] As in Fig. As shown in Figure 1, a longitudinal acceleration a can occur when a motor vehicle is cornering, supported by the wheels of the motor vehicle10 against a surface. x and lateral acceleration a y occur, which result in an overall acceleration a ges is composed of. The relationship between the longitudinal acceleration a x and lateral acceleration a y can be simplified and represented in a Kamm friction circle 12.
[0027] As in Fig. As shown in Figure 2, in normal driving operation the motor vehicle is not operated up to a friction limit 18 of the wheels on the surface, so that a historical driving profile 14 in the Kamm friction circle 12 only covers a small sub-area, around which a tolerance range 16 can easily be provided, which does not exceed that of a friction limit 18 of the Kamm friction circle 12 and should still be considered an acceptable driving situation for a driver due to the historical driving profile 14.
[0028] As in Fig.As shown in Figure 3, the inventive method 20 can incorporate an artificial intelligence with an MLLM 22, to which measured time series data 24 can be supplied after tokenization 26. The time series data 24 can also be used in a physical model to represent the current driving state in the form of a graphical representation, in particular as a plot or latent representation 28 of a Kamm friction circle. Using a time series of these plots or latent representations 28, the historical driving profile 14 can be graphically represented. The plot / latent representation 28 of the current driving state and / or the historical driving profile 14 can be supplied to the MLLM 22 as data via a graphical tokenization 30. In addition, contextualization 32 of the input data for the MLLM 22 can be performed via suitable prompting.Preferably, the MLLM can undergo fine-tuning 34, particularly in a resource-efficient manner based on LoRA or QLoRA. Based on this, the MLLM can then evaluate the sensor data supplied in this way and, if necessary, perform an indirect driving intervention 36, for example in the form of a warning to the driver, and / or a direct driving intervention 38, for example in the form of an autonomously initiated emergency braking maneuver. QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] EP 4 145 242 A1
[0002]
Claims
[1] Method (20) for, in particular autonomous and / or assisted, driving of a motor vehicle (10), in which Sensor data is fed into a driver assistance system. based on the sensor data, a possible driving intervention (36,38) is determined, where, for the determination of the driving intervention (36, 38) using an MLLM (22), several measurement series (24) of vehicle data are taken into account, where, for training the MLLM (22), such measurement series (24) of vehicle data were linked with a physical model of driving stability, the MLLM (22) plausibly validates the driving intervention (36, 38) determined on the basis of the considered measurement series (22) of vehicle data using the physical model used in the training and In the event that the plausibility check using the physical model does not determine sufficient driving stability, the driving intervention (36, 38) is corrected until the plausibility check using the physical model determines sufficient driving stability. [2] Method (20) according to claim 1, wherein longitudinal forces and lateral forces acting on a wheel of the motor vehicle (10) are taken into account in the physical model, wherein insufficient driving stability is detected if a skidding of the motor vehicle (10) is likely. [3] Method (20) according to claim 1 or 2, wherein a vehicle dynamics model based on the Kamm friction circle (12) and / or the Krempel friction ellipse is used in the physical model. [4] Method (20) according to one of claims 1 to 3, wherein training and prompting (30) of the MLLM (22) with the physical model in the form of a graphical or latent representation (28) of parameters relevant to driving stability, in particular the longitudinal forces and lateral forces acting on the wheel of the motor vehicle, is carried out. [5] Method (20) according to claim 4, wherein a time series of graphical / latent representations (28) of the physical model generated from measurement data of the motor vehicle (10) are taken into account. [6] Method (20) according to claim 5, wherein, in the event that sufficient driving stability for the driving intervention (36, 38) has been determined, the driving intervention (36, 38) is corrected until the driving intervention (36, 38) lies within a predefined tolerance window (16) to the historical data (14) of the time series of the physical model, provided that sufficient driving stability is possible within the predefined tolerance window (16). [7] Method (20) according to any one of claims 1 to 6, wherein a loss function was applied for the training of the MLLM (22) in which a detected driving intervention that violates the physical model is punished disproportionately. [8] Method (20) according to any one of claims 1 to 7, wherein the measurement series (24) on vehicle data comprise time series data on a lateral acceleration, a longitudinal acceleration, a steering angle, a steering angle rate, a driving speed, an estimated coefficient of friction of a road surface, weather data influencing the estimated coefficient of friction, road cleaning data influencing the estimated coefficient of friction and / or measurement data from an ESP system. [9] Driver assistance system for autonomous and / or assisted driving of a motor vehicle (10), wherein the driver assistance system is configured to carry out the method according to any one of claims 1 to 8. [10] Motor vehicle (10) with a driver assistance system according to claim 8.
Citation Information
Patent Citations
Perception field based driving related operations
EP4145242A1