Method for training a machine learning model and application of the machine learning model in the trajectory planning of a vehicle, and vehicle

WO2026189782A1PCT designated stage Publication Date: 2026-09-17MERCEDES BENZ GROUP AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/054340
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-13
Filing Date
2026-02-18
Publication Date
2026-09-17

Smart Images

  • Figure EP2026054340_17092026_PF_FP_ABST
    Figure EP2026054340_17092026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for training a machine learning model (1) of a vehicle control system having a camera for object identification, and to an application in an inference mode for trajectory planning of a vehicle, wherein; A: starting images (2) of objects (3) with an annotation (4) are provided to the machine learning model (1) for training, comprising a distance (5) up to a maximum distance and an object height (6); B: scaled-down starting images (7) with an annotation (4) of the object height (6) known from step A are provided to the machine learning model (1) for training; and the following steps are carried out in the inference mode of the machine learning model (1): C: processing, by means of the machine learning model (1), camera images (8) captured by the camera, said processing comprising: (i) identifying objects (3) corresponding to unscaled starting images (2) in the camera images (8) and determining the associated distance (5); and (ii) identifying objects (3) corresponding to scaled starting images (7) in the camera images (8) and determining, from the pixel height (9) of the object (3) depicted in the camera images (8), a known image width (10) of the camera and the annotated object height (6), the associated distance (5') lying above the annotated maximum distance of the unscaled starting images (2); and D: transmitting the objects (3) identified in step C and associated distances (5, 5') to the vehicle control system set up for trajectory planning.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Mercedes-Benz Group AG

[0002] Methods for training a machine learning model and application of the machine learning model in the trajectory planning of a vehicle and vehicle

[0003] The invention relates to a method for training a machine learning model of a vehicle control system and an application of the trained machine learning model in an inference mode for trajectory planning of a vehicle, as well as a vehicle in which the machine learning model is applied.

[0004] The level of vehicle automation is constantly increasing. To safely navigate a vehicle through traffic, it is necessary for the vehicle to detect static and dynamic objects in its environment and determine their relative positions. Various sensor systems are used for this environmental perception, such as cameras, radar sensors, LiDAR sensors, ultrasonic sensors, and the like.

[0005] Camera images are usually analyzed using artificial intelligence. Convolutional Neural Networks (CNNs) have proven particularly effective in this process. Machine vision is also known as computer vision. One task is to recognize vehicles ahead and mark them in the camera images. This is typically achieved by drawing a bounding box around each vehicle. Beyond simple object recognition, machine learning-based models can perform other tasks, such as classifying objects, determining their dimensions, or calculating the distance between an object and the camera, or between a vehicle ahead and the driver's own vehicle.

[0006] For such a machine learning model to handle its assigned tasks, extensive training is necessary. The fundamental process involves feeding the machine learning model camera images annotated with so-called annotations. Such an annotation can be defined by or encompass a bounding box. Additional information can also be included, such as object classification, object height, and / or distance. These annotations, or "labels," can be added to camera images manually or automatically. For example, a bounding box can be set manually by a user, although this is usually quite time-consuming.

[0007] Distance information can be gathered using active environmental sensors, such as radar and / or LiDAR, and supplemented for the respective bounding box. However, with increasing distance, objects become more difficult to detect even with radar-based or LiDAR-based sensor systems.

[0008] The farther away an object is from the vehicle, the smaller it appears in the camera images. This means that objects can only be detected up to a certain maximum distance using machine vision. The risk increases that the distance to an object ahead cannot be determined using any sensor modality. Therefore, if a machine learning model is to be able to detect objects up to a distance of 300 meters, for example, the training material must also contain camera images showing objects at that distance. If the required detection range exceeds the detection range of the LiDAR, it is unclear how the distance should be provided for annotation.

[0009] Image annotation for deep neural networks is known from DE 102022 114047 A1. This publication describes the acquisition of a vehicle's surroundings using cameras and the evaluation of corresponding camera images using artificial neural networks. Vehicles ahead are detected and marked with a bounding box. Based on longitudinal and lateral movement data of the vehicle itself, the system estimates how the image content of the camera images changes over time, which is used to adjust the bounding box. The resulting section of the camera image can be cropped and enlarged. Object recognition is then performed for this image section. By filtering out irrelevant information, the artificial neural network is better able to recognize relevant image content even in situations with challenging visibility conditions.The present invention is based on the objective of providing an improved method for training a machine learning model and an application of the trained machine learning model for trajectory planning of a vehicle, with the help of which trajectory planning is made possible taking into account objects located further ahead of the vehicle.

[0010] According to the invention, this problem is solved by a method for training a machine learning model and its application in the trajectory planning of a vehicle with the features of claim 1. Advantageous embodiments and further developments as well as a vehicle for carrying out the method are described in the dependent claims.

[0011] An inventive method for training a machine learning model of a vehicle control system with a camera for object recognition and an application in an inference mode for trajectory planning of a vehicle provides that A: the machine learning model is provided with initial images of objects with an annotation, each comprising a distance distance extending up to a maximum distance and an object height;

[0012] B: The machine learning model is provided with downscaled source images for training, annotated with the object height known from step A;

[0013] and the following steps are performed in the inference mode of the machine learning model:

[0014] C: Processing camera images captured by the camera, wherein the processing includes:

[0015] - (i) Identifying objects in the camera images that correspond to unscaled source images and determining the associated distance; and

[0016] (ii) Identifying objects in the camera images that correspond to scaled source images and determining the associated distance above the annotated maximum distance of the unscaled source images from the pixel height of the object depicted in the camera images, a known image distance of the camera and the annotated object height; and

[0017] D: Transmitting the detected objects and their associated distances from step C to the vehicle control system configured for trajectory planning. Using the method according to the invention, a computer vision model is enabled to detect objects at distances exceeding the maximum distance and, in addition, to estimate their distance to the vehicle itself. The method utilizes the fact that objects or object types are recognized as such and that the object height is already known from annotated images of these objects. Based on the geometric relationships between object height, the pixel height of the object depicted in the camera image, and the image distance (i.e., the distance between the lens and the sensor), the distance to objects not annotated with respect to distance can be determined. Examples of object types include passenger cars or trucks, particularly of a specific design and make, people, animals, traffic signs, etc.

[0018] Based on the geometric relationships, the distance to unannotated objects can be reliably determined using methods based on the intercept theorem.

[0019] The inventive method is based on two core ideas. Firstly, the machine learning model is enabled by the execution of step B to recognize objects located far ahead, and secondly, taking into account fixed distance relationships, the desired distance is determined from the height of an object output by the machine learning model.

[0020] According to an advantageous embodiment of the method according to the invention, the object height is divided by the pixel height and the result is multiplied by the image distance to calculate the distance in step C(ii). This mathematical relationship is based on the lens equation, or the so-called image scale. This equation describes a relationship between the actual height of a real object and its distance to a lens plane, the height of an image of the object, and the distance of this image to the lens plane. The three quantities "object height," "image height," and "image distance" are known in the context of the invention, so the "object distance" can be calculated from them. The distance (object distance) is calculated from the object height, the pixel height (image height), and the image distance.

[0021] In step A, the initial training of the machine learning model takes place, enabling the model to recognize generally defined objects and determine their height and distance from the camera. For this purpose, the aforementioned annotated source images are used as input data during training. This procedure is well known to those skilled in the art.

[0022] Objects located so far from the camera that the maximum distance is reached or exceeded cannot be detected by the appropriately trained machine learning model in this way, let alone have their distance estimated. This is primarily because no corresponding distance can be determined as an annotation. According to the invention, step B is now performed. The source images used in step A can be used for this step. The source images are now scaled down, i.e., reduced in size, using a scaling factor. Established scaling methods can be used for this purpose. By scaling down the source images, the object to be detected, in particular a vehicle driving ahead, shrinks so that it appears to the machine learning model to be at a greater distance.The object marker, specifically in the form of a bounding box, and the object height are inherited from the annotation in step A. Through training, the machine learning model is enabled to recognize and classify objects at a distance increased by the scaling factor and to estimate the object height. The applicant has recognized that the object height is scale-invariant and therefore does not need to be adjusted. This allows for a robust and reliable determination of the distance based on the relationship between pixel height, image distance, and object height.

[0023] In inference mode, in step C, the camera images are captured and processed by a computing unit in the vehicle. The machine learning model processes these images to recognize objects at different distances. This processing step comprises the following sub-steps:

[0024] According to step C(i), objects are detected up to the annotated maximum distance. The distance is estimated directly by the machine learning model. Optionally, the object height can also be determined by the machine learning model.

[0025] According to step C(ii), the machine learning model recognizes objects whose distance exceeds the annotated maximum distance. Their object height is determined, and the distance (5') is calculated based on the image scale. The image distance (10) depends on the camera configuration and is therefore known. The pixel height (9) corresponds to the number of pixels representing the object in the camera image in the vertical direction. The vertical direction typically runs orthogonally upwards from the background.

[0026] In step D, information about the presence of corresponding objects and their associated distances is then transmitted to the downstream trajectory planning algorithm. This information is then considered in the trajectory planning process during inference mode, and the vehicle's control behavior is adjusted accordingly. Based on the trajectories determined in this way, control parameters for the vehicle's actuators can be calculated, and the actuators can be controlled to follow the trajectory. This allows the vehicle to be accelerated or decelerated in a controlled manner, as well as to perform steering maneuvers.

[0027] According to a further advantageous embodiment of the method according to the invention, the scaling factor for downscaling the source images in step B is a fixed value. Varying scaling factors can disrupt the training of machine learning models. For example, if objects up to a size of 60 pixels are annotated in the source image, this will be 60 pixels with a scaling factor of 1.540 pixels and with a factor of 230 pixels. These different scaling factors make it ambiguous for the machine learning model whether objects up to 30, 40, or 60 pixels should be detected. Without special techniques, such as so-called "ignore region handling," which themselves require complex annotations, only a faulty training of the machine learning model is possible.According to the invention, however, the scaling factor is kept constant, so that no separate measures are required for trouble-free training. The maximum possible distance at which objects can still be visually recognized thus increases according to the product of the maximum distance and the scaling factor.

[0028] A further advantageous embodiment of the method according to the invention provides that the distance determined in step C(ii) is transmitted, along with the corresponding camera image, to a central computing unit for further training of the machine learning model. The central computing unit can be a cloud server or server cluster. Such a server is often also referred to as a backend. For this purpose, the vehicle can be connected to the internet via a telecommunications unit using a mobile network. Corresponding camera images can also be scaled down.

[0029] According to a further advantageous embodiment of the method according to the invention, the annotation of the source images for improving trajectory planning also includes a classification of the respective objects. This enables the machine learning model to classify objects. Corresponding classifications, such as cars, trucks, pedestrians, cyclists, or the like, can thus be considered as an additional parameter in trajectory planning. For example, different typical movement behaviors can be assumed for different object classes, such as the highly erratic riding style of a cyclist uphill or the absence of emergency stops for a truck when turning sideways at an intersection due to blind spots. The vehicle can thus be controlled even more safely.

[0030] A further advantageous embodiment of the method according to the invention provides that the distances determined in inference mode in steps C(i) and / or C(ii) are validated using sensor signals of different modalities. The vehicle can include additional sensor systems, such as, in particular, radar sensors, LiDAR sensors, and / or ultrasonic sensors. Depth information can be obtained using such sensors, and thus corresponding distances to objects can also be determined. These findings obtained with other sensors can then be used to validate the visually determined distances, as long as the respective objects are within the detection range of said sensors.

[0031] In a vehicle comprising a camera for environmental perception and a vehicle control system, a machine learning model trained using a method described above is implemented in the vehicle control system, in particular in the form of an artificial neural network. The camera has, in particular, the same resolution as the source images used in the training of the machine learning model. The vehicle can be any road vehicle such as a car, truck, van, bus, or the like. It is also conceivable that it could be a rail vehicle, watercraft, or aircraft. The objects in question do not necessarily have to be vehicles ahead or other road users. Static environmental objects can also be taken into account.

[0032] Further advantageous embodiments of the inventive method for training a machine learning model of a vehicle control system with a camera for object recognition and an application in an inference mode for trajectory planning of a vehicle and the vehicle according to the invention also result from the exemplary embodiments which are described in more detail below with reference to the figures.

[0033] This shows:

[0034] Fig. 1 is a schematic representation of a camera image processed by a vehicle in the course of machine vision according to the prior art; Fig. 2 is a schematic representation of the process of a method according to the invention for training a machine learning model usable for machine vision;

[0035] Fig. 3 is a schematic representation of the relationship of the image scale; and Fig. 4 is a schematic representation of a camera image processed by a vehicle according to the invention in the course of machine vision.

[0036] Figure 1 shows a camera image 8 captured by a surround-view camera of a vehicle known from the prior art. The camera image 8 is processed by a machine learning model (not shown) to detect vehicles ahead. These vehicles represent objects 3. The machine learning model is trained to annotate each object 3 with an annotation 4. This annotation includes a corresponding object marker, here in the form of a bounding box, as well as a distance 5 and an object height 6. The distance 5 describes the distance of the object 3 to the vehicle itself or to the camera that captured the camera image 8. The object height 6 corresponds to the height of the object 3.

[0037] Objects 3 can only be detected up to a certain maximum distance using the machine learning model trained according to the state of the art, because the training dataset only contains annotated distance values ​​5 up to this maximum distance. In Figure 1, detectable objects 3 are symbolized by a checkmark and unrecognizable objects 3 by a cross.

[0038] Figure 2 shows the procedure of a method according to the invention for training a machine learning model 1. The camera image 8 shown in Figure 1 can represent an initial image 2 for the learning method. The annotation 4 can be inserted manually by a user or at least partially automatically, for example, taking into account distance information generated by an active measuring system, such as a distance radar and / or a LiDAR sensor system.

[0039] As Figure 2 shows, the original image 2 is now scaled down by a fixed scaling factor. This results in a scaled original image 7, which is then fed into the machine learning model 1 as input data during training. The annotation 4, already present in the original image 2, is partially reused. Annotation 4 serves to mark objects 3 to be recognized. Furthermore, the object height 6 is retained, allowing the machine learning model 1 to learn to recognize correspondingly small objects 3 and estimate their object height 6. By scaling down the original image 2 to the scaled original image 7, the effect is created that the object 3 is located at a greater distance. This allows not only the use of an existing training dataset of camera images 8, but also because the respective camera images 8 are already annotated with their respective annotations 4, thus eliminating the need for additional labeling.Through the training shown in Figure 2, the machine learning model 1 is enabled to recognize objects 3 and determine their object height 6 that are at a greater distance from the camera or the vehicle than the maximum distance.

[0040] However, the machine learning model 1 is unsure of the distance 5. To determine this, the relationship of the image scale shown in Figure 3 is used. Figure 3 shows the relationship between the distance 5' (denoted here by dw), the object height 6 (denoted here by hw), the pixel height 9 (denoted here by hß), and the image distance 10 (denoted here by dß). The distance 5' can thus be calculated from the object height 6, the pixel height 9, and the image distance 10. Specifically, the object height 6 is divided by the pixel height 9, and the result is multiplied by the image distance 10. It should be noted that the distances 5 from Figure 1 or 2 (i.e., the distance E directly determined by the machine learning model) and the distance 5' from Figure 3 both represent the distance to an object. The distinction between 5 and 5' is intended to clarify the different methods of determination.Distance 5 in step C(i) of claim 1 is the distance that the trained machine learning model directly determines. The model has learned to recognize objects up to the maximum distance and to derive their distance from the image features. Distance 5' in step C(ii) of claim 1 is the distance that is mathematically calculated as described above. As Figure 4 shows, in an inference mode of the machine learning model 1, camera images 8 acquired by the machine learning model 1 can be processed, and thus objects 3 can also be correctly recognized that are located ahead of the vehicle at a distance corresponding to the product of the scaling factor and the maximum distance. Furthermore, the distance 5' can also be determined for such objects 3 located far ahead.

[0041] The information about detected objects 3 and their respective distances 5, 5' is now fed into a downstream algorithm for trajectory planning for the vehicle. This allows for the determination of a particularly safe and long-term planned trajectory for the vehicle. The vehicle can thus be controlled with particular safety, at least semi-automatically, and especially autonomously.

[0042] Using the method according to the invention, the camera-based detection range for sensor-based environmental sensing of vehicles can be increased, while reducing the effort and costs of labeling training data.

Claims

Mercedes-Benz Group AG Patent claims 1. Method for training a machine learning model (1) of a vehicle control system with a camera for object recognition and an application in an inference mode for trajectory planning of a vehicle, wherein A: The machine learning model (1) is provided with initial images (2) of objects (3) with an annotation (4), each comprising a distance (5) up to a maximum distance and an object height (6); B: The machine learning model (1) is provided with downscaled initial images (7) with an annotation (4) of the object height (6) known from step A; and the following steps are performed in the inference mode of the machine learning model (1): C: Processing camera images captured by the camera (8) by the machine learning model, wherein the processing includes: - (i) Identifying objects (3) in the camera images (8) that correspond to unscaled source images (2) and determining the associated distance (5); and - (ii) Identifying objects (3) in the camera images (8) that correspond to the scaled source images (7) and determining the associated distance (5') above the annotated maximum distance of the unscaled source images (2) from the pixel height (9) of the object (3) depicted in the camera images (8), a known image distance (10) of the camera and the annotated object height (6); and D: Transmitting the detected objects (3) and associated distances (5, 5') from step C to the system set up for trajectory planning Vehicle control system.

2. Method according to claim 1. characterized by the fact that To calculate the distance (5) in step C (ii) the object height (6) is divided by the pixel height (9) and the result is multiplied by the image distance (10).

3. Method according to claim 1 or 2, characterized by the fact that The scaling factor for downscaling the source images (2) in step B is a fixed value.

4. Method according to claim one of claims 1 to 3, characterized by the fact that The distance (5) determined in step C (ii) is transmitted with the corresponding camera image (8) to a central computing unit for further training of the machine learning model (1).

5. Method according to any one of claims 1 to 4, characterized by the fact that The annotation (4) of the source images (2) for improving trajectory planning includes a classification of the respective objects (3).

6. Method according to any one of claims 1 to 5, characterized by the fact that The distances determined in inference mode in steps C(i) and C(ii) (5) are made plausible with sensor signals of different modalities.

7. Vehicle comprising a camera for environmental sensing and a vehicle control system, characterized by the fact that in the vehicle control system a machine learning model (1) trained using a method according to one of claims 1 to 6 is implemented, in particular in the form of an artificial neural network.