Computer-implemented method and system for object detection using a generative deep learning model, and training method
By employing a generative deep learning model pre-trained on natural language data and retrained with a small dataset of sensor data, the method addresses the inefficiencies of existing object detection methods, achieving reduced training effort and effective object detection in automated driving systems.
Patent Information
- Application Number
- PCT/EP2024/086976
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-12-17
- Publication Date
- 2025-06-26
AI Technical Summary
Existing methods for object detection in highly automated driving require extensive annotation and training efforts, which are resource-intensive and inefficient.
A computer-implemented method and system using a generative deep learning model pre-trained on natural language data, which is then retrained using a small dataset of individual images and/or point clouds from a vehicle environment, reducing the need for extensive annotation and training.
This approach significantly reduces the training effort required for object detection, allowing for efficient adaptation of the model to specific object detection tasks in various vehicle environments, while maintaining effective performance in object detection and prediction.
Smart Images

Figure EP2024086976_26062025_PF_FP_ABST
Abstract
Description
[0001] Description
[0002] title
[0003] Computer-implemented method and system for object detection using a generative deep learning model and training procedures
[0004] The present invention relates to a computer-implemented method for providing a generative deep learning model for object detection.
[0005] Furthermore, the invention relates to a computer-implemented method for object detection using a generative deep learning model.
[0006] The invention further relates to a system for providing a generative deep learning model for object detection. Furthermore, the invention relates to a system for object detection using a generative deep learning model.
[0007] The invention further relates to a computer program with program code for carrying out the method according to the invention and to a computer-readable data carrier with program code of a computer program for carrying out the method according to the invention when the computer program is executed on a computer.
[0008] State of the art
[0009] Algorithms for object detection for highly automated driving can be provided using various training methods. EP 3446281 A1 discloses a training method for object recognition, wherein the training method comprises providing at least one training image in plan view, aligning a training object in the training image along a predetermined direction, annotating at least one training object from the at least one training image using a predefined annotation scheme, extracting at least one feature vector for describing the content of the at least one labeled training object and at least one feature vector for describing at least one background scene, and training a classifier model based on the extracted feature vectors.
[0010] Traditionally, a neural network for object detection is trained on a standard dataset consisting of several individual frames. To achieve satisfactory performance in object detection, the training dataset usually requires annotation or labeling. The annotation step is very resource-intensive, regardless of whether it is performed manually or at least partially automated.
[0011] Accordingly, there is a need to improve existing methods for object detection for highly automated driving in such a way that the training effort for the machine learning algorithm is reduced. It is therefore an object of the invention to provide an improved method and system for providing a machine learning algorithm for object detection that requires less training effort.
[0012] Disclosure of the invention
[0013] The object is achieved according to the invention by a computer-implemented method for providing a generative deep learning model for object detection with the features of patent claim 1.
[0014] The object is further achieved according to the invention by a computer-implemented method for object detection using a generative deep learning model having the features of patent claim 6.
[0015] The object is further achieved according to the invention by a system for providing a generative deep learning model for object detection with the features of patent claim 12.
[0016] Furthermore, the problem is solved by a system for object detection using a generative deep learning model with the features of patent claim 13.
[0017] The object is further achieved according to the invention by a computer program having the features of patent claim 14 and a computer-readable data carrier having the features of patent claim 15. The invention relates to a computer-implemented method for providing a generative deep learning model for object detection, comprising providing a generative deep learning model pre-trained on the basis of natural language data, in particular a pre-trained generative transformer model.
[0018] The method further comprises retraining the generative deep learning model using a training data set of individual images and / or point clouds based on sensor data of a vehicle environment.
[0019] The invention further relates to a computer-implemented method for object detection using a generative deep learning model. The method comprises providing a first data set based on sensor data of a vehicle's surroundings, comprising a plurality of individual images, in particular time series images and / or point clouds.
[0020] Furthermore, the method comprises applying the generative deep learning model trained according to the invention, in particular a generative transformer model, to the first data set comprising the plurality of individual images and / or point clouds for object defect detection.
[0021] The method further comprises outputting a , a
[0022] Result of the object detection representing, in particular annotated objects in the majority of individual images and / or point clouds having the second data set.
[0023] The invention further relates to a computer-implemented method for validating an automated driving function of a motor vehicle using the result of the computer-implemented method for object detection according to the invention.
[0024] The invention further relates to a system for providing a generative deep learning model for object detection. The system comprises a first training calculation unit configured to provide a generative deep learning model pre-trained on the basis of natural language data, in particular a pre-trained generative transformer model.
[0025] In addition, the system comprises a second training calculation unit which is configured to retrain the generative deep learning model using a training data set of individual images and / or point clouds based on sensor data of a vehicle environment.
[0026] The invention further relates to a system for object detection using a generative deep learning model. The system comprises a data provision unit configured to provide a first data set based on sensor data of a vehicle's surroundings, comprising a plurality of individual images, in particular time series images and / or point clouds.
[0027] In addition, the system comprises a calculation unit which is configured to apply a generative deep learning model trained according to one of claims 1 to 5, in particular a generative transformer model, to the first data set comprising the plurality of individual images and / or point clouds for object detection, and a data output unit which is configured to output a second data set representing a result of the object detection, in particular comprising annotated objects in the plurality of individual images and / or point clouds.
[0028] The invention further relates to a computer program with program code for carrying out the inventive method for object detection when the computer program is executed on a computer, as well as to a computer-readable data carrier with program code of a computer program for carrying out the inventive method when the computer program is executed on a computer.
[0029] Machine learning algorithms are based on the use of statistical methods to train a data processing system to perform a specific task without having been explicitly programmed to do so. The goal of machine learning is to construct algorithms that can learn from data and make predictions. These algorithms create mathematical models that can be used, for example, to classify data—in this case, to detect objects.
[0030] Image data refers to data that can be reproduced as an image or graphic using a special program. The fact that an object is represented in image data also means that the corresponding image data shows the object or contains a representation of the object.
[0031] Image data can be, for example, video image data or radar image data. Point cloud data can be, for example, LiDAR point cloud data. Another possible data type is position data from a GPS sensor.
[0032] One idea of the present invention is to use a generative deep learning model pre-trained on a large data set, which is already capable of processing sensor data from a vehicle-mounted environment detection sensor.
[0033] Generative deep learning models trained on natural language data have the ability to predict the next token for a corresponding query. In the context of a natural language query, a token can be, for example, a letter. Context vectors are then calculated from this token. These represent not only the sequence of letters, but also their position in the text and their context.
[0034] This principle is also applicable to sensor data, where individual image pixels or points of a point cloud can be considered as tokens.
[0035] An image is a list of points with X and Y values. A point cloud is a list of points with X, Y, and Z values. This allows the model to predict pixels or points.
[0036] The retraining of the generative deep learning model is advantageously carried out using a comparatively small dataset of sensor data from the vehicle's surroundings compared to the initial training data used to train the generative deep learning model. This allows the model to be retrained for a specific task, such as traffic sign recognition, using the smaller dataset of sensor data.
[0037] During retraining, for example, parts of the generative deep learning model can be frozen and not trained. The subsequent parts are then retrained based on the data set of sensor data.
[0038] The combination of the pre-trained model and the retrained model can thus be used as a new model for inference. The main advantages are that the training strategy allows a base model to be trained on large, diverse, non-sequential data sets and then extended and optimized with a small model for application to sequential data.
[0039] Further embodiments of the present invention are the subject of the further subclaims and the following description with reference to the figures.
[0040] According to a preferred development of the invention, it is provided that the training data set comprises individual images and / or point clouds of a predetermined driving situation and / or a predetermined environmental condition of the vehicle surroundings.
[0041] This advantageously makes it possible to retrain the pre-trained model with specific data of the given driving situation and / or the given environmental conditions of the vehicle environment in order to achieve improved performance of the model with regard to a specific object detection task in the vehicle environment.
[0042] According to a further preferred development of the invention, the predetermined driving situation comprises an urban traffic environment, a non-urban traffic environment and / or a motorway traffic environment, and wherein the predetermined environmental condition comprises a time of day, a weather condition and / or a road surface condition. The model can thus advantageously be trained using driving situation data from respective traffic environments. In this way, for example, a model can be trained exclusively with driving situation data from an urban traffic environment, so that this model is used in the inference exclusively in an urban traffic environment. Further models can, for example, each be trained for a non-urban traffic environment and / or a motorway traffic environment as well as respective environmental conditions.
[0043] Country-specific data from the respective traffic environment can also be considered or trained during training. Country-specific data can include, for example, country-specific traffic signs, road width, and country-specific road markings.
[0044] According to a further preferred development of the invention, it is provided that the training data set for retraining (S2) the generative deep learning model (A) has between 50 and 500 individual images (10) and / or point clouds (12), in particular between 100 and 300 individual images (10) and / or point clouds (12).
[0045] Due to the use of a very small number of individual images (10) and / or point clouds (12) for retraining the pre-trained model, training to adapt the model to a specific object detection task of an automated driving function can be implemented with very little effort. According to a further preferred development of the invention, the training data set (TD) for retraining (S2) the generative deep learning model (A) comprises unannotated raw sensor data.
[0046] Since no annotated data is required for training generative deep learning models, this can advantageously lead to a significant gain in efficiency when training the model compared to conventional training methods for other models.
[0047] According to a further preferred development of the invention, it is provided that the generative deep learning model carries out a prediction of a position of the road markings, traffic signs and / or road users for a predetermined future point in time.
[0048] Thus, the model can advantageously perform not only static object detection, but also dynamic object detection, in which the position of respective objects in future frames or time series images can be predicted. In this context, for example, a speed of an ego vehicle, which represents the reference point for the respective time series images, as well as a speed of other road users can be taken into account. According to a further preferred development of the
[0049] The invention provides that the predetermined future point in time is determined relative to the acquisition time of the respective individual image and / or the respective point cloud of the vehicle's surroundings.
[0050] The future time indicates in how many future frames or individual images or point clouds a corresponding object should be detected.
[0051] According to a further preferred development of the invention, it is provided that the generative deep learning model processes multimodal input data, in particular text and image data, and outputs multimodal output data, in particular text and / or image data.
[0052] For the object detection task, it is necessary that the model receives, in addition to the image or point cloud data, at least initially text data as input, which specifies the corresponding object detection task.
[0053] Furthermore, the model is able to output both text data and image or point cloud data.
[0054] For example, the model can set a bounding box around a detected object or specify the coordinates of the detected object in the image. Furthermore, the model can output additional metadata regarding detected objects, such as the speed of other road users. Furthermore, the model can include, for example, lane markings and / or the lane of the ego vehicle in the output image or video.
[0055] Label point cloud data .
[0056] The features of the computer-implemented method for providing a machine learning algorithm for object detection described herein are equally applicable to the system for providing a machine learning algorithm for object detection and vice versa.
[0057] Likewise, the features of the computer-implemented method for object detection by a machine learning algorithm described herein are applicable to the system for object detection by a machine learning algorithm and vice versa.
[0058] Short description of the drawings
[0059] For a better understanding of the present invention and its advantages, reference is now made to the following description in conjunction with the accompanying drawings.
[0060] The invention is explained in more detail below with reference to exemplary embodiments which are shown in the schematic illustrations of the drawings.
[0061] It shows :
[0062] Fig. 1 is a flow diagram of a computer-implemented method for providing a generative deep learning model for object detection according to a preferred embodiment of the invention;
[0063] Fig. 2 is a flow diagram of a computer-implemented method for object detection using a generative deep learning model according to the preferred embodiment of the invention;
[0064] Fig. 3 is a schematic representation of a system for providing a generative deep learning model for object detection according to the preferred embodiment of the invention; and
[0065] Fig. 4 is a schematic representation of a system for object detection using a generative deep learning model according to the preferred embodiment of the invention.
[0066] Unless otherwise indicated, like reference symbols refer to like elements in the drawings.
[0067] Detailed description of the embodiments
[0068] This is Fig . l shown computer-implemented method for providing a generative deep learning model for object detection according to a preferred
[0069] Embodiment of the invention comprises providing
[0070] S 1 of a generative deep learning model A pre-trained on the basis of natural language data, in particular a pre-trained generative transformer model. Furthermore, the method comprises retraining S2 of the generative deep learning model A using a training data set TD of individual images 10, in particular time series images and / or point clouds 12, based on sensor data SD of a vehicle environment.
[0071] The training data set TD comprises individual images 10 and / or point clouds 12 of a given driving situation and / or a given environmental condition of the vehicle's surroundings. The given driving situation comprises an urban traffic environment, a non-urban traffic environment, and / or a motorway traffic environment. The given environmental condition comprises a time of day, for example, a time of day or night, as well as a weather condition and / or a road condition, for example, a coefficient of friction of the road.
[0072] The training data set TD for retraining S2 of the generative deep learning model A preferably comprises between 50 and 500 individual images 10 and / or point clouds 12, in particular between 100 and 300 individual images 10 and / or point clouds 12.
[0073] The training dataset TD for retraining S2 of the generative deep learning model A also contains unannotated raw sensor data.
[0074] Fig. 2 shows a flow diagram of a computer-implemented method for object detection using a generative deep learning model according to the preferred embodiment of the invention. The method comprises providing S 1 ' a first data set Dl based on sensor data SD of a vehicle's surroundings, comprising a plurality of individual images 10, in particular time series images and / or point clouds 12.
[0075] Furthermore, the method comprises applying S2 ' the generative deep learning model A trained according to the invention, in particular a generative transformer model, to the first data set Dl comprising the plurality of individual images 10 and / or point clouds 12 for object defect detection.
[0076] The method further comprises outputting S3 ' a second data set D2 representing a result of the object detection, in particular annotated objects in the plurality of individual images 10 and / or point clouds 12.
[0077] The generative deep learning model A performs a detection of road markings, traffic signs and / or road users based on the plurality of provided individual images 10 and / or point clouds 12.
[0078] Furthermore, or alternatively, the generative deep learning model A predicts the position of the lane markings, traffic signs, and / or road users for a specified future point in time. The specified future point in time is determined relative to the acquisition time of the respective individual image and / or the respective point cloud 12 of the vehicle's surroundings.
[0079] The generative deep learning model A processes multimodal input data, in particular text and image data. Furthermore, the generative deep learning model A produces multimodal output data, in particular text and / or image data.
[0080] Fig. 3 shows a schematic representation of a system 1 for providing a generative deep learning model for object detection n according to the preferred embodiment of the invention.
[0081] The system 1 comprises a first training calculation unit 16 which is configured to provide a generative deep learning model A pre-trained on the basis of natural language data, in particular a pre-trained generative transformer model.
[0082] Furthermore, the system 1 comprises a second training calculation unit 18 which is configured to retrain the generative deep learning model A using a training data set TD of individual images 10 and / or point clouds 12 based on sensor data SD of a vehicle environment.
[0083] Fig. 4 shows a schematic representation of a system 2 for object detection using a generative deep learning model according to the preferred embodiment of the invention. The system 2 comprises a data provision unit 20, which is configured to provide a first data set D1 based on sensor data SD of a vehicle's surroundings, comprising a plurality of individual images 10, in particular time series images, and / or point clouds 12.
[0084] In addition, the system comprises a calculation unit 22 which is configured to apply a generative deep learning model A trained according to the method according to the invention, in particular a generative transformer model, to the first data set D1 comprising the plurality of individual images 10 and / or point clouds 12 for object detection.
[0085] The system 2 further comprises a data output unit 24 which is configured to output a second data set D2 representing a result of the object detection, in particular annotated objects in the plurality of individual images 10 and / or point clouds 12.
[0086] Although specific embodiments have been illustrated and described herein, it will be understood by those skilled in the art that numerous alternative and / or equivalent implementations exist. It should be noted that the example embodiment or example embodiments are only examples and are not intended to limit the scope, applicability, or configuration in any way. Rather, the foregoing summary and detailed description will provide one skilled in the art with a convenient road map for implementing at least one example embodiment, it being understood that various changes in the functionality and arrangement of elements may be made without departing from the scope of the appended claims and their legal equivalents. In general, this application is intended to encourage changes orTo cover adaptations or variations of the embodiments presented herein. For example, the order of the method steps may be modified. Furthermore, the method may be carried out sequentially or in parallel, at least in sections.
[0087] Reference symbol list
[0088] 1 system
[0089] 2 systems 10 individual images
[0090] 12 point clouds
[0091] 16 first training calculation unit
[0092] 18 second training calculation unit
[0093] 20 Data provision unit 22 Calculation unit
[0094] 24 output unit
[0095] A generative deep learning model
[0096] The first data set
[0097] D2 second data set SD sensor data
[0098] TD training data
[0099] S 1-S3 procedural steps
[0100] S l ' -S3 ' process steps
Claims
Claims 1. Computer-implemented method for providing a generative deep learning model (A) for object detection, comprising the steps of: providing (S1) a generative deep learning model (A) pre-trained on the basis of natural language data, in particular a pre-trained generative transformer model; and retraining (S2) the generative deep learning model (A) using a training data set (TD) of individual images (10), in particular time series images and / or point clouds (12), based on sensor data (SD) of a vehicle environment.
2. Computer-implemented method according to claim 1, wherein the training data set (TD) comprises individual images (10) and / or point clouds (12) of a given driving situation and / or a given environmental condition of the vehicle surroundings.
3. Computer-implemented method according to claim 2, wherein the specified driving situation comprises an urban traffic environment, a non-urban traffic environment and / or a motorway traffic environment and wherein the specified Environmental condition includes a time of day, a weather condition and / or a road condition.
4. Computer-implemented method according to one of the preceding claims, wherein the training data set (TD) for retraining (S2) the generative deep learning model (A) comprises between 50 and 500 individual images (10) and / or point clouds (12), in particular between 100 and 300 individual images (10) and / or point clouds (12).
5. Computer-implemented method according to one of the preceding claims, wherein the training data set (TD) for retraining (S2) the generative deep learning model (A) comprises unannotated raw sensor data.
6. Computer-implemented method for object detection using a generative deep learning model (A), comprising the steps of: providing (Sl') a first data set (Dl) based on sensor data (SD) of a vehicle environment, comprising a plurality of individual images (10), in particular time series images and / or point clouds (12); Applying (S2') a generative deep learning model (A) trained according to one of claims 1 to 5, in particular a generative transformer model, to the first data set (Dl) comprising the plurality of individual images (10) and / or point clouds (12) for object detection; and Outputting (S3') a second data set (D2) representing a result of the object detection, in particular comprising annotated objects in the plurality of individual images (10) and / or point clouds (12).
7. Computer-implemented method according to claim 6, wherein the generative deep learning model (A) performs a detection of road markings, traffic signs and / or road users based on the plurality of provided individual images (10) and / or point clouds (12).
8. Computer-implemented method according to claim 7, wherein the generative deep learning model (A) performs a prediction of a position of the lane markings, traffic signs and / or road users for a given future time.
9. Computer-implemented method according to claim 8, wherein the predetermined future time is determined relative to the acquisition time of the respective individual image and / or the point cloud (12) of the vehicle surroundings.
10. Computer-implemented method according to one of claims 6 to 9, wherein the generative deep learning model (A) processes multimodal input data, in particular text and image data, and outputs multimodal output data, in particular text and / or image data.
11. A computer-implemented method for validating an automated driving function of a motor vehicle using the result of the computer-implemented method for object detection according to one of claims 6 to 10.
12. System (1) for providing a generative deep learning model (A) for object detection, comprising: a first training calculation unit (16) which is configured to provide a generative deep learning model (A) pre-trained on the basis of natural language data, in particular a pre-trained generative transformer model; and a second training calculation unit (18) which is configured to retrain the generative deep learning model (A) using a training data set (TD) of individual images (10) and / or point clouds (12) based on sensor data (SD) of a vehicle environment.
13. System (2) for object detection using a generative deep learning model (A), comprising: a data provision unit (20) which is configured to provide a first data set (Dl) based on sensor data (SD) of a vehicle environment and comprising a plurality of individual images (10), in particular time series images, and / or point clouds (12); a calculation unit (22) which is configured to apply a generative deep learning model (A) trained according to one of claims 1 to 5, in particular a generative transformer model, to the first data set (D1) comprising the plurality of individual images (10) and / or point clouds (12) for object detection; and a data output unit (24) which is configured to output a second data set (D2) representing a result of the object detection, in particular comprising annotated objects in the plurality of individual images (10) and / or point clouds (12).
14. A computer program product comprising a computer program comprising software means for carrying out one of the methods according to any one of claims 1 to 5 and 6 to 10, wherein the computer program is executed on a computer.
15. A computer-readable data carrier with program code of a computer program for carrying out at least parts of a method according to any one of claims 1 to 5 and 6 to 10 when the computer program is executed on a computer.
Citation Information
Patent Citations
Training method and detection method for object recognition
EP3446281A1
Map and environment based activation of neural networks for highly automated driving
US20190213451A1
Pretraining framework for neural networks
US20230019211A1