Method for training a machine learning model for object detection in a vehicle environment

The method combines generated and operational data for iterative training of machine learning models, enhancing detection accuracy and adaptability to sensor changes, addressing data scarcity and retraining challenges in vehicle object detection.

DE102024120801A1Inactive Publication Date: 2025-08-21BAYERISCHE MOTOREN WERKE AG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE102024120801
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2025-08-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing machine learning models for object detection in vehicle environments face challenges in achieving high detection accuracy due to insufficient training data, especially for rare scenarios, and require significant retraining when sensor types change, leading to high development effort.

Method used

A method combining generated object detection data from a generative machine learning model with operational data from vehicle tests, using an iterative training process to adapt and improve the model, allowing for efficient training and adaptation to new sensors and rare scenarios without complete retraining.

Benefits of technology

Enhances training efficiency and reliability by expanding the training dataset with artificially generated data, effectively addressing rare scenarios and adapting to sensor changes, reducing the need for extensive retraining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for training a machine learning model for object detection in a vehicle environment comprises the following steps: generating first generated object detection data using a generative machine learning model based on first multimodal input data; determining first operational object detection data in a first test operation of a vehicle; training a second machine learning model based on the first generated object detection data and the first operational object detection data; determining second operational object detection data in a second test operation of the vehicle; and training a third machine learning model based on the first generated object detection data and the second operational object detection data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present disclosure relates to a method for training a machine learning model for object detection in a vehicle environment.

[0002] Efficient and reliable object detection is particularly important for automated and autonomous vehicle systems. It must always be ensured that all relevant objects in the vehicle's surroundings can be reliably detected within the shortest possible time using the vehicle's operating sensor data and incorporated into the vehicle control system. Object detection must also ensure that no spurious objects (e.g., fog, spray, pollen, optical reflections) or irrelevant objects (e.g., leaves, insects, and birds) are classified as traffic-relevant objects, thereby causing unnecessary and unexpected driving maneuvers.

[0003] Machine learning models such as neural networks have generally proven suitable for evaluating extensive vehicle sensor data. One disadvantage, however, is that high detection accuracy can usually only be achieved if the machine learning model has previously been subjected to a complex training process based on extensive training data. Particular difficulties arise, on the one hand, in providing a sufficient amount of training data, which, for safety reasons, must also contain a sufficient number of rare detection scenarios (so-called "corner cases"). For example, the quality of the optical vehicle sensor data can be severely impaired by smoke from a burning vehicle involved in an accident. Such a scenario rarely occurs under real traffic conditions and is therefore usually underrepresented in available training data.However, obtaining high-quality training data with sufficient coverage of rare detection scenarios is associated with considerable effort.

[0004] A further challenge is that the sensor types in vehicles are constantly evolving, for example, with regard to resolution accuracy and maximum range of the sensors or the data formats used. However, a minor change in a sensor type can significantly impact the detection accuracy of a trained machine learning model. Existing machine learning models, such as deep neural networks, therefore always have to be retrained whenever the optical vehicle sensor system is modified. This involves considerable effort in vehicle development, which is undesirable.

[0005] One object of the present disclosure is to reduce the effort required to train a machine learning model for object detection in a vehicle environment, particularly in connection with new sensors for monitoring the vehicle environment and the sufficient consideration of rare detection scenarios. At the same time, the reliability of the trained networks should be increased wherever possible.

[0006] The problem is solved by the features of the independent claims. The subclaims contain further developments of the disclosure.

[0007] According to one aspect of the disclosure, the object is achieved by a method for training a machine learning model for object detection in a vehicle environment, the method comprising at least the following steps: - generating first generated object detection data by means of a generative machine learning model based on first multimodal input data; - determining first operational object detection data in a first test operation of a vehicle, wherein the first operational object detection data comprise first operational sensor data of the vehicle and first object data, wherein the first operational sensor data represent an environment of the vehicle in the first test operation and the first object data are determined by means of a first machine learning model based on the first operational sensor data;- Training a second machine learning model based on the first generated object detection data and the first operational object detection data; - Determining second operational object detection data in a second test operation of the vehicle, wherein the second operational object detection data comprises second operational sensor data of the vehicle and second object data, wherein the second operational sensor data represents the surroundings of the vehicle in the second test operation and the second object data is determined by means of the second machine learning model based on the second operational sensor data; and - Training a third machine learning model based on the first generated object detection data and the second operational object detection data.

[0008] The method according to the invention is characterized, on the one hand, by the combination of generated object detection data and operational object detection data. The latter are obtained during a respective test operation of the vehicle and thus represent real training data that reflect the actual application situation when performing object detection in the vehicle. The generated object detection data, on the other hand, are artificially generated using a generative machine learning model. Using the generated data, the training basis can be significantly expanded both in terms of scope and content in order to accelerate the training process and improve object detection performance. Controlling the generated object detection data based on the multimodal input data is particularly beneficial.By combining input data representing multiple sensor types or other data types, particularly complex and rare detection scenarios can be addressed in a targeted and efficient manner. In this way, specific detection scenarios can also be artificially generated.

[0009] A further aspect of the invention consists in an advantageous iterative training process in which the object detection data from multiple test operations is used. The machine learning model used for object detection in the test operation is updated between test operations in order to directly use the improved new training status for generating improved object detection data. Machine learning models as such are generally available in various predefined types, e.g., as neural networks or vision transformers, including powerful training algorithms, and can be used in the described method to obtain a trained machine learning model that can be described by a set of parameters and operated for object detection on this basis.

[0010] According to one embodiment, the second machine learning model is trained on the basis of the first machine learning model, and the third machine learning model is trained on the basis of the second machine learning model. Thus, the machine learning models do not need to be completely retrained each time, but can be reliably and efficiently further trained and improved accordingly in an iterative training process. In this way, for example, an existing machine learning model can be adapted to new sensor types or rare detection scenarios.

[0011] According to a further embodiment, the third machine learning model is provided in the vehicle for object detection. The vehicle can then be controlled in an automated driving mode based on object data determined using the third machine learning model. The third machine learning model can be used, in particular, in a real customer operation to offer fully autonomous or partially automated driving functions. It should be understood that, within the scope of the described training method, multiple training loops can also be performed, with the last machine learning model preferably being used for practical use in customer operation after a convergence criterion has been reached.

[0012] According to a further embodiment, the input data comprises sensor data of a first sensor type and / or a second sensor type, wherein the first sensor type and / or the second sensor type is adapted to at least one vehicle sensor type of the vehicle that is used to determine the first and / or second operating sensor data. For example, the sensor types represented by the multimodal input data can correspond to a vehicle sensor type. The performance of object detection can be particularly increased in this way. It is also possible to use the input data to consider sensor types that are only used in the vehicle during the training process. The training basis can thus be specifically expanded to include one or more sensor types that are only used during successive test operations to generate the respective operating sensor data.Nevertheless, training of the machine learning models can already begin.

[0013] According to a further embodiment, the first operating sensor data comprises sensor data of a first vehicle sensor type, and the second operating sensor data comprises sensor data of a second vehicle sensor type. The sensor data of the first vehicle sensor type are preferably also part of the first multimodal input data. In this way, the generative data can benefit from the operational data of the first vehicle sensor type and ensure particularly good adaptation. The first vehicle sensor type can, for example, be a radar sensor or another vehicle-specific sensor type with a first resolution and / or a first data format.

[0014] Preferably, within the scope of the method, second generated object detection data are additionally generated by means of the generative machine learning model on the basis of second multimodal input data. The second multimodal input data preferably comprises sensor data of the second vehicle sensor type. The third machine learning model can then be trained in addition to or alternatively to the first generated object detection data on the basis of the second generated object detection data. The second vehicle sensor type can be, for example, a radar sensor or another vehicle-typical sensor type with a second resolution that differs from the first resolution. The sensor data of the second vehicle sensor type can also be defined in a second data format that differs from the first data format of the first vehicle sensor type.By training the third machine learning model based on data that also represents the second vehicle sensor type, the second machine learning model can be adapted to the characteristics of a new vehicle sensor type. This does not require a completely new training process. This approach is particularly advantageous when the first and second vehicle sensor types do not differ fundamentally from each other, but only with regard to certain parameters. For example, the second resolution can be higher than the first resolution without the content properties of the data as such fundamentally differing. However, adaptation to completely new sensor types is also possible.

[0015] According to a further embodiment, the multimodal input data represents several object detection scenarios that occur less frequently than other object detection scenarios in a customer's vehicle operation. Using the comparatively easily obtainable generated object detection data, the rare scenarios can be specifically considered without having to laboriously capture them using real data from vehicle operation. For example, the rare object detection scenarios can each be assigned a probability value that is lower than a threshold value, e.g., lower than 1% of all object detection scenarios. In this way, rare object detection scenarios can be objectively classified.

[0016] According to a further embodiment, the first multimodal input data comprises image data, text data, and / or audio data. The image data can, in particular, comprise 3D pixel data and / or 2D pixel data from a laser scanner and / or a radar sensor. Other image data can be 2D image data from an image sensor. Rare object detection scenarios can be specifically taken into account using text data, e.g., by the text data specifying plumes of smoke in the vehicle's surroundings. The multimodal input data generally represents properties of expected or desired operating sensor data of the vehicle. The input data generally comprises several different data types.

[0017] The second and / or third machine learning model is preferably trained using an automated machine learning algorithm. Corresponding algorithms are generally known, e.g., AutoAl. It should be understood that, after training, machine learning models are each represented by a predetermined set of parameters that forms a calculation rule. The operating sensor data is processed using the calculation rule to determine the object data representing the detected objects in the vehicle's environment, e.g., using position data and / or a contour to describe the object shape.

[0018] According to a further embodiment, the generative machine learning model comprises a first neural subnetwork that generates a generation dataset based on the multimodal input data. The generative machine learning model preferably comprises a second neural subnetwork that generates the generated object detection data based on the generation dataset. Generative machine learning models, particularly of the neural network type, are known per se, although the described method is not limited to a specific type of this model.

[0019] The individual steps of the method are preferably computer-implemented, particularly with regard to determining the data and training the machine learning models. The individual machine learning models are preferably each based on neural networks, which can in particular be formed by convolutional neural networks (CNNs). In particular, some or all of the machine learning models of the described method can each be formed by a neural network, in particular a convolutional network. Alternatively, however, other types of (training-based) machine learning models can also be used, such as Vision Transformer.

[0020] According to a further aspect of the disclosure, a data processing device is provided. The data processing device is configured to perform one or more steps of the described method. Optionally, the data processing device is configured to perform a method step described as advantageous or optional and / or to implement a method feature in order to achieve an associated technical effect.

[0021] According to a further aspect of the disclosure, a computer program and / or a computer-readable medium is provided. The computer program and / or the computer-readable medium comprise instructions which, when the program or instructions are executed by a data processing device, cause the device to perform the method according to the disclosure and / or steps thereof. Optionally, the computer program and / or the computer-readable medium comprises instructions which, when the program or instructions are executed by a data processing device, cause the device to perform the method steps described as advantageous or optional in order to achieve an associated technical effect.

[0022] According to one aspect of the disclosure, a motor vehicle is provided with a data processing device. The data processing device of the motor vehicle and / or the motor vehicle are configured to perform a method step described as advantageous or optional and / or to implement a method feature in order to achieve an associated technical effect. The motor vehicle preferably has one or more sensors for detecting the surroundings of the motor vehicle, such as a laser scanner (e.g., a lidar scanner), a 2D camera, and / or a radar sensor. During operation of the vehicle, the sensors output operational sensor data, which is processed by means of a machine learning model to obtain object data.The object data represents one or more objects in the vehicle's surroundings and serves as the basis for an automated driving function of the vehicle, in particular for automated vehicle control while avoiding a collision with an object. The machine learning model used in the vehicle can in particular be the first, second, or third machine learning model described above in connection with the method.

[0023] In the following, the aspects of the disclosure are further described by way of example only with reference to the figures, which show in detail the following: Fig. 1 is a schematic flow diagram of a method according to one aspect of the disclosure; Fig. 2 an exemplary detection scenario in a vehicle environment; Fig. 3 is a schematic block diagram illustrating steps for training a machine learning model for object detection in a vehicle environment; and Fig. 4 a schematic data processing device for carrying out the method of Fig. 1.

[0024] Fig. 1 shows a schematic flow diagram of a computer-implemented method 100 for training a machine learning model adapted for detecting objects in a vehicle environment. The method is described with further reference to Fig. 2 and Fig. 3 described.

[0025] An example vehicle environment is shown in Fig. 2. In the driver's field of vision are, among others, a passenger car 24, a truck 26 and another passenger car 28. The vehicles 24, 26 and 28 form objects that are detected using several sensors. In the example of Fig. 2, the vehicle has a laser scanner 50 and a radar sensor 52. The laser scanner 50 and the radar sensor 52 each generate operational sensor data, which is fed to a data processing device 54 of the vehicle in order to detect the vehicles 24, 26, and 28. The training of the machine learning model used to evaluate the operational sensor data is the subject of method 100.

[0026] The method 100 comprises the following steps in detail. In method step 110, first generated object detection data are created using a generative machine learning model based on first multimodal input data. The input data is multimodal, i.e., the input data comprises different data types, including sensor data of a first sensor type and a second sensor type. In method step 120, first operational object detection data are determined during a first test operation of a vehicle. The first operational object detection data comprise first operational sensor data of the vehicle (e.g., sensor data of the laser scanner 50 and the radar sensor 52). The first operational sensor data thus represent the surroundings of the vehicle during the first test operation. To detect the objects in the surroundings (e.g., the vehicles 24, 26, and 28), the first operational sensor data are processed using a first machine learning model.Initial object data is obtained as input data, representing the positions and / or other object properties, such as object classification and object contour. The first machine learning model can, in particular, be an initial network trained exclusively on generated object detection data.

[0027] In method step 130, a second machine learning model is trained based on the first generated object detection data and the first operational object detection data. The training is performed on the basis of the first machine learning model in order to improve its detection performance. Upon completion of this training step, the second machine learning model is therefore available with a detection performance that is improved compared to the first machine learning model.

[0028] In method step 140, second operational object detection data are determined in a second test operation of the vehicle. The second operational object detection data comprise second operational sensor data of the vehicle and second object data. The second operational sensor data represent the environment of the vehicle in the second test operation (e.g., with other vehicles, their position and type of Fig. 2). The second object data is determined using the second machine learning model based on the second operational sensor data.

[0029] To further improve detection performance, the second machine learning model is further trained in method step 150 based on the first and / or further generated object detection data as well as the second operational object detection data. As a result, a third machine learning model is obtained whose detection performance is further improved compared to the second machine learning model.

[0030] The procedure 100 of Fig. 1 is characterized by an iterative training scheme that can efficiently and reliably generate a powerful machine learning model for object detection.

[0031] To further clarify process aspects, the following schematic flow diagram of Fig. 3. The multimodal input data for the generated object detection data includes 2D image data 70, video data 72, 3D pixel data 74, and text data 76. The 3D pixel data 74 (“point cloud data”) may, in particular, include sensor data corresponding to that of the laser scanner 50 and the radar sensor 52. Blocks 78 and 82 of Fig. 3 together represent the generation of generated object detection data 84 using a generative machine learning model. In block 78, a generation data set 80 is first generated using a subnetwork (not shown in detail), which is then processed by another subnetwork to obtain the generated object detection data 84.

[0032] Block 86 represents the step of training a machine learning model for object detection. In addition to the generated object detection data 84, operational object detection data 88 are used, which are determined during a respective test operation of the vehicle (block 90). After training a respective machine learning model, a parameter set 92 is available that defines the underlying machine learning model, e.g., a neural network. As described in connection with Fig. 1, after execution of method step 130, a parameter set 92 for the second machine learning model is available. This parameter set 92 forms the basis for a second test operation of the vehicle. For this purpose, the first machine learning model in the vehicle is updated on the basis of the parameter set 92. Then, during the execution of the second test operation, the second operational object detection data are determined (see method step 140), which are used to train the third machine learning model (see block 86). The machine learning models described in connection with the figures, i.e., the first, second, and third machine learning models, are preferably each designed as a neural network.

[0033] The method 100 is preferably carried out with a data processing device 55 which is Fig. 4 is shown schematically. The data processing device 55 comprises, as hardware components, at least a processor 56, a non-volatile memory 58 (e.g., an SD card), and a memory 60 (e.g., a main memory, RAM). A computer program for executing the steps of the method 100 can be permanently stored in the non-volatile memory 58 and loaded into the memory 60 for execution. The data processing device 55 can, in particular, be formed in a central server. The data processing device 54 of Fig. 2 can be constructed correspondingly to the data processing device 55 and connected to the data processing device 54 via a wireless connection in order to efficiently and securely transmit the operational object detection data 88 and the parameter set 92 between the devices 54 and 55 and in this way to facilitate the practical implementation of the method 100. List of reference symbols 24 passenger cars 26 trucks 28 passenger cars 50 laser scanners 52 radar sensor 54 Data processing device 55 Data processing device 56 processor 58 Non-volatile memory 60 storage 70 2D image data 72 video data 74 3D pixel data 76 text data 78 Creating a generation data record 80 Generation data set 82 Creating generated object detection data 84 Generated object detection data 86 Training a machine learning model for object detection 88 Operational object detection data 90 Determining operational object detection data 92 parameter set 100 procedures 110 Process step 120 process steps 130 process steps 140 process steps 150 process steps

Claims

[1] A method for training a machine learning model for object detection in a vehicle environment, the method comprising: - generating first generated object detection data by means of a generative machine learning model based on first multimodal input data (110); - Determining first operational object detection data in a first test operation of a vehicle, wherein the first operational object detection data comprise first operational sensor data of the vehicle and first object data, wherein the first operational sensor data represent an environment of the vehicle in the first test operation and the first object data are determined by means of a first machine learning model on the basis of the first operational sensor data (120); - training a second machine learning model based on the first generated object detection data and the first operational object detection data (130); - Determining second operational object detection data in a second test operation of the vehicle, wherein the second operational object detection data comprise second operational sensor data of the vehicle and second object data, wherein the second operational sensor data represent the environment of the vehicle in the second test operation and the second object data are determined by means of the second machine learning model on the basis of the second operational sensor data (140); and - training a third machine learning model based on the first generated object detection data and the second operational object detection data (150). [2] The method of claim 1, wherein the second machine learning model is trained based on the first machine learning model and the third machine learning model is trained based on the second machine learning model. [3] The method according to claim 1 or 2, wherein the machine learning model is provided in the vehicle for object detection, and wherein the vehicle is controlled in an automated driving mode on the basis of object data determined by means of the third machine learning model. [4] Method according to one of the preceding claims, wherein the input data comprises sensor data of a first sensor type and / or a second sensor type, wherein the first sensor type and / or the second sensor type are adapted to at least one vehicle sensor type of the vehicle which is used for determining the first and / or second operating sensor data. [5] Method according to one of the preceding claims, wherein the first operational sensor data comprises sensor data of a first vehicle sensor type and wherein the second operational sensor data comprises sensor data of a second vehicle sensor type, wherein the first multimodal input data comprises sensor data of the first vehicle sensor type and wherein the method further comprises: Generating second generated object detection data by means of the generative machine learning model based on second multimodal input data, wherein the second multimodal input data comprises sensor data of the second vehicle sensor type, and wherein the training of the third machine learning model takes place in addition to or alternatively to the first generated object detection data on the basis of the second generated object detection data. [6] Method according to one of the preceding claims, wherein the multimodal input data represents a plurality of object detection scenarios that occur less frequently in a customer operation of the vehicle than other object detection scenarios. [7] Method according to one of the preceding claims, wherein the first multimodal input data comprise image data, text data and / or audio data, in particular wherein the image data comprise 3D pixel data and / or 2D pixel data of a laser scanner and / or a radar sensor. [8] Method according to one of the preceding claims, wherein the training of the second and / or third machine learning model is carried out by means of an automated machine learning algorithm. [9] Method according to one of the preceding claims, wherein the generative machine learning model comprises a first neural subnetwork that generates a generation data set based on the multimodal input data, and wherein the generative machine learning model comprises a second neural subnetwork that generates the generated object detection data based on the generation data set. [10] Data processing device (54) which is configured to at least partially carry out the steps of the method (100) according to one of claims 1 to 9. [11] Computer program and / or computer-readable medium comprising instructions which, when the program or instructions are executed by a data processing device (54), cause the device (54) to carry out the method (100) and / or steps of the method (100) according to one of the preceding claims.

Citation Information

Patent Citations

  • COMPUTER-IMPLEMENTED METHOD FOR TRAINING A COMPUTER VISION MODEL

    DE102021200348A1

  • Ml-based automatic recognition of new and relevant data sets

    EP3961511A1