Computer-implemented method and device for creating at least one synthetic image data set

By using an image classifier to detect and modify dynamic objects in real datasets, the method generates synthetic datasets that enhance neural network robustness and generalizability, addressing the challenge of distinguishing between static and dynamic objects in autonomous vehicle training.

DE102023212620A1Pending Publication Date: 2025-06-18ZF FRIEDRICHSHAFEN AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102023212620
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-18

AI Technical Summary

Technical Problem

Existing neural networks for autonomous vehicles struggle to distinguish between static and dynamic objects in real-world datasets, leading to overfitting and reduced generalizability, as they cannot modify the pose of dynamic objects in real-world data.

Method used

A computer-implemented method using an image classifier to detect dynamic objects in real datasets and a generative AI to create synthetic datasets by modifying or removing these objects, ensuring the neural network is trained on diverse scenarios.

Benefits of technology

Enhances the robustness and generalizability of neural networks by creating synthetic datasets that mimic real-world environments, reducing overfitting and improving the vehicle's driving performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method (100) and a device (200) for creating at least one synthetic image data set (SB). The method (100) comprises: providing (110) a real image data set (RB) characterizing a captured image of a real environment; providing (120) an image classifier for recognizing at least one dynamic object (O1; O2) in a real image data set (RB) and creating object information characterizing the dynamic object (O1; O2) recognized in the real image data set (RB); recognizing (130), by means of the image classifier, at least one first dynamic object (O1) in the real image data set (RB); creating (140), by means of the image classifier, object information characterizing the recognized first object (O1); providing (150) a trained generative artificial intelligence, AI;Creating (160) at least one synthetic image data set (SB) based on the real image data set (RB) and the object information by changing a position or removing the detected first object (O1) in the real image data set (RB);
Need to check novelty before this filing date? Find Prior Art

Description

The invention relates to a computer-implemented method and a device for creating at least one synthetic image data set, a computer-implemented method and a device for training a neural network for controlling a vehicle, a vehicle and respective computer program products.For at least partially, in particular fully autonomously driving vehicles, a neural network for controlling the vehicle must be trained in advance in order to achieve reliable driving performance. While methods such as mimic learning have an immense potential for mimicking observed driver behavior, the ultimate performance correlates directly with the quality of the recorded data set.On the one hand, in end-to-end learning, the need for a costly labeling process for processing a real dataset is dispensed with, and on the other hand, the missing labels do not allow the underlying model to explicitly distinguish between a type of entity for any recording scenes.For example, a fully data-controlled AI model or neural network is unable to distinguish between a building and a parked vehicle if the position of the vehicle is kept constant throughout the acquisition. In contrast to a synthetic dataset, it is not possible to change the pose of all dynamic objects in a real dataset in order to reduce this risk. Instead, it is preferable to expand the recorded data set during data preprocessing to improve the generalization and robustness of the final AI model.It is an object of the present invention to provide a computer implemented method and apparatus which at least improve one or more of the aforementioned disadvantages. In particular, it is an object of the present invention to expand the dataset pool in such a way that the robustness of AI models can be improved.According to a first aspect, the object is achieved by a computer-implemented method for creating at least one synthetic image data set. The method comprising:providing a real image data set characterizing a recorded image of a real environment;providing an image classifier for recognizing at least one dynamic object in a real image data set and creating object information characterizing the dynamic object recognized in the real image data set;detecting, by means of the image classifier, at least one first dynamic object in the real image data set;creating, by means of the image classifier, object information characterizing the recognized first object;providing a trained generative artificial intelligence, AI;creating at least one synthetic image data set based on the real image data set and the object information by changing a position or removing the recognized first object in the real image data set.The proposed method improves the quality of real and synthetic data sets by creating new synthetic image data sets by means of the generative AI, which have not yet been acquired in the original real image data set. The inventors have recognized that existing generative AI models not only offer the possibility of changing existing image datasets, but also creating completely new image datasets, in particular based on text descriptions and / or instructions. For the automation of the process, the image classifier is used, which is designed for the detection of dynamic objects in the real image data set. This step prevents augmentation requests for non-existent or even static objects (e.g., buildings). By recognizing dynamic objects, they can be changed, for example, in their position and / or length for a synthetic image data set. This ensures that data about the actually observed scene, here the real environment, are not lost or are corrupted. This achieves that a neural network to be trained is less susceptible to overfitting to a specific state and / or a specific scene, i.e. a time instance of the environment.The first dynamic object as well as any further dynamic object can be a movable object. By contrast, non-dynamic objects are buildings, road signs, plants, trees, barriers and the like. The dynamic objects may be non-static objects.The provision of the image classifier can comprise a storage of the image classifier, in particular on a storage unit. Alternatively or additionally, the provision of the image classifier can comprise communicating with the image classifier stored in a distributed network, a cloud and / or an external computer in such a way that the image classifier receives the real image data set as input, recognizes the at least one first object on the basis thereof and generates the associated object information. The object information can be sent back to a transmitter of the real image data set and received by the latter. Consequently, the image classifier can be included or provided externally.The provision of the real image data set can comprise a reception of the real image data set and / or a storage of the real image data set. The real image data set can be stored on a storage unit. The real image data record can be sent and / or provided by at least one vehicle, a camera, an external unit, a cloud and / or a storage unit.Providing the trained generative AI may include storing the generative AI, in particular on a storage unit. Alternatively or additionally, the provision of the generative AI can comprise communicating with the generative AI stored in a or the distributed network, a or the cloud and / or an or the external computer, such that the generative AI receives the real image data set and the object information as input and generates the at least one synthetic image data set based thereon. The generated image data set can be sent back to a transmitter of the real image data set and / or of the object information and received by the latter. Consequently, the generative AI may be included with or provided externally.The method can furthermore comprise storing the at least one synthetic image data set, in particular on a storage unit.The generative AI is known per se to the person skilled in the art. Generative artificial intelligence (also called generative AI or GenAI) is artificial intelligence capable of generating texts, images, or other media using generative models. Generative AI learns the patterns and structure of its input data and then generates new data having similar characteristics.The image classifier can be trained in such a way that it is designed to recognize dynamic objects in image data sets. This can be done, for example, by means of supervised learning. Furthermore, the image classifier can be trained in such a way that, as shown below, no classifiers but rather suitable text descriptions and / or instructions can be generated on the basis of the recognized dynamic object in the image dataset. The results of the image classifier may then be passed to the generative AI to create synthetic image datasets without requiring manual queries. The generative AI may be or include a generative AI model.A real image data record can be a recording of a real environment, for example a traffic situation of a vehicle. By contrast, a synthetic image data record can be a non-real recorded environment and / or recording of a invented environment, for example a changed traffic situation of the vehicle. The synthetic image data set can be based on the real image data set.The object information may further characterize the position, a size, a volume, a perspective distortion, a variability of the appearance, for example color, and / or a theoretical direction of movement of the detected first object.The object information may further characterize a text description of the detected first object, the position, the size, the volume, and / or the theoretical direction of movement of the detected first object. The text description can be a heading and / or signature of the real image data set and / or of the recognized first object.The text description can characterize a type, a theoretical target position to be reached and / or further properties of the first object, which can be used in particular by the generative AI for creating the at least one synthetic image data set.Features described with respect to the first object may be performed for other recited objects.The object information may further include an instruction to the generative AI, the instruction characterizing the changing of the position to a changed position and / or the removal of the detected first object.The method may further comprise:adding, by means of the image classifier, at least one second dynamic object into the real image data set,wherein the object information further characterizes the second object in the real image data set,wherein the creation of the synthetic image data set is further based on the added second object.The creation of the at least one synthetic image data set can be a creation of a plurality of synthetic image data sets that are different from one another.The plurality of mutually different synthetic image data sets can be created by respectively changing the position and / or removing the first object and / or changing a position and / or removing the second object.The method can be carried out fully automatically. Once a real image data set with at least one dynamic object has been provided, a plurality of synthetic images can be automatically created based on this real image data set. In this case, the image classifier can automatically recognize the dynamic object, provide the object information to the generative AI and generate the generative AI in accordance with synthetic image datasets.According to a second aspect, the object is achieved by a device for creating at least one synthetic image data set, comprising means for carrying out the method according to the first aspect. The means can comprise a storage unit for storing the real and / or synthetic image data set and / or for storing a computer program product mentioned below. The means may further comprise a processor for executing the computer program product. Furthermore, the means can comprise a communication unit for receiving the real image data set and / or for transmitting the synthetic image data set.According to a third aspect, the object is achieved by a computer program product comprising instructions which, when the program is executed by a processor, cause the processor to execute the method according to the first aspect.According to a fourth aspect, the object is achieved by a computer program product comprising a computer-readable database having at least one synthetic image data set which has been created by means of the method according to the first aspect.According to a fifth aspect, the object is achieved by a computer-implemented method for training a neural network for controlling a vehicle, comprising:providing at least one synthetic image data set produced by means of the method according to the first aspect;training the neural network based on the at least one synthetic image data set such that the neural network is configured to control the vehicle.The training can further comprise training based on the real image data set.According to a sixth aspect, the object is achieved by a device for training a neural network for controlling a vehicle, comprising means for carrying out the method according to the fifth aspect. The means can comprise a storage unit for storing the real and / or synthetic image data set and / or for storing a computer program product mentioned below. The means may further comprise a processor for executing the computer program product. Furthermore, the means can comprise a communication unit for receiving the real image data set and / or for transmitting the synthetic image data set. The vehicle can be an at least partially, in particular completely autonomously driving vehicle.According to a seventh aspect, the object is achieved by a computer program product comprising instructions which, when the program is executed by a processor, cause the processor to execute the method according to the fifth aspect.According to an eighth aspect, the object is achieved by a vehicle comprising the device according to the second aspect and / or the device according to the sixth aspect. The vehicle can be an at least partially, in particular completely autonomously driving vehicle.Method features which have been described with respect to the method according to the first aspect can be implemented as method and / or device features of the second to eighth aspects.Preferred exemplary embodiments are explained by way of example with reference to the enclosed figures. The following are shown: FIG. 1 shows a schematic representation of a computer-implemented method for creating at least one synthetic image data set; FIG. 2 shows a schematic representation of a real image data set and a synthetic image data set; FIG. 3 shows a schematic illustration of a device for creating at least one synthetic image data set; FIG. 4 is a schematic illustration of a computer-implemented method for training a neural network for controlling a vehicle; FIG. 5 shows a schematic illustration of an apparatus for training a neural network for controlling a vehicle; and FIG. 6 shows a vehicle with such a device.In the figures, identical or substantially functionally identical or similar elements are denoted by the same reference numerals.FIG. 1 shows a computer-implemented method 100 for creating at least one synthetic image data set. The method 100 can be stored in the form of a computer program product, for example on a memory unit and / or a processor.The method 100 comprises providing 110 a real image dataset RB characterizing a recorded image of a real environment, see for example FIG. 2. FIG. 2 shows an image and / or a snapshot from an eoperssive of a vehicle on a roadway in the real image dataset RB. In front of the vehicle, a first object O 1, here a further vehicle O 1, travels in the right lane. On the right side of the roadway there is a second object O2, here a pedestrian O2-The method 100 further comprises providing 120 an image classifier for recognizing at least one dynamic object O 1 in a real image data set RB and creating object information characterizing the dynamic object O 1 recognized in the real image data set RB.Furthermore, the method 100 comprises a detection 130, by means of the image classifier, of at least one first dynamic object O 1 in the real image dataset RB. According to FIG. 2, the image classifier is designed to recognize the further vehicle O 1 and the pedestrian O 2 as dynamic objects.The method 100 further comprises creating 140, by means of the image classifier, object information characterizing the recognized first object O 1, in FIG. 2 the two objects O 1, O 2. The object information further characterizes the position, a size, a volume and / or a theoretical direction of movement of the detected first object O 1. For example, the image classifier can recognize that the further vehicle O 1 is oriented forward in the direction of travel. Also, the image classifier can recognize that the pedestrian O 2 is looking forward.The image classifier can further adapt the object information based on this information such that the object information further characterizes a text description of the recognized objects O 1, O 2, the positions, the sizes, the volumes and / or the theoretical movement directions of the recognized objects O 1, O 2. For example, the text description may be: "The vehicle O 1 moves forward in the traveling direction. The pedestrian O 2 moves forward in parallel to the traveling direction.".The object information may further include an instruction to a generative artificial intelligence, AI, mentioned below. This instruction may be based on the text description. For example, the instruction may be: "create a synthetic image data set by moving the vehicle O 1 by 4 meters forward along the direction of travel and by moving the pedestrian O 2 by 1 meter forward parallel to the direction of travel.".The method 100 further comprises providing 150 the trained generative artificial intelligence, KI, and creating 160 at least one synthetic image data set SB based on the real image data set RS and the object information by changing a position or removing the recognized first object, here the first and second objects O 1, O 2 in the real image data set RB. The synthetic image data set can be stored and used for training a neural network. Alternatively or additionally, the synthetic image data set can be output directly to the neural network.As shown in FIG. 2, the generative KI can generate the synthetic image data set SB, in which the objects O 1, O 2 can be found at new positions, based on the real image data set RB and the object information. The remaining information of the real image data set RB can be retained in real terms.Alternatively or additionally, at least one of the objects O 1, O 2 can be removed from the real image data set RB, so that the synthetic image data set SB only has one of the objects O 1, O 2.Alternatively or additionally, a third dynamic object can be added by the image classifier, in particular the object information. The position of the third object can also be changed.Consequently, it is possible to generate a plurality of synthetic image data sets SB based on at least one real image data set RB.FIG. 3 shows a device 200 for creating at least one synthetic image data set SB. The apparatus 200 comprises a storage unit 210 and a processor 220. The storage unit 210 is configured to store a computer program product comprising instructions which, when the program is executed by the processor 220, cause the processor to execute the method 100.FIG. 4 shows a computer-implemented method 300 for training a neural network for controlling a vehicle. The method 300 may be stored in the form of a computer program product, for example on a memory unit and / or a processor.The method 300 comprises providing 310 at least one synthetic image data set SB produced by means of the method 100 and training 320 the neural network on the basis of the real image data set RB and / or the at least one synthetic image data set SB in such a way that the neural network is formed for controlling the vehicle.FIG. 5 shows an apparatus 400 for training a neural network. The apparatus 400 comprises a memory unit 410 and a processor 420. The storage unit 410 is configured to store a computer program product comprising instructions which, when the program is executed by the processor 420, cause the latter to execute the method 300.FIG. 6 shows an at least partially, in particular fully autonomously driving vehicle 500 comprising the device 400. Alternatively or additionally, the vehicle 500 may include the device 200.Reference numerals denote reference numerals100 computer-implemented method for creating at least one synthetic image data set; 110 Providing a real image data set 120 Providing an image classifier 130 Recognizing, by means of the image classifier, at least one first dynamic object 140 Creating, by means of the image classifier, object information 150 Providing a trained generative artificial intelligence, KI 160 Creating at least one synthetic image data set RB real image data set SB synthetic image data set O1 first object O2 second object 200 Device for creating at least one synthetic image data set 210 Storage unit 220 Processor 300 Computer-implemented method for training a neural network for controlling a vehicle 310 Providing at least one synthetic image data set 320 created by means of the method 100 Training the neural network 400 Device for training a neural network for controlling a vehicle 410 Storage unit 420 Processor 500 Vehicle

Claims

Computer-implemented method (100) for creating at least one synthetic image dataset (SB), comprising: providing (110) a real image dataset (RB) characterizing a recorded image of a real environment; providing (120) an image classifier for recognizing at least one dynamic object (O1; O2) in a real image dataset (RB) and creating object information characterizing the dynamic object (O1; O2) recognized in the real image dataset (RB); recognizing (130), by means of the image classifier, at least one first dynamic object (O1) in the real image dataset (RB); creating (140), by means of the image classifier, object information characterizing the recognized first object (O1); providing (150) a trained generative artificial intelligence, AI; creating (160) at least one synthetic image data set (SB) based on the real image data set (RB) and the object information by changing a position or removing the recognized first object (O1) in the real image data set (RB).The method (100) according to claim 1, wherein the object information further characterizes the position, a size, a volume and / or a theoretical direction of movement of the detected first object (O1).Method (100) according to claim 1 or 2, wherein the object information further characterizes a text description of the detected first object (O1), the position, the size, the volume and / or the theoretical direction of movement of the detected first object (O1).The method (100) according to any of the preceding claims, wherein the object information further comprises an instruction to the generative AI, wherein the instruction characterizes changing the position to a changed position and / or removing the detected first object (O1).Method (100) according to one of the preceding claims, further comprising: adding, by means of the image classifier, at least one second dynamic object into the real image data set (RB), wherein the object information further characterizes the second object in the real image data set (RB), wherein the creation (160) of the synthetic image data set (SB) is further based on the added second object.Method (100) according to one of the preceding claims, wherein the creation of the at least one synthetic image data set (SB) is creation of a plurality of synthetic image data sets different from one another, wherein the plurality of synthetic image data sets different from one another is created by a respective change of the position and / or the removal of the first object (O1) and / or a change of a position and / or a removal of the second object.Method (100) according to one of the preceding claims, wherein the method (100) is carried out fully automatically.Apparatus (200) for creating at least one synthetic image data set, comprising means (210; 220) for carrying out the method (100) according to one of the preceding claims.A computer program product comprising instructions which, when the program is executed by a processor, cause the processor to perform the method (100) of any one of claims 1 to 7.Computer program product, comprising a computer-readable database with at least one synthetic image data set (SB) which has been created by means of the method (100) according to one of claims 1 to 7.Computer-implemented method (300) for training a neural network for controlling a vehicle (500), comprising: providing (310) at least one synthetic image data set (SB) produced by means of the method (100) according to one of Claims 1 to 7; training (320) the neural network based on the at least one synthetic image data set (SB) such that the neural network is formed for controlling the vehicle (500).Apparatus (400) for training a neural network for controlling a vehicle (500), comprising means (410; 420) for carrying out the method (300) according to claim 11.A computer program product comprising instructions which, when the program is executed by a processor, cause the processor to perform the method (300) of claim 11.A vehicle (500) comprising the apparatus (200) of claim 8 and / or the apparatus (400) of claim 12.