Use of synthetically generated outliers for improved vehicle detection
The method of training an artificial neural network using domain-specific images and synthetically generated outliers addresses the robustness and efficiency issues in existing object recognition systems, resulting in improved object detection and classification for autonomous driving applications.
Patent Information
- Application Number
- DE102023131995
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-16
- Publication Date
- 2025-05-22
AI Technical Summary
Existing object recognition systems for autonomous and automated driving are not sufficiently robust and efficient, leading to missed object detections, false alarms, and interruptions in autonomous driving, which can result in safety hazards and increased costs.
A method for training an artificial neural network that includes providing domain-specific images, generating synthetically generated outliers using these images, and training the network based on both the domain-specific images and the outliers. This approach enhances the network's robustness and ability to distinguish between known and unknown objects.
The proposed solution results in a more robust artificial neural network capable of reliably detecting and classifying objects, reducing false alarms and interruptions in autonomous driving, thereby enhancing safety and operational efficiency.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method for detecting objects, a method for training an artificial neural network for detecting objects, as well as a corresponding computer program product, a computer-readable data carrier, a control unit for carrying out one of the methods and a corresponding vehicle with a corresponding control unit.
[0002] Artificial neural networks are known for classifying and / or localizing objects in input images. This can be used, for example, to assign objects in an input image to a category based on classification. In the context of autonomous and / or at least partially automated driving, such artificial neural networks could be used to detect the environment (road situation) and enable vehicle operation based on the classification.
[0003] However, the known methods and systems from the state of the art have various disadvantages. For example, they may not be sufficiently robust and / or efficient. It is possible that objects are not detected. In particular, this can lead to these objects not being taken into account during operation (particularly during driving, e.g. in autonomous and / or automated driving). In the worst case, this can result in accidents with drastic consequences (e.g. personal injuries). It is also possible that an error, in particular a false alarm, occurs when unknown objects are detected. This can lead to the system being switched off and thus prevent a greater degree of autonomous and / or automated driving. It can also trigger an undesired driving maneuver. In particular, it is therefore possible that even with supposedly unimportant or harmless objects, the autonomous orautomated driving is interrupted, at least temporarily. There may be a high false positive and / or false negative rate of known systems and / or methods. In particular, this can be due to outliers that cannot be detected, for example, due to a lack of features identifiable by the artificial neural network. This can be disadvantageous because it requires the driver to be ready to intervene at all times, particularly since it is never known when the system will shut down and intervention might be necessary. This can lead to potential hazards and / or increased costs. In particular, this can (unnecessarily) reduce safety and lead to inadequate driving comfort and / or reduced sales figures. This can also be disadvantageous because the system cannot learn from failure to detect and / or errors (or crashes).Furthermore, known systems and methods often require too much training data and / or input from users (machine learning experts and / or drivers). Training with purely or partially synthetic data, especially in or based on edge cases, can lead to inadequate generalization of the artificial neural network, especially since these data may be too dissimilar to actually occurring real-world outliers. In particular, sufficient adjustability (or control) with regard to outliers may be insufficient. Probabilistic recognition models often use Gaussian assumptions (especially distributions), which are inadequate, especially since these may inadequately represent real distributions. With post-hoc methods, especially without training or a training goal, which are optimized especially after training a neural network, the stability and / or generalization may be inadequate.
[0004] It is therefore an object of the present invention to at least partially overcome at least one of the disadvantages described above. In particular, the object of the invention is to provide a more robust artificial neural network.
[0005] The above object is achieved by a method having the features of the first independent patent claim, a method having the features of the second independent patent claim, a computer program product having the features of the independent computer program product claim, a computer-readable data carrier having the features of the independent patent claim relating to a computer-readable data carrier, a control unit having the features of the independent patent claim relating to a control unit, and a vehicle having the features of the independent vehicle claim. Further features and details of the invention emerge from the subclaims, the description, and the drawings. Features and details which are related to the invention orthe methods according to the invention are described, of course, also in connection with the computer program product according to the invention and / or in connection with the computer-readable data carrier according to the invention and / or in connection with the control unit according to the invention and / or in connection with the vehicle according to the invention and vice versa, so that with regard to the disclosure of the individual aspects of the invention, reference is or can always be made to each other. In particular, advantages described in the context of the first, second, third, fourth, fifth and / or sixth aspect also apply to the first, second, third, fourth, fifth and / or sixth aspect.
[0006] The above object is achieved according to a first aspect by a method for detecting objects, in particular for a vehicle, comprising: - Providing an artificial neural network, in particular according to the second aspect, - capturing an input image by a capturing device, - Providing the input image to the artificial neural network, - Identifying an image area in the input image by the artificial neural network, wherein at least one object is detected in the image area, - Classifying, by a classifier, the image area by the artificial neural network to obtain a classification, wherein during the classification at least one known object and / or unknown object is detected in the image area, - Outputting output information based on the classification.
[0007] The method according to the first and / or second aspect can preferably be used in the context of autonomous and / or automated driving, in particular to enable automatic detection of objects.
[0008] The method according to the first and / or second aspect can be computer-implemented and / or performed repeatedly. In particular, it is conceivable that the method according to the second aspect precedes the method according to the first aspect. It can also be provided that the method according to the second aspect is used to enable readjustment (in particular relearning or further learning) of the artificial neural network, in particular during operation of a vehicle and / or the method according to the first aspect.
[0009] The artificial neural network according to the first aspect can preferably be generated and / or provided by a method for training according to the second aspect.
[0010] The method according to the first aspect can be used for the (automatic) detection of objects.
[0011] An input image can be captured by a capture device, in particular an optical system, e.g., at least one camera. This can be integrated on and / or in a vehicle, for example, in the front area. The capture device can provide the input image to the artificial neural network via a data connection.
[0012] The identification of image regions and / or objects in the input image can comprise the identification of several and / or a large number of, in particular rectangular, image regions. For example, the input image can be scanned in a raster-like manner. In this case, object detection can occur, which is performed by the artificial neural network. For example, an image region can be detected which, in particular, contains a (known) street sign. An image region can contain at least one object, in particular a known and / or unknown object.
[0013] The classification can be carried out by a classifier which was trained according to the second aspect. During the classification, the image region can be processed by the artificial neural network in order to obtain a classification of the image region and / or the at least one object. In this case, a classification can be two-stage, preferably "known" and "unknown". A known object can be, for example, a street sign or a pedestrian. An unknown object can be, for example, a grizzly bear or a passerby disguised as a grasshopper. Preferably, an unknown object cannot have been directly represented by the training or the training data. Thus, it can be provided that the method according to the first aspect recognizes an (unknown) object, but cannot assign it to a known object (from the training). In other words, a determination can be made that an (unknown) object has been detected.However, this cannot be classified as known. Classifying the image area may therefore include classifying the at least one object in the image area.
[0014] Output information can include the classification, for example, "known" and / or "unknown." Alternatively or additionally, this can include the calculated and / or estimated location of the object, in particular the distance. Closer objects may pose a greater danger. This can be incorporated into a (quantitative) danger factor, which is greater the closer the object is, in particular to the vehicle. Conversely, the influence on the danger factor can be smaller the further away the object is. Alternatively or additionally, the determined and / or estimated weight, shape, mass and / or material of the object can be incorporated accordingly.
[0015] The output of information can be based on the classification and, in particular, can be output to the driver, for example, to provide a warning. A "known" or "unknown" message can be output. This can occur in addition to or as an alternative to the (particularly visual and / or acoustic) output of the detected object. Output can occur via a control console, a head-up display, and / or a mobile device (such as a smartphone or smartwatch).
[0016] Alternatively or additionally, output information can be provided to a control unit. The control unit can then initiate a driving action, for example, at least partial (preferably complete) braking, evasive maneuvering, and / or adjusting the speed (in particular, a reduction and / or increase).
[0017] Alternatively or additionally, output information can be output to a method according to the second aspect. This can be combined with output to a driver, a cloud, and / or another algorithm, which can result in a more precise classification. Through (re-)training with the method according to the second aspect and / or an alternative method, the unknown object can thus become a known object. In other words, the artificial neural network can be readjusted and / or retrained.
[0018] Alternatively or additionally, output information can be sent to a cloud. This can be used for further classification (see below). This allows for a ground truth to be established for certain (rare) initially unknown objects.
[0019] Alternatively or additionally, the output information can be output via a data connection, in particular mobile data, Bluetooth, Wi-Fi, and / or the Internet. This allows classification by at least one, preferably several, other drivers and / or users, similar to a CAPTCHA process.
[0020] It may also be provided that a verification function is initiated. For example, the driver and / or an algorithm can determine whether the object is unknown (only to the existing artificial neural network) and / or, if applicable, what type of object it is.
[0021] It may also be planned to request updates and / or adjustments. This may be planned to ensure that other systems, especially those of the same design, have already classified the unknown object. This allows learning from the results of other systems.
[0022] It may also be provided to initiate a query as to whether readjustment, particularly by a driver, is possible and / or desirable for unknown objects. The driver can then classify the unknown object. This can be initiated, in particular, after a trip and / or before the next trip.
[0023] Within the scope of the invention, it may be advantageous that the classification is carried out based on an uncertainty value, wherein a known object is detected if the uncertainty value is below or identical to an uncertainty threshold, wherein an unknown object is detected if the uncertainty value is above an uncertainty threshold.
[0024] In addition or alternatively, the uncertainty value may be included in the hazard factor (see above). Thus, a high uncertainty value can lead to a comparatively high hazard factor. Conversely, a low uncertainty value can lead to a comparatively lower hazard factor.
[0025] It can be provided that, depending on the uncertainty value and / or the hazard factor, the control unit triggers a corresponding driving action, with the extent of the driving action preferably being dependent on the uncertainty value and / or the hazard factor. For example, an (emergency) braking action may be more severe with a higher uncertainty value and / or hazard factor, particularly because the underlying hazard and / or worst-case scenario may be (statistically) greater.
[0026] Within the scope of the invention, it is conceivable that if a known object is recognized in the image area during identification, further classification takes place after classification and / or output, wherein the known object is assigned to at least one known category.
[0027] This can also be implemented as or within the framework of a "perception task loss." Known objects can be divided into categories, for example, "road sign," "vehicle," "car," "truck," "bicycle," "pedestrian," and / or "wild animal." Further classification steps can be performed, for example, to perform an even finer classification into (sub-)categories and / or to extract information, e.g., by recognizing a speed limit on a road sign. Thus, the process can be structured in multiple stages, which can lead to better results overall due to the specific tasks of the respective components, which can be trained individually. Thus, a more robust artificial neural network can be provided to realize the above advantages (or reduce the discussed disadvantages).
[0028] Within the scope of the invention, it can be provided that the image area has a rectangular section of the input image and / or an object-specific section, wherein in particular an object-specific section is created based on the contours of the object.
[0029] A rectangular section allows for systematic scanning of the input image. This allows for efficient yet thorough examination of all areas of the input image. This can also be used for output, for example, when displaying the captured objects, especially visually.
[0030] Alternatively and / or additionally, an object-specific section can be used. This can create a particularly easily recognizable and / or intuitive display. Thus, a driver can identify the object and / or type of object based on the section, similar to how a glance at a tachometer provides an indication of the (approximate) current rpm without actually recording the exact number. This can be particularly helpful when a large number of objects appear simultaneously and / or in rapid succession.
[0031] The above object is further achieved according to a second aspect by a method for training an artificial neural network for detecting objects, in particular for a vehicle, comprising: - Providing an artificial neural network to be trained, - Providing training data for training the artificial neural network, wherein the training data comprises domain-specific images, in particular for the application area of the artificial neural network, - Generating synthetically generated outliers using the domain-specific images, - Training the artificial neural network based on the domain-specific images and the synthetically generated outliers.
[0032] The training method according to the second aspect can generate an artificial neural network which can advantageously be used for a method according to the first aspect.
[0033] The method according to the second aspect can be (temporarily) preceded by the method according to the first aspect. Thus, an initial training of the artificial neural network can take place before it is deployed.
[0034] In particular, the method for training can be used, in particular started, when the method for application (first aspect) detects an unknown object. In other words, readjustment can be performed. This can, in particular, also include data from other vehicles and / or vehicle fleets. This can transform a previously unknown object into a known object that can be recognized and / or classified in the future.
[0035] The application area can include autonomous and / or automated driving. This can advantageously involve the detection and / or differentiation of known and unknown objects, in particular to realize the above-mentioned advantages.
[0036] Domain-specific images can comprise, in particular, unmodified images from the corresponding domain. Domain-specific images can include known objects, which can in particular be known as such within the framework of a ground truth, which is preferably used within the framework of training. A domain can in particular be road traffic relevant for vehicles. For example, in the case of vehicles, the (input) images can be used which are collected from the environment (e.g., road traffic situations) during normal operation, e.g., during at least partially autonomous driving. These can be captured by a capture device, e.g., a camera. Accordingly, large, known data sets from many vehicles and / or an entire vehicle fleet can also be used.Preferably, the domain-specific images can be specific to a (planned) application location, in particular a vehicle, for example, country-specific and / or continent-specific. For example, images can be specific to road traffic in Europe, the USA, and / or China. This can allow for consideration of local conditions, such as left- or right-hand traffic, known and / or unknown objects (in particular vehicles, road signs, vegetation, animals), weather conditions, and / or visibility conditions (e.g., lighting, fog, smog).
[0037] The training can be based on at least partially unaltered training data. These can, for example, contain only known objects. Furthermore, synthetically generated outliers can be included, which in particular comprise (altered) training data into which synthetically generated outliers have been integrated. This can enable a more robust artificial neural network, especially in application, to advantageously realize the aforementioned advantages.
[0038] It is also preferably conceivable that the generation of the synthetically generated outliers is carried out by a diffusion model, wherein in particular the diffusion model uses structured text blocks to integrate synthetically generated outliers using the domain-specific images based on domain-foreign objects.
[0039] The diffusion model can also be referred to as a "diffusion inpainting model." The diffusion model can generate particularly realistic images that advantageously approximate and / or simulate potentially actually occurring unknown objects during the application of the artificial neural network. In particular, a higher degree of similarity can be achieved than with known methods. Accordingly, the artificial neural network can be trained particularly robustly and / or reliably, and preferably later deployed accordingly. When using the artificial neural network, this can reduce susceptibility to errors or false alarms (false positive results). Furthermore, it can prevent unknown objects from being detected or only poorly detected. This can increase security and / or robustness.Advantageously, this can prevent shutdown, making autonomous and / or at least partially automated driving (at all) completely possible.
[0040] In this case, objects from outside the domain can contain unexpected information which is particularly unlikely to be expected in the domain in question, e.g. in road traffic (e.g. a grizzly bear or a Ferris wheel). Objects from outside the domain can, at least in part, comprise unknown objects. In this case, it can be provided that these unknown objects are not “learned” as such. Preferably, it can be provided that the artificial neural network is trained to detect unknown objects (as a whole) in order to advantageously distinguish these from known objects as reliably as possible. This can advantageously lead to a more robust artificial neural network through training. During application, the artificial neural network can preferably receive an unexpected, unknown and / or improbable input image ora corresponding image area from it, can be better recognized (in particular classified as unknown or known). In particular, an outlier and / or unknown object can be recognized more reliably. This means that the artificial neural network can be used without an interruption, for example due to a malfunction. Furthermore, information that lies substantially outside the corresponding image section (in particular comprising the unknown object) can be used, since this information is relatively highly likely to be reliable and can in particular include known objects, which are preferably nevertheless recognized. This can advantageously allow substantially uninterrupted, or in particular at least less frequently interrupted, control by an autonomous driving system of a vehicle.For example, a detected immobile unknown object may cause a vehicle to simply navigate around it (avoid) instead of applying the brakes and / or shutting down.
[0041] Additionally or alternatively, it can be provided that further unknown objects are taken into account during training. Accordingly, for example, within the framework of the method according to the first aspect, an unknown object can be detected. This can be classified as unknown, for example by a human. Subsequently, in the method according to the second aspect, e.g., as part of a readjustment, the unknown object, in particular as a "real" outlier, can be used. This can lead to a further increase in robustness, especially compared to training based exclusively on objects from outside the domain.
[0042] In the simplest case, structured text blocks can consist of letters strung together. These can include descriptions, particularly machine-readable descriptions, which particularly relate to certain (particularly domain-alien) circumstances and / or objects. For example, a domain-alien object can include a grizzly bear. A structured text block could therefore read: "Bear runs towards the vehicle on the right side of the road." Based on the structured text block, the diffusion model can integrate the domain-alien object of a grizzly bear into a domain-specific image, e.g., a traffic situation, in order to generate a synthetically generated outlier. The structured text blocks can be compiled manually, particularly by humans. Alternatively or additionally, they can be linked to the domain-alien objects. It can also be provided that text blocks of the domain-specific images and / or objects (e.g.,"Bicycle approaching from the front at the right side of the road") can be combined and / or mixed with text blocks of domain-alien objects (e.g., "Grizzly bear," which could result in a structured text block: "Grizzly bear approaching from the front at the right side of the road"). This makes it possible, during training, to obtain an artificial neural network that is more robust and, in particular, less or not at all susceptible to the disadvantages mentioned above.
[0043] Within the scope of the invention, it is optionally possible for the generation to comprise at least a partial modification of the domain-specific images, and in particular of the non-domain objects.
[0044] This modification (also known as "image impainting") can result in particularly realistic integration into a domain-specific image. This can also enable better recognition and / or reliable discrimination (known / unknown) when using an artificial neural network.
[0045] For example, the non-domain objects can be subjected to transformation and / or morphing, such as translation, rotation, scaling, distortion, and / or coloring. This allows for particularly realistic integration into the domain-specific image.
[0046] In particular, the integration can comprise inserting the non-domain object into a domain-specific image. Preferably, only a portion of the domain-specific image can be replaced. Preferably, the non-domain object can represent a synthetically generated outlier in and / or with the domain-specific image. The non-domain object can be integrated into a domain-specific image, in particular a predefined one. This can comprise, for example, scaling, whereby, in particular, the size of the non-domain object is adjusted as realistically as possible. This can lead to a particularly realistic representation of the synthetically generated outlier.
[0047] The domain-specific images can comprise a number between 10 and 10^8 images, preferably between 100 and 10^7 images, preferably between 1000 and 10^6 images, particularly preferably between 10^4 and 10^5 images.
[0048] Furthermore, it can be provided within the scope of the invention that the generation of synthetically generated outliers is carried out based on a mask of the domain-specific images, wherein in particular an integration of the domain-foreign objects is carried out within the mask.
[0049] In this case, it can be provided that the mask is placed on the domain-specific image(s), or that the mask produces a section of the domain-specific image, wherein the section is preferably at least partially replaced by the non-domain object. In this case, it can be provided that the mask is used as a surface section (e.g., rectangular or corresponding to a shape) for an area within which a non-domain object is inserted, in order to advantageously create a synthetically generated outlier.
[0050] With regard to the present invention, it is conceivable that the generation, in particular based on a mask, comprises the assignment of features, in particular to the domain-alien objects and / or the synthetically generated outliers. Alternatively or additionally, it can also be provided that the features are assigned to the known objects and / or unknown objects. Based on the features, a classifier can then be trained. This can be carried out based on a (known) ground truth of the corresponding images. The classifier can be trained to recognize known objects and / or, in particular, to distinguish them from unknown objects.
[0051] Furthermore, it is conceivable that the synthetically generated outliers, and in particular the non-domain objects, are based at least partially on at least one publicly available image dataset. The non-domain objects may contain at least one unknown object.
[0052] In the simplest case, the domain-extraneous objects are at least partially taken from a publicly available image dataset, for example, the CIFAR-10 dataset. This allows a particularly broad spectrum of domain-extraneous objects to be used, which can advantageously lead to greater robustness of the artificial neural network, especially after training.
[0053] Within the scope of the invention, it may be advantageous that the training comprises the optimization of an uncertainty function of a classifier: - where known objects, in particular based on domain-specific images, are assigned a low uncertainty value by the uncertainty function, - where unknown objects, in particular based on objects outside the domain and / or synthetically generated outliers, are assigned a high uncertainty value by the uncertainty function.
[0054] Optimization can involve adaptation or training. The classifier can then perform a classification based on the uncertainty value through training. This can, for example, lead to classification into the categories "known" (in particular, known objects) and "unknown" (in particular, unknown objects).
[0055] The uncertainty function can correspond to a so-called cost function (especially a "loss function"), which is optimized during training. Features of the images used for training can be learned, particularly to enable and / or train for differentiation.
[0056] It can be provided, in particular subsequently, that further classification is trained, which enables a more precise specification (in particular division into categories) of known objects. It can furthermore be provided that the method according to the first aspect is carried out, in particular after the method according to the second aspect. In other words, the artificial neural network can be applied to the training. In this case, known objects can in particular be differentiated from unknown objects, in particular by classifying image regions. It can also be provided that known objects are divided into more precise categories, preferably by further classification (e.g. vehicle, bicycle, pedestrian, as well as their actions, e.g. "standing", "coming towards", etc.).
[0057] The above object is further achieved according to a third aspect by a computer program product according to the invention, comprising instructions which, when the computer program product is executed by a computer, cause the computer to implement the method according to the first and / or second aspect.
[0058] This results in the same advantages with regard to a computer program product according to the invention as have already been described with regard to a method according to the invention.
[0059] The above object is further achieved according to a fourth aspect by a computer-readable data carrier in which instructions are stored which, when executed by a computer, cause the computer to carry out the method according to the first and / or second aspect.
[0060] This results in the same advantages with regard to a computer-readable data carrier according to the invention as have already been described with regard to the first, second and / or third aspect.
[0061] The above object is further achieved according to a fifth aspect by a control unit comprising a computing unit and a memory unit in which instructions are stored which, when at least partially executed by the computing unit, carry out a method according to the first and / or second aspect.
[0062] The control unit can be designed to enable (at least partially) autonomous and / or automated driving, in particular of the vehicle.
[0063] This results in the same advantages with regard to a control unit according to the invention as have already been described with regard to the first, second, third and / or fourth aspect.
[0064] The above object is further achieved according to a sixth aspect by a vehicle comprising a control unit according to the fifth aspect.
[0065] The vehicle can be designed to enable (at least partially) autonomous and / or automated driving.
[0066] This results in the same advantages with regard to a vehicle according to the invention as have already been described with regard to the first, second, third, fourth and / or fifth aspect.
[0067] Further advantages, features, and details of the invention will become apparent from the following description, in which several embodiments of the invention are described in detail with reference to the drawings. The features mentioned in the claims and in the description may be essential to the invention individually or in any combination. These schematically show Fig. 1 a method for detecting objects, Fig. 2 a method for training an artificial neural network, Fig. 3 a vehicle, and Fig. 4 a method for training an artificial neural network.
[0068] In the following figures, identical reference numerals are used for the same technical features, even for different embodiments.
[0069] Fig. 1 shows a method 100 for detecting objects 10, 20, in particular for a vehicle 300, comprising: - Providing 110 an artificial neural network 1, - capturing 120 an input image 121 by a capturing device 301, - Providing 130 the input image 121 to the artificial neural network 1, - Identifying 140 an image area 141 in the input image 121 by the artificial neural network 1, wherein at least one object 10, 20 is recognized in the image area 141, - Classifying 150, by a classifier 241, the image area 141 by the artificial neural network 1 to obtain a classification 151, wherein in the classification 151 at least one known object 10 and / or unknown object 20 is recognized in the image area 141, - Outputting 160 an output information 161 based on the classification 151.
[0070] Fig. 2 shows a method 200 for training an artificial neural network 1 for detecting objects 10, 20, in particular for a vehicle 300, comprising: - Providing 210 an artificial neural network 1 to be trained, - Providing 220 training data 221 for training the artificial neural network 1, wherein the training data 221, in particular for the application area of the artificial neural network 1, comprise domain-specific images 222, - generating 230 synthetically generated outliers 223 using the domain-specific images 222, - Training 240 of the artificial neural network 1 based on the training data 221 and the synthetically generated outliers 223.
[0071] Fig. 3 shows a vehicle 300 comprising a detection device 301, which, in particular, receives input images comprising a known object 10 and / or an unknown object 20. The vehicle has a control unit ECU comprising a computing unit CU and a data storage device MU. The detection device 301 can be connected to the control unit via a data connection, in particular to transmit input images to the control unit, preferably the computing unit CU. The vehicle 300, in particular the control unit ECU, can have an artificial neural network 1.
[0072] Fig. 4 shows, by way of example, a method 200 for training an artificial neural network 1. Structured text blocks 225 can be used to generate synthetic outliers 223, preferably in the domain-specific images 222, for example based on objects 226 that are foreign to the domain. This can be done by a diffusion model 224, which preferably receives and / or processes the structured text blocks. Images with known objects 10 and images with unknown objects 20 can therefore be provided as training data 221 used. Images with unknown objects 20 can be generated by the diffusion model 224. Images with known objects 10 can be provided based on domain-specific images 222 (a street is indicated as an example). An assignment 229 of features can be carried out, which, for example, can be Fig. 4 are represented as circles, with white circular areas symbolizing the features of synthetically generated outliers 223 or unknown objects 20, and black circular areas representing the features of known objects 10.
[0073] The artificial neural network 1 can then be trained 240. This can include optimizing an uncertainty function 242 of a classifier 241. In other words, the classifier 241 can be trained to recognize unknown objects 20 and / or to distinguish them from known objects 10.
[0074] A further classification can also be trained, which allows a more precise specification of the known objects 10. This is shown, merely as an example, in the lower right area of Fig. 4 shown.
[0075] It may further be provided that method 100 is carried out, particularly following method 200. In other words, the artificial neural network 1 may be applied to the training. In particular, known objects 10 can be distinguished from unknown objects 20, in particular by classifying image regions 150. It may also be provided to classify known objects 10 into more precise categories, preferably by further classification 170. List of reference symbols 1 artificial neural network 10 known objects 11 Category 20 unknown object 30 Uncertainty value 31 Uncertainty threshold 100 procedures (application) 110 Provision 120 Capturing an input image 121 Input image 130 Providing the input image 140 Identifying an image area 141 image area 142 rectangular section of the image area 143 object-specific section of the image area 150 Classifying the image area (known / unknown) 151 Classification 160 Outputting output information 161 Output information 170 further classification (of known objects) 171 Assigning known objects to categories 200 procedures (training) 210 Providing an artificial neural network to be trained 220 Providing training data 221 training data 222 domain-specific images 223 synthetically generated outliers 224 Diffusion model 225 structured text blocks 226 non-domain objects 227 Change 228 Mask 229 Assigning features 230 Generating synthetically generated outliers 240 Training the artificial neural network 241 Classifier 242 Uncertainty function 242.1 low uncertainty value 242.2 high uncertainty value 300 vehicles 301 Detection device CU computing unit ECU control unit MU storage unit
Claims
[1] Method (100) for detecting objects (10, 20), in particular for a vehicle (300), comprising: - Providing (110) an artificial neural network (1) according to one of claims 5 to 11, - capturing (120) an input image (121) by a capturing device (301), - providing (130) the input image (121) to the artificial neural network (1), - identifying (140) an image area (141) in the input image (121) by the artificial neural network (1), wherein at least one object (10, 20) is recognized in the image area (141), - Classifying (150), by a classifier (241), the image area (141) by the artificial neural network (1) to obtain a classification (151), wherein in the classification (151) at least one known object (10) and / or unknown object (20) is recognized in the image area (141), - Outputting (160) output information (161) based on the classification (151). [2] Method (100) according to claim 1, characterized by that the classification (150) is carried out based on an uncertainty value (30), wherein a known object (10) is detected if the uncertainty value (30) is below or identical to an uncertainty threshold value (31), wherein an unknown object (20) is detected if the uncertainty value (30) is above an uncertainty threshold value (31). [3] Method (100) according to claim 1 or 2, characterized by that, if a known object (10) is recognized in the image area (141) during identification (140), a further classification (170) is carried out after classification (150) and / or output (160), wherein an assignment (171) of the known object (10) to at least one known category (11) is carried out. [4] Method (100) according to one of the preceding claims, characterized by that the image area (141) has a rectangular section (142) of the input image (121) or an object-specific section (143), wherein in particular an object-specific section (143) is created based on the contours of the object (10, 20). [5] Method (200) for training an artificial neural network (1) for detecting objects (10, 20), in particular for a vehicle (300), comprising: - providing (210) an artificial neural network (1) to be trained, - Providing (220) training data (221) for training the artificial neural network (1), wherein the training data (221), in particular for the application area of the artificial neural network (1), comprise domain-specific images (222), - generating (230) synthetically generated outliers (223) using the domain-specific images (222), - Training (240) the artificial neural network (1) based on the domain-specific images (222) and the synthetically generated outliers (223). [6] Method (200) according to the preceding claim, characterized by that the generation (230) of the synthetically generated outliers (223) is carried out by a diffusion model (224), wherein in particular the diffusion model (224) uses structured text blocks (225) to integrate synthetically generated outliers (223) using the domain-specific images (222) based on domain-foreign objects (226). [7] Method (200) according to the preceding claims 5 or 6, characterized by that the generation (230) comprises an at least partial modification (227) of the domain-specific images (222), and in particular of the domain-foreign objects (226). [8] Method (200) according to one of the preceding claims 5 to 7, characterized bythat the generation (230) of synthetically generated outliers (223) is carried out based on a mask (228) of the domain-specific images (222), wherein in particular an integration of the domain-foreign objects (226) is carried out within the mask (228). [9] Method (200) according to one of the preceding claims 5 to 8, characterized by that the generation (230), in particular based on a mask (228), comprises the assignment (229) of features, in particular to the domain-alien objects (226) and / or the synthetically generated images (223). [10] Method (200) according to one of the preceding claims 5 to 9, characterized by that the synthetically generated outliers (223), and in particular the domain-alien objects (226), are based at least in part on at least one publicly available image dataset. [11] Method (200) according to one of the preceding claims 5 to 10, characterized bythat the training (240) comprises optimizing an uncertainty function (242) of a classifier (241): - wherein known objects (10), in particular based on domain-specific images (222), are assigned a low uncertainty value (242.1) by the uncertainty function (242), - wherein unknown objects (20), in particular based on domain-foreign objects (226) and / or synthetically generated outliers (223), are assigned a high uncertainty value (242.2) by the uncertainty function (242). [12] A computer program product comprising instructions which, when executed by a computer, cause the computer to implement the method (100, 200) according to any one of the preceding claims. [13] Computer-readable data carrier in which instructions are stored which, when executed by a computer, cause the computer to carry out the method (100, 200) according to one of the preceding claims 1 to 11. [14] Control unit (ECU), comprising a computing unit (CU) and a memory unit (MU) in which instructions are stored which, when at least partially executed by the computing unit (CU), carry out a method (100, 200) according to one of the preceding claims. [15] Vehicle (300) comprising a control unit (ECU) according to the preceding claim.
Citation Information
Patent Citations
System for detection and management of uncertainty in perception systems, for new object detection and for situation anticipation
WO2022243337A2