Computer-implemented method for generating annotations for an unannotated image for training a neural network

DE102024202152A1Pending Publication Date: 2025-09-11ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102024202152
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2025-09-11

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method for generating annotations for an unannotated image for training a neural network, wherein the method is carried out using an annotated image dataset, wherein the image dataset comprises a plurality of annotated images, wherein the method comprises the following steps: - determining an annotation mask (S10, S12) for each annotated image, wherein the annotation mask describes an area in the respective annotated image which is mapped to the non-annotated image by a transformation; - determining a mask weight (S10, S14) for each annotation mask; and - generating at least one annotation (S16) for the unannotated image using the annotation masks, their mask weights and the annotations of the image data set.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to the generation of annotations for images and image data to be used for training a neural network.

[0002] A common challenge in many machine learning applications is the problem of insufficient training data when training AI models. A training dataset that is too small can lead to several problems.

[0003] If a KL model is trained on a dataset that is too small, there is a risk of overfitting. This means that the model memorizes the existing data too much and is unable to generalize properly to new, unknown data. Essentially, the model overfits the noise and outliers of the limited training dataset.

[0004] A training dataset that is too small can further limit the model's ability to recognize patterns and relationships in the data. The model may have difficulty learning relevant features and applying them to new data, leading to inaccurate predictions. In many real-world applications, data naturally varies. Therefore, if the training dataset is too small, the model may have difficulty adequately covering and capturing the diversity of this data. In particular, if rare events occur in the data, they may be underrepresented in a small training dataset. The model may not be able to correctly detect and predict these rare events.

[0005] The availability of sufficient data is crucial for training a neural network, especially for demanding tasks such as image classification or segmentation. If the amount of data is too small, various problems can arise that significantly impact the network's performance and capabilities.

[0006] In addition to the problems already mentioned, too little data for training a neural network on an image processing task can lead to the training not representing the full range of possible input data. The network could then have difficulty dealing with different scenarios, lighting, viewing angles, etc.

[0007] If certain classes or categories are underrepresented in the images, the network may have difficulty correctly detecting or segmenting them. This leads to a bias in model performance. Furthermore, neural networks benefit from the ability to extract relevant features from images. With insufficient data, important features may not be sufficiently learned, which impacts the network's performance.

[0008] Furthermore, a lack of diversity in the training data can make the network vulnerable to perturbations that may arise in real-world use cases. The network might have difficulty dealing with noise, bias, or other unforeseen variables.

[0009] To address these problems, it is important to use high-quality and sufficient training data to ensure robust and generalized performance of the neural network. State of the art

[0010] The "Refign" method for training neural networks for semantic image segmentation is known from "Refign: Align and Refine for Adaptation of Semantic Segmentation to Adverse Conditions," Bruggemann et al., 2023. The goal is to develop a model that learns from annotated images from a source domain S and unannotated images from the target domain T, as well as a reference domain R. The model is designed to predict semantic segmentation maps for images in the target domain. During training, only ground-truth segmentation annotations for the source domain are available.

[0011] The "Refign" method is an extension of self-learning-based methods for unsupervised domain adaptation (UDA). It describes a training algorithm that uses a pre-trained alignment module and a non-parametric refinement module. The alignment module ensures that the target and reference images are spatially accurately aligned. The refinement module improves the class probabilities of the target images using the aligned reference probabilities and a confidence map. The goal is to use high-quality reference predictions to self-train the data to the target domain.

[0012] The alignment module is based on an extended Warp Consistency (WarpC) approach. The refinement module uses adaptive pseudo-label refinement using confidence values ​​and a strategy for large static classes.

[0013] The invention is therefore based on the object of proposing a method with which annotations for non-annotated images can be generated from annotated image data.

[0014] The problem is solved by the subject matter of the independent claims. Disclosure of the invention

[0015] According to a first aspect of the invention, this object is achieved by a computer-implemented method for generating annotations for an unannotated image for training a neural network. The method is carried out using an annotated image dataset, wherein the image dataset comprises a plurality of annotated images.

[0016] The procedure includes the following steps: - Determining an annotation mask for each annotated image, wherein the annotation mask describes an area in the respective annotated image that is mapped to the non-annotated image by a transformation; - Determining a mask weight for each annotation mask; and - Generating at least one annotation for the unannotated image using the annotation masks, their mask weights and the annotations of the image dataset.

[0017] The annotated image dataset is a collection of images containing annotations. For example, the images may contain data that provides information about the scenes and objects depicted in the images. This could include position, distance, and other data.

[0018] For example, the image dataset may include the images of a video sequence taken with a camera moving along a street.

[0019] The unannotated image should be similar to the images in the annotated image dataset.

[0020] For example, the unannotated image may have been recorded under different conditions, such as a different time of day, different weather, or similar, than the images in the annotated image dataset.

[0021] The method can be used to extend annotations from already annotated sources to unannotated data, thus creating a larger database for training a neural network. The method can therefore be used in particular for annotating images that are similar to one another.

[0022] For example, if a photograph of a person driving through a street has been annotated one morning in sunny weather, images of the same street from other times of day and / or in different weather conditions can be automatically annotated using the already annotated photograph.

[0023] The images should be similar for the process. This means that the unannotated images should contain scenes or elements that are already included in the annotated image dataset. This could be an entire street, as mentioned in the example above. However, the similarity can also be lower, for example, if the annotated image dataset contains images of elements, in particular certain vehicles, vehicle types, buildings, building parts such as windows and doors, traffic signs, people, vegetation, or other everyday objects. Images of a motorway section, for example, can be used to annotate images of other motorway sections, provided the motorway sections are similar in some way. This can be the case, for example, if the number of lanes is the same, there are no construction sites, and / or the lighting conditions are at least similar.

[0024] The procedure comprises three steps, two of which are performed multiple times.

[0025] First, an annotation mask is created for each previously annotated image. The annotation mask describes a region in the unannotated image—which may also include the entire unannotated image—that can be recognized with sufficient similarity in the respective annotated image.

[0026] The image sections don't have to be congruent. A certain degree of transformation is permissible and, to a certain extent, to be expected. If the images were congruent, creating annotations would provide no added value for training the neural network.

[0027] In particular, the transformation can be a Euclidean transformation of the images or individual image sections of the annotated image. The Euclidean transformation is a mathematical mapping that preserves certain geometric properties in Euclidean space. These transformations include translations (translations), rotations (rotations), reflections (symmetries), and scalings (stretchings). Applying a Euclidean transformation to objects in space changes their position, orientation, and size while preserving basic geometric properties such as distances between points, angles between lines, and parallel relationships.

[0028] For example, the unannotated image shows a street scene. A road runs between two buildings. A vehicle is parked on one side, obscuring the view of the building behind it.

[0029] The annotated images show the street at a different time of day, meaning the lighting conditions are slightly different and the vehicle is not in front of the building in question.

[0030] The annotation masks now describe the sections in the annotated images that are recognized in the unannotated image. For example, this could be the side of the street where the vehicle is not visible, and on the other side, at least the part of the building that is not obscured by the vehicle in the unannotated image. The sky between the buildings is not included in the annotation mask because it has a different color due to the different time of day.

[0031] In the entire annotated image dataset, there may also be images showing vehicles that are similar to the vehicle in the unannotated image.

[0032] When merging the annotation masks from the annotated images, at least one annotation for the not yet annotated image can receive the annotations from the annotated images where similar elements can be found.

[0033] The annotations are weighted with a mask weight. The mask weight can be used to determine which annotation from which annotated image should be used to generate the annotation for the unannotated image. Various metrics can be used, such as image quality, measures of the quality of the annotations in the annotated images, and / or statistical information.

[0034] The mask weight can affect the entire image or parts of it. Annotated images in which no element from the unannotated image is found can receive a total mask weight of 0. Images in which elements can be detected receive a high weight factor for the positions of the corresponding element(s), and a lower weight factor in other areas. The transitions can be smooth or discrete.

[0035] In the final step, the annotation masks and mask weights of the annotated images are combined, and an annotation for the unannotated image is generated. For this purpose, the entire set of annotated images is considered for each image area of ​​the unannotated image. The annotation masks, if available, indicate which sections of the annotated images the annotations can be transferred from to the unannotated image. The mask weights determine how strong the annotation of an individual image should be compared to the other annotated images.

[0036] For example, two annotated images may provide different annotations for the same region in the unannotated image. The mask weights can be used to determine which of the annotated images should be used to annotate the region in the unannotated image.

[0037] This proposes a method for transferring annotations from already annotated image data to unannotated images. This solves the problem of the invention.

[0038] In one embodiment, the mask weight scales with the extent of the mapping of the respective annotated image to the unannotated image.

[0039] A pixel in the annotated image can be mapped to a pixel of the same color in the unannotated image. The mapping does not have to be unique. However, it is quite possible that a pixel is mapped out of context. For example, if a blue vehicle and a blue sky are present in the annotated image, a blue vehicle in the unannotated image can receive the annotation from the blue vehicle or the blue sky in the annotated image. It is very likely that every pixel in the unannotated image will find a counterpart in the annotated image in this way, which complicates or even hinders the correct generation of annotations.

[0040] To account for this, the mask weight can consider the size or extent of the mapping of the annotated image to the unannotated image. For example, a particularly contiguous area can be determined that is mapped for annotation. The weight for the annotation could then be scaled proportionally to the area of ​​the mapping. This would mean that the annotation would be weighted more heavily the more evidence for the annotation already exists in the annotated image.

[0041] By linking the extent of the figure and the weight for the annotation, the accuracy of generating annotations for the unannotated image is increased because the annotations are generated with more evidence based on the figure and its extent.

[0042] In one embodiment, the annotation mask for all or part of the pixels in the unannotated image comprises a transformation vector originating from a pixel in the respective annotated image.

[0043] In this embodiment, the annotation mask can be represented by a tensor. The tensor consists of a matrix where each pixel of the unannotated image contains entries that are 0 if no matching region was found in the annotated image, or a vector that includes the corresponding pixel or pixel range including the annotation for that region.

[0044] This embodiment allows an annotation to be assigned to the pixels in the unannotated image in a particularly simple manner.

[0045] In one embodiment, the region from the annotated image mapped to the unannotated image comprises at least the same color value, brightness value and / or contrast as the corresponding region in the unannotated image.

[0046] Images consist of pixels that have different properties, such as their color, brightness, or, if surrounding pixels are taken into account, their contrast. In addition, additional parameters can be defined that can be used to search for corresponding pixels in the unannotated image using the pixels from the annotated image.

[0047] The term "same" is to be interpreted broadly in this embodiment. Two pixels have the "same" property—color value, brightness, etc.—if they are within a certain tolerance.

[0048] For example, if brightness is defined as a value between 0 for black and 1 for white, two pixels with brightness values ​​of 0.552 and 0.514 can have the same brightness as long as their difference is below a threshold. In the example above, a threshold of 0.03 would be sufficient to consider the pixels' brightnesses the same.

[0049] The threshold can be chosen depending on many factors. For example, the quality of the annotated images, the quality of the unannotated image, the task of the neural network for which the annotations are being trained, and / or other factors and conditions can be taken into account.

[0050] In one embodiment, the at least one annotation comprises a plurality of annotations, wherein different subregions of the unannotated image have different annotations.

[0051] An image can contain multiple elements that are annotated differently. A simple example is an image of a traffic scene with several road users recognizable in the image. The road users can have different annotations, such as the type of road user, their position, and / or their speed.

[0052] This embodiment advantageously allows different annotations to be assigned to different elements in the unannotated image.

[0053] In one embodiment, the at least one annotation for the unannotated image or a portion of the unannotated image is determined using a mapping probability, wherein the mapping probability is a measure of the quality of the mapping of the annotated image to the unannotated image.

[0054] This embodiment can be used in particular together with the embodiment described above in which the mask weights scale with the extent of the image.

[0055] The mapping probability can be determined as a measure of the quality of the mapping, for example, by subjecting the transformation of the mapped element to a plausibility check using its annotation. In particular, a Euclidean transformation can be strong or weak. The stronger the transformation, the more the mapping of the element is deformed compared to its representation in the annotated image. The greater the deformation, the less likely it is that the annotation from the annotated image can be transferred to the unannotated image. Instead, it may be that the element in the annotated image does not correspond to the element in question in the unannotated image.

[0056] This advantageously improves the accuracy of annotation generation for the unannotated image, since difficult or at least questionable images do not have less annotation weight.

[0057] In one embodiment, the method is performed for a plurality of unannotated images.

[0058] This embodiment can advantageously be used to annotate an entire, unannotated image data set.

[0059] The unannotated image dataset can, for example, be a video sequence consisting of unannotated images. For example, if a video sequence is recorded and annotated for training the neural network, and it turns out during training that the amount of training data is insufficient, the method can be applied to a subsequently recorded video to reduce or automate the effort required for the annotation of the subsequently recorded images.

[0060] In one embodiment, the annotated image dataset comprises images of a first domain and wherein the unannotated image is an image of a second domain.

[0061] In this embodiment, in particular, sufficient training data is available from the first domain, but not from the second domain.

[0062] A domain is characterized by certain characteristics, properties, and rules that are typical for it. In the context of the present invention, different domains may have different types of data, features, and patterns. For example, the data may have been recorded with different sensors or measuring devices, or different settings may have been used during recording, or different conditions may have prevailed.

[0063] The proposed method allows annotations from the first domain to be transferred to the second domain. This does not necessarily apply to all annotations in the images of the first domain. For example, only individual data may be transferable.

[0064] In one example, a video of a traffic scene was recorded during the day. A second, unannotated video was recorded at night. The annotations in the video of the vehicles visible in the images, such as their positions and speeds, may be transferable, but the data on the time of day and / or weather conditions cannot be transferred from the annotated images to the unannotated images. Other methods may be more suitable for this purpose.

[0065] In one embodiment, the method further comprises the step of validating the at least one generated annotation by processing the unannotated image using a neural network and comparing the processing result with the at least one generated annotation.

[0066] Preferably, several neural networks are trained with the annotated image dataset, whereby the decision as to whether the generated annotations are valid is then made by majority vote.

[0067] The neural network used for this embodiment is to be distinguished from the neural network for which the unannotated images are to be annotated. The neural network used in this embodiment is preferably trained to assign the annotations present in the annotated images. Later, when the neural network for which the method proposed here is applied is trained, additional training data with additional annotations can be used. Thus, the architecture of the neural network used for this embodiment can be kept very simple, so that even the comparatively few images of the annotated image dataset are sufficient for its training.

[0068] In one embodiment, the neural network was trained using the annotated image dataset.

[0069] Advantageously, training the neural network with the annotated image dataset can simplify the validation of the generated annotations. The neural network only knows the annotations used in the annotated image dataset. Analogous to the need-to-know principle, the neural network can therefore be specialized for its task, validation.

[0070] In one embodiment, the at least one generated annotation is discarded if the comparison of the generated annotation with the processing result results in a difference that is greater than a defined threshold.

[0071] The objective of the invention is to generate annotations for unannotated images. However, this must not be done at the expense of the data's authenticity. Generating incorrect annotations can be more damaging than not having sufficient data to train a neural network.

[0072] In a further aspect, the invention relates to a computer program with program code for carrying out a method as described above when the computer program is executed on a computer.

[0073] In a further aspect, the invention relates to a computer-readable data carrier with program code of a computer program for carrying out a method as described above when the computer program is executed on a computer.

[0074] In a further aspect, the invention relates to a system for generating annotations for an unannotated image, wherein the system is designed to carry out a method as described above.

[0075] In summary, the present invention provides a method for generating annotations for an unannotated image, a computer program, a computer-readable program code, and a system for carrying out the method.

[0076] The described designs and further training courses can be combined as desired.

[0077] Further possible embodiments, developments and implementations of the invention also include combinations of features of the invention described previously or below with regard to the embodiments that are not explicitly mentioned. Short description of the drawings

[0078] The accompanying drawings are intended to provide a further understanding of embodiments of the invention. They illustrate embodiments and, in conjunction with the description, serve to explain principles and concepts of the invention.

[0079] Other embodiments and many of the aforementioned advantages will become apparent upon review of the drawings. The elements illustrated in the drawings are not necessarily drawn to scale.

[0080] It shows: Fig. 1 schematically shows the sequence of the method according to one embodiment.

[0081] In the figures of the drawings, the same reference symbols designate the same or functionally equivalent elements, parts or components, unless otherwise stated.

[0082] Fig. 1 shows schematically the sequence of the method according to an embodiment.

[0083] The method begins with step S10, in which an annotation mask and a mask weight are determined for each image of the annotated image data set.

[0084] In step S12, the annotation mask is determined. To do this, an attempt is made to map the respective annotated image or parts of it onto the unannotated image. The elements present in the annotated images can be transformed in various ways, in particular stretched, compressed, or scaled. Those parts of the annotated image for which this is successful are masked.

[0085] In step S14, a mask weight is determined for each annotation mask. The mask weight indicates how heavily the annotation mask or its regions are used in determining the annotation for the unannotated image.

[0086] Preferably, the mask weight is determined in such a way that the size or extent of the contiguous annotated regions is taken into account in the annotation mask. If these regions are larger, more pixels can be mapped to the corresponding region of the unannotated image, thereby making the assumption that the annotation is transferable from one image to the other more robust.

[0087] Once an annotation mask and a mask weight have been determined for all images in the annotated image dataset, an annotation for the unannotated image can be generated from this in step S16. The generated annotation is composed of the annotations of the corresponding regions in the annotated images. The generated annotation can apply to the entire, unannotated image and / or a sub-region thereof. For example, the annotation can include information about a weather condition, a season, and / or a time of day and would thus be valid for the entire image. Furthermore, the annotation can relate to individual objects recognizable in the image, such as pedestrians, vehicles, traffic signs, or scenery. These annotations would be limited to the recognized objects. When generating the annotation in step S16, multiple annotations can therefore also be generated. QUOTES CONTAINED IN THE DESCRIPTION

[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited non-patent literature

[0000] Refign: Align and Refine for Adaptation of Semantic Segmentation to Adverse Conditions", Bruggemann et al., 2023

[0010]

Claims

[1] Computer-implemented method for generating annotations for an unannotated image for training a neural network, wherein the method is carried out using an annotated image data set, wherein the image data set comprises a plurality of annotated images, the method comprising the following steps: - determining an annotation mask (S10, S12) for each annotated image, wherein the annotation mask describes an area in the respective annotated image which is mapped to the non-annotated image by a transformation; - determining a mask weight (S10, S14) for each annotation mask; and - generating at least one annotation (S16) for the unannotated image using the annotation masks, their mask weights and the annotations of the image data set. [2] The computer-implemented method of claim 1, wherein the mask weight scales with the extent of the mapping of the respective annotated image to the unannotated image. [3] A computer-implemented method according to any one of the preceding claims, wherein the annotation mask for all pixels or a portion of the pixels in the unannotated image comprises a transformation vector originating from a pixel in the respective annotated image. [4] Computer-implemented method according to one of the preceding claims, wherein the region from the annotated image mapped onto the unannotated image comprises at least the same color value, brightness value and / or contrast as the corresponding region in the unannotated image. [5] Computer-implemented method according to one of the preceding claims, wherein the at least one annotation comprises a plurality of annotations, wherein different subregions of the unannotated image have different annotations. [6] Computer-implemented method according to one of the preceding claims, wherein the at least one annotation for the unannotated image or a portion of the unannotated image is determined using a mapping probability, wherein the mapping probability is a measure of the quality of the mapping of the annotated image to the unannotated image. [7] A computer-implemented method according to any one of the preceding claims, wherein the method is performed for a plurality of unannotated images. [8] A computer-implemented method according to any one of the preceding claims, wherein the annotated image data set comprises images of a first domain and wherein the unannotated image is an image of a second domain. [9] A computer-implemented method according to any one of the preceding claims, the method further comprising the step of: - Validating the at least one generated annotation by processing the unannotated image using a neural network and comparing the processing result with the at least one generated annotation. [10] The computer-implemented method of claim 9, wherein the neural network was trained using the image dataset. [11] Computer-implemented method according to one of claims 9 or 10, wherein the at least one generated annotation is discarded if the comparison of the generated annotation with the processing result results in a difference that is greater than a defined threshold. [12] A computer program comprising program code for carrying out a method according to any one of the preceding claims when the computer program is executed on a computer. [13] A computer-readable data carrier comprising program code of a computer program for carrying out a method according to any one of claims 1 to 11 when the computer program is executed on a computer. [14] A system for generating annotations for an unannotated image, the system being configured to carry out a method according to any one of claims 1 to 11.