Method and apparatus for cleaning image data sets
By cleaning the image dataset with an intersection-over-union (IoU) filter and using an IoU threshold to identify and remove redundant images, the redundancy problem in the image dataset is solved, and the training and recognition accuracy of machine learning models in automated unpacking machines and autonomous vehicles is improved.
Patent Information
- Application Number
- CN202510663349.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-24
- Filing Date
- 2025-05-22
- Publication Date
- 2025-11-25
AI Technical Summary
Existing technologies struggle to effectively remove redundancy from image datasets, especially in automated unpacking machines and autonomous vehicles, impacting the training and recognition accuracy of machine learning models.
Intersection-over-Union-Filters are used to clean the image dataset. Redundant images are identified and removed by comparing the Intersection-over-Union (IoU) ratio between images. In particular, key images are compared with the remaining images, and a threshold is set to determine the degree of redundancy.
It effectively reduces semantic redundancy in image datasets, improves the training and recognition accuracy of machine learning models, and enhances the quality of datasets, especially in scenarios with static backgrounds and dynamic foregrounds.
Smart Images

Figure CN121010844A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to methods and apparatus for cleaning image datasets used to train and / or validate and / or test machine learning models. The invention also relates to a method for training and / or validating and / or testing machine learning models that can be used for the classification and / or segmentation of image data, particularly in automated unpacking machines and / or vehicles with at least one autonomous driving capability and / or automated optical inspection. Background Technology
[0002] In automation technology, the development of automated package unpacking machines presents a challenge, primarily due to the need to handle large volumes of data. This data typically originates from camera data streams, which rapidly and continuously provide a large number of images. One issue is dealing with the inherent redundancy in these image sequences. Therefore, semantic redundancy filtering is desired to develop efficient and accurate content recognition system prototypes. This technique helps identify and eliminate unnecessary repetitions in the data, stemming not only from sequentiality but also from overlapping content features within the images.
[0003] Semantic redundancy filtering, previously primarily used in image classification, now employs methods such as agglomerative clustering in latent spaces. This approach allows for the analysis of the importance and information content of each data sample within the context of the entire dataset, and corresponding weighting. Such techniques are crucial for reducing data volume without losing relevant information.
[0004] Furthermore, methods for handling inter-pixel redundancy are known, such as methods for buffer allocation, which aim to minimize pixel-level overlap. Additionally, content protection techniques such as perceptual hashing and depth-aware hashing are also known, designed to protect the key features of an image while remaining independent of visual changes such as rotation, scaling, or compression.
[0005] Furthermore, nonmaximum suppression, a technique known in object recognition, is typically combined with the analysis of validation metrics that consider different overlap thresholds at a fixed classification confidence level. These techniques are particularly useful for neural networks applied to single-image object recognition and enable accurate object classification during training on selected regions of interest.
[0006] Although some methods are already known, there is still potential for further development. Summary of the Invention
[0007] Therefore, the objective of this invention is to provide an improved method and / or apparatus in this respect.
[0008] The task is solved by a method according to the features of claim 1. The task is also solved by a device according to the features of claim 10.
[0009] Based on the first aspect, a method for cleaning image datasets used to train and / or validate and / or test machine learning models is proposed, the method comprising the following steps: - Provides image datasets with a large number of images; - By applying Intersection-over-Union-Filters, a large number of images, especially the predetermined contrast images, are compared with at least a portion of the remaining images in the large number of images; Based on this comparison, at least one redundant image relative to the comparison image is determined from at least a portion of the remaining images in the large number of images; and - Clean up the image dataset by removing at least one redundant image from the large number of images.
[0010] It should be understood that the steps and further optional steps according to the invention need not be performed in the order shown, but may be performed in other orders. Furthermore, further intermediate steps may be provided. Each step may also include one or more sub-steps without departing from the scope of the method according to the invention.
[0011] According to the second aspect, an apparatus for cleaning an image dataset used to train and / or validate and / or test a machine learning model is proposed, wherein the apparatus has an evaluation and computation unit configured to perform the following steps: - Provides image datasets with a large number of images; - By applying an intersection-over-union filter, a large number of images, especially predetermined contrast images, are compared with at least a portion of the remaining images in the large number of images; Based on this comparison, at least one redundant image relative to the comparison image is determined from at least a portion of the remaining images in the large number of images; and - Clean up the image dataset by removing at least one redundant image from the large number of images.
[0012] The descriptions provided for this method apply accordingly to this device. It should be understood that, according to language conventions, features expressed in the form of a method can be rewritten in the language for the device, without the need to explicitly list such expressions here.
[0013] Even when using active learning on unlabeled data, semantically redundant data can be cleaned up using computationally inexpensive additional importance classification. Under the assumption that the data stream does not originate from an open-world application scenario (image data is part of a closed system) and that images are detected within a scene with a fixed or unchanging image background, potentially changing image foregrounds can be examined and their redundancy analyzed. Leveraging a simplified prediction space where only the image foreground changes (e.g., a simplified prediction space for object block annotations), intersection-over-union (IoU) comparisons become available in image data preprocessing.
[0014] This method and device enable the creation of arbitrary, semantically redundant datasets for training and / or validating and / or testing machine learning models. Here, the intersection-over-union (IoU) metric is used to clean the initial image dataset.
[0015] For example, two images with the highest IoU metric within the same category can be compared. In this way, preferably, if at least one object in the compared image pair does not meet the IoU filtering criteria, the two images can be considered similar and therefore redundant.
[0016] Here, relevant key images or contrasting images can be selected, and comparisons should be made based on these key images or contrasting images. In this way, semantically similar and therefore redundant relevant key images can be selected so that they can be automatically excluded from the dataset or dataset split, in particular. This makes the metric more robust.
[0017] Compared to proactive solutions based on data control, the most significant innovation of IoU-based semantic image redundancy filters can be seen as their application to datasets with static or quasi-static image backgrounds, allowing the filtering process to focus on the dynamics of objects located in the foreground of the images being compared. Under this assumption, IoU-based semantic redundancy filters have proven to be more efficient and effective in recognition, for example, than introducing new labels for semantic redundancy.
[0018] Intersection over Union (IoU) is a metric used in computer vision, particularly in tasks such as object recognition and / or segmentation. IoU measures the overlap between two regions—specifically, the overlap between the predicted region (by the model) and the ground truth region (the actual location of the object)—to evaluate the accuracy of the prediction. An IoU value of 1 indicates a perfect match, while a value of 0 indicates that the predicted and ground truth regions do not overlap. Here, for example, an IoU threshold (e.g., 0.5) can be specified to determine whether a prediction should be considered correct. In this case, it is preferable to use this threshold to determine whether an image is considered similar to a comparison image and therefore redundant, or dissimilar to a comparison image and therefore non-redundant.
[0019] On the other hand, the proposed application of the cross-union filter includes comparing at least one keyframe in at least one comparison image that includes at least one object with keyframes in each of the remaining images to be compared (the keyframes are preferably set at the same position as in the comparison image).
[0020] In computer vision, a keyframe is preferably defined as a representative frame within a sequence of images or videos that contains information important or meaningful to the processing task. In object tracking and similar tasks, keyframes can be used to specify the important locations of objects in a video. The relationship between keyframes and Intersection over Union (IoU) in computer vision (especially in video analysis or object tracking) arises from evaluating the accuracy of object recognition and tracking across multiple frames. Here, keyframes help identify important points in a video sequence where the IoU metric can be applied to evaluate the accuracy of object tracking. IoU is calculated to measure how well the tracked object matches manually labeled ground truth data in these keyframes.
[0021] On the other hand, it is proposed that an image dataset be provided, which includes pre-selected key images used as at least one comparison image, where these key images may be redundant.
[0022] Selecting visually relevant key images helps clean the image dataset and preserve potentially redundant images. From a provided (preferably labeled or annotated) image dataset—which can be used for object recognition and where images are labeled by setting bounding boxes or keyframes that identify objects in a scene—relevant key images that may be redundant are preferably selected. In the simplest case, this might mean manually selecting some unwanted redundant images, but other redundancy measures can also be used to select relevant key images. Here, a cleaned image dataset can preferably be considered. Furthermore, subsets can be created from the cleaned dataset in any proportion, for example, dividing the data into 60% training data, 20% validation data, and 20% test data. The preferably selected relevant key images are compared as contrast images to the remaining images of the preferred pre-cleaned image dataset or any subset of that dataset. The images in the image dataset preferably each include a static background image followed by a foreground object that dynamically changes between consecutive images.
[0023] On the other hand, it is proposed that identifying at least one redundant image relative to the comparison image in at least a portion of the remaining images of a large number of images includes comparing the corresponding image with a predetermined threshold or similarity metric.
[0024] In the current case, it is preferable to apply an IoU filter to the relevant key image and all remaining images for each common category. The IoU filter is preferably set fixedly, in particular with a threshold between 0.0 and 1.0, for example, 0.9. Here, the relevant key images are preferably used. Then, all categories are iterated through, where for each category and preferably for each relevant key image, a similarity metric with each other image is calculated. Here, it is preferable to compare the objects included in each key image with the corresponding objects in the image to be compared. This can also be performed on only a subset of the images. If at least one overlapping object is found, each compared image is considered semantically redundant.
[0025] On the other hand, it is proposed to select at least one contrast image from a subset of the image dataset and compare the contrast image with images from another subset of the image dataset.
[0026] Particularly preferred is to compare the pre-filtered images with the cleaned image dataset. This step preferably involves assigning redundant images to a subset different from the subset under investigation of the image dataset.
[0027] By applying an intersection-over-union (IoU) filter to compare, specifically predetermined contrast images from a large dataset, with at least a portion of the remaining images in the large dataset; and by determining, based on this comparison, at least one redundant image relative to the contrast images from the at least portion of the remaining images in the large dataset, this process can preferably be performed individually or comparatively on multiple subsets of the image dataset. Thus, the comparison and determination can preferably be performed on pairs of subsets of image data, for example, in the case of a tripartite pair consisting of training-validation data, training-test data, and validation-test data. For example, if subsets of training image data and validation image data are compared with each other, and relevant key images are selected from the training subset, images that are similar to each other and therefore redundant can be removed from the validation subset to clean up semantic similarity in the validation subset.
[0028] A cutoff threshold can also be applied to divide the image dataset, for example, to form a new subset from the image dataset. The newly formed subset may deviate from the original proportion, for example, a cutoff proportion of 60%-20%-20%. Therefore, before using non-redundant dataset samples for training and / or testing and / or validating machine learning models, it is preferable to use this step to re-filter the non-redundant dataset samples according to the principle of randomness using the original cutoff proportion. For inaccurate image sequences, i.e., for non-contiguous images, this method can still function as a redundancy filter.
[0029] On the other hand, it is proposed that a large amount of image data be detected sequentially by an imaging sensor. "Sequential" refers to the image data being detected in a specific order or sequence. "Sequential processing" describes processing the image data according to the order in which it arrives or is created. At least one imaging sensor may include a camera, an ultrasonic sensor, a lidar sensor, a radar sensor, etc.
[0030] On the other hand, a method is proposed for training and / or validating and / or testing a machine learning model that can be used for the classification and / or segmentation of image data, particularly (but not limited to) in automated unpacking machines and / or in vehicles with at least one autonomous driving function and / or in automated optical inspection, the method comprising: training and / or validating and / or testing the machine learning model based on an image dataset cleaned according to the method claimed in this application.
[0031] It should be understood that the methods used to train and / or validate and / or test machine learning models for classifying and / or segmenting image data can also be used in other technical fields. Examples include medical technology, security technology, surveillance technology, authentication technology, automation technology, robotics, and so on.
[0032] The method used to clean up datasets can, in principle, be applied to places where large amounts of image data may be processed in a fast sequence. This method is particularly advantageous when the image background remains constant while the foreground or foreground plane changes.
[0033] On the other hand, protection is also claimed for a control device that includes a vehicle and / or robotic system and / or industrial machine with autonomous driving capabilities, and on said control device, a machine learning model trained and / or verified and / or tested in accordance with the method in one of its aspects.
[0034] On the other hand, a computer program having program code is claimed, which, when executed on a computer, is used to perform at least a portion of the method in one of its aspects. In other words, a computer program (product) is claimed that includes instructions, which, when executed by a computer, cause the computer to perform the method / steps of the method in one of its aspects.
[0035] On the other hand, a computer-readable data carrier containing program code of a computer program is proposed, which, when executed on a computer, is used to perform at least a portion of the method in one of its aspects. In other words, the present invention relates to a computer-readable (storage) medium comprising instructions that, when executed by a computer, cause the computer to perform the method / steps of the method in one of its aspects.
[0036] The described design and improvement schemes can be combined with each other arbitrarily.
[0037] Other possible designs, improvements, and implementations of the present invention also include combinations of features of the invention not explicitly mentioned in the preceding or following descriptions of the embodiments. Attached Figure Description
[0038] The accompanying drawings are intended to provide a further understanding of embodiments of the invention. They illustrate the embodiments and, in conjunction with the description, serve to explain the principles and concepts of the invention.
[0039] Other embodiments and many of the mentioned advantages are derived with reference to the accompanying drawings. The elements shown in the drawings are not necessarily shown to scale.
[0040] Figure 1 A schematic flowchart of one embodiment of the method is shown; Figure 2 A schematic block diagram of one embodiment of the method is shown; Figure 3 A schematic visualization of redundant images is shown; and Figure 4A schematic visualization of a non-redundant image is shown.
[0041] In the accompanying drawings, the same reference numerals denote the same or functionally identical elements, parts or components, unless otherwise specified. Detailed Implementation
[0042] Figure 1 A schematic flowchart of a method for cleaning an image dataset used to train and / or validate and / or test a machine learning model is shown.
[0043] This method can be performed at least in part by device 100 in any embodiment, which may include a number of components not shown in detail, such as one or more providing means and / or at least one evaluation and calculation means. It should be understood that the providing means may be incorporated into or separate from the evaluation and calculation means. Furthermore, device 100 (which may be part of a system) may include storage means and / or output means and / or display means and / or input means.
[0044] The computer-implemented method includes at least the following steps: In step S1, an image dataset with a large number of images is provided.
[0045] In step S2, by applying an intersection-over-union filter, a predetermined comparison image in a large number of images is compared with at least a portion of the remaining images in the large number of images.
[0046] In step S3, based on the comparison, at least one redundant image relative to the comparison image is determined from at least a portion of the remaining images of the large number of images.
[0047] In step S4, the image dataset is cleaned by removing at least one redundant image from the large number of images.
[0048] Figure 2 A schematic block diagram of one embodiment of the method is shown. In step S200, a subset of the image dataset to be used for training is provided. In step S202, a subset of the image dataset to be used for validation is provided. Preferably, in step S204, irrelevant keyframes are visually selected to provide a pre-cleaned subset in this manner. In step S206, for both subsets, an IoU filter is applied to at least one comparison image and all the remaining images of each subset to be compared, for each common category.
[0049] In optional step S208, the filtered images are compared with images in the cleaned subset of the image dataset. In optional step S210, a cutoff threshold is applied to thus form a new subset partition. In this way, new training data subsets and new validation data subsets can be provided in steps S212 and S214. Furthermore, a subset of image data identified as redundant images can be provided in optional step S216.
[0050] Figure 3 and Figure 4 A schematic illustration shows images in an image dataset that should be cleaned, filtered, or have redundant images removed according to this method. The images in the image dataset are preferably collected sequentially. These images come, for example, from a camera that provides a series of images of a scene at a predetermined frame rate. One possible redundancy factor is that, for example, due to the characteristics of the scene (e.g., sampling the scene at a frame rate of 30 images per second), there are no significant changes between consecutive images.
[0051] In the example shown, an image of a scene (in this case, a warehouse) is detected. This scene includes packaging materials (e.g., boxes, cartons, parcels, etc.) in different states (e.g., open, half-open, closed, etc.), photographed from above in a static environment. Here, the background of the scene remains unchanged in the current context. To ensure the completeness of the dataset as much as possible, packaging materials of different locations, orientations, and / or sizes are detected in the current context.
[0052] This method allows for the evaluation of the generalization ability of object recognition. Furthermore, it avoids sample redundancy. Now, as previously described, images that are similar to each other should be filtered from the dataset to clean up unwanted redundancy. Here, similarity at the image feature level is not considered in the current case, because, for example, the packaging materials look very similar when viewed from above at the detection scale.
[0053] Instead, the contents of the packaging materials should be considered, preferably only volume information. As mentioned earlier, the IoU index is used to compare images, thereby finding the arrangement of packaging materials identified as similar to each other in the image pairs being compared.
[0054] Figure 3A schematic visualization of redundant image data is shown. In the left image 300, relevant keyframes 302 and 303 are shown as circles. Keyframes 302 and 303 represent the current object, which is compared to the corresponding dashed line 304 in the right image 306. This is a non-redundant variation because, in this example, there is no object in the area of another schematically marked keyframe 308 in the right image 306. On the other hand, keyframe 303 overlaps between the two images 300 and 306, which is indicated by the reference numeral 303 drawn in both the left image 300 and the right image 306. Thus, a redundant object is found here, making the entire right image 306 marked as redundant.
[0055] The smaller circles 310 shown in images 300 and 306 represent smaller objects, which can also be used for redundancy comparison. The right image 306, identified as redundant, is filtered out from the image dataset due to the dashed keyframe 303 or objects using the IoU-based algorithm described herein or by this method.
[0056] Figure 4 A schematic visualization of a non-redundant image is shown. In the left image 400, relevant keyframes 402 and 403 are displayed as circles, and comparisons of all objects are indicated by dashed lines 404 shown in the right image 406. Smaller frames 408, or smaller circles, represent smaller objects that can also be used for redundant comparisons. In this example, the right image 406 is identified as a non-redundant image because there is no overlap between the left image 400 and the right image 406.
Claims
1. A method for cleaning an image dataset used to train and / or validate and / or test a machine learning model, the method comprising the following steps: - Provide (S1) an image dataset with a large number of images; - By applying an intersection-over-union (IoU) filter with a set threshold, the IoU images, especially predetermined comparison images, are compared with at least a portion of the remaining images in the IoU images (S2), wherein applying the IoU filter includes: At least one keyframe (302, 303, 308, 402, 403) in at least one comparison image, which includes at least one object, is compared with keyframes (302, 303, 308, 402, 403) in each of the remaining images to be compared, the keyframes being set at the same positions as in the comparison image; - Based on the threshold, determine (S3) at least one redundant image (306, 310) relative to the comparison image from at least a portion of the remaining images of the large number of images; and - The image dataset is cleaned (S4) by removing at least one redundant image from the large number of images.
2. The method of claim 1, wherein providing (S1) the image dataset includes pre-selected key images for use as the at least one comparison image, wherein the key images can be redundant.
3. The method according to any one of the preceding claims, wherein determining (S3) at least one redundant image relative to the comparison image in at least a portion of the remaining images of the large number of images comprises comparing the corresponding image with a predetermined threshold.
4. The method according to any one of the preceding claims, wherein the at least one comparison image is selected from a subset of the image dataset, and the comparison image is compared with images from another subset of the image dataset.
5. The method according to any one of the preceding claims, wherein a large amount of image data is detected sequentially by an imaging sensor.
6. A method for training and / or validating and / or testing a machine learning model, said machine learning model being capable of classifying and / or segmenting image data in an automated unpacking machine and / or in a vehicle with at least one autonomous driving capability and / or in automated optical inspection, said method comprising: The machine learning model is trained and / or validated and / or tested based on the cleaned image dataset cleaned according to any one of claims 1 to 5.
7. A computer program having program code, which, when executed on a computer, performs at least a portion of the method according to any one of claims 1 to 6.
8. A computer-readable data carrier having program code of a computer program, which, when executed on a computer, performs at least a portion of the method according to any one of claims 1 to 6.
9. An apparatus (100) for cleaning an image dataset used to train and / or validate and / or test a machine learning model, wherein the apparatus (100) includes an evaluation and computation unit configured to perform the following steps: - Provides image datasets with a large number of images; - By applying an intersection-over-union (IoU) filter with a set threshold, comparing, in particular, predetermined contrast images from the large number of images with at least a portion of the remaining images from the large number of images, wherein applying the IoU filter includes: At least one keyframe (302, 303, 308, 402, 403) in at least one comparison image, which includes at least one object, is compared with keyframes (302, 303, 308, 402, 403) in each of the remaining images to be compared, the keyframes being set at the same positions as in the comparison image; - Based on the threshold, at least one redundant image relative to the comparison image is determined from at least a portion of the remaining images of the large number of images; as well as - The image dataset is cleaned by removing at least one redundant image (306, 310) from the large number of images.