Target category assignment of objects
By using a multi-category classifier to determine the confidence value of the intermediate category and map it to the target category, the problem of class changes in the prior art requires retraining, and efficient and interpretable image data classification is achieved.
Patent Information
- Application Number
- CN202510149318.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-14
- Filing Date
- 2025-02-11
- Publication Date
- 2025-08-15
AI Technical Summary
In the prior art, in image data classification, especially in multi-label classification, the cumbersome process of retraining the neural network when facing changes in category definitions, resulting in inefficiency.
The confidence value of the intermediate category is determined by using a multi-category classifier, and the intermediate category is converted into the target category through mapping rules, avoiding training directly for the target category, and using simple mapping rules or parameterization algorithms for conversion.
This enables no need to retrain the neural network when the category changes, simplifies the target category allocation process of image data, and improves efficiency and interpretability.
Smart Images

Figure CN120495612A_ABST
Abstract
Description
[0001] The invention relates to an image acquisition device and a method for assigning object classes to objects according to the preambles of claims 1 and 11.
[0002] In many image processing applications, particularly in logistics and automation, it is necessary to identify objects or their properties. In addition to classical methods, machine learning or artificial intelligence methods have long been used to classify image data or the objects recorded therein. Since the seminal article "Imagenet classification with deep convolutional neural networks" by Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton (AlexNet) in Advances in Neural Information Processing Systems 25 (2012), deep neural networks (deep learning) have almost undisputedly dominated the field. However, significant progress has also been made in the intervening period.
[0003] One type of classification is so-called multi-label classification. In this case, multiple attributes or categories are assigned to an object. Below, the German term "multi-class classification" will be used to refer to this situation, although strictly speaking, it was originally the opposite of binary classification. For example, Jesse Read and Fernando Perez-Cruz's article "Deep learning for multi-label classification" published in arXiv preprint arXiv:1502.05988 (2014) deals with multi-class classification in the sense defined above.
[0004] The article “ML-decoder: Scalable and versatile classification head” published by Tal Ridnik et al. in the Winter Conference on Applications of Computer Visio, 2023 describes a particularly powerful multi-class classification. In it, the attention mechanism known from the Transformer architecture is used, where only linear workload is achieved by adaptation (abandoning the self-attention layer). The ML decoder is regarded as a supplementary classification head, which is previously preprocessed with another neural network (for example, ResNet or TResNet). For TResNet, see the article “TResNet: High performance gpu-dedicated architecture” published by Tal Ridnik et al. in the Proceedings of the IEEE / CVF International Conference on Computer Vision, 2021. An alternative to this is MobileViT, from the article “Separable self-attention for mobile vision transformers” published by Sachin Mehta and Mohammad Rastegari in the arXiv preprint arXiv:2206.02680(2022). Finally, it is worth mentioning that it may be helpful to use asymmetric error evaluation (loss function) in multi-class classification, see the article "Asymmetric loss for multi-label classification" published by Tal Ridnik et al. in Proceedings of the IEEE / CVF International Conference on Computer Vision, 2021.
[0005] In practice, training a neural network for a specific classification task is challenging because it requires a tremendous amount of work. This becomes particularly problematic if the category definitions change over time. Until now, such adjustments have only been possible through retraining, or at least large-scale retraining.
[0006] It is therefore an object of the present invention to further improve the class assignment of image data.
[0007] This object is achieved by an image capture device and a method for assigning a target class to an object according to claims 1 and 11. The target class is the actual result of the classification and is referred to as a target class in order to distinguish it from the intermediate classes that will be introduced, since according to the invention, the classification is carried out in two stages. Image data with the object to be classified are recorded using an image sensor. A control and evaluation unit evaluates the image data in order to assign the target class, and for this purpose a machine learning method, in particular a neural network, is implemented in the control and evaluation unit. The control and evaluation unit is understood to be at least one arbitrary hardware component that can be arranged inside the image capture device and / or connected to the image capture device and that provides the necessary computing power and storage capacity.
[0008] The present invention is based on the following basic idea: First, an intermediate class is determined using a multi-class classifier. The result of this intermediate step is a confidence value that indicates the degree of certainty with which the image data is assigned to the corresponding intermediate class. Multiple assignments to several or even all intermediate classes are expressly desired. Numerical, informal confidence values are conceivable, which use a "yes / no" message to determine whether an intermediate class has been assigned. However, quantitative confidence values, for example in the interval [0, 1], are preferred, as they can always be implemented using simple scaling. The multi-class classifier operates using machine learning methods.
[0009] Subsequently, an assignment to at least one target class, preferably precisely one target class, is performed based on the intermediate class. To this end, the confidence values of the intermediate class are mapped to the target class. The mapping is therefore a function or an assignment rule, which assigns one or more target classes to a tuple of confidence values. Preferably, the tuple has as many elements as the intermediate class, and any dimensional differences with tuples having a different number of elements can in any case be restored to this situation by padding with zeros or by projection mapping. The intermediate class and the target class are not identical to one another, so that at least one target class is not found in the intermediate class. The mapping reconstructs the target class from the intermediate class in a very intuitive and simplified manner, taking into account pre-given rules.
[0010] Therefore, the machine learning method is trained for the intermediate categories instead of for the target category. The latter is conceivable in principle, but according to the present invention, this approach is precisely avoided so that tedious (re)training can be omitted when the target category changes. Preferably, the mapping after the multi-category classification does not require training, but is the result of a relatively simple optimization. For this purpose, it is also conceivable to use machine learning methods, in particular a second neural network. The effort required for its training is very low, because only a very small number of intermediate categories compared to the pixels of the image need to be taken into account as input data. However, preferably, the mapping is a simple, deterministic assignment rule or a simple algorithm parameterized according to the rules of the target category, without the need for machine learning or neural network methods.
[0011] The advantage of the present invention is that it allows for a very simple assignment of object classes to image data or the objects recorded therein, specifically those designated for applications in logistics or automation. Thanks to the optimized mapping of intermediate classes to target classes, changing requirements can be accommodated and adapted to new challenges. This eliminates the need to repeat the most complex step, namely, training a multi-class classifier.
[0012] Preferably, none of the target classes is an intermediate class. As mentioned, the intermediate classes and the target classes are not identical, otherwise the mapping would be redundant or could be done by an extremely simple rule (such as identity mapping The mapping is achieved by using a combination of methods (e.g., projections) to map the target class and the intermediate classes. According to this embodiment, the target class and the intermediate classes should even be disjoint. Therefore, each target class depends on more than one intermediate class, the target class is a mixture of the intermediate classes, and the mixture rule of this connection is the mapping.
[0013] Preferably, the intermediate class is defined by at least one of the following properties of the recorded object: material (in particular plastic, Styrofoam, wood, or metal), strength (in particular rigid or flexible), or shape (in particular cuboid, cylindrical, toroidal, or irregular). These are only some examples of possible intermediate classes; a multi-class classifier can be trained for any intermediate class. Other conceivable intermediate classes involve color, reflectivity, or size, provided that a comparison scale is provided, for example, by a fixed recording situation.
[0014] Preferably, exactly two target classes are provided, in particular carton or non-carton. Although there are multiple intermediate classes, according to this embodiment, only one of the two target classes is ultimately assigned binary. An example is the distinction related to logistics applications, namely whether the object is a carton or not. In one embodiment, the carton properties are completely derived from the other properties of the intermediate classes. However, even if there is an intermediate class with carton as outer packaging material, the target classes can be distinguished by further rules, for example, because cartons surrounded by plastic strips or plastic sleeves should no longer be classified into the target class "carton". Such additional conditions can be taken into account in the mapping of intermediate classes to target classes.
[0015] Preferably, the multi-class classifier has an attention mechanism. This enables a particularly efficient assignment to intermediate classes.
[0016] Preferably, the multi-class classifier has a first stage that generates an embedding from the image data and a second stage that determines the intermediate class from the embedded features. In the first stage (backbone network), features are extracted, especially in the form of embeddings, and then used by the classification head to determine the intermediate class. This architecture (for example, with TResNet or MobileViT as the backbone network and an ML decoder as the classification head) is particularly suitable for reliably determining the intermediate class.
[0017] Preferably, the mapping uses thresholds to evaluate intermediate categories separately. In general, the mapping is an arbitrary function from an m-dimensional space to an n-dimensional space, where there are m intermediate categories and n target categories. In this embodiment, by first evaluating each of the m intermediate categories using a threshold, the space of possible mappings is significantly reduced. This produces a result equivalent to an m-bit binary word, so only these binary words need to be assigned the corresponding target category. This greatly simplifies the mapping required. It is not even necessary to distinguish all m-bit binary words. For example, for a "cardboard" decision, it is sufficient that the intermediate categories "wood" or "plastic" are above a threshold, and the mapping should then be assigned to "not cardboard." Simplifying the mapping to thresholds has the additional advantage that these thresholds are intuitively understandable to the user. Thus, there are interpretable intermediate categories or interpretable assessments of the impact of intermediate categories on the target category. In contrast, intermediate results or feature maps from traditional training of machine learning methods are often opaque and incomprehensible to the observer ("black box"). In particular, interpretability makes it very easy for the user to readjust. For example, it might be the case that a package should not be classified as the target class "carton" even if it contains only a small amount of plastic. However, the mapping was initially parameterized so that cartons with low confidence values for plastic are still assigned the target class "carton." Field technicians can now simply adjust the threshold for plastic so that packages are no longer classified as carton due to their plastic content, as desired. This is accomplished by simply resetting the parameters and requires neither retraining the multi-class classifier nor reoptimizing the mapping.
[0018] Preferably, the mapping is learned by a multi-class classifier determining confidence values for a plurality of example images annotated with the desired target class and, during optimization, determining a mapping that, given the confidence values found for each example image, reproduces the mapping of the associated annotated target class as best as possible. The desired rule for the target class is thus specified in the form of example images and the target class derived from the rule for the corresponding example images, for example, in a manual labeling process in which a human observer annotates the example images according to the rule. The source of the annotated example images is irrelevant to the present invention. If the example images annotated in this manner are evaluated by the multi-class classifier, the intermediate classes are then known, and the target classes corresponding thereto are known based on the labels of the example images. Processing the plurality of example images generates a plurality of tuples of the form ((intermediate_1, ..., intermediate_m), (target_1), ..., (target_n)), from which a mapping that reproduces the plurality of tuples as comprehensively as possible can be determined, for example, by means of function fitting or other optimization methods. Other machine learning methods, in particular a second neural network, are conceivable as an alternative to function fitting. Its training is no longer based on massive data of original image data, but only on the mentioned multiple tuples, and is therefore less complex.
[0019] Preferably, the mapping is first initialized using arbitrary thresholds for each intermediate class, and the optimization varies only the thresholds. This corresponds to the simplified mapping discussed above, which uses thresholds to evaluate intermediate classes individually. This may not lead to a global optimum, but it can lead to a sufficiently efficient mapping, with the advantage that the optimization problem is significantly simplified.
[0020] The image detection device is preferably installed on a conveyor, on which the objects to be sorted are transported through the field of view of the image sensor. In particular, multiple cameras are provided, and the control and evaluation unit is designed to combine the image data recorded by the cameras into a common image. For example, the conveyor may be part of a production line or sorting system in the automation or logistics industry, and the conveyor may transport objects one by one to the inspection area. In some cases, the field of view of a single camera is too small for the objects or the conveyor. In this case, an image detection device with two or more cameras can be used, whose image data is combined (stitched).
[0021] The method according to the invention is a computer-implemented method, which runs, for example, in a camera or another computing unit, whether in real time in a computing unit at least indirectly connected to the camera or in a delayed manner in any computing unit.
[0022] Preferably, the mapping is learned by the multi-class classifier determining confidence values for a plurality of example images annotated with the desired target class and, during optimization, determining a mapping that, given the confidence values found for each example image, reproduces the mapping of the associated annotated target class as best as possible. This corresponds to the process described above. Prior to this, the multi-class classifier is trained, for example, in supervised learning, based on example images annotated with intermediate classes. Training the multi-class classifier and determining the mapping are separate steps. Training the multi-class classifier can be performed in completely different locations, at different times, with different equipment, and using different example images, or at least using a different training dataset that has at least its own annotations, i.e., annotations with intermediate classes rather than the target class. As repeatedly emphasized, after training for the intermediate classes, the multi-class classifier is not trained or retrained again for the target class; the target class is determined by the mapping of the intermediate classes of the multi-class classifier.
[0023] Furthermore, the method according to the invention can be further developed in a similar manner to the image acquisition device and at the same time exhibit similar advantages. Such advantageous features are described exemplarily but not exhaustively in the dependent claims which are dependent on the independent claim. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Other features and advantages of the present invention will be described in more detail below based on exemplary embodiments and with reference to the accompanying drawings. In the accompanying drawings:
[0025] Figure 1 An overview diagram of using a camera to classify objects being transported on a conveyor belt through the camera's field of view is shown;
[0026] Figure 2 An exemplary flow chart for first classifying into an intermediate category and then mapping the intermediate category to a target category is shown;
[0027] Figure 3 A diagram showing mapping of intermediate categories to target categories; and
[0028] Figure 4 An exemplary flow chart for training a multi-class classifier to classify into an intermediate class and for finding a mapping of the intermediate class to the target class is shown.
[0029] Figure 1The camera 10 is shown mounted above a conveyor belt 12, which conveys objects 14 through a detection area 18 of the camera 10, as indicated by the arrow 16. The fixed mounting of the camera 10 at the conveyor belt is a common application in practice, for example for logistics tasks or automation tasks or quality control. However, the present invention relates to the classification of images or objects 14 recorded with the images, in particular in order to initiate subsequent processing steps based on the classification, such as sorting, requesting manual further processing, etc. Therefore, this example should not be understood as restrictive, and the objects 14 can also be presented to the camera 10 in other ways. Figure 1 In the example, the objects 14 differ in presentation only by their shape, although the classification may also involve other properties.
[0030] The camera 10 uses an image sensor 20 to detect image data of the object 14 being fed in. This image data is further processed by a control and evaluation unit 22. For example, the control and evaluation unit 22 includes at least one computing module, such as a microprocessor or CPU (Central Processing Unit), FPGA (Field Programmable Gate Array), DSP (Digital Signal Processor), ASIC (Application-Specific Integrated Circuit), AI processor, NPU (Neural Processing Unit), GPU (Graphics Processing Unit), VPU (Video Processing Unit), etc. The control and evaluation unit 22 is shown as an internal computing module. Alternatively, there may be multiple computing modules, which may also be at least partially external to the camera 10. The external computing unit can be any type of computer, including a laptop, smartphone, tablet, or controller, or it can be a local network, edge device, or cloud. Furthermore, the control and evaluation unit 22 responsible for classification or inference may include hardware that is completely different from the hardware used for training, learning, or parameterizing the classification.
[0031] Furthermore, the specific imaging method is not important for the present invention, so the camera 10 can be constructed according to any known principle. For example, only one line is recorded at a time, and the control and evaluation unit 22 combines the lines detected during the conveying movement into image data. Using a matrix image sensor 20, a larger area can be recorded in one recording, wherein the recording in the conveying direction and the recording transverse to the conveying direction can also be combined here. Unlike Figure 1 Such multiple recordings can also be recorded simultaneously or in a temporally overlapping manner using a camera 10 with multiple cameras. The camera 10 can output information such as image data or specific classification results related to the image data or the recorded objects 14 via the interface 24.
[0032] Figure 2 An exemplary flow chart for classifying image data or objects 14 recorded in the image data is shown. In step S1, image data is recorded. In step S2, features related to the image data are generated in a backbone network. For example, TResNet or MobileViT or MobileViTv2 can be used as the backbone network. In multi-class classification, it may be advantageous to use an asymmetric error function (loss function).
[0033] In step S3, the features of the backbone network are provided to a classification head using an attention mechanism. This classification head replaces ("drop-in replacement") an alternative but still usable pooling (GAP, global average pooling). Attention mechanisms, especially those made famous by the Transformer architecture, lead to better results by taking context into account. The ML decoder in the literature cited in the introduction is particularly suitable, and details about the backbone network and the asymmetric error function can also be found in the literature. In the ML decoder, the attention mechanism of the Transformer is modified to help reduce the workload, and reference can also be made to the literature on this point.
[0034] The intermediate class and its associated confidence value (score) are output as the classification result. The procedure, divided into two parts in steps S2 and S3: the backbone network and the subsequent classification head, is a preferred embodiment of a multi-class classifier with a particularly modern, high-performance architecture. However, various other classifications are also possible, as long as the result is an intermediate class with a confidence value.
[0035] The intermediate categories determined by the multi-class classifier in steps S2 and S3 are not yet the expected results of the classification, which is why they are called intermediate categories. In step S4, the actual target category is determined based on the intermediate categories and their confidence values. This is done by Figure 3 This is achieved using the simple allocation rules or mapping shown. Figure 3The left side of the diagram shows some exemplary intermediate categories. For example, these intermediate categories relate to material, shape, and strength, but the present invention can also handle any even more refined definitions of intermediate categories, i.e., it is not limited to these categories nor to the specific examples of plastic, polystyrene foam, wood, metal, cuboid, cylinder, toroid, irregular, rigid, or flexible. The height of the bar indicates how clearly the multi-class classifier found the intermediate category, i.e., the confidence value. Here, it is explicitly allowed that multiple intermediate categories are identified very clearly at the same time, i.e., multiple intermediate categories are identified simultaneously with high confidence values.
[0036] The mapping, indicated by the arrows, assigns at least one target class to the confidence values of the intermediate classes. The mapping takes into account the distribution of confidence values across the intermediate classes using the assignment rules defined in the mapping in order to assign a specific target class. In a preferred embodiment, there is only one target class in each case, but similar to the intermediate classes, multiple classifications can also be envisioned in the final result of the target classes, so that the mapping is correspondingly multidimensional in both its definition and its value range.
[0037] Then in step S5, the target category is determined. Figure 3 In the example shown, mapping initially produces not only the object class but also the confidence value for that object class. Thresholding or determining a maximum value can further reduce the number of object classes, typically to just one. As mentioned previously, mapping can also directly produce just one object class with or without a confidence value. In this example, the intermediate classes "Plastic" and "Irregular" are strongly represented, so the latter is assigned to the binary object class "Carton / Not Carton."
[0038] Figure 4 An exemplary flowchart for training a multi-class classifier to classify as an intermediate class and for finding a mapping from the intermediate class to the target class is shown. In step T1, the multi-class classifier is trained. This is accomplished, for example, in supervised learning based on training images annotated with the intermediate class. Training a neural network for a given classification task based on example images and associated labels is well known and will not be described in detail here. Further explanation can also be found in the literature cited in the introduction.
[0039] In step T2, a multi-class classifier is used to evaluate the example image. Figure 3As shown on the left, confidence values for the intermediate categories are obtained for the corresponding example images. These example images are annotated with the target category, not with the intermediate category. Therefore, these example images are not the images used to train the multi-category classifier in step T1, and the multi-category classifier has already been trained at this stage. Here, the images in step T1 are allowed to be repeated in step T2, but the images are at least annotated in a different way. Figure 3 As shown, various intermediate categories related to material, shape, or strength can be provided, while "carton / not carton" can be determined as an object class, for example. For example, example images are manually annotated with associated object classes, particularly following a predefined set of rules. This annotation takes the form of dividing the image into partitions belonging to the respective object class. While it is possible to train a classifier for the object class directly based on the example images, this effort is precisely avoided according to the present invention.
[0040] In step T3, after executing step T2 multiple times, multiple allocation examples are obtained. Figure 3 (Its combination Figure 2 The application of mapping intermediate categories to target categories has been explained) can also be understood as Figure 4 Thus, the corresponding assignment example is equivalent to an m-tuple of confidence values for m intermediate categories associated with the confidence values for n target categories, which can be written as ((intermediate category_1, ..., intermediate category_m), (target category_1), ... (target category_n)). As mentioned above, alternatively, the confidence value for the target category can be omitted by simply specifying whether the target category exists in a binary manner.
[0041] In step T4, an optimization method is now used to determine a mapping that is as compatible as possible with the allocation examples of step T3, or reproduces these allocation examples. Of course, the mapping mentioned here does not refer to a mapping that corresponds point by point to the allocation examples and outputs arbitrary results for deviated input values, but rather a mapping that is generally best adapted to the allocation examples in the selected error metric and satisfies predetermined regulations such as smoothness and other secondary conditions. Ultimately, this is a function fitting, and all known methods can be used for it. One method is to use a hyperparameter optimization (HPO) tool such as Optuna (Takuya Akiba et al. published an article "Optuna: A next-generation hyperparameter optimization framework" in Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, 2019).
[0042] Optimizing any mapping from the m-dimensional space of intermediate categories to the n-dimensional space of target categories over all possible mappings requires a certain amount of work and may not converge to a useful optimum. Therefore, one can imagine mappings that only allow for certain categories, in particular those that are initially evaluated only for each intermediate category. For example, Figure 2 The mapping in step S4 can only compare the intermediate categories with the threshold, and infer the intermediate categories from the intermediate categories with confidence values higher than the threshold. Figure 4 Step T4 is simplified to finding the optimal threshold.
[0043] Therefore, by Figure 4 In step T4, a mapping is found, learned, or parameterized to reclassify the original intermediate categories into new target categories. The target categories can then be specified using rules or example images annotated based on the rules. The effort required to find the mapping in step T4 is significantly less than the effort required to train or retrain a multi-class classifier based on the target categories.
Claims
1. An image acquisition device (10) for assigning an object class to an object (14), wherein the image acquisition device (10) comprises: an image sensor (20) for recording image data with the object (14); and a control and evaluation unit (22) designed to evaluate and classify the image data using a machine learning method, in particular a neural network, and to assign an object class to the image data, It is characterized in that The control and evaluation unit (22) is further designed to use a multi-class classifier as a machine learning method for classification into multiple intermediate categories, the multi-class classifier determining corresponding confidence values for assigning the image data to the corresponding intermediate categories, and subsequently determining the target category by applying a mapping of the confidence values to the target category.
2. The image detection device (10) according to claim 1, wherein None of the target categories are intermediate categories.
3. The image detection device (10) according to claim 1 or 2, wherein: The intermediate class is defined by at least one of the following properties of the object recorded: material, in particular plastic, Styrofoam, wood or metal; strength, in particular rigid or flexible; or shape, in particular cuboid, cylindrical, toroidal or irregular.
4. The image detection device (10) according to any one of the preceding claims, wherein Provide exactly two target classes, specifically carton or non-carton.
5. The image detection device (10) according to any one of the preceding claims, wherein The multi-class classifier has an attention mechanism.
6. The image detection device (10) according to any one of the preceding claims, wherein The multi-class classifier has a first stage of generating an embedding from the image data and a second stage of determining the intermediate class from features of the embedded data.
7. The image detection device (10) according to any one of the preceding claims, wherein The mapping uses a threshold to evaluate the intermediate categories individually.
8. The image detection device (10) according to any one of the preceding claims, wherein The mapping is learned in that the multi-class classifier determines confidence values for a plurality of example images annotated with the desired target class and determines, in an optimization, a mapping that reproduces the associated annotated target class as best as possible given the confidence values found for each example image.
9. The image detection device (10) according to claim 8, wherein: The mapping is first initialized using arbitrary thresholds for each intermediate class, and the optimization only varies the thresholds.
10. An image detection device (10) according to any one of the preceding claims, which is installed at a conveyor device (12) on which the objects (14) to be classified are conveyed through the field of view (18) of the image sensor (20), wherein in particular a plurality of cameras are provided and the control and evaluation unit (22) is designed to combine the recordings of the cameras in the image data to form a common image.
11. A method for assigning an object class to an object (14), wherein image data with the object are evaluated and classified using a machine learning method, in particular a neural network, and the object class is assigned to the image data. It is characterized in that A multi-class classifier is used as a machine learning method to classify into multiple intermediate categories, the multi-class classifier determines corresponding confidence values for assigning the image data to corresponding intermediate categories, and then determines the target category by applying a mapping of the confidence values to the target category.
12. The method according to claim 11, wherein The mapping is learned in that the multi-class classifier determines confidence values for a plurality of example images annotated with the desired target class and determines in an optimization a mapping that reproduces the associated annotated target class as best as possible given the predefined confidence values found for each example image.