Sensitivity of a measurement image classifier to input image variations

By training an image classifier to map input images to intermediate images and generate realistic distorted images, the problem of identifying defects in product quality control is solved, the accuracy of identification is improved and the scrap rate is reduced.

CN113807384BActive Publication Date: 2026-05-29ROBERT BOSCH GMBH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ROBERT BOSCH GMBH
Filing Date
2021-06-11
Publication Date
2026-05-29

Smart Images

  • Figure CN113807384B_ABST
    Figure CN113807384B_ABST
Patent Text Reader

Abstract

The invention relates to a method (100) for measuring a sensitivity (2*) of an image classifier (2) to a change of an input image (1), the image classifier assigning the input image (1) to one or more classes (3a-3c) of a pre-given classification, the method having the following steps: mapping (110) the input image (1) by at least one pre-given operator (4) to an intermediate image (5) having a lower information content and / or a worse signal-to-noise ratio than the input image (1); providing (120) at least one generator (6) trained to produce a realistic image assigned by the image classifier (2) to a specific class (3a-3c) of the pre-given classification; producing (130) a variant (7) of the input image (1) from the intermediate image (5) using the generator (6).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the behavior control of a trainable image classifier, which can be used, for example, for quality control of products produced in batches. Background Technology

[0002] In mass production, continuous quality control is typically required. The goal is to identify quality issues as quickly as possible so that the causes can be eliminated promptly and minimal losses of affected units as scrap.

[0003] The optical control of the product's geometry and / or surface is rapid and non-destructive. WO 2018 / 197 074 A1 discloses a testing apparatus in which an object can be placed under a large number of lighting conditions, wherein an image of the object is recorded by a camera under each of these lighting conditions. The morphology of the object is analyzed from these images.

[0004] The product image can also be directly assigned to one of several pre-defined categories using an image classifier based on an artificial neural network. Based on this, the product can be assigned to one of several pre-defined quality categories. In the simplest case, this classification is binary ("qualified" / "unqualified"). Summary of the Invention

[0005] Within the scope of this invention, a method for measuring the sensitivity of an image classifier to changes in an input image is provided.

[0006] The image classifier assigns an input image to one or more categories pre-defined by a given classification. The input image can, for example, be an image of a batch-produced, nominally identical product. The image classifier can be trained, for instance, to assign the input image to one or more of at least two possible categories representing a quality assessment of the corresponding product.

[0007] For example, products can be binary classified as "qualified" or "unqualified" (NOK) based on images. It is also possible and meaningful to classify them into categories that include more intermediate levels between "qualified" and "unqualified".

[0008] The concept of an image, in principle, encompasses any distribution of information arranged as a two-dimensional or multi-dimensional grid. For example, this information could be the intensity values ​​of image pixels recorded using any imaging modality (e.g., using an optical camera, a thermal imaging camera, or ultrasound). However, any other data, such as audio data, radar data, or lidar data, can also be converted into an image and then classified in the same way.

[0009] In this method, the input image is mapped to an intermediate image by at least one pre-given operator, the intermediate image having less information content and / or a worse signal-to-noise ratio compared to the input image.

[0010] For example, the information content can be understood in the sense of Shannon's information theory as the minimum number of independent parameters (e.g., bits) required to characterize an image, such as those needed to transmit the image through a channel.

[0011] Thus, for example, scaling down the input image to a lower pixel resolution results in the intermediate image being described entirely by these fewer pixels. If the intermediate image is then scaled back up to the original pixel resolution of the input image after being scaled down, no change is made. A portion of the information originally contained in the input image has been irretrievably lost.

[0012] The same applies if an encoder using an autoencoder structure converts the input image into a reduced-dimensional representation. This applies whether the reduced-dimensional representation is used directly as an intermediate image or converted back to an image with the same pixel resolution as the original input image. An autoencoder structure can, for example, include an encoder and a decoder, where the encoder produces the reduced-dimensional representation, and the decoder converts the representation back. The encoder and decoder can then be trained together with the goal of producing an image as similar as possible to the original input image upon conversion back.

[0013] Blurring at least a portion of the input image causes the information in the image to change more slowly in space; that is, specific spatial frequencies are removed from the image. Therefore, the image is locally or globally described by fewer spatial frequencies. Thus, the blurring also reduces the information content.

[0014] For example, by adding additive and / or multiplicative noise to the intermediate image, the signal-to-noise ratio of the intermediate image can be made worse compared to the input image.

[0015] Affine transformations can also reduce the information content of intermediate images compared to the input image. Furthermore, the input image can also be mapped to an intermediate image with less information content and / or a poorer signal-to-noise ratio by a trained intermediate image generator. For example, the intermediate image generator can be trained to produce images from the input image with interference that is difficult to describe analytically. This generator can learn to retain this interference on completely different images based on example images with such interference.

[0016] At least one generator is provided, trained to produce images that are realistic on the one hand and assigned to a specific pre-given category by the image classifier on the other. Using the generator, variations of the input image are generated from intermediate images. Here, in addition to the intermediate image, random noise may be fed to the generator to produce many variations from the same intermediate image, which are then classified into specific pre-given categories by the image classifier.

[0017] In this way, for example, an input image assigned to the "qualified" category by the image classifier can be selectively transformed into a variant assigned to the "unqualified" category by the same image classifier. If there are other intermediate levels as other categories between "qualified" and "unqualified," variants belonging to one of these other categories can also be selectively generated from the same input image. Here, a separate generator can be responsible for each target category. However, a unique "conditional" generator can also be used, for example, to which the desired target category can be pre-defined.

[0018] In principle, the input image can also be directly fed to the generator to obtain the variant. The consideration behind bypassing intermediate images is that a significantly stronger change must be made to the original input image to guide the image classifier away from its initial judgment and assign the variant to a different category. These significantly stronger changes are difficult to align with the goal that the variant should be a realistic image.

[0019] For the purposes mentioned in quality control, it is important that the variation is realistic. What is desired here is precisely the effect that might occur during real image recording and cause the input image to be classified differently.

[0020] The classification uncertainty performed by the image classifier now increases significantly by first generating intermediate images with reduced information content and / or a degraded signal-to-noise ratio. Meanwhile, the intermediate images still retain the essential features of the input image, particularly those inserted into the low-frequency components of the input image. From this state of increased uncertainty, only relatively small interventions are needed on the image to obtain the variant assigned to another category by the image classifier. This small intervention is thus still within the range that can be introduced into the image without stripping it of its realistic features.

[0021] In a particularly advantageous design, the intermediate image generator is trained to produce realistic images that are as similar as possible to the input image and are simultaneously classified by the image classifier with the greatest possible uncertainty. As described above, starting from this, it is possible to change the category to which the intermediate image is classified by the image classifier with high probability through small variations. For example, the similarity to the input image can be measured using the cosine distance between activations induced by these images in the deep feature maps of a pre-trained neural network.

[0022] The variants provide directly understandable information about which specific changes in the input image cause the image classifier to assign the image to a category different from the current category. If the specific category to which the variant generated by the generator belongs is a different category from the category to which the input image is assigned by the image classifier, then a comparison of one or more variants with the input image directly provides the changes that drive the image beyond the decision boundary between categories. Thus, the sensitivity of the image classifier to changes in the input image can be analyzed both qualitatively and quantitatively. Alternatively or in combination, any comprehensive statistics about multiple variants can also be used for sensitivity evaluation. Furthermore, for example, the variant that is closest to the input image based on cosine distance, feature clustering, or any other metric can also be identified.

[0023] For example, it can be concluded that product contamination during the manufacturing process or light reflection on the product surface under specific lighting conditions can cause defective products to be assigned to the "non-conforming" category and thus discarded as waste.

[0024] Conversely, this analysis could also reveal, for example, that specific defects in the physical image recordings during quality control—such as poor depth of field or insufficient lighting in the camera used—cause images of defective products to be assigned to the "acceptable" category, and thus the defective product incorrectly passes through the quality control process. In a production process that produces a large number of perfectly matched products, quality defects in individual products are indicated with a high probability by clearly demarcated individual damage. For example, a scratch might run across a surface, or a coating might peel off at a specific location. If such damage cannot be identified due to poor image quality, a defective product may not be recognized as defective.

[0025] To specifically investigate problems during image recording, in another particularly advantageous design, the operator used to generate the intermediate image is additionally modified, and the sensitivity's dependence on the operator is analyzed. For example, the operator can be characterized by parameters that can be changed. In this analysis, it can be found, for example, that even minimal variation in depth of field significantly worsens many erroneous recognitions, while the recognition still functions even under significantly worse lighting conditions.

[0026] Finally, it can be identified, for example, whether the image classifier might be making its judgments based not on relevant features of the image at all, but on other features. For instance, if it is difficult to identify the features of the actual product during quality control, the image classifier could bypass this and instead examine the robotic arm holding the product in the camera.

[0027] In another particularly advantageous design, in response to the image classifier's determined sensitivity meeting pre-given criteria, the product involved in the input image is pre-labeled for manual post-production control, and / or the conveyor is manipulated to sort the product out of the production process. This saves the significant additional technical costs required for recording and analyzing images within the scope of automated quality control, costs that would otherwise be necessary to automatically clarify all suspicious and boundary cases. Manual post-production control of a few examples of mass-produced products is significantly more economically advantageous than increasing the hit rate to the point of completely eliminating suspicious cases requiring post-production control in automated quality control.

[0028] The variations produced in the described manner can not only be used to measure the current sensitivity of the image classifier to changes in the input image, but more precisely, they can also be used to improve the training of the image classifier and thus, in particular, sharpen its recognition performance near the decision threshold. Therefore, in another advantageous design, the variations are used as additional training images for training the image classifier. For example, these training images can be manually labeled with the category that the image classifier should assign to them. This does not necessarily have to be the same category to which the input image is also assigned.

[0029] The present invention also relates to a method for training a generator used in the above-described method. In this method, a large number of training images are provided, wherein it is not necessary to know, for these training images, which training image should nominally be classified into which category by the image classifier. As previously described, intermediate images are generated from these training images respectively.

[0030] The parameters characterizing the generator's behavior are now optimized for the following objective: the variants generated by the generator are assigned to a pre-given category by the image classifier and are similar to the training images. This similarity can be measured selectively for all training images, or only for training images assigned by the image classifier to the same category as the variants to be generated by the generator.

[0031] Here, any metric can be used to quantitatively measure the assignment to a specific pre-given category to which the variant should be assigned by the image classifier. For example, any standard can be used to measure the difference between the actual classification score and the expected classification score associated with the pre-given category.

[0032] The degree to which a discriminator, trained simultaneously or alternately with the generator, can distinguish the variant from the training images is particularly advantageous for measuring the similarity between the variant and the training images. This comparison can be based optionally on the diversity of all training images or only the diversity of training images assigned by the image classifier to the same category as the variant to be generated by the generator. During training, the generator and discriminator form a Generative Adversarial Network (GAN), where only the generator continues to be used when measuring sensitivity after training has ended.

[0033] In another particularly advantageous design, a generator that produces images from categories is additionally trained such that the variants are as similar as possible to the training images, to which the image classifier also classifies the training images. In this way, it is ensured that the generator not only learns to produce variants “from nothing” (i.e., from noise), but also learns to reconstruct specific input images. For this training, “batches” of input images assigned to the same category by the image classifier can be aggregated, for example.

[0034] If an intermediate image generator is used to produce intermediate images, this intermediate image generator can be trained, for example, simultaneously or alternately with the generator. Thus, for example, it can be expected in the common cost function used for training that...

[0035] • Each generator produces a variant that is also actually classified into the desired target category by the image classifier;

[0036] • Each generator produces variants that are generally difficult to distinguish from the training images or training images of the desired target category;

[0037] • Each generator reconstructs the input image classified into the desired target category by the image classifier in the most optimal way;

[0038] • The intermediate image generator produces a realistic image that is as similar as possible to the input image or at least retains the basic features of the input image; and

[0039] The images generated by the intermediate image generator are simultaneously classified by the image classifier with the greatest possible uncertainty.

[0040] The cost function (loss function) is synthesized with any weights (optionally specified by hyperparameters). The generator can then be trained based on this cost function using any optimization method (e.g., ADAM or gradient descent), and optionally the intermediate image generator can also be trained.

[0041] This method can be implemented entirely or partially by a computer. Therefore, the invention also relates to a computer program having machine-readable instructions that, when executed on one or more computers, cause the computers to perform one of the described methods. In this sense, embedded systems of vehicle control devices and technical equipment capable of executing machine-readable instructions should also be considered as computers.

[0042] Similarly, the present invention also relates to machine-readable data carriers and / or downloadable products having computer programs. Downloadable products are digital products that can be transmitted via a data network, i.e., digital products that can be downloaded by users of said data network, and said digital products may be available for sale, for example, in online stores for immediate download.

[0043] In addition, the computer may be equipped with the computer program, the machine-readable data carrier, or the download product.

[0044] Other measures that improve the invention are shown in more detail below, together with the accompanying drawings, in conjunction with the description of preferred embodiments of the invention.

[0045] Other measures that improve the invention are shown in more detail below, together with the accompanying drawings, in conjunction with the description of preferred embodiments of the invention. Attached Figure Description

[0046] in:

[0047] Figure 1 An embodiment of a method 100 for measuring the sensitivity 2* of an image classifier 2 is shown;

[0048] Figure 2 An exemplary reclassification of variant 7 of input image 1 is shown;

[0049] Figure 3 An embodiment of a method 200 for training generator 6 is shown. Detailed Implementation

[0050] Figure 1This is a schematic flowchart of an embodiment of a method 100 for measuring the sensitivity 2* of an image classifier 2 to changes in an input image 1. According to step 105, the input image 1 is, in particular, an image of a nominally identical product that can be selected from mass production. The image classifier 2 can then be trained to classify the input image 1 into pre-given categories 3a-3c, which represent quality assessments of the corresponding products.

[0051] In step 110, the input image 1 is mapped to an intermediate image 5 using a pre-given operator 4, which may in particular be a trained intermediate image generator 4a. Compared to the input image 1, the intermediate image 5 has less information content and / or a worse signal-to-noise ratio. In step 120, at least one generator 6 is provided. The generator 6 is trained to produce realistic images of specific categories 3a-3c assigned by the image classifier 2 to a pre-given category.

[0052] Specifically, a separate generator 6 can be set up for each category 3a-3c. However, there are also applications where a single generator 6 is sufficient. For example, it is possible to simply ask why a product is classified into the worst quality category 3a-3c and is intended to be sorted out. Then, the single generator 6 can generate a realistic image that has been classified into that worst quality category 3a-3c by the image classifier 2.

[0053] In step 130, generator 6 generates a variant 7 of input image 1 from intermediate image 5. Specifically, for example, according to block 131, this variant can here be variant 7 classified by image classifier 2 into categories 3a-3c, different from the original input image 1. Variant 7 is then still identifiable based on the original input image 1, but is modified to exceed the decision threshold of image classifier 2 between the two categories. This directly reflects the sensitivity 2* of image classifier 2 to changes in input image 1.

[0054] Sensitivity 2* can be analyzed, for example, from comparisons 132 of one or more variants 7 with the input image 1 and / or from comprehensive statistics 133 of multiple variants 7, according to block 134. Then, additionally, the operator 4 used to generate the intermediate image 5 can be modified according to block 111, such that the dependence of the sensitivity 2* of the image classifier 2 on operator 4 can be analyzed according to block 135. As described above, this allows for the evaluation, in particular, of which measures to reduce information content or which types and intensities of noise have a particularly detrimental effect on the classification performed by the image classifier 2.

[0055] However, variant 7 can not only be used as a measure of sensitivity 2*, but can also be additionally used in step 140 as other training images for training image classifier 2. Training images close to the decision threshold of image classifier 2 are particularly suitable for sharpening that decision threshold.

[0056] In step 150, it can be checked whether the determined sensitivity 2* of the image classifier 2 (which may also be expressed as the dependence of sensitivity 2* on operator 4 2*(4)) meets a pre-given criterion. If it does (truth value 1), the product involved in the input image 1 can be pre-labeled, for example, in step 160, for manual post-processing control. Alternatively or in combination with this, the conveyor 8 can be manipulated in step 170 to sort the product out of the production process.

[0057] Figure 2 The example illustrates how a variant 7 of input image 1 can be classified by image classifier 2 into categories 3a-3c that are different from the original input image 1 by generating an intermediate image 5 with reduced information and then reconstructing it using generator 6.

[0058] Input image 1 shows a nut 10 in the form of a regular hexagon with an internal thread 11 in the middle. Due to a material or manufacturing defect, a crack 12 extends from the outer periphery of the internal thread 11 to the outer edge of the nut 10. If this input image 1 is fed to image classifier 2, the input image is classified into category 3a, which corresponds to the quality assessment "NOK".

[0059] However, if the input image is converted into an intermediate image 5 with reduced information using the blurring operator 4, the corners of the nut 10 are rounded. The basic shape of the hexagon can still be identified. However, this blurring has resulted in the crack 12 being visible only in a strongly reduced manner.

[0060] Generator 6 is trained to produce realistic images classified into category 3b by image classifier 2, which corresponds to the quality criterion "OK". When intermediate image 5 is fed to generator 6, the last remnants of crack 12 that were still identifiable in intermediate image 5 also disappear.

[0061] A comparison of variant 7 with the original input image 1 shows that crack 12 causes a difference between classifying it into category 3a = "unqualified" and classifying it into category 3b = "qualified".

[0062] Figure 3 This is a schematic flowchart of an embodiment of a method 200 for training generator 6. In step 210, a large number of training images 1# are provided. In step 230, intermediate images 5# are generated from these training images 1# using at least one pre-given operator 4.

[0063] In step 230, the parameter 6* representing the behavior of generator 6 is optimized for the following objectives: the variant 7 generated by generator 6 is assigned to a pre-given category 3a-3c by image classifier 2 and is similar to the training image 1#.

[0064] Here, according to block 231, the similarity between variant 7 and training image 1# can be measured by the degree to which a discriminator trained simultaneously or alternately with generator 6 can distinguish variant 7 from training image 1#.

[0065] According to block 232, generator 6 can be additionally trained to produce images from categories 3a-3c such that variant 7 is as similar as possible to training image 1#, and image classifier 2 also classifies training image 1# into that category.

[0066] According to block 233, an intermediate image generator 4a that generates intermediate images 5 can be trained simultaneously or alternately with generator 6.

[0067] Optimization of parameter 6* can continue until any termination criterion is met. The state representation of the completed generator 6 after training of parameter 6* is completed.

Claims

1. A method (100) for measuring the sensitivity (2*) of an image classifier (2) to changes in an input image (1), the image classifier assigning the input image (1) to one or more categories (3a-3c) of a pre-given category, the method comprising the following steps: • The input image (1) is mapped (110) to an intermediate image (5) by at least one pre-given operator (4), the intermediate image having less information content and / or a worse signal-to-noise ratio compared to the input image (1); • Provide (120) at least one generator (6), said at least one generator being trained to produce realistic images of a specific category (3a-3c) assigned by said image classifier (2) to the pre-given category; • The generator (6) generates (130) a variant (7) of the input image (1) from the intermediate image (5). The specific category (3a-3c) to which the variant (7) generated by the generator (6) belongs is a different category (3a-3c) from the category (3a-3c) to which the input image (1) is assigned (131) by the image classifier (2), and wherein • Comparison (132) of one or more variants (7) with the input image (1), and / or • From the comprehensive statistics on multiple variants (7) (133) Analyze (134) the sensitivity (2*) of the image classifier (2) to changes in the input image (1).

2. The method (100) according to claim 1, wherein the pre-given operator (4) comprises: • Reduce the input image (1) to a lower pixel resolution; and / or • A dimension-reduced representation of the input image (1) is generated using an encoder with an autoencoder structure, and / or • Add additive and / or multiplicative noise; and / or • Blur at least a portion of the input image; and / or • At least one affine transformation, and / or • Other trained intermediate image generators (4a).

3. The method (100) according to claim 2, wherein the intermediate image generator (4a) is trained to generate a realistic image that is as similar as possible to the input image (1) and is classified by the image classifier (2) with the greatest possible uncertainty.

4. The method (100) according to any one of claims 1 to 3, wherein the operator (4) is modified (111) in order to generate the intermediate image (5), and wherein the sensitivity (2*) of the image classifier (2) is analyzed (135) for its dependence on the operator (4).

5. The method (100) according to any one of claims 1 to 4, wherein the variant (7) is used (140) as another training image for training the image classifier (2).

6. The method (100) according to any one of claims 1 to 5, wherein (105) images of nominally identical products produced in batches are selected as input images (1), and wherein the image classifier (2) is trained to assign the input images (2a-3c) to one or more of at least two possible categories (3a-3c), the at least two possible categories representing quality assessments of the corresponding products.

7. The method (100) according to claim 6, wherein in response to the determined sensitivity (2*) of the image classifier (2) meeting a pre-given criterion (150), the product involved in the input image (1) is pre-labeled (160) for manual post-production control, and / or the conveyor (8) is manipulated (170) to sort the product out of the production process.

8. A method (200) for training a generator (6) used in the method (100) according to any one of claims 1 to 7, comprising the following steps: • Provides (210) a large number of training images (1#); • Use at least one pre-given operator (4) to generate (220) intermediate images (5#) from the training image (1#); • Optimize (230) the parameters (6*) characterizing the behavior of the generator (6) for the following objective: the variant (7) generated by the generator (6) is assigned to a pre-given category (3a-3c) by the image classifier (2) and is similar to the training image (1#).

9. The method (200) according to claim 8, wherein the similarity between the variant (7) and the training image (1#) is measured (231) by the degree to which the discriminator, trained simultaneously or alternately with the generator (6), can distinguish the variant (7) from the training image (1#).

10. The method (200) according to any one of claims 8 to 9, wherein the generator (6) that generates images from categories (3a-3c) is additionally trained (232) such that the variant (7) is as similar as possible to the training image (1#), and the image classifier (2) also classifies the training image (1#) into the category (3a-3c).

11. The method (200) according to any one of claims 8 to 10, wherein, An intermediate image generator (4a) is trained (233) simultaneously or alternately with the generator (6) to generate the intermediate image (5).

12. A computer program product containing machine-readable instructions, which, when executed on one or more computers, cause the one or more computers to perform the method (100, 200) according to any one of claims 1 to 11.

13. A machine-readable data carrier having a computer program containing machine-readable instructions that, when executed on one or more computers, cause the one or more computers to perform the method (100, 200) according to any one of claims 1 to 11.

14. A downloadable product having a computer program containing machine-readable instructions that, when executed on one or more computers, cause the one or more computers to perform the method (100, 200) according to any one of claims 1 to 11.

15. A computer equipped with a computer program product according to claim 12 and / or a machine-readable data carrier according to claim 13 and / or a download product according to claim 14.