Anomaly detection in digital images

The method improves anomaly detection in digital images by calculating semantic similarity scores between object classes using a knowledge graph and metrics, addressing the challenge of distinguishing normal from abnormal scenes in autonomous driving and manufacturing.

JP2026501872APending Publication Date: 2026-01-16ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025541859
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-18
Filing Date
2024-01-11
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing methods for anomaly detection in digital images, such as those in complex driving scenes and visual object classification, struggle to accurately distinguish between normal and abnormal object combinations, particularly in autonomous driving and manufacturing scenarios.

Method used

A computer-implemented method that determines semantic similarity scores between classes of objects in digital images using a knowledge graph, employing metrics like averages, weighted scores, and thresholds to enhance anomaly or normality detection, leveraging classifiers and neural networks for improved decision-making.

Benefits of technology

Enhances the accuracy of anomaly detection by reliably identifying normal and abnormal scenes or object combinations, enabling effective operation in autonomous driving and manufacturing applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026501872000001_ABST
    Figure 2026501872000001_ABST
Patent Text Reader

Abstract

The present invention relates to an apparatus and a computer-implemented method for processing digital images for the detection of anomalies or normalities, the method including providing a digital image (402), determining a first class for a first object and a second class for a second object depicted in the digital image (404) dependent on the digital image, determining a score dependent on a semantic similarity between the first class and the second class (406), and detecting anomalies or normalities dependent on the score (408, 410).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] background The present invention relates to a method and apparatus for processing digital images for the detection of abnormalities or normalities. [Background technology]

[0002] Biase, GD, Blum, H., Siegwart, R., Cadena, C.: Pixel-wise anomaly detection in complex driving scenes. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021. pp. 16918-16927. Computer Vision Foundation / IEEE (2021) discloses determining whether a given image is anomalous at the pixel level.

[0003] Eiter, T., Kaminski, T.: Exploiting contextual knowledge for hybrid classification of visual objects. In: JELIA. Lecture Notes in Computer Science, vol. 10021, pp. 223-239 (2016) discloses classifying images based on manually specified constraints, ensuring that none of the constraints are violated. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Biase, GD, Blum, H., Siegwart, R., Cadena, C.: Pixel-wise anomaly detection in complex driving scenes. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021. pp. 16918-16927. Computer Vision Foundation / IEEE (2021) [Non-patent document 2] Eiter, T., Kaminski, T.: Exploiting contextual knowledge for hybrid classification of visual objects. In: JELIA. Lecture Notes in Computer Science, vol. 10021, pp. 223-239 (2016) Summary of the Invention [Means for solving the problem]

[0005] Disclosure of the Invention A computer-implemented method for processing a digital image includes providing a digital image, determining a first class for a first object and a second class for a second object depicted in the digital image depending on the digital image, determining a score depending on a semantic similarity between the first class and the second class, and detecting anomalies or normalities depending on the score. The score indicates a pairwise semantic similarity between the classes. Information about the semantic similarities that influence the score enhances the detection of anomalies or normalities, thereby improving detection results.

[0006] The method may include determining a score for a plurality of class pairs of objects shown in the digital image, determining a metric dependent on the plurality of scores determined for the plurality of class pairs, and detecting abnormalities or normalities dependent on the metric.

[0007] The metric may include an average of a first score of the plurality of scores and a second score of the plurality of scores, and basing the detection on an average of the plurality of scores further improves the detection results.

[0008] The metric may include an average of multiple weighted scores, and basing the detection on an average of these scores further improves the detection results.

[0009] The method may include determining a probability that the digital image includes an object of a first class, determining a first weight depending on the probability that the digital image includes an object of the first class, and weighting scores for pairs including the first class by the first weight, which further improves detection results.

[0010] The metric may include an extreme score among the scores, in particular the minimum or maximum score. This metric is based on the extreme scores of the class pairs considered in the detection. This further improves the detection. Depending on the definition of the score calculation, the extreme score may indicate the most relevant class pair or the least relevant class pair. Depending on the definition, this improves the detection of normality or abnormality.

[0011] Determining the metric may include determining that the first score is less than the second score and determining the metric depending on the first score. By definition, a larger score may indicate a more related class pair or a less related class pair. Depending on the definition, considering smaller scores improves the detection of normality or abnormality.

[0012] Detecting abnormality or normality may include comparing the metric to a threshold and detecting the abnormality or normality depending on the result of comparing the metric to the threshold.

[0013] The method may include determining a parameter to indicate the confidence depending on the difference between the metric and a threshold, which provides additional information to explain the detection.

[0014] Determining the metric may include determining a list containing the scores, the scores being ordered in the list in particular ascending or descending order, which facilitates calculation.

[0015] Detecting anomalies or normalities may include classifying the list using a classifier having an output to indicate anomalies and / or an output to indicate normalities, among other things, which further improves the performance of the detection.

[0016] The method may include determining a behavior or output of the device depending on the detection of normality or abnormality.

[0017] An apparatus for processing digital images for the detection of abnormalities or normalities comprises at least one processor and at least one memory, the at least one processor configured to execute instructions that, when executed by the at least one processor, cause the apparatus to perform the method, and the at least one memory configured to store instructions, the apparatus having the advantages as described for the method.

[0018] The computer program is a computer program comprising instructions for causing a computer to carry out the method when the computer program is executed by a computer, the computer program having the advantages as described for the method.

[0019] Further advantageous embodiments can be derived from the following description and drawings. [Brief explanation of the drawings]

[0020] [Figure 1] 1 shows a schematic diagram of an apparatus for processing digital images; [Figure 2] FIG. 2 is a diagram illustrating a first digital image. [Figure 3] FIG. 2 is a diagram schematically illustrating a second digital image. [Figure 4] FIG. 1 shows a schematic diagram of a method for processing a digital image. DETAILED DESCRIPTION OF THE INVENTION

[0021] FIG. 1 shows a schematic diagram of an apparatus 100 for processing digital images.

[0022] The device 100 comprises at least one processor 102 and at least one memory 104 .

[0023] The device 100 comprises an interface 106 for a sensor 108 and / or the sensor 108. In the example shown in FIG.

[0024] The sensor 108 may be, for example, a camera, a radar sensor, a LiDAR sensor, a motion sensor, an infrared sensor, or an ultrasonic sensor.

[0025] The sensors 108 are configured to capture digital images. In one example, the sensors 108 are configured to capture sensor data from one or more sensors 108, for example, to determine a digital image of an environment of the device 100. The digital image may be determined by the device 100 depending on the sensor data.

[0026] The digital image may represent visual data, radar data, LiDAR data, ultrasound data, or a combination thereof.

[0027] The device 100 may be configured to detect objects in digital images. In one example, the device 100 is configured to detect objects from a set of digital images captured by the sensor 108.

[0028] The device 100 may be configured to detect classes of objects in digital images. In one example, the device 100 is configured to detect classes from a collection of digital images captured by a sensor 108.

[0029] The present disclosure relates to the problem of semantic anomaly detection for determining whether a particular measured scene or scenario is normal or abnormal. Given some input data, e.g., a collection of images, the device 100 is configured to determine whether the input data represent a realistic or normal combination of objects or an unrealistic or abnormal combination. As an example, a combination of objects such as a car, a pedestrian, and a road corresponds to a normal situation, while a combination of a car, a tiger, and a road is considered an abnormal situation in the present disclosure.

[0030] The problems considered are important and relevant in many applications, for example in autonomous driving or in the visual inspection of products assembled by robots.

[0031] For example, in the field of autonomous driving, the device 100 is a self-driving vehicle or part thereof configured to reliably distinguish between normal and abnormal combinations of objects. The device 100 is configured to base its determination on the normal combination of objects. In one example, the device 100 is configured to detect a scene having an abnormal combination of objects as a normal scene. For example, the device 100 is configured to detect a scene including a ride-on toy, a pedestrian, and a road as objects as a normal scene. In one example, the device 100 is configured to detect a scene having an abnormal combination of objects as an abnormal scene. In one example, the device 100 is configured to determine whether the abnormal combination of objects indicates a normal scene or an abnormal scene.

[0032] For example, in the manufacturing field, the device 100 may be an assembly or part thereof configured to detect the combination of parts in a product, particularly an automated manufactured product. The device 100 may be configured to determine whether the detected combination of parts indicates a normal combination of parts or an abnormal combination of parts. The device 100 may, for example, be configured to detect that the product includes a plastic top and a metal bottom. The device 100 may, for example, be configured to detect that this combination is abnormal for a particular electronic control unit architecture. The device 100 may be configured to identify a problem with the product if an abnormal combination is detected.

[0033] The device 100 may include an output unit 110. The output unit 110 is configured to, for example, output a result of the detection of normality or abnormality. The output unit 110 may be configured to control the operation of the device 100, for example, an action taken by the device 100.

[0034] 2 shows a first digital image 202. The first digital image 202 shows a portion of a road 204 and a vehicle 206. The vehicle 206 is traveling on the road 204 toward a cross street 208. A pedestrian 210 is crossing the cross street 208, while the vehicle 206 indicates its intention to turn onto the cross street 208. The first digital image 202 is a normal digital image. A normal digital image refers to a digital image that captures a scene that is likely to occur in a real-world scenario.

[0035] 3 shows a second digital image 302. The second digital image 302 shows a portion of a road 304 and a vehicle 306. The vehicle 306 is traveling on the road 304 toward a tiger 308 sitting on the road 304. The second digital image 302 is an abnormal digital image. An abnormal digital image refers to a digital image that captures a scene that is unlikely to occur in a real-world scenario.

[0036] The at least one memory 104 is configured to store instructions that, when executed by the at least one processor 102, cause the device 100 to perform steps in a method for processing digital images.

[0037] At least one memory 104 is configured to store a knowledge graph. The knowledge graph includes nodes that represent classes for objects present in the digital image. The knowledge graph includes edges that pairwise connect the nodes of the knowledge graph. The edges indicate relationships between the nodes and, therefore, the classes that these nodes represent.

[0038] A knowledge graph specifically represents a collection of factual information linked together as a directed graph.

[0039] A knowledge graph is encoded as a set of triples, for example (subject; predicate; object). In one triple, the subject of the triple corresponds to a node, the object corresponds to a node, and the predicate corresponds to an edge.

[0040] FIG. 4 shows the steps of the method.

[0041] The method processes a digital image to detect anomalies or normalities in the digital image, where anomalies refer to non-normal digital images, and normalities refer to normal digital images.

[0042] In step 402, a digital image is provided.

[0043] The digital image is captured, for example, by the sensor 108 .

[0044] The digital image includes an object.

[0045] In step 404, the class of the object is determined.

[0046] A plurality of classes are provided, and for at least one of the plurality of classes, a probability is determined that the digital image contains an object of that class, the probability indicating the likelihood that the digital image contains an object of the at least one class.

[0047] According to one example, the probability for at least one class is determined depending on the digital image.

[0048] According to one example, multiple objects are detected and a class is determined for at least one of the objects.

[0049] According to one example, a probability that an object is of one of the multiple classes is determined for each class of the multiple classes.

[0050] The objects are detected, for example, using an object detector. The probabilities or classes are determined, for example, using a classifier. An existing classifier or object detector that has already been pre-trained can be used.

[0051] In one example, the classifier assigns multiple probabilities to an object, one of the multiple probabilities indicating, for a respective class of multiple classes, the likelihood that the object is of that class.

[0052] The method may use object detection or may determine the probability of an object's class without using object detection.

[0053] An example for classifying detected objects using a classifier or object detector is the CLIP model disclosed in Radford, A., Kim, JW, Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. In: ICML. Proceedings of Machine Learning Research, vol. 139, pp. 8748-8763. PMLR (2021).

[0054] The method will be further described by considering a first class, a second class and a third class of the plurality of classes.

[0055] The first class, the second class, and the third class are determined depending on the digital image. For the first class, a first probability that an object of the first class is shown in the digital image is determined. For the second class, a second probability that an object of the second class is shown in the digital image is determined. For the third class, a third probability that an object of the third class is shown in the digital image is determined.

[0056] Other probabilities may be determined for other classes.

[0057] In step 406, a score s is calculated for a class pair including class i and class j. ij is determined.

[0058] Score ij is determined for a number of class pairs of objects shown in the digital image.

[0059] Score s for one pair ij is determined depending on the semantic similarity between class i and class j, which is determined depending on the knowledge graph.

[0060] According to one example, the higher the score, the more related the classes used to determine the score are. That is, the higher the score for a class pair, the more likely the combination of objects from the two classes in that pair is indicative of a normal digital image. That is, the lower the score for a class pair, the more likely the combination of objects from the two classes in that pair is indicative of a non-normal digital image.

[0061] The score definition may be reversed and the following metrics and detection decisions may be adjusted accordingly.

[0062] The method may include determining a score for any pair of classes of the plurality of classes.

[0063] The knowledge graph is for example a ConceptNet or WebChild to take into account the semantic relationships of objects that may occur in a scene.

[0064] ConceptNet is described in Speer, R., Chin, J., Havasi, C.: Conceptnet 5.5: An open multilingual graph of general knowledge. In: AAAI. pp. 4444-4451. AAAI Press (2017).

[0065] WebChild is described in Tandon, N., de Melo, G., Suchanek, FM, Weikum, G.: Webchild: harvesting and organizing commonsense knowledge from the web. In: WSDM. pp. 523-532. ACM (2014).

[0066] A score can be determined for pairs of classes that are assigned a higher probability than the other classes. In one example, a score is determined for the k classes that are assigned the highest probabilities.

[0067] According to one example, the classes for determining the score are selected depending on the probabilities of those classes.

[0068] In one example, the classes for determining the score are selected because the probabilities assigned to them are higher than the probabilities assigned to other classes.

[0069] A score is determined for each of the two classes, and the score depends on the semantic similarity between the classes.

[0070] In step 408, a metric is determined depending on the first score and the second score.

[0071] In one example, determining the metric includes determining an average of multiple scores.

[0072] In one example, an average of the first score and the second score is determined.

[0073] In one example, the k scores s with the highest probabilities ij The average of is determined.

number

[0074] In one example, determining the metric includes determining an average of a number of weighted scores.

[0075] In one example, the average of the weighted scores is the weight p for class i. i and the weight p for class j j and the k scores s with the highest probability ij is determined using

number

[0076] In one example, the weights are determined depending on the probabilities assigned to the classes, and these weights determine a weighted score.

[0077] In one example, the average is determined depending on the weighted scores, for example, the first weighted score is determined depending on the first weight and the first score, and the second weighted score is determined depending on the second weight and the second score.

[0078] For example, the first weight is determined depending on the probability that the digital image contains an object of a first class and the probability that the digital image contains an object of a second class, and the second weight is determined depending on the probability that the digital image contains an object of the first class and the probability that the digital image contains an object of a third class.

[0079] The weighting may be based on predetermined human rules, for example, a human may define that the class pair of infant and car should have a higher weight than the original unweighted semantic score because it may represent a dangerous / unnormal situation. The weighting may be determined by predetermined rules or by a field of use. In some embodiments, not all abnormalities or scores are equal. Rules may be used to identify class pairs that appear normal based on the original unweighted semantic scores as abnormal.

[0080] In one example, determining the metric includes determining an extreme score, which may be the minimum or maximum score in a plurality of scores determined for the digital image.

[0081] In the case of the minimum score, for example, the k scores s with the highest probability are ij The smallest of these is determined.

number

[0082] The method is not limited to using a minimum or maximum score. The method may include finding a small score that is not necessarily a minimum score. The method may include finding a large score that is not necessarily a maximum score. Determining the metric may include determining that a first score of the plurality of scores is less than a second score of the plurality of scores, and determining the metric dependent on the first score.

[0083] Instead of determining a scalar metric as described above, determining the metric may include determining a list containing multiple scores, with the multiple scores being ordered in ascending or descending order in the list, among other things.

[0084] In step 410, abnormalities or normalities are detected depending on the metric.

[0085] According to one example, detecting the anomaly includes determining that the metric is below a threshold.

[0086] According to one example, detecting normality includes determining that a metric is above or equal to a threshold.

[0087] The threshold may be determined during training.

[0088] For example, a scalar metric is determined for a plurality of digital images, including images known to depict normal scenes and digital images known to depict abnormalities.

[0089] The threshold value is selected to be the largest possible value such that an abnormality is correctly detected for a predetermined amount or percentage of the plurality of digital images, an abnormality is incorrectly detected for a predetermined amount or percentage of the plurality of digital images, normality is correctly detected for a predetermined amount or percentage of the plurality of digital images, and / or normality is incorrectly detected for a predetermined amount or percentage of the plurality of digital images.

[0090] Preferably, the threshold is determined from digital images known to represent normal scenes. The threshold can be adjusted based on experimental scores. If most normal pairs exceed 0.2, a threshold of 0.2 can be selected. The threshold is selected to be the largest possible value such that normality is correctly detected for a predetermined amount or percentage of digital images in the plurality of images, or abnormality is incorrectly detected for a predetermined amount or percentage of digital images in the plurality of digital images. This avoids the use of digital images containing anomalies.

[0091] Different class lists can be defined for examples of abnormal and normal combinations. The class names are used to calculate the scores and corresponding scalar metrics.

[0092] The appropriate threshold is selected as above from the class name instead of the digital image. The advantage of this approach is that no input data for the digital image classifier is required.

[0093] The method may include determining a parameter for indicating the confidence level depending on the difference between the metric and a threshold to which the metric is compared.

[0094] Detecting anomalies or normalities may include classifying the list using a classifier having an output to indicate anomalies and / or an output to indicate normalities, among other things. The classifier may be a neural network configured to process the list and determine an output.

[0095] An example of such a neural network includes two fully connected layers. More complex neural networks are possible. When properly trained, the output of the neural network's activation function, e.g., softmax, can be used as a confidence score for the detection. Alternatively, the output can be calibrated by an uncertainty method and used as a confidence score for the detection.

[0096] The results of the normality or abnormality detection may be stored or output. The results may indicate the abnormality or normality of the digital image. The results may also include a confidence level.

[0097] During training, it is not necessary to use all possible classes seen during inference. Appropriate thresholds or appropriate network weights are learned during training. Through information in the knowledge graph, new class pairs with similar high or low similarity scores will still yield correct normal or abnormal classification results.

[0098] In step 412, the operation of the device 100, for example an output or action by the device 100, may be controlled depending on the result.

[0099] For example, if the results indicate the normality of the digital image, the digital image can be used to determine the behavior of the device 100 .

[0100] For example, if the results indicate an abnormality in the digital image, the digital image may be ignored for purposes of determining the behavior of the device 100 .

[0101] In the field of autonomous driving, the action may be driving a vehicle.

[0102] In the manufacturing sector, the action may be the use or selection of a product.

Claims

1. 1. A computer-implemented method for processing a digital image for detection of abnormalities or normalities, comprising: The method comprises: Providing a digital image (402); determining (404) a first class for a first object and a second class for a second object shown in the digital image dependent on the digital image; determining a score depending on the semantic similarity between the first class and the second class (406); Detecting abnormalities or normalities depending on the scores (408, 410); A method comprising:

2. The method comprises: determining (406) scores for a plurality of class pairs of objects shown in the digital image; determining 408 a metric dependent on the scores determined for the plurality of class pairs; Detecting the abnormality or normality depending on the metric (410); The method of claim 1 , comprising:

3. The method of claim 2 , wherein the metric comprises an average of the plurality of scores.

4. The method of claim 2 , wherein the metric comprises an average of multiple weighted scores.

5. The method comprises: determining (404) a probability that the digital image contains an object of the first class; determining (408) a first weight depending on the probability that the digital image contains an object of the first class; weighting the scores for pairs that include the first class by the first weight; The method of claim 4, comprising:

6. The method of claim 2 , wherein the metric comprises an extreme score, in particular a minimum or maximum score, of the plurality of scores.

7. Determining the metric (408) includes: determining that a first score of the plurality of scores is less than a second score of the plurality of scores; determining the metric dependent on the first score; The method of claim 2 , comprising:

8. Detecting the abnormality or normality (410) includes: comparing the metric to a threshold; detecting said abnormality or said normality depending on a result of comparing said metric with said threshold; The method of any one of claims 1 to 7, comprising:

9. The method comprises: determining (410) a parameter for indicating a confidence level depending on the difference between the metric and the threshold; The method of claim 8, comprising:

10. Determining the metric (408) includes determining a list including the plurality of scores; the scores are in particular ordered in ascending or descending order in the list, The method of claim 2.

11. 11. The method of claim 10, wherein detecting (410) the anomaly or the normality comprises classifying the list using a classifier having, inter alia, an output for indicating an anomaly and / or an output for indicating normality.

12. The method comprises: Determining (412) the behavior or output of the device (100) depending on the detection of normality or abnormality.

12. The method of claim 1, comprising:

13. An apparatus (100) for processing digital images for the detection of abnormalities or normalities, comprising: The device (100) comprises: at least one processor (102); At least one memory (104); Equipped with The at least one processor (102) is configured to execute instructions that, when executed by the at least one processor (102), cause the apparatus (100) to perform the method of any one of claims 1 to 12; the at least one memory (104) is configured to store the instructions; 1. An apparatus (100) comprising:

14. 13. A computer program comprising instructions for causing a computer to carry out the method of any one of claims 1 to 12 when the computer program is run by the computer.

Citation Information

Patent Citations

  • Method and system for anomaly detection using multimodal knowledge graph

    EP3816856A1

  • Identification apparatus and identification method

    JP2020087037A