Method for determining whether an object represented in an input image has an anomaly
Patent Information
- Application Number
- DE102024201098
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-07
- Publication Date
- 2025-08-07
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The present disclosure relates to methods for determining whether an object represented in an input image has an anomaly.
[0002] The detection of anomalous regions on objects in images in an industrial environment is a common problem, often referred to as "industrial anomaly detection." Such image-based detection of anomalies such as damage, missing, rearranged, or altered parts on objects in an industrial environment can be relevant in various contexts where objects are handled automatically and can be an essential component of factory automation.
[0003] Accordingly, efficient methods for determining whether an object represented in an input image has an anomaly are desirable.
[0004] According to various embodiments, a method for determining whether an object represented in an input image has an anomaly is provided, comprising: • Generating, for each reference image of one or more reference images, each reference image showing at least one reference object of one or more reference objects, a respective reference image descriptor image, • Generating an input image descriptor image for the input image, • Determining a region of the input image that shows an object that corresponds to one of the reference objects by comparing descriptor values of the input image descriptor image with descriptor values of at least one of the reference image descriptor images, • Determining an assessment of the correspondence of the object shown by the determined area with the reference object to which the object shown by the determined area corresponds, based on a comparison of the descriptor values contained in the input image descriptor image for the determined area with the descriptor values contained in the respective reference image descriptor image for the area of the respective reference image showing the reference object, and • Marking the object as having an anomaly depending on the assessment determined.
[0005] The method described above enables anomaly detection based on dense visual descriptors, thus leveraging a framework that can also be used for other tasks. For example, dense visual descriptors could be used for object detection or for the recognition of keypoints on objects. These keypoints could, for example, be used for grasping an object (i.e., a suitable location on the object where it is grasped). If such a dense descriptor network (or "dense object net") has already been trained for another such task (e.g., for a set of reference objects), it can be used for anomaly detection (and thus, for example, for the detection of broken objects during robot manipulation) without further training.
[0006] Furthermore, the method proposed above uses dense descriptors to detect a reference object on the test image and does not require object annotations (i.e., labels) indicating which objects are visible on the respective test image. Furthermore, it can also be applied to a test image containing multiple object instances (i.e., the same object multiple times).
[0007] Various examples of implementation are given below.
[0008] Embodiment 1 is a method for determining whether an object represented in an input image has an anomaly (ie, an anomaly detection method) as described above.
[0009] Embodiment 2 is a method according to embodiment 1, comprising generating the reference image descriptor image by feeding the respective reference image to a dense object mesh and generating the input image descriptor image by feeding the input image to the dense object mesh.
[0010] A dense object network can be efficiently trained to produce high-quality descriptor images (i.e., assigning similar descriptors to object locations that correspond to each other and assigning very different descriptors to object locations that do not correspond to each other).
[0011] Embodiment 3 is a method according to Embodiment 1 or 2, wherein determining the region of the input image comprises determining one of the reference images showing the reference object to which the object shown by the region corresponds, and performing an affine transformation between the region and a region of the determined reference image showing the reference object to which the object shown by the region corresponds.
[0012] This assumes and exploits the rigidity of the object, allowing for high accuracy in determining the area.
[0013] Embodiment 4 is a method according to any one of embodiments 1 to 3, wherein the anomaly is damage.
[0014] In particular, damaged objects can be detected and, for example, sorted out. However, other anomalies can also be detected, such as missing or altered parts of an object.
[0015] Embodiment 5 is a method according to any one of embodiments 1 to 4, further comprising, if the object is identified as an object having an anomaly, determining a location of the object having the anomaly based on the or a further comparison of the descriptor values contained in the input image descriptor image for the determined area with the descriptor values contained in the respective reference image descriptor image for the area of the respective reference image showing the reference object.
[0016] The descriptor images can also be used to more accurately localize anomalies (e.g. damage) on objects by comparing the descriptors.
[0017] Embodiment 6 is a method for controlling a technical system, comprising • Receiving an input image; • Checking whether an object shown in the input image has an anomaly according to any one of embodiments 1 to 5 • Treating (e.g., manipulating, grasping, etc.) the object depending on whether it has been marked as having an anomaly.
[0018] Embodiment 7 is a method according to embodiment 6, comprising sorting out the object if it has been marked as an object having an anomaly.
[0019] Embodiment 8 is a method according to embodiment 6 or 7, further comprising determining a location and / or pose for capturing or processing the object using the input image descriptor image.
[0020] This allows the input image descriptor (and a corresponding tool for generating descriptor images, e.g., a suitably trained machine learning model such as a dense object network) to be used both for anomaly detection and for determining a location and / or pose for picking up or processing the object (e.g., a grasping pose). Thus, if a corresponding tool for determining a location and / or pose for picking up or processing the object is already available, no additional tool for anomaly detection is required.
[0021] Embodiment 9 is a data processing device (in particular control device) which is configured to carry out a method according to one of the embodiments 1 to 8.
[0022] Embodiment 10 is a computer program including instructions that, when executed by a processor, cause the processor to perform a method according to any one of embodiments 1 to 8.
[0023] Embodiment 11 is a computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform a method according to any one of embodiments 1 to 8.
[0024] In the drawings, like reference characters generally refer to the same parts throughout the several views. The drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the invention. In the following description, various aspects are described with reference to the following drawings. Fig. 1 shows a robot. Fig. 2 illustrates a flow for detecting anomalies according to one embodiment. Fig. 3 shows a flowchart illustrating a method for determining whether an object represented in an input image has an anomaly, according to one embodiment.
[0025] The following detailed description refers to the accompanying drawings, which, by way of illustration, show specific details and aspects of this disclosure in which the invention may be practiced. Other aspects may be utilized, and structural, logical, and electrical changes may be made without departing from the scope of the invention. The various aspects of this disclosure are not necessarily mutually exclusive, as some aspects of this disclosure may be combined with one or more other aspects of this disclosure to form new aspects.
[0026] Various examples are described in more detail below.
[0027] Fig. 1 shows a robot 100.
[0028] The robot 100 includes a robot arm 101, for example, an industrial robot arm for handling or assembling a workpiece (or one or more other objects). The robot arm 101 includes manipulators 102, 103, 104 and a base (or support) 105 by means of which the manipulators 102, 103, 104 are supported. The term "manipulator" refers to the movable components of the robot arm 101, the actuation of which enables physical interaction with the environment, e.g., to perform a task. For control, the robot 100 includes a (robot) controller 106 designed to implement the interaction with the environment according to a control program. The last component 104 (furthest from the support 105) of the manipulators 102, 103, 104 is also referred to as the end effector 104 and may include one or more tools, such as a welding torch, a gripping instrument, a painting device or the like.
[0029] The other manipulators 102, 103 (located closer to the support 105) can form a positioning device, so that, together with the end effector 104, the robot arm 101 is provided with the end effector 104 at its end. The robot arm 101 is a mechanical arm that can provide functions similar to a human arm (possibly with a tool at its end).
[0030] The robot arm 101 may include joint elements 107, 108, 109 that connect the manipulators 102, 103, 104 to each other and to the support 105. A joint element 107, 108, 109 may have one or more joints, each of which can provide rotational movement (i.e., rotary movement) and / or translational movement (i.e., displacement) for associated manipulators relative to each other. The movement of the manipulators 102, 103, 104 may be initiated by actuators controlled by the controller 106.
[0031] The term "actuator" can be understood as a component configured to effect a mechanism or process in response to its drive. The actuator can implement instructions generated by the controller 106 (the so-called activation) into mechanical movements. The actuator, e.g., an electromechanical transducer, can be configured to convert electrical energy into mechanical energy in response to its activation.
[0032] The term "controller" can be understood as any type of logic-implementing entity, which may, for example, include a circuit and / or a processor capable of executing software, firmware, or a combination thereof stored in a storage medium, and which can issue instructions, e.g., to an actuator in the present example. The controller can, for example, be configured by program code (e.g., software) to control the operation of a system, a robot in the present example.
[0033] In the present example, the controller 106 includes one or more processors 110 and a memory 111 that stores code and data based on which the processor 110 controls the robot arm 101. According to various embodiments, the controller 106 controls the robot arm 101 based on a machine learning model 112 stored in the memory 111.
[0034] One task of robot 100 may be to sort out defective objects from a group of objects 113, e.g., before they are packaged and sold. To do this, robot 100 must be able to recognize defective objects—that is, more generally, detect anomalies.
[0035] Accordingly, according to various embodiments, the machine learning model 112 is configured and trained to enable the robot 100 to detect anomalies in the objects 113, e.g., to determine whether an object 113 is defective.
[0036] The robot 100 may, for example, be equipped with one or more cameras 114 that enable it to take images of its workspace and thus of the objects 113.
[0037] According to various embodiments, the machine learning model 112 is a neural network 112 and the controller 106 provides input data to the neural network 112 based on the one or more digital images (color images, depth images, or both) of one or more of the objects 113.
[0038] According to various embodiments, anomaly detection is based on learned dense visual descriptors, i.e., the neural network 112 is a dense object network according to various embodiments. A dense object network (DON) maps an image to an arbitrarily dimensional (D dimension) descriptor space image. The dense object network is a neural network that can be trained using self-supervised learning to output a descriptor space image for an input image. This allows images of known objects to be mapped to descriptor images containing descriptors that identify locations on the objects regardless of the perspective of the image. A respective DON can be provided for each object type. The descriptors make it possible, for example, to find grasping poses or general locations for manipulating an object (e.g., grasping, welding, etc.), for example as follows: (1) For each object type of interest, a (reference) image of a reference object of the object type is mapped from the respective DON to a descriptor image. (2) A user selects pixels (e.g., by clicking) on the image of the object where the object is to be grasped, and the descriptors of the selected pixels (i.e., the descriptor value to which the pixel was mapped by the DON 115) are registered. The user can also select pixels, for example, by selecting regions, e.g., by clicking on the vertices of a polygon or pixels within a corresponding convex hull. This is repeated until all regions or locations on the surface desired by the user have been marked (and the associated descriptors have been registered). Analogously, the user can also select locations that are to be avoided when picking up (e.g., grasping) the object. (3) (1) and (2) can be repeated for multiple reference images to capture all sides of an object. This results in a known (and registered) set of descriptors corresponding to preferred (or avoided) locations for capturing (these can be considered "keypoints"). Using the DON, these locations of an object can be detected in the input image by searching for the registered descriptors in the input image. Accordingly, the following procedure is followed: (4) For the input image a. The DON determines a descriptor image from the input image b. The descriptor image for the input image is searched for the registered descriptors, and a manipulation preference image (e.g., with pixel values in the interval [0, 1]) is generated such that the pixel value of a pixel indicates how well the pixel's descriptor matches one of the registered descriptors, illustratively in the form of a "(descriptor match) heatmap" with respect to the match with the registered descriptors and thus with the locations selected by the user. This can be done, for example, by generating a respective heatmap for each registered descriptor, and forming the manipulation preference image as a (pixel-wise) maximum across all these heatmaps.
[0039] A dense (visual) descriptor network thus transforms, for example, RGB values view-invariantly into a multidimensional 1-1 representation (i.e., pixel-by-pixel) with the same output width and height, but possibly different channel dimensions. Such a representation space promises that descriptors located at the same location of an object on two different images will be very close to each other, even under lighting changes or physical transformations, while descriptors that the descriptor network assigns to different locations of an object will be different. When the descriptor images of two objects are overlaid, the difference between the descriptor images indicates how similar the objects are. It is to be expected that different regions (i.e., regions showing different objects) will also exhibit a large difference in the descriptor space.Therefore, if one of the two objects is a normal (reference) object, the difference in descriptor values can be interpreted as an anomaly score using a specific metric (e.g., I2 norm). Areas with a high anomaly score likely contain anomalies, while areas with a low anomaly score are likely to be very similar to the reference object.
[0040] To detect anomalies in an input image (hereinafter, the image on which possible anomalies are to be detected is referred to as the “test image”), a DON (representing a learned dense representation of reference objects) can be used to detect a reference object on the test image, an affine transformation (rotation, translation and shear) of the object from the respective reference image to the test image (i.e., the representation of the object in the test image) can be estimated (and thus the area in the test image that represents the object that corresponds to the reference object (i.e., e.g., is equal to) can be determined), and the descriptor values of the representation of the object in the reference image and in the test image can be compared. Based on the differences in the descriptor values (e.g., the descriptor value ranges) of the area in the reference image that shows the reference object (i.e.,An anomaly score is calculated based on the region in the test image containing the (matching) object. If this score exceeds a threshold, the region in the test image containing the object is labeled as an abnormal region (and thus the object is also labeled as having an anomaly).
[0041] Fig. 2 illustrates a flow for detecting anomalies according to one embodiment.
[0042] It is assumed that a trained dense object network 201 (possibly for each object type) is available. Using the trained dense object network 201, a respective (reference) descriptor image 204 is generated for one or more reference images 202, each of which shows one or more reference objects 203 in a normal (undamaged) state.
[0043] For a test image 205 containing a potentially damaged (test) object 206, the object 206 is first identified based on the reference objects 203. Each reference object 203 is a "normal" version (i.e., a reference version) of an object (or, in other words, a reference instance of a specific object or object type).
[0044] The test image 205 is also mapped to a (test) descriptor image 207 using the dense object network 201 (if necessary per object type).
[0045] To detect the object(s) present in the test image 205, the reference descriptor image(s) 204 can be used, e.g., by tracking keypoints from the respective reference image 202 to the test image 205 via the descriptors as described above. This means that a matching reference image 202 is searched for the test image 205 that shows a reference object 203 of the same object type as the test object 206, i.e., a region in the test image 205 is searched for that shows an object corresponding to the reference object 203 shown in the reference image 202. This is done by ensuring that the descriptor values in a region of the reference image 202 match the descriptor values in the region of the test image 205 well, e.g. the I2 norm of the pixel-wise difference of the descriptor values of the regions is below a certain threshold (whereby, if several DONs are present for several object types, such a comparison of descriptor values may be necessary).is only performed for descriptor images generated by the DON for the same object type).
[0046] If such a reference image 202 or such an area in the test image 205 has been found, an affine transformation of the recognized reference object from the reference image to the test image, ie from the area of the reference image 202 to the area of the test image 205, whose descriptor values match well, is determined.
[0047] If several test objects 206 are detected in the test image 205, a respective affine transformation can be estimated for each of these object instances (possibly from different reference images 202 if they are test objects that match different reference objects).
[0048] The affine transformation takes into account, for example, rotation, translation, and shear. It is assumed that the test object 206 is a rigid body. The affine transformation can be estimated, for example, using the RANSAC algorithm to robustly adapt a transformation based on a series of detected keypoints. Another possibility is to use an optimization method that finds the best adaptation (by transformation) of the region of the reference image 202 showing the respective object to the region of the test image 206 that was detected to represent the object, with respect to the descriptor values (i.e., such that the region of the reference image 202 is superimposed onto the region of the test image 206 by the transformation in such a way that the pixel-by-pixel difference between the (correspondingly superimposed) descriptor values is as small as possible on average (e.g., in the sense of the I2 norm)).
[0049] If such a transformation has been found, a difference descriptor image 208 is formed, which contains, for each pixel (at least in the area 209 corresponding to the found object), the difference between the reference image descriptor image 204 transformed according to the transformation and the test descriptor image 207. More generally, the descriptor values contained in the descriptor image 207 of the test image 205 for the area 209 are compared with the descriptor values contained in the reference image descriptor image for the area of the reference image 202 showing the (matching) reference object (i.e., the area mapped to the area 209 by the transformation).
[0050] The magnitude of the pixel values (i.e., the difference between the descriptor values of the reference image descriptor image 204 and the test descriptor image 207) in the region 209 in the difference descriptor image 208 corresponding to the detected object, e.g., determined in the form of the I2 norm of the pixel values of the difference descriptor image 208 in the region 209, is used as (or as the basis for) an anomaly assessment. Thus, the pixel-by-pixel difference between the descriptor values for the reference object 203 and the test object 206 serves as the basis for the anomaly assessment.
[0051] The anomaly score is then compared, for example, with a threshold value. If it exceeds the threshold value, this is interpreted as meaning that the test object 206 has an anomaly (e.g., is damaged). One or more anomalous locations 210 of the test object 206 can also be found based on the local distribution of the difference values in the area 209. If multiple objects were found in the test image 205, this (determination of the transformation, determination of the anomaly score, and, if necessary, identification of anomalous locations of the test object 206) can be performed for each object found.
[0052] In summary, according to various embodiments, a method is provided as described in Fig. 3 shown.
[0053] Fig. 3 shows a flowchart 300 illustrating a method for determining whether an object represented in an input image (also referred to herein as a test image) has an anomaly, according to one embodiment.
[0054] In 301, for each reference image of one or more reference images, each reference image showing at least one reference object of one or more reference objects, a respective reference image descriptor image is generated (which contains descriptors for the respective reference image, for example, pixel-by-pixel, although subsampling may occur). The reference images can be captured in advance of the reference objects.
[0055] At 302, an input image descriptor image is generated for the input image (which contains descriptors for the input image, again, for example, pixel-by-pixel, whereby subsampling may also occur). According to one embodiment, the reference image descriptor images and the input descriptor image are generated according to the same image-to-descriptor image mapping rule (e.g., the same machine learning model, e.g., DON).
[0056] In 303, based on a comparison of descriptor values of the input image descriptor image with descriptor values of at least one of the reference image descriptor images, an area of the input image is determined that shows an object that corresponds to one of the reference objects (e.g., there is a reference object for each object type of interest and a reference object is searched that matches an area of the input image).
[0057] In 304, an assessment of the correspondence of the object shown by the determined region with the reference object to which the object shown by the determined region corresponds is determined based on a comparison of the descriptor values contained in the input image descriptor image for the determined region with the descriptor values contained in the respective reference image descriptor image for the region of the respective reference image showing the reference object.
[0058] In 305, the object is marked as having an anomaly depending on the determined rating. For example, it is marked as having an anomaly if the rating is above a predetermined threshold.
[0059] In other words, according to various embodiments, a dense (pixel-wise) representation of one or more reference objects is learned, which is view-invariant (i.e., invariant to rotations or translations of the respective reference object with respect to a camera), and for each object type for which possible anomalies are to be detected, one or more images containing normal (i.e., e.g., undamaged) instances of the object type (i.e., reference objects) are collected as a set of reference images. By comparing descriptor images of the normal objects (i.e., the reference images) and the descriptor image of a test image, anomalies (e.g., damaged or missing areas) are discovered on an object represented in the test image. In general, anomalies in a technical system can be detected in this way.
[0060] The (dense visual) ones are generated, for example, using a neural network (e.g., with a ResNet architecture), specifically a dense descriptor network. However, other approaches are also conceivable, as long as descriptors can be generated that are unambiguous, robust, and view (and light) invariant for surface points of the respective (reference) objects.
[0061] Anomalies can be identified in an input image by segmenting the input image accordingly (e.g., into regions showing normal objects, objects showing anomalies, and background).
[0062] Color and depth images, for example, serve as input data for the machine learning models. These can also be supplemented by sensor signals from other sensors such as radar, LiDAR, ultrasound, motion, thermal images, etc. For example, an RGB and depth image is captured in a robot cell, and the image (or several such images, e.g., a video) is used for anomaly detection.
[0063] The procedure of Fig. 3 can be performed by one or more computers having one or more data processing units. The term “data processing unit” can be understood as any type of entity that enables the processing of data or signals. The data or signals can, for example, be treated according to at least one (i.e., one or more than one) specific function performed by the data processing unit. A data processing unit can comprise or be formed from an analog circuit, a digital circuit, a logic circuit, a microprocessor, a microcontroller, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a programmable gate array (FPGA) integrated circuit, or any combination thereof.Any other way of implementing the respective functions described in more detail herein may also be understood as a data processing unit or logic circuit arrangement. One or more of the method steps described in detail herein may be performed (e.g., implemented) by a data processing unit through one or more specific functions performed by the data processing unit.
[0064] According to various embodiments, the method is therefore particularly computer-implemented.
[0065] The approach of Fig. 3 serves, for example, to generate a control signal for a robotic device. The term “robotic device” can be understood as referring to any technical system (with a mechanical part whose movement is controlled), such as a computer-controlled machine, a vehicle, a household appliance, a power tool, a manufacturing machine, a personal assistant, or an access control system. Data (e.g. scalar time series), in particular from a sensor, e.g. a camera, can be evaluated and then the technical system can be controlled accordingly. For example, depending on anomalies found (e.g. on an object), a signal can be generated for the technical system, e.g. to treat damaged objects in a certain way.
[0066] According to one embodiment, the method of Fig.3, for example, in a robotic application for removing objects from containers. In a combined application, a dense descriptor network can also be used to define specific gripping points on objects and simultaneously detect damage to the objects. This detected damage could then be used to automatically sort out the damaged objects and exclude them from further processing.
Claims
[1] A method for determining whether an object (113, 206) represented in an input image (205) has an anomaly, comprising: generating (301), for each reference image (202) of one or more reference images, each reference image (202) showing at least one reference object of one or more reference objects, a respective reference image descriptor image (204), generating (302) an input image descriptor image (207) for the input image (205), Determining (303) a region (209) of the input image (205) showing an object (113, 206) corresponding to one of the reference objects based on a comparison of descriptor values of the input image descriptor image (207) with descriptor values of at least one of the reference image descriptor images, Determining (304) an assessment of the correspondence of the object (113, 206) shown by the determined area (209) with the reference object to which the object (113, 206) shown by the determined area (209) corresponds, based on a comparison of the descriptor values contained in the input image descriptor image (207) for the determined area (209) with the descriptor values contained in the respective reference image descriptor image for the area of the respective reference image (202) showing the reference object; and Marking (305) the object (113, 206) as an object having an anomaly depending on the determined assessment. [2] The method of claim 1, comprising generating the reference image descriptor image (204) by feeding the respective reference image (202) to a dense object network (112, 201) and generating the input image descriptor image (207) by feeding the input image (205) to the dense object network (112, 201). [3] The method of claim 1 or 2, wherein determining the region (209) of the input image (205) comprises determining one of the reference images (202) showing the reference object to which the object (113, 206) shown by the region (209) corresponds, and an affine transformation between the region (209) and a region of the determined reference image (202) showing the reference object to which the object (113, 206) shown by the region (209) corresponds. [4] A method according to any one of claims 1 to 3, wherein the anomaly is damage. [5] Method according to one of claims 1 to 4, further comprising, if the object (113, 206) is identified as an object having an anomaly, determining a location (210) of the object (113, 206) having the anomaly based on the or a further comparison of the descriptor values contained in the input image descriptor image (207) for the determined area (209) with the descriptor values contained in the respective reference image descriptor image for the area of the respective reference image (202) showing the reference object. [6] Method for controlling a technical system (101), comprising: Receiving an input image (205); Checking whether an object (113, 206) represented in the input image (205) has an anomaly according to one of claims 1 to 5; and treating the object (113, 206) depending on whether it has been marked as an object (113, 206) having an anomaly. [7] The method of claim 6, comprising discarding the object (113, 206) if it has been identified as an object (113, 206) having an anomaly. [8] The method of claim 6 or 7, further comprising determining a location and / or pose for capturing or processing the object (113, 206) using the input image descriptor image (207). [9] Data processing device (106) which is arranged to carry out a method according to one of claims 1 to 8. [10] A computer program comprising instructions which, when executed by a processor, cause the processor to perform a method according to any one of claims 1 to 8. [11] A computer-readable medium storing instructions which, when executed by a processor, cause the processor to perform a method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Automatic Defect Classification Without Sampling and Feature Selection
US20160163035A1
Cited By
Manipulator motion point position testing device
CN120902014A