Method for detecting object in image data
The method enhances robotic object detection by segmenting and grouping similar image regions using machine learning, addressing the challenge of distinguishing multiple instances of the same object type and identifying anomalies.
Patent Information
- Application Number
- CN202510040190.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-12
- Filing Date
- 2025-01-10
- Publication Date
- 2025-07-15
AI Technical Summary
The prior art is difficult to reliably detect multiple instances of the same object type present in a container in a robot and distinguish objects from foreign objects, especially under different viewing angles and light conditions.
Using a machine learning-based object segmentation model, an object type model is formed through image area similarity measurement and local feature analysis, and instance combination and recognition of the same object type is achieved.
It improves the accuracy of object detection under different viewing angles and light conditions, can effectively distinguish objects from foreign objects, and improves the robot's grasping ability in open-world scenarios.
Smart Images

Figure CN120318483A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method for detecting objects in image data. Background Art
[0002] Picking (i.e., grasping) an object from a container is an important problem in robotics. For this to be performed automatically, a robot must in particular be able to detect the object to be grasped, for example, select the correct object if different objects (such as screws and nuts) are present in the container, or also distinguish an object from undesired objects (such as remaining packaging material). Correspondingly, reliable object detection methods are desired especially for scenarios in which there are multiple instances of the same object type (i.e., the same object, e.g., detecting the same screw multiple times). Summary of the Invention
[0003] According to various embodiments, there is provided a method for detecting an object in image data, having:
[0004] · Segmenting an input image into a plurality of image regions, where each image region shows a corresponding instance of a corresponding object type or the image background;
[0005] · Determining at least one set of image regions that show instances of the same object type according to an image region similarity metric;
[0006] · Combining, for each determined set, the instances of the object type shown by the set of image regions into a model for the object type; and
[0007] · Detecting instances of the object type or other objects in the input image or one or more additional input images according to the model.
[0008] The above method achieves reliable object detection for application scenarios in which there are multiple identical objects (i.e., multiple instances of the same object type).
[0009] For example, the segmentation is performed by means of a machine learning model trained for segmenting image data.
[0010] By using, in a first step, an (up-to-date, machine learning-based) object segmentation model (such as a neural network, e.g., a convolutional neural network), object understanding is achieved, which is generally not the case in purely local feature-based methods. This improves usability in real-world scenarios in which objects can be detected from different perspectives or under different lighting conditions.
[0011] Various embodiments are described below.
[0012] Embodiment 1 is the method for detecting an object in image data as described above.
[0013] Example 2 is a method according to Example 1, in which it is determined that at least one group has: for each pair of image regions, a value of an image region similarity measure is determined and a maximum clique is determined, where if the image region similarity measure between two image regions is higher than a preset threshold, the two regions are regarded as associated.
[0014] Therefore, groups of image regions showing instances of the same object type can be effectively determined. The similarity measure is, for example, a distance measure between features of the image regions and / or a measure of the consistency of a set of local features (keypoints in English), such as in SIFT (Scale-Invariant Feature Transform).
[0015] Example 3 is a method according to Example 1 or 2, having: determining multiple groups of image regions, the image regions showing corresponding instances of the same object type according to an image region similarity measure; comparing the number of image regions belonging to the groups; and determining an instance of the object type to be manipulated, the object to be manipulated being shown by the image regions of the group containing the most image regions.
[0016] Thereby, for example, an object to be manipulated (e.g., an object to be extracted from a container) can be distinguished from foreign objects (e.g., packaging remnants).
[0017] Example 4 is a method according to any one of Examples 1 to 3, where the model is a two-dimensional or three-dimensional model of the shape of an object of the object type.
[0018] Then, other possible instances of the object type can be compared thereby, in order to improve the detection accuracy for other instances thereby. In particular, instances that were first "ignored" during segmentation can also be found.
[0019] Example 5 is a method according to any one of Examples 1 to 4, where detecting other objects in the input image or one or more additional input images according to the model is identifying objects that deviate from the object type according to the model.
[0020] Therefore, "other objects" are, for example, anomalies or outliers, such as defective objects of that object type. By creating a model, such anomalies (e.g., defective objects) can be effectively identified without prior knowledge of the object type.
[0021] Example 6 is a method for controlling an engineering system, having detecting one or more objects according to any one of Examples 1 to 5 and controlling the engineering system to manipulate one or more detected objects.
[0022] Embodiment 7 is a data processing device (particularly a control device), which is designed to execute the method according to any one of Embodiments 1 to 6.
[0023] Embodiment 8 is a computer program with instructions that, when executed by a processor, cause the processor to execute the method according to any one of Embodiments 1 to 6.
[0024] Embodiment 9 is a computer-readable medium storing instructions that, when executed by a processor, cause the processor to execute the method according to any one of Embodiments 1 to 6. Description of the Drawings
[0025] In the drawings, in the overall different views, similar reference numerals generally refer to the same components. The drawings are not necessarily to scale, and instead, the emphasis is usually placed on the description of the principles of the present invention. In the following description, various aspects are described with reference to the following drawings.
[0026] Figure 1 A robot is shown.
[0027] Figure 2 Illustrate creating an object model from an input image.
[0028] Figure 3 A flowchart showing a method for detecting an object in image data according to one embodiment is shown. Detailed Description of the Embodiments
[0029] The following detailed description refers to the accompanying drawings, which show the specific details and aspects in which the present disclosure can implement the present invention for explanation. Without departing from the scope of protection of the present invention, other aspects can be used and structural, logical, and electrical changes can be made. The various aspects of the present disclosure are not necessarily mutually exclusive, because some aspects of the present disclosure can be combined with one or more other aspects of the present disclosure to form new aspects.
[0030] The following describes each example in more detail.
[0031] Figure 1 A robot 100 is shown.
[0032] The robot 100 includes a robotic arm 101, such as an industrial robotic arm, for manipulating or assembling workpieces (or one or more other objects). The robotic arm 101 includes manipulators 102, 103, 104 and a base (or support) 105, and the manipulators 102, 103, 104 are supported by the base. The term "manipulator" refers to a movable element of the robotic arm 101, and the actuation of the movable element enables physical interaction with the environment, for example in order to perform tasks. For control, the robot 100 includes a (robot) control device 106, which is configured to: implement interaction with the environment according to a control program. The last element 104 of the manipulators 102, 103, 104 (farthest from the support 105) is also referred to as an end effector 104 and may include one or more tools, such as a welding torch, a fixture, a painting tool, etc.
[0033] The other manipulators 102, 103 (closer to the support 105) may form a positioning device such that the robotic arm 101 is provided with an end effector 104 at its end together with the end effector 104. The robotic arm 101 is a robotic arm capable of performing functions similar to those of a human arm (possibly having tools at its end).
[0034] The robotic arm 101 may include joint elements 107, 108, 109 that connect the manipulators 102, 103, 104 to each other and to the support 105. The joint elements 107, 108, 109 may have one or more joints, and each such joint may provide a rotatable movement (i.e., rotational movement) and / or translational movement (i.e., translation) of the respective manipulators relative to each other. The movement of the manipulators 102, 103, 104 may be caused by actuators controlled by the control device 106.
[0035] The term "actuator" may be understood as a component designed to affect a mechanism or process in response to being driven. The actuator may implement a command (so-called activation) output by the control device 106 as a mechanical movement. An actuator, such as an electromechanical transducer, may be configured to: convert electrical energy into mechanical energy in response to its activation.
[0036] The term "control device" may be understood as any type of logic of an implementing entity, which may include, for example, a circuit and / or a processor, firmware, or a combination thereof capable of executing software stored in a storage medium, and the control device may output instructions, for example to an actuator in the current example. The control device may be configured, for example, by program code (such as software), to control the operation of the robotic device.
[0037] In the current example, the control device 106 includes one or more processors 110 and a memory 111 that stores code and data, and the processor 110 controls the robotic arm 101 based on the code and data. According to various embodiments, the control device 106 controls the robotic arm 101 based on a machine learning model 112 stored in the memory 111 and sensor data (such as image data of the camera 114, which can be a color image (RGB), but can also have depth information (RGB-D)). For example, the robot 100 is intended to manipulate the object 113. For example, the manipulation task is, for example, to pick up the object 113 from the container 115 (so-called bin picking).
[0038] In many application scenarios (especially in logistics), there are repeated objects of the same type, that is, multiple instances of the same object type (or in other words, multiple occurrences of the same object). For example, the container 115 contains multiple identical objects (i.e., instances of the same object type, such as multiple identical screws) that must be picked up for further processing.
[0039] The information that an object repeats in the corresponding scene (i.e., multiple instances of the same object type) can be used for more robust perception. For example, if the container 115 also contains another object (e.g., remaining packaging material), it can be concluded based on the presence of only one instance (or a few instances) of the other object that this is not the object that should be picked up.
[0040] According to various embodiments, a method for object detection is provided, in which the detection accuracy is improved by obtaining information about an object from multiple instances of the object.
[0041] To this end, according to various embodiments, a foundation model (e.g., the foundation model corresponds to the machine learning model 112) is used to find object proposals in the input data (e.g., in the input image), such as a segmentation model, such as Segment Anything Model (SAM). Based on the object proposals (i.e., the determined image regions that each show a corresponding object), instances of the same object type are then searched for by finding object proposals with similar appearances, for example, by searching for patterns present in multiple object proposals (e.g., based on low-level features such as edges or pixel values). This enables better object understanding and ultimately an improvement in detection accuracy in open-world scenarios.
[0042] According to various embodiments, thus in the first step, object proposals are obtained by means of an ML (machine learning) base model, and in the second step, the object proposals are compared with each other, for example, by searching for corresponding transformations that mutually map for the object proposals, such that clusters of similar object proposals are formed, that is, groups of instances of the same object type are formed, and the instances, for example, show in the input image from different perspectives or are exposed to different light conditions. In the third step, different instances of the object can be combined to create a model (or "template") for the object.
[0043] For the first step, any (e.g., open-world) object segmentation method can be used, that is, for example, a list of object masks is provided in an RGB input image or an RGB-D input image, ideally for each instance of one or more (physical) objects in the input image. It is assumed that the object segmentation method can: provide a mask for each object in the corresponding scene of interest.
[0044] The second step is used to: filter out object instances belonging to the same object type (i.e., the same object). Typically, the object segmentation method provides many object proposals, also proposals for the background. For example, in the second step, a local feature-based method and homography estimation are used (i.e., for a pair of object proposals, the possible transformation from one object proposal to another is calculated) to obtain an evaluation of the similarity of two object proposals pairwise. For example, the evaluation can be represented as a graph (i.e., the nodes correspond to the object proposals, and each edge between two corresponding object proposals of a pair has the corresponding evaluation of the pair). In such a graph, the maximum clique can be searched, that is, a group or cluster of objects is searched, that is, a group is searched for which there is a transformation (with certain limitations) between every pair of object proposals from the group (i.e., the image region similarity metric is higher than a preset threshold). The object proposals of such a maximum clique are considered to be instances of the same object type (i.e., the same object).
[0045] In the third step, the individual object instances in a group are combined into a model or "object template". Here, for example, by combining, such as stitching together, multiple parts (object proposals), the complete shape of the object is established from different perspectives provided by the object proposals of the corresponding group. Depending on the input modality, this can be done in the 2D-RGB space or in 3D space, for example, using a point cloud. Different methods can be used for both modalities to fit the multiple parts of the object.
[0046] Figure 2 Illustrate creating an object model 205 from an input image 201.
[0047] The input image 201 is first passed to an object segmentation method (e.g., SAM). The output of the method is a list 202 of N object masks in the image (e.g., a list of binary images with 1s at the positions of the corresponding objects).
[0048] Then, the list of object masks is forwarded to the "matching stage". Here, pairwise "matching scores" (i.e., values of a similarity metric, i.e., values of a similarity assessment) are calculated between all N 2 object proposals of the object proposals. For this purpose, methods for finding local features can be used, such as classical computer vision methods like SIFT (Scale-Invariant Feature Transform) or other local feature-based methods. Based on the local features, a homography between two object proposals can be estimated for each pair of object proposals. Here, a transformation (translation, rotation, scaling, shearing) from one object proposal to another is estimated. For example, the estimated transformation is stored together with the similarity assessment in an N×N similarity matrix, which, as explained above, can be presented as a graph 203.
[0049] One or more relevant clusters are calculated from the similarity matrix, i.e., one or more groups of objects that are all pairwise similar. In practice, for this purpose, a threshold of the similarity assessment (i.e., the edge weights of the graph) can first be confirmed in the similarity matrix, and the similarity assessment is converted into a binary matrix containing 1s (for similarity assessments greater than the threshold) and 0s (for similarity assessments less than the threshold) (by definition, equality with the threshold can be evaluated as 1 or 0). A 1 at a matrix location means: there is a transformation between the two object proposals with the said index, which in turn means that the index shows instances of the same object type, e.g., from different perspectives (shown as solid edges in Figure 2 as graph 203). A zero means: this is not the case (shown as dashed edges in Figure 2 as graph 203).
[0050] The maximum clique can be found in the matrix or graph 203, i.e., a set of object proposals that are all pairwise similar. The result of this stage is a list of clusters (in this example, there is only one cluster 204), where each cluster is composed of object proposals (or object masks) of the same object type (i.e., the same object) (where, for example, only such clusters containing multiple object proposals are considered).
[0051] Now, for example, a 2D image stitching method or a 3D point cloud matching method can be used to create an object model 205 from each cluster 204. The method can be used for downstream tasks, such as calculating grasping poses or identifying anomalies.
[0052] In summary, according to various embodiments, a method as shown in Figure 3 is provided.
[0053] Figure 3 FIG. 300 is a flowchart showing a method for detecting an object in image data according to one embodiment.
[0054] In 301, an input image is segmented into a plurality of image regions, where each image region shows a corresponding instance of a corresponding object type or the image background (where a particular object may be attributed to the image background, for example, objects that are not of interest in a corresponding application scenario or objects that exist only individually).
[0055] In 302, at least one set of image regions showing instances of the same object type is determined according to an image region similarity metric (i.e., the similarity between the image regions is determined, and if the similarity is higher than a specific threshold, it is interpreted as: the image regions show instances of the same object type, i.e., show the same object).
[0056] In 303, for each determined set, the instances of the object type shown by the set of image regions are combined into a model (pattern, template) of the object type.
[0057] In 304, instances of the object type or other objects are detected (or identified) in the input image or one or more additional input images according to the model.
[0058] As described above, Figure 3 the processing manner can be used, for example, in a situation where there are multiple instances of the same object type in container 115 and robot 100 has to find each instance for extraction. However, the processing manner can also be applied in other application scenarios, such as, for example, for counting instances, for finding anomalies, and for quality assurance, etc., where in these other application scenarios, sensor data is presented, especially sensor data that can be represented as image data, i.e., generally sensor data in the form of a matrix having one or more channels, and multiple instances of the same object type of interest are presented. In addition to color images and depth images, the sensor data can also be or include image data from various other sensors (i.e., data arranged in a matrix form), such as radar, lidar, ultrasonic sensors, motion sensors, thermal imaging sensors, etc.
[0059] For example, anomalies in an engineering system can be identified in the following way: By knowing multiple instances of an object in a scenario and the model obtained therefrom (i.e., the "correct" object appearance), anomalies can be identified based on one or more objects deviating from the model (e.g., a damaged screw missing a head in container 115 does not match the model, i.e., the expected appearance image). Thus, for example, an undamaged part can be used to create a model of how the object should look, and a damaged part can be identified based on the part not matching the model.
[0060] Figure 3 The processing can be part of a pipeline in which the processing provides object (model) information, such as the (expected) geometry or (expected) appearance of each object. This information can then be used for further processing, such as for controlling a robot or other engineering systems, such as computer-controlled machines, vehicles, household appliances, power tools, manufacturing machines, personal assistants, or access control systems. For this purpose, for example, one of the object instances is located based on the corresponding sensor data (e.g., determining where the screw is in container 115 and controlling the robot arm 101 to grasp the screw).
[0061] Individual instances can also be tracked based on low-level features in the sensor data (e.g., in a video, i.e., a sequence of images).
[0062] Figure 3 The method can be performed by one or more computers having one or more data processing units. The term "data processing unit" can be understood as any type of entity that implements data or signal processing. For example, data or signals can be processed based on at least one (i.e., one or more than one) specific function performed by the data processing unit. The data processing unit can include analog circuits, digital circuits, logic circuits, microprocessors, microcontrollers, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), field-programmable gate array integrated circuits (FPGAs), or any combination thereof, or be constituted by them. Any other way of implementing the corresponding functions described more precisely herein can also be understood as a data processing unit or a logic circuit device. One or more of the method steps described in detail herein can be performed (e.g., implemented) by the data processing unit through one or more special functions performed by the data processing unit.
[0063] Thus, according to various embodiments, the method is particularly implemented by a computer.
Claims
1. A method for detecting an object (113) in image data, the method comprising: segmenting (301) an input image (201) into a plurality of image regions, wherein each image region shows a respective instance of a respective object type or the image background; determining (302) at least one set (204) of image regions that show instances of the same object type according to an image region similarity metric; combining (303) for each determined set (204) the instances of the object type shown by the image regions of the set (204) into a model (205) for the object type; and detecting (304) instances of the object type or other objects in the input image (201) or one or more additional input images according to the model (205).
2. The method according to claim 1, wherein determining the at least one set (204) comprises: determining a value of the image region similarity metric for each pair of image regions and determining a maximum clique, wherein two image regions are considered to be associated if the image region similarity metric between the two image regions is higher than a preset threshold.
3. The method according to claim 1 or 2, comprising: determining multiple sets (204) of image regions that show instances of respective same object types according to the image region similarity metric; comparing the number of image regions belonging to the sets (204); and determining an instance of the object type to be manipulated, the object to be manipulated being shown by the image regions of the set (204) containing the largest number of image regions.
4. The method according to any one of claims 1 to 3, wherein the model (205) is a two-dimensional or three-dimensional model of the shape of an object of the object type.
5. The method according to any one of claims 1 to 4, wherein detecting other objects in the input image (201) or the one or more additional input images according to the model (205) is identifying objects that deviate from the object type according to the model (205).
6. A method for controlling an engineering system (101), comprising detecting one or more objects (113) according to any one of claims 1 to 5 and controlling the engineering system (101) to manipulate the one or more detected objects (113).
7. A data processing device (106) designed to execute the method according to any one of claims 1 to 6.
8. A computer program having instructions which, when executed by a processor, cause the processor to execute the method according to any one of claims 1 to 6.
9. A computer-readable medium storing instructions which, when executed by a processor, cause the processor to execute the method according to any one of claims 1 to 6.