Failure detection and failure recovery for ai depalletizing

JP2023113572A5Pending Publication Date: 2025-07-17FANUC LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022208010
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-02-03
Filing Date
2022-12-26
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing neural networks struggle to accurately segment and identify boxes of varying sizes and shapes due to insufficient training data, leading to missed detections and partial segmentations, which can cause robotic failures during pick-and-place operations.

Method used

A system utilizing a 3D camera to acquire RGB and depth map images, followed by a modified deep learning mask R-CNN for image segmentation, combined with fault detection modules to identify and correct segmentation errors, and a manual labeling process to refine the neural network training.

Benefits of technology

Enhances the accuracy of box identification by robots, reducing robotic failures and enabling the handling of diverse box sizes and shapes through iterative training with user-corrected data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a system and a method for detecting and correcting a failure of a processed image.SOLUTION: A method for identifying inaccurately depicted boxes in an image obtains a 2D RGB image of the boxes and a 2D depth map image of the boxes using a 3D camera, where pixels in the depth map image are assigned a value identifying the distance from the camera to the boxes. The method generates a segmentation image of the boxes using a neural network by performing an image segmentation process that extracts features from the RGB image and segments the boxes by assigning a label to pixels in the RGB image so that each box in the segmentation image has the same label and different boxes in the segmentation image have different labels. The method analyzes the segmentation image to determine if the image segmentation process has failed to accurately segment the boxes in the segmentation image.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to systems and methods for detecting and correcting defects in processed images, and more specifically, to systems and methods for detecting and correcting defects in segmentation images generated by a neural network. The segmentation images have a specific application for identifying boxes lifted by a robot from a stack of boxes.

Background Art

[0002] Robots perform a number of commercial tasks, including pick-and-place operations. In the pick-and-place operation, the robot grasps and moves an object from one position to another. For example, the robot can grasp a box from a pallet and place the box on a conveyor belt. Here, the robot often employs an end effector having a suction cup to hold the box. In order for the robot to effectively grasp the box, the robot needs to recognize the width, length, and height of the box to be grasped, and the width, length, and height are input to the robot controller before the pick-and-place operation. However, in many cases, the sizes of the boxes on the same pallet are different, so it is inefficient to input the box size to the robot during the pick-and-place operation. The boxes can also be arranged side by side at the same height. In this case, it is difficult to distinguish whether these boxes are separate boxes or a single large box.

[0003] U.S. Patent Application No. 17 / 015,817, filed on September 9, 2020, assigned to the assignee of this application and incorporated herein by reference, discloses a system and method for identifying boxes to be picked up by a robot from a pile of boxes. The method includes using a 3D camera to acquire a 2D red-green-blue (RGB) color image of the boxes and a 2D depth map image of the boxes, where pixels in the depth map image are assigned values ​​that identify the distance from the camera to the boxes. The method employs a modified deep learning mask R-CNN (convolutional neural network) to generate a segmented image of the boxes by performing an image segmentation process. The image segmentation process extracts features from the RGB image, combines the extracted features in the image, and assigns labels to pixels in the feature image such that pixels of each box in the segmented image have the same label and pixels of different boxes in the segmented image have different labels. Next, the method uses segmented images to identify the position where the box is to be picked up.

[0004] The method disclosed in Application No. 817 employs a deep learning neural network for the image filtering step, the region proposal step, and the binary segmentation step. Deep learning is a particular type of machine learning that provides better learning performance by representing a particular real-world environment as a hierarchy of increasingly complex concepts. Deep learning typically employs a software structure that includes several layers of neural networks that perform nonlinear processing, where each successive layer receives the output from the previous layer. The layers generally include an input layer that receives raw data from a sensor, several hidden layers that extract abstract features from the data, and an output layer that identifies a particular thing based on the feature extraction from the hidden layers.

[0005] A neural network contains neurons or nodes, each having a "weight" which is multiplied by the input to the node to determine the probability of something being true or false. More specifically, each node has a weight, which is a floating-point number multiplied by the input to the node to produce an output that is a certain ratio of the inputs. The weights are initially "trained" or set by having the neural network analyze a set of known data under supervision, and by minimizing a cost function so that the network obtains the highest probability of the correct output.

[0006] Deep learning neural networks are often employed to provide image feature extraction and transformation for the visual detection and classification of objects in images. Here, a video or image stream may be analyzed by the network to identify and classify objects and learn through a process to better recognize objects. The number of layers and nodes within a neural network determines the network's complexity, computation time, and performance accuracy. The complexity of a neural network can be reduced by reducing the number of layers, the number of nodes within a layer, or both. However, reducing the complexity of a neural network reduces the accuracy of the neural network's learning. Here, it has been shown that reducing the number of nodes within a layer has an accuracy advantage over reducing the number of layers within the network.

[0007] This type of deep learning neural network requires significant data processing and is data-driven, meaning it needs a lot of data to train. For example, for a robot to pick up boxes, the neural network is trained to identify boxes of a specific size and shape from a training dataset of boxes. Here, more boxes are used and needed to train the network. However, the robot's end-user may need to pick up boxes of different sizes and shapes that were not part of the training dataset. Furthermore, different textures and labeling on a single box can generate bad detections. Thus, the neural network may not have the segmentation performance necessary to identify such boxes. This can result in different types of bad detections, such as missed detection bad detections and partial segmentation bad detections. A missed detection occurs when a smaller box is on top of a larger box, and the smaller box is not detected by the segmentation process, causing the robot to collide with the top box. Here, the smaller box is not part of the training dataset. Partial segmentation occurs when there is only one box, and the label or other features on the box cause the neural network to segment the box into two, which can result in the robot picking up the box off-center. In this case, the box is tilted, which can lead to obvious problems.

[0008] A possible solution to the aforementioned problem is to provide specific end users with more boxes or specific types of boxes to use in training the neural network. However, the segmentation model used in the neural network is imperfect, which could lead to damage to the robot and other objects, thus discouraging its use. Without using an imperfect neural network, bad samples that could be used to improve the model are not obtained; in other words, data is not provided to further train and fine-tune the neural network. [Overview of the project]

[0009] The following discussion discloses and describes a system and method for identifying boxes that are inaccurately depicted within an image of a group of boxes, such as detecting overlooked boxes and partially detected boxes. The method uses a 3D camera to acquire a 2D red-green-blue (RGB) color image of the boxes and a 2D depth map image of the boxes. Here, pixels in the depth map image are assigned values ​​that identify the distance from the camera to the boxes. The method generates a segmented image of the boxes using a neural network, for example, by performing an image segmentation process. The image segmentation process extracts features from the RGB image and segments the boxes by assigning labels to pixels in the RGB image such that pixels in each box in the segmented image have the same label and pixels in different boxes in the segmented image have different labels. The method analyzes the segmented image to determine whether the image segmentation process failed to accurately segment the boxes in the segmented image.

[0010] Additional features of this disclosure will become apparent from the following description and the accompanying claims, in conjunction with the attached drawings. [Brief explanation of the drawing]

[0011] [Figure 1] Figure 1 shows a robotic system that includes a robot that picks up boxes from a pallet and places the boxes onto a conveyor belt. [Figure 2] Figure 2 is a schematic block diagram of a system that detects and corrects defects in a processed image used to identify an object to be grasped by a robot. [Figure 3] Figure 3 shows a top-view RGB segmentation image of a group of boxes randomly placed on a palette, illustrating missed detection errors. [Figure 4] Figure 4 is a graph with distance on the horizontal axis and the number of pixels on the vertical axis, showing the bounding box formed around a larger box that has a smaller box at its top. [Figure 5] Figure 5 shows a top-view RGB segmentation image of a group of boxes randomly placed on a palette, exhibiting partial segmentation defects. [Modes for carrying out the invention]

[0012] The following discussion of embodiments of the present disclosure relating to a system and method for detecting and correcting defects in processed images is essentially illustrative and is not intended in any way to limit the present invention or any use or application of the present invention. For example, the system and method may have an application in identifying a box to be picked up by a robot. However, the system and method may have other applications.

[0013] Figure 1 shows a robot system 10 including a robot 12 having an end effector 14 configured to pick up a box 16 from a pile 18 of boxes 16 arranged on a pallet 20 and place the box 16 onto a conveyor belt 22. System 10 is intended to represent any type of robot system that can be benefited from the discussion herein, and robot 12 can be any robot suitable for the purpose. A 3D camera 24 is positioned to capture a 2D top RGB image and a depth map image of the pile 18 of boxes 16 and to provide these images to a robot controller 26 that controls the movement of robot 12. The boxes 16 may have different orientations on the pallet 20, may be stacked in multiple layers on the pallet 20, and may be of different sizes.

[0014] Figure 2 is a schematic block diagram of a system 30 that detects and corrects defects in a processed image used to identify an object, such as a box 16, to be picked up by a robot 12. The system 30 includes a 3D camera 32 representing a camera 24 that acquires a top RGB image and a 2D depth map image of the object. The RGB image and depth map image are provided to an analysis module 34. The analysis module 34 processes the RGB image and the 2D depth map image in any way suitable for providing the processed image to identify the object in a format that allows the robot 12 to pick it up. In a non-limiting example, the analysis module 34 performs a segmentation process as described in Patent Application No. 817, employing a modified deep learning mask R-CNN to generate a segmented image of the box 16 by performing an image segmentation process. The image segmentation process extracts features from an RGB image and a depth map image, combines the extracted features in the image, and assigns labels such as color to pixels in the feature image such that pixels in each box in the segmentation image have the same label and pixels in different boxes in the segmentation image have different labels.

[0015] Figure 3 is a top-level RGB segmentation image 36 of a group of boxes 38 randomly placed on a palette 40, output from the analysis module 34, where one of the boxes 38 is a smaller box 42 placed on top of a larger box 44. The analysis module 34 identifies the segmented position and orientation of the boxes 38 in the segmentation image 36 and selects a candidate box, e.g., the best box, as the next box to be picked up by the robot 12. Figure 3 shows the segmentation bounding boxes 46 around the boxes 38 at the highest point of the pile of boxes 38 in the segmentation image 36, where each bounding box 46 is defined around pixels with the same label. As is clear, there is no bounding box 46 around the smaller box 42, thus indicating a missed detection.

[0016] The analysis module 34 transmits the processed image to the defect detection module 48 to identify defects in the image. For example, in the particular embodiment described above, module 34 transmits position and orientation information about a box 38 having a bounding box 46 in the image 36 to the defect detection module 48, which determines whether the position and orientation of the box 38 are correct or incorrect. The defect detection module 48 also receives an RGB image and a depth map image from the camera 32. The defect detection module 48 includes several defect detection submodules 50, each running in parallel and each operating to detect one of the defects discussed herein.

[0017] One of the fault detection submodules 50 can identify missed detections. The depth map image allows the fault detection module 48 to determine the distance of each pixel in the image 36 from the camera 32. The distance of each pixel within a segmentation bounding box 46 in the image 36 should be approximately the same as all other pixels within that segmentation bounding box 46. By individually examining the distances of each group of pixels within each bounding box 46, it is possible to determine whether there are any missed box detections.

[0018] Figure 4 is a graph with distance on the horizontal axis and the number of pixels on the vertical axis for bounding boxes 46 around a larger box 44, and such a graph is generated for each bounding box 46. Higher peaks 60 with a number of pixels further away from the camera 32 identify the distance of most of the pixels within the bounding box 46 around the larger box 44, and thus provide the distance of the larger box 44 from the camera 32. Lower peaks 62 with a number of pixels closer to the camera 32 identify the distance of pixels for smaller boxes 42 closer to the camera 32. Thus, boxes where detection was missed are identified. If the graph has only one peak, only one of the boxes 38 is within its bounding box 46. Note that a threshold is used to identify a second peak so that a smaller peak that is not the box 38 where detection was missed is not identified as a missed detection.

[0019] One of the defect detection submodules 50 within the defect detection module 48 may identify a partial segmentation in which there are multiple separate bounding boxes 46 for a single box 38. Each of the bounding boxes 46 output from the analysis module 34 is determined by the analysis module 34 to generally have a predetermined high probability, e.g., 95%, of indicating the boundary of box 38 in the image 36. However, the analysis module 34 identifies many bounding boxes 46 that are less likely to indicate the boundary of box 38. These are discarded in the segmentation process disclosed in application 817. For system 30, the analysis module 34 retains some or all of the less likely bounding boxes 46 and outputs them to the defect detection module 48. This is illustrated by Figure 5, which shows a top-view RGB segmentation image 70 of a group of boxes 72 randomly placed on a palette 74, output from the analysis module 34. Image 70, in the top layer, includes a left box 76 enclosed by a bounding box 78, a right box 80 enclosed by a bounding box 82, and an intermediate box 84 showing several bounding boxes 86. Bounding boxes 78 and 82 are actual multiple bounding boxes that overlap each other, indicating that bounding boxes 78 and 82 likely define the boundaries of boxes 76 and 80, respectively. Thus, the separated bounding box 86 indicates that the boundary of box 84 is not identified with high confidence and can be segmented as many boxes as possible within the segmentation image in the segmentation process disclosed in Application No. 817. Therefore, by looking at the cross ratio, i.e., how well the bounding boxes overlap for a particular box, it is possible to identify boxes that may have been partially segmented.

[0020] The fault detection submodule 50 within the fault detection module 48 can also detect other faults. One of these other faults may be the detection of an empty pallet. Before a box is placed on an empty pallet, a depth map image of the empty pallet is acquired and compared to a real-time depth map image in one of the submodules 50. If a sufficient number of pixels above a threshold indicate the same distance, module 48 recognizes that all boxes have been picked up.

[0021] Another type of defect may be the detection of whether a segmented box in a segmentation image is identified as being larger or smaller than the largest or smallest known box size, which is the detection of a defective box in the segmentation image. One of the defect detection submodules 50 can see the size of each bounding box in the segmentation image, and if a bounding box is larger than the largest box size or smaller than the smallest box size, an alarm indicating a defect may be sent.

[0022] Returning to Figure 2, if none of the fault detection submodules 50 detect any of the faults described above, the position and orientation of the next box to be selected and picked up is transmitted to the actuation module 90. The actuation module 90 operates the robot 12 to pick up the box and place it on the conveyor belt 22. Next, the system 30 causes the camera 32 to take the next top RGB image and depth map image of the pile 18 in order to pick up the next box 16 from the pallet 20. If any or some of the fault detection submodules 50 detect a particular fault, no signal is sent to the actuation module 90, and the robot 12 stops. Next, one or more alarm signals are sent to the manual labeling module 92, which is a user interface computer and screen monitored by a human user. The human user views the processed image generated by the analysis module 34 on the screen to visually identify the problem causing the fault. For example, a human user may use a computer-based algorithm to draw or redraw appropriate bounding boxes around a box or multiple boxes that are causing a problem, and then send the adjusted RGB segmentation image to the actuation module 90. The robot 12 then picks up the manually labeled boxes in the image.

[0023] The processed images including defects from the manual labeling box 92, and the correct or corrected images from the operation module 90 are sent to the database 94 and stored in such a way that the system 30 includes defective images, images obtained by correcting defective images, and appropriately processed images from the defect detection module 48. Next, at a selected time, these images are sent to the fine-tuning module 96. The fine-tuning module 96 uses the corrected images, good images, and defective images to correct the processing in the analysis module 34, such as further training the neural network nodes in the analysis module 34. Thus, the analysis module 34 is further trained to pick up boxes in the user's equipment that may not have been used to train the neural network nodes using the original dataset, and is trained to prevent defective images from being generated when the configuration of the box that previously generated defective images occurs next.

[0024] As will be fully understood by those skilled in the art, several of the various steps and processes discussed herein for describing the present disclosure may refer to operations performed by a computer, processor, or other electronic computing device that manipulates or transforms data or both using electrical phenomena. Such computers and electronic devices may employ various volatile memories or non-volatile memories or both, including a non-transitory computer-readable medium storing an executable program including various codes or executable instructions that can be performed by the computer or processor. The memory or computer-readable medium or both may include all forms and types of memory and other computer-readable media.

[0025] The foregoing discussion merely discloses and describes the preferred embodiments of the present disclosure. Those skilled in the art will readily recognize that various changes, modifications, and variations can be made without departing from the spirit and scope of the present disclosure as defined in the following claims, from such discussion and the accompanying drawings and claims.

Claims

1. A method for identifying and correcting an object depicted inaccurately within an image of a group of objects, the method comprising: using a 3D camera to obtain a 2D red-green-blue (RGB) color image of the object; using the 3D camera to obtain a 2D depth map image of the object, wherein a value identifying the distance from the camera to the object is assigned to pixels within the depth map image; processing the RGB image and the depth map image to generate a processed image of the object; analyzing the processed image to determine whether the object is accurately depicted within the processed image; when the analysis of the processed image determines that the processing of the RGB image and the depth map image was unable to accurately depict the object within the processed image, transmitting the processed image to a user interface, the user interface enabling a user to correct the processed image; storing accurate processed images, defective processed images, and corrected processed images; using the stored accurate processed images, defective processed images, and corrected processed images to train the processing of the RGB image and the depth map image; applying the trained processed image to new RGB and depth map images.

2. The object is a box, and the processing of the RGB image and the depth map image includes generating a segmented image of the box using a neural network by performing an image segmentation process, the image segmentation process extracting features from the RGB image and segmenting the box by assigning labels to pixels within the RGB image such that each box within the segmented image has the same label and different boxes within the segmented image have different labels. The method according to claim 1.

3. The analysis of the processed image includes identifying boxes that were missed in detection by analyzing the labels of the boxes within the segmented image to determine whether boxes within the segmented image are not identified with different labels. The method according to claim 2.

4. The generation of the segmentation image includes providing a bounding box around each labeled box in the segmentation image, wherein the boxes in the segmentation image that are not identified with different labels do not have a bounding box, the method according to claim 3.

5. The analysis of the labels of the boxes in the segmentation image includes counting the pixels within each bounding box having the same distance value from the depth map image, and determining that there are multiple boxes within the bounding box if there are multiple pixel counts higher than a predetermined threshold, the method according to claim 4.

6. The analysis of the processed image includes identifying partially segmented boxes, the method according to claim 2.

7. The generation of the segmentation image includes providing various degrees of confidence to each bounding box around each box in the segmentation image that the bounding box identifies a box in the segmentation image, and the analysis of the segmentation image includes observing the intersection over union of the multiple bounding boxes around each box in the segmentation image, wherein a predetermined low intersection ratio indicates a partially segmented box in the image, the method according to claim 6.

8. The generation of the segmentation image includes providing a bounding box around each box in the segmentation image, and the analysis of the segmentation image includes looking at the size of each bounding box in the segmentation image and determining whether the segmented box in the segmentation image is larger or smaller than a box of a predetermined maximum size or smaller than a box of a predetermined minimum size by determining whether the bounding box is larger than a box of a predetermined maximum size or smaller than a box of a predetermined minimum size, the method according to claim 2.

9. The analysis of the processed image includes identifying an empty pallet by obtaining a depth map image of the empty pallet before the object is placed on the empty pallet, comparing the depth map image of the empty pallet with the depth map image from the 3D camera, and identifying the empty pallet if a sufficient number of pixels higher than a threshold value indicate the same distance. The method according to claim 1.

10. A method for identifying boxes inaccurately depicted within an image of a group of boxes, the method being used by a robot to identify which boxes to pick up, the method comprising: Using a 3D camera to obtain a 2D red-green-blue (RGB) color image of the box; Using the 3D camera to obtain a 2D depth map image of the box, wherein a value identifying the distance from the camera to the box is assigned to the pixels within the depth map image; Generating a segmentation image of the box using a neural network by performing an image segmentation process, the image segmentation process extracting features from the RGB image and segmenting the box by assigning labels to the pixels within the RGB image such that each box within the segmentation image has the same label and different boxes within the segmentation image have different labels, and the generation of the segmentation image including providing a bounding box around each labeled box within the segmentation image; Identifying boxes missed in detection within the segmentation image by analyzing the labels of the boxes within the segmentation image to determine whether the boxes within the segmentation image are not identified with different labels, wherein the boxes within the segmentation image not identified with different labels do not have a bounding box. The method comprising.

11. The analysis of the label of the box in the segmentation image includes counting the pixels within each bounding box having the same distance value from the depth map image, and determining that there are multiple boxes within the bounding box if there are multiple pixel counts higher than a predetermined threshold. The method according to claim 10.

12. The method according to claim 10, further comprising transmitting the segmentation image to a user interface if the boxes in the segmentation image are not identified with different labels, and the user manually identifies the boxes with different labels.

13. The method according to claim 10, further comprising storing an accurate segmentation image, a defective segmentation image, and a corrected segmentation image, and training the neural network using the stored accurate segmentation image, defective segmentation image, and corrected segmentation image.

14. A method for identifying a box inaccurately depicted within an image of a group of boxes, the method being used by a robot to identify which box to pick up, the method comprising: using a 3D camera to obtain a 2D red-green-blue (RGB) color image of the box; using the 3D camera to obtain a 2D depth map image of the box, wherein a value identifying the distance from the camera to the box is assigned to the pixels in the depth map image. By performing an image segmentation process, a neural network is used to generate a segmented image of the box, wherein the image segmentation process extracts features from the RGB image and assigns labels to the pixels in the RGB image such that each box in the segmented image has the same label and different boxes in the segmented image have different labels, thereby segmenting the box, and the generation of the segmented image includes providing various degrees of confidence to a plurality of bounding boxes around each box in the segmented image, wherein each bounding box identifies the box in the segmented image. Identifying a partially segmented box in the segmented image by observing the intersection over union of the plurality of bounding boxes around each box in the segmented image, wherein a predetermined low intersection ratio indicates a partially segmented box in the image. Claim 15 The method according to claim 14, further comprising transmitting the segmented image to a user interface when a partially segmented box in the segmented image is identified, and the user manually identifying the partially segmented box. Claim 16 The method according to claim 15, further comprising storing an accurate segmented image, a defective segmented image, and a corrected segmented image, and training the neural network using the stored accurate segmented image, defective segmented image, and corrected segmented image. Claim 17 A robot system for identifying an object inaccurately depicted in an image of a group of objects, the system being used by a robot to identify which object to pick up, the system A 3D camera that provides a 2D red-green-blue (RGB) color image and a 2D depth map image of the object. Means for processing the RGB image and the depth map image to generate a processed image of the object. means for analyzing the processed image to determine whether the object is accurately depicted within the processed image; means for transmitting the processed image to a user interface if the analysis of the processed image determines that the processing of the RGB image and the depth map image was unable to accurately depict the object within the processed image, wherein the user interface enables a user to modify the processed image; means for storing accurate processed images, defective processed images, and corrected processed images; means for training the processing of the RGB image and the depth map image using the stored accurate processed images, defective processed images, and corrected processed images; means for applying the trained processed image to new RGB and depth map images, a system comprising.

18. The object is a box, and the means for analyzing the processed image employs a deep learning convolutional neural network that performs an image segmentation process to generate a segmentation image of the box, the image segmentation process extracts features from the RGB image and the depth map image, combines the extracted features within the image, and assigns labels to the pixels within the segmentation image such that each box within the segmentation image has the same label and different boxes within the segmentation image have different labels. The system according to claim 17.

19. The means for analyzing the processed image determines whether a box has been overlooked by analyzing the labels of the boxes within the segmentation image to determine whether the boxes within the segmentation image are not identified with different labels. The deep learning convolutional neural network provides a bounding box around each labeled box within the segmentation image, and the boxes within the segmentation image that are not identified with different labels do not have a bounding box. The system according to claim 18.

20. The means for analyzing the processed image identifies a partially segmented box, and the deep learning convolutional neural network provides various degrees of confidence to a plurality of bounding boxes around each box in the segmentation image, with each bounding box identifying a box in the segmentation image. The means for analyzing the processed image observes the intersection over union of the plurality of bounding boxes around each box in the segmentation image, and a predetermined low intersection over union indicates a partially segmented box in the image. The system according to claim 18.