SYSTEM AND METHOD FOR PICKING ITEMS - Patent application
The method and system address recognition errors in robotic depalletizing by leveraging human operator input to automate recognition improvements, reducing costs and time through automated data updates and neural network training.
Patent Information
- Application Number
- JP2024215128
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-12-10
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2044-12-10
Smart Images

Figure 0007813861000001 
Figure 0007813861000002 
Figure 0007813861000003
Abstract
Description
[Technical Field]
[0001] The present disclosure generally relates to methods and systems for performing object recognition and unloading. [Background technology]
[0002] Depalletizing goods from pallets is performed by retailers, wholesalers, and other third-party logistics vendors as part of their business operations. Some pallets may only have one type of parcel or a single stock-keeping unit (SKU), while other pallets may contain different types of parcels, also known as mixed SKUs. Robotic devices, such as robotic arms, are used to perform palletizing and depalletizing of both single and mixed SKUs.
[0003] In the related art, recognition software utilizing computer vision techniques (e.g., traditional rule-based methods, machine learning-based methods, etc.) has been developed and utilized to recognize objects and their locations. After performing object recognition, the robotic device then proceeds to pick / grasp and move the identified object based on the recognition results. However, challenges remain in that the recognition software cannot accurately recognize the article / object, leading to palletizing and depalletizing failures. This is particularly problematic when it comes to depalletizing mixed SKUs, as many types of objects are involved.
[0004] In the related art, a method is disclosed for performing remote perception assistance and object identification correction, where remote assistance is desired to verify the object identification and provide a correction, based on which additional processing operations on the object are then performed by the robot.
[0005] In the related art, a method for training a machine learning model to identify objects for picking is disclosed, where data used in training the machine learning model is collected from edge cases where the machine learning model fails to detect the object. Summary of the Invention [Problem to be solved by the invention]
[0006] Currently, remote recovery is required when a false recognition result or error occurs. Remote recovery is performed by having the robotic device send images of objects on a pallet to a human operator, who then performs object area selection (e.g., creating a rectangular area for the object) for objects that have not been detected or that have been incorrectly grouped with other objects, and commands the robotic device to pick / grasp the now-identified object. However, this remote recovery process is insufficient for two reasons: 1) manual object selection when performed by a human can be time-consuming due to the complexity of object selection and tedious annotation, and 2) the recognition accuracy / performance of recognition software is not improved when useful information, such as object boundaries, is not provided or utilized.
[0007] Figures 1(A)-1(C) illustrate an exemplary process flow of a conventional recovery process. As shown in Figure 1(A), a recognition error occurs and is identified. Various recognition errors are associated with the objects shown in Figure 1(A). In particular, the upper-left and upper-right objects are not recognized, the middle object has a pose error, the lower-right object is not correctly fitted, and the lower-left object is not segmented and is incorrect. Using the middle object on pallet 100 as an example, a pose error has been identified, requiring remote recovery assistance. As shown in Figure 1(B), the middle object is remotely selected by a human operator under the recovery approach. Determining which object to select and highlighting that object requires specialized knowledge and can be time-consuming. As shown in Figure 1(C), the information required to improve object recognition is manually annotated, resulting in additional costs and time. [Means for solving the problem]
[0008] Aspects of the present disclosure involve an innovative method for recognizing and unloading a plurality of objects. The method may include repeatedly performing the following until the plurality of objects are unloaded: collecting object data by a processor through a vision sensor; performing object recognition by determining object pose and object position based on the object data; if an object of the plurality of objects is recognized and determined to be available for picking, picking up and unloading the object using a robotic device; if no object is determined to be available for picking, determining, by the processor, an occurrence of an object recognition error; if an object recognition error is detected, performing a recovery process to address the object recognition error; and if no object recognition error is detected, recognizing, by the processor, completion of unloading of the plurality of objects.
[0009] Aspects of the present disclosure involve an innovative non-transitory computer-readable medium storing instructions for recognizing and unloading a plurality of objects. The instructions may include repeatedly performing the following until the plurality of objects are unloaded: collecting object data by a processor through a vision sensor; performing object recognition by determining object pose and object position based on the object data; if an object of the plurality of objects is recognized and determined to be available for picking, picking up and unloading the object using a robotic device; if no object is determined to be available for picking, determining that an object recognition error has occurred; if an object recognition error is detected, performing a recovery process to address the object recognition error; and if no object recognition error is detected, recognizing completion of unloading of the plurality of objects.
[0010] An aspect of the present disclosure involves an innovative server system for recognizing and unloading a plurality of objects, which may include repeatedly performing the following until the plurality of objects are unloaded: collecting object data through a vision sensor, performing object recognition by determining object pose and object position based on the object data, picking up and unloading the object using a robotic device if an object of the plurality of objects is recognized and determined to be available for picking, determining if an object recognition error has occurred if no object is determined to be available for picking, performing a recovery process to address the object recognition error if an object recognition error is detected, and recognizing completion of unloading of the plurality of objects if no object recognition error is detected.
[0011] Aspects of the present disclosure involve an innovative system for recognizing and unloading a plurality of objects, which may include repeatedly performing the following until the plurality of objects are unloaded: collecting object data through a vision sensor; performing object recognition by determining object pose and object position based on the object data; picking up and unloading the object using a robotic device if an object of the plurality of objects is recognized and determined to be available for picking; determining that an object recognition error has occurred if no object is determined to be available for picking; performing a recovery process to address the object recognition error if an object recognition error is detected; and recognizing completion of unloading of the plurality of objects if no object recognition error is detected.
[0012] A general architecture embodying various features of the present disclosure is described below with reference to the drawings. The drawings and related description are provided to illustrate example embodiments of the present disclosure and are not intended to limit the scope of the disclosure. Throughout the drawings, reference numbers are also used again to indicate correspondence between referenced elements. [Brief explanation of the drawings]
[0013] [Figure 1(A)] FIG. 1 illustrates an exemplary process of a conventional recovery process. [Figure 1(B)] FIG. 1 illustrates an exemplary process of a conventional recovery process. [Figure 1(C)] FIG. 1 illustrates an exemplary process of a conventional recovery process. [Figure 2] FIG. 2 illustrates an exemplary system 200, according to one exemplary implementation. [Figure 3] FIG. 3 illustrates an example process flow 300 for performing object recognition and unloading, according to one example implementation. [Figure 4] FIG. 1 illustrates an exemplary remote recovery process, according to an exemplary embodiment. [Figure 5]FIG. 2 illustrates an exemplary recognition error detection process, according to an exemplary implementation. [Figure 6] FIG. 6 illustrates an alternative process flow 600 for performing object recognition and unloading, according to an example implementation. [Figure 7] FIG. 1 illustrates an exemplary confidence threshold update process, according to one exemplary implementation. [Figure 8] FIG. 8 illustrates an alternative process flow 800 for performing object recognition and unloading, according to an example implementation. [Figure 9] FIG. 1 illustrates an exemplary ML model training and testing process, according to an exemplary embodiment. [Figure 10] FIG. 10 illustrates an exemplary process flow 1000 for performing object recognition and unloading using trainable object recognition capabilities, according to one exemplary implementation. [Figure 11] FIG. 11 illustrates an exemplary graphic user interface (GUI) 1100 for performing area selection and error type display, according to an exemplary implementation. [Figure 12] FIG. 1 illustrates an exemplary computing environment having an exemplary computing device suitable for use in some exemplary implementations. DETAILED DESCRIPTION OF THE INVENTION
[0014] The following detailed description provides details of the figures and exemplary embodiments of the present application. Reference numbers and descriptions of elements that are duplicated between figures are omitted for clarity. Terms used throughout the description are provided by way of example and are not intended to be limiting. For example, use of the term "automatic" can include fully automatic or semi-automatic implementation, including user or administrator control over particular aspects of the implementation, depending on the desired implementation of those skilled in the art practicing the embodiments of the present application. Selection can be performed by a user via a user interface or other input means, or can be implemented via a desired algorithm. The exemplary embodiments as described herein can be used either alone or in combination, and the functionality of the exemplary embodiments can be implemented via any means according to the desired embodiment.
[0015] The present exemplary embodiment relates to a method and system for recognizing and unloading multiple objects. The exemplary embodiment utilizes simple yet informative input from a human operator during remote recovery that can be used to automatically generate useful information that improves the overall recognition function / process.
[0016] FIG. 2 illustrates an exemplary system 200 according to one exemplary embodiment. The system 200 may be used for palletizing / depalletizing single or mixed SKUs. The system 200 may include components such as, but not limited to, a robotic device 202, a vision sensor 204, a processor 206, and a memory 208. The robotic device 202 may be a robotic arm that performs functions of picking / grabbing and moving objects. The vision sensor 204 may be a camera directed toward a pallet 212 containing several objects 214. The vision sensor 204 captures images or video of the objects 214 on the pallet 212, which are then used to measure the objects 214. In some exemplary embodiments, the vision sensor captures the objects 214 in a data format other than images or video, such as a point cloud.
[0017] The processor 206 performs data processing on data (e.g., images, video, etc.) collected from the visual sensor 204 and issues commands to the robotic device 202 to control its movement. In particular, the processor 206 performs object recognition using the collected data. The memory 208 stores the data collected from the visual sensor 204, the recognition data as generated by the processor 206, and instructions / programs used by the various components of the system 200. A request to perform remote recovery may be sent to a user / operator through a graphic user interface (GUI) 210. The collected data and recognition results as generated from object recognition may be sent to the user / operator for review on the GUI 210, and a user response to the request may be entered through the GUI 210 and received by the processor 206.
[0018] 3 shows an example process flow 300 for performing object recognition and unloading, according to one example implementation. Process flow 300 begins in step S302, where object data is collected / measured using the vision sensor 204. In step S304, object recognition is performed to use the collected data to recognize the object's location and pose (e.g., object dimensions, etc.) and identify the object's availability for picking (e.g., an object may be recognized but not maneuverable due to object overlap, obstructions, etc.). Object recognition may be performed using at least one of a learning-based method and / or a rule-dependent method.
[0019] In step S306, the object is recognized and executed. objectBased on the recognition, a determination is made as to whether the object is available for picking. If the object is recognized as available for picking, the process proceeds to step S308, where the recognized object is picked up / grasped and moved. In step S318, a determination is made as to whether the object was successfully picked up. If the answer is "yes" in step S318, the process then returns to step S302, where object data is collected again. If the answer is "no" in step S318, the process then proceeds to step S310, which is described in more detail below.
[0020] If no objects are recognized as available for picking in step S306, the process continues to step S310, where a determination is made as to whether a recognition error has occurred. In some exemplary implementations, recognition error detection involves scanning the current object This is performed by comparing the recognition results with previous object recognition results (past results). FIG. 5 illustrates an exemplary recognition error detection process according to one exemplary implementation. As shown in FIG. 5, a recognition error is detected when an object 502 was correctly recognized in a previous object recognition cycle / step but is incorrectly recognized as a single object in the current step. In an alternative exemplary implementation, a recognition error may be detected when depth data in one image area indicates the presence of an object when the object is not recognized. In an alternative exemplary implementation, a failure to pick or unload a recognized object may indicate a recognition error. In an alternative exemplary implementation, a recognition error may exist when none of the recognized objects are suitable for picking and unloading. Whether an object is suitable for picking can be determined in many ways, such as minimum / maximum object size, likelihood of collision with other objects, etc. If the answer is "no" in step S310, the process ends.
[0021] If a recognition error is detected in step S310, a remote recovery request is sent / issued to the user / operator in S312 along with the collected data / objects for review. In step S314, the user / operator then generates a correction response including a selected area in the collected object data where the recognition error occurred and a specified error type of the recognition error. FIGS. 4(A) and 4(B) show an exemplary remote recovery process according to an exemplary embodiment. As shown in FIG. 4(A), a recognition error is detected and a remote recovery request is sent / issued to the user / operator. As shown in FIG. 4(B), the user / operator makes a selection on the provided object data and indicates an error type associated with the selected object (e.g., error type 1, error type 2, error type 3, etc.). In some exemplary embodiments, area selection and error type display may be made using a GUI. FIG. 11 shows an exemplary GUI 1100 for performing area selection and error type display according to an exemplary embodiment. The remote recovery request is received on the user device 1102 and displayed through a GUI 1104. The information displayed on the GUI 1104 may include the current object placement and recognition results. As shown in Figure 11, the user / operator can provide annotations such as boundary selection and error type identification.
[0022] If user input is received, the collected data is then updated based on the user input in step S316, and object recognition is performed in step S304 using the updated data from step S316. Process flow 300 is repeated until no objects remain on the pallet for processing.
[0023] FIG. 6 illustrates an alternative process flow 600 for performing object recognition and unloading according to an exemplary embodiment. Process flow 600 is similar to process flow 300 of FIG. 3, except for several additional steps. Process flow 600 begins in step S602, where object data is collected / measured using the vision sensor 204. In step S604, using the collected data, object recognition is performed to detect the object's location and pose, and an object confidence calculation for the object is performed based on the determined location and object pose of the object to generate a confidence value. Alternatively, object confidence may be calculated using other methods, such as, but not limited to, object probabilities output from a neural network during object recognition. Each confidence value is associated with a corresponding object. Object recognition may be performed using at least one of an object recognition learning-based method or a rule-based method.
[0024] In step S606, a confidence threshold is applied and used in determining whether an object has been detected / recognized. In step S608, a determination is made as to whether an object has been recognized as available for picking based on a confidence value comparison with the applied confidence threshold. In particular, the object's confidence value is compared against the confidence threshold. If the object's confidence value is equal to or greater than the confidence threshold, the object is determined to have been detected / recognized. If the object is identified as available for picking, the process then continues to step S610, where the recognized object is picked up / grasped and moved. In step S624, a determination is made as to whether the object was successfully picked up. If the answer is "yes" in step S624, the process then returns to step S602, where object data collection occurs again. If the answer is "no" in step S624, the process proceeds to step S612, which is described in more detail below.
[0025] However, in step S608, no objects are recognized as available for picking (e.g., all objects are below the confidence threshold). value, the object is determined as not detected / recognized, and the process then proceeds to step S612, where a determination is made as to whether a recognition error has occurred. In some exemplary implementations, recognition error detection is performed by checking the current object This is done by comparing the results of the recognition with previous object recognition results (past results).If the answer is "no" in step S612, the process ends.
[0026] If a recognition error is detected in step S612, a remote recovery request is collected for review in step S614. data / The object recognition information is sent / issued to the user / operator along with the object. In step S616, the user / operator then selects an area for the collected object data where a recognition error occurred and indicates the error type of the recognition error. In some exemplary implementations, area selection and error type indication may be done using a GUI. The process then continues to step S618, where the user / operator determines whether there is an undetected object caused by an inadequate confidence threshold. If the answer to step S618 is "no," the process continues to step S620, where the collected data is then updated based on the user input, and object recognition is performed again in step S604 using the updated data from step S620.
[0027] If the answer to step S618 is "yes," the process continues to step S622, where areas associated with non-detected objects are created and confidence threshold adjustments are performed accordingly. Once the confidence threshold has been updated, the process then returns to step S606, where the updated confidence threshold is applied. Process flow 600 continues repeatedly until no objects remain on the pallet for processing.
[0028] 7(A) and 7(B) illustrate an exemplary confidence threshold update process according to one exemplary implementation. As shown in FIG. 7(A), undetected objects and detected objects are identified. Undetected objects are objects with confidence values below the confidence threshold, and detected objects are objects with confidence values equal to or greater than the confidence threshold. As shown in FIG. 7(B), a user / operator makes an adjustment or update to the confidence threshold, thereby leading to the detection of a previously undetected object.
[0029] 8 illustrates an alternative process flow 800 for performing object recognition and unloading, according to an exemplary embodiment. Process flow 800 utilizes a neural network in performing object recognition. Process flow 800 begins at step S802, where object data is collected / measured using the vision sensor 204. At step S804, object edge / boundary detection is performed using red, green, and blue (RGB) images. At step S806, object edge / boundary detection is performed using depth images.
[0030] Object edge / boundary detection through RGB and depth images is performed using a machine learning (ML) algorithm. FIG. 9 illustrates an exemplary ML model training and testing process, according to an exemplary embodiment. During the model training phase, a sensor 902, such as a vision sensor 204, is used to capture an RGB image 904 and a depth image 910 of an object. In some exemplary embodiments, two or more sensors 902 may be utilized in capturing the RGB image 904 and the depth image 910 of the object. For example, a first sensor 902 may be utilized in capturing the RGB image 904, and a second sensor 902 may be utilized in capturing the depth image 910. The RGB image 904 and the depth image 910 are then separately utilized to train two ML models. In some exemplary embodiments, a deep neural network (DNN) is utilized as the ML model. As illustrated in FIG. 9 , the RGB image 904 is utilized to train an RGB DNN 906 to generate a trained RGB DNN 908. The depth image 910 is used to train a depth DNN 912 to generate a trained depth DNN 914. The trained DNN is utilized to detect edges / boundaries of objects in the RGB image 904 and the depth image 910. The output of the trained DNN may include a set of edge / boundary probability images with pixel values representing the likelihood that a pixel belongs to an edge / boundary.
[0031] During the implementation / testing phase, a sensor 902 captures an RGB image 916 and a depth image 920 of an object to be processed. In some example implementations, more than one sensor 902 may be utilized in capturing the RGB image 916 and the depth image 920 of the object. The RGB image 916 is sent to a trained RGB DNN 908 to generate a set of edge / boundary images indicating edge probabilities 918. The depth image 920 is sent to a trained depth DNN 914 to generate a set of edge / boundary images indicating edge probabilities 922. Weights are then assigned to the edge probabilities 918 and 922, and the two are combined to generate a combined edge probability 924.
[0032] The pixel values of the combined edge probabilities 924 are then compared to a threshold to determine whether the pixel belongs to an edge / boundary. The result is a binarized edge map (binarized edge image 926) where each pixel is either an edge / boundary or not. Objects can then be recognized (recognized objects 930) using methods such as image segmentation 928 on the detected edges / boundaries.
[0033] Referring again to FIG. 8, the edges / boundaries from the RGB image and the depth image are combined using weight assignments in step S808. In step S810, objects are recognized using the combined edges / boundaries. The process then continues to step S812, where the objects are recognized and the resulting image is processed. object Based on the recognition, a determination is made as to whether the object is recognized as available for picking. If the object is recognized as available for picking, the process then proceeds to step S814, where the recognized object is picked up / grasped and moved. In step S826, a determination is made as to whether the object was successfully picked up. If the answer is "yes" in step S826, the process then returns to step S802, where object data is collected again. If the answer is "no" in step S826, the process proceeds to step S816, which is described in more detail below.
[0034] If no objects are recognized as available for picking in step S812, the process continues to step S816, where a determination is made as to whether a recognition error occurred. If the answer is "no" in step S816, the process ends. If a recognition error is detected in step S816, a remote recovery request is sent to the object collected for review in step S818. data / The object recognition error information is sent / issued to the user / operator along with the object. In step S820, the user / operator then selects an area of the collected object data where a recognition error occurred and indicates the error type of the recognition error. In some exemplary implementations, the area selection and error type indication can be done using a GUI.
[0035] In step S822, the user / operator determines the presence of an incorrectly segmented object or a failed object segmentation (an incorrectly unsegmented object). The process then continues to step S824, where adjustments are made to the weights of the edge probabilities (edge probabilities 918 and edge probabilities 922). If an incorrect object segmentation is present, the weight associated with the trained RGB DNN 908 (edge probabilities 918) is increased. Otherwise, the weight associated with the trained RGB DNN 908 is decreased. Alternatively, the weight associated with the trained depth DNN 914 (edge probabilities 922) is decreased if an incorrect object segmentation is detected or increased if a failed object segmentation is present. Once step S824 is completed, the process then returns to step S808, where the updated weights are applied. The process flow 800 is repeated until no objects remain in the palette for processing.
[0036] 10 shows an example process flow 1000 for performing object recognition and unloading using trainable object recognition functionality, according to one example implementation. Process flow 1000 begins in step S1002, where object data is collected / measured using the vision sensor 204. In step S1004, object recognition is performed on the collected data using the trainable object recognition functionality. As part of object recognition, identification of the object's availability for picking is also performed.
[0037] In step S1006, the object is recognized and executed. objectBased on the recognition, a determination is made as to whether the object is available for picking. If the object is recognized as available for picking, the process then proceeds to step S1008, where the recognized object is picked up / grasped and moved. In step S1020, a determination is made as to whether the object was successfully picked up. If the answer is "yes" in step S1020, the process then returns to step S1002, where object data is collected again. If the answer is "no" in step S1020, the process proceeds to step S1010, which is described in more detail below.
[0038] If no objects are recognized as available for picking in step S1006, the process continues to step S1010, where a determination is made as to whether a recognition error has occurred. In some exemplary implementations, recognition error detection involves scanning the current object This is done by comparing the results of the recognition with previous object recognition results (past results).If the answer is "no" in step S1010, the process ends.
[0039] If a recognition error is detected in step S10I0, a remote recovery request is sent / issued to the user / operator in S1012 along with the collected data / objects for review. In step S1014, the user / operator then selects an area of the collected object data where a recognition error occurred and indicates the error type of the recognition error. If user input is received, the collected data is then updated based on the user input in step S1016, and object recognition is re-run in step S1004 using the updated data. In addition to re-running object recognition in step S1004, the trainable object recognition function can be trained online or offline using updated data to improve the function's recognition accuracy rate in step S1018. For example, the trained DNNs (trained RGB DNN 908 and trained depth DNN 914) of FIG. 9 can be further trained using the updated results. Process flow 1000 is repeated until no objects remain on the pallet for processing.
[0040] The exemplary embodiments may have various benefits and advantages. For example, the exemplary embodiments may enable a reduction in costs and processing time associated with object recognition. In particular, costs and processing time are reduced as a result of reduced annotation complexity. Furthermore, the use of annotations and updated recognition results for online / offline function training achieves improved recognition accuracy rates in object recognition.
[0041] 12 illustrates an exemplary computing environment having an exemplary computing device suitable for use in some exemplary implementations. The computing device 1205 in the computing environment 1200 can include one or more processing units, cores, or processors 1210, memory 1215 (e.g., RAM, ROM, and / or the like), internal storage 1220 (e.g., magnetic, optical, solid-state storage, and / or organic), and / or IO interface 1225, any of which can be coupled over a communication mechanism or bus 1230 for communicating information or can be incorporated into the computing device 1205. The IO interface 1225 can be further configured to receive images from a camera or provide images to a projector or display, depending on the desired implementation.
[0042] Computing device 1205 may be communicatively coupled to input / user interface 1235 and output device / interface 1240. Either or both of input / user interface 1235 and output device / interface 1240 may be wired or wireless interfaces and may be detachable. Input / user interface 1235 may include any device, component, sensor, or interface, physical or virtual, that can be used to provide input (e.g., buttons, touchscreen interface, keyboard, pointing / cursor control, microphone, camera, Braille, motion sensor, accelerometer, optical reader, and / or the like). Output device / interface 1240 may include a display, television, monitor, printer, speaker, Braille, or the like. In some exemplary implementations, input / user interface 1235 and output device / interface 1240 may be incorporated with or physically coupled to computing device 1205. In other implementations, other computing devices may function as or provide the functionality of input / user interface 1235 and output device / interface 1240 for computing device 1205 .
[0043] Examples of computing devices 1205 may include, but are not limited to, highly mobile devices (e.g., smartphones, devices mounted on vehicles and other machines, devices carried by people or animals, and the like), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, and the like), and devices not designed for mobility (e.g., desktop computers, other computers, information kiosks, televisions with one or more processors embedded and / or televisions with one or more processors coupled thereto, radios, and the like).
[0044] Computing device 1205 may be communicatively coupled (e.g., via IO interface 1225) to external storage 1245 and to a network 1250 for communication with any number of networked components, devices, and systems, including one or more computing devices of the same or different configurations. Computing device 1205 or any other connected computing device may function as, provide services to, or be referred to as a server, client, thin server, general-purpose machine, special-purpose machine, or otherwise.
[0045] IO interface 1225 may include, but is not limited to, wired and / or wireless interfaces using any communication or IO protocol or convention (e.g., Ethernet, 802.11x, Universal System Bus, WiMax, modem, cellular network protocols, and the like) for communicating information to and / or from at least all connected components, devices, and networks in computing environment 1200. Network 1250 may be any network or combination of networks (e.g., the Internet, a local area network, a wide area network, a telephone network, a cellular network, a satellite network, and the like).
[0046] The computing device 1205 can use and / or communicate using computer-usable or computer-readable media, including transitory and non-transitory media. Transitory media include transmission media (e.g., metallic cables, fiber optics), signals, carrier waves, and the like. Non-transitory media include magnetic media (e.g., disks and tape), optical media (e.g., CD-ROM, digital video disks, Blu-ray® disks), solid-state media (e.g., RAM, ROM, flash memory, solid-state storage), and other non-volatile storage or memory.
[0047] The computing device 1205 can be used to implement techniques, methods, applications, processes, or computer-executable instructions in some exemplary computing environments. The computer-executable instructions can be retrieved from a transitory medium and stored on and retrieved from a non-transitory medium. The executable instructions can be from one or more of any programming language, scripting language, and machine language (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, etc.).
[0048] The processor 1210 can run under any operating system (OS) (not shown) in a native or virtual environment. One or more applications can be deployed, including a logic unit 1260, an application programming interface (API) unit 1265, an input unit 1270, an output unit 1275, and an inter-unit communication mechanism 1295 for different units to communicate with each other, with the OS, and with other applications (not shown). The above-mentioned units and elements can vary in design, function, configuration, or implementation and are not limited to the above description. The processor 1210 can have the form of a hardware processor, such as a central processing unit (CPU), or can be a combination of hardware and software units.
[0049] In some exemplary implementations, when information or instructions for execution are received by API unit 1265, it may be communicated to one or more other units (e.g., logic unit 1260, input unit 1270, output unit 1275). In some examples, logic unit 1260 may be configured to control the flow of information between units and, in some exemplary implementations described above, direct the services provided by API unit 1265, input unit 1270, and output unit 1275. For example, the flow of one or more processes or implementations may be controlled by logic unit 1260 alone or in conjunction with API unit 1265. Input unit 1270 may be configured to obtain inputs for the calculations described in the exemplary implementations, and output unit 1275 may be configured to provide outputs based on the calculations described in the exemplary implementations.
[0050] The processor 1210 may be configured to collect object data through a visual sensor, as shown in FIG. 3. The processor 1210 may also be configured to perform object recognition by determining an object pose and an object position based on the object data, as shown in FIG. 3. The processor 1210 may also be configured to pick up and unload an object using a robotic device if an object of the plurality of objects is recognized, as shown in FIG. 3. The processor 1210 may also be configured to determine the occurrence of an object recognition error, as shown in FIG. 3. The processor 1210 may also be configured to execute a recovery process to address the object recognition error if an object recognition error is detected, as shown in FIG. 3. The processor 1210 may also be configured to recognize the completion of unloading of the plurality of objects if no object recognition error is detected, as shown in FIG. 3.
[0051] Some portions of the detailed description are presented in terms of algorithms and symbolic representations of operations within a computer. These algorithmic descriptions and symbolic representations are the means used by those skilled in the data processing arts to convey the substance of their innovations to others skilled in the art. An algorithm is a series of defined steps leading to a desired end state or result. In exemplary implementations, the performed steps require physical manipulations of tangible quantities to achieve a tangible result.
[0052] Unless otherwise specified, and as will be apparent from the description, throughout this specification, descriptions utilizing words such as "processing," "calculating," "computing," "determining," "displaying," or the like, are understood to include the actions and processes of a computer system or other information processing device that manipulates and converts data represented as physical (electronic) quantities in the registers and memory of the computer system into other data similarly represented as physical quantities in the memory or registers of the computer system or other information storage, transmission, or display devices.
[0053] Exemplary embodiments may also relate to apparatuses for performing the operations herein. This apparatus may be specially constructed for the desired purposes, or may include one or more general-purpose computers selectively activated or reconfigured by one or more computer programs. Such computer programs may be stored on a computer-readable medium, such as a computer-readable storage medium or a computer-readable signal medium. Computer-readable storage media may include tangible media, such as, but not limited to, optical disks, magnetic disks, read-only memory, random-access memory, solid-state devices and drives, or any other type of tangible or non-transitory medium suitable for storing electronic information. Computer-readable signal media may include media such as carrier waves. The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. A computer program may include a pure software implementation containing instructions for performing the operations of a desired embodiment.
[0054] Various general-purpose systems may be used with the programs and modules according to the examples herein, or it may eventually be convenient to construct specialized apparatus to perform the desired method steps. Moreover, the example embodiments are not described with reference to any particular programming language. It will be understood that a variety of programming languages may be used to implement the teachings of the example embodiments as described herein. Instructions of the programming language may be executed by one or more processing devices, such as, for example, a central processing unit (CPU), a processor, or a controller.
[0055] As is known in the art, the operations described above may be performed by hardware, software, or some combination of software and hardware. Various aspects of the exemplary embodiments may be implemented using circuits and logic devices (hardware), while other aspects may be implemented using instructions stored on a machine-readable medium (software) that, when executed by a processor, cause the processor to perform methods that implement the present application. Furthermore, some exemplary embodiments of the present application may be implemented solely in hardware, while other exemplary embodiments may be implemented solely in software. Furthermore, the various functions described may be performed in a single unit or may be distributed across multiple components in any number of ways. When implemented by software, the methods may be executed by a processor, such as a general-purpose computer, based on instructions stored on a computer-readable medium. If desired, the instructions may be stored on the medium in compressed and / or encrypted format.
[0056] Additionally, other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the teachings herein. Various aspects and / or components of the described exemplary embodiments may be used singly or in any combination. It is intended that the specification and exemplary embodiments be considered exemplary only, with the true scope and spirit of the present application being indicated by the following claims. [Explanation of symbols]
[0057] 208 memory 206 processors 204 Visual Sensor 214 Object 212 palettes 202 Robot Devices 1205 Computer Devices 1210 processor 1215 memory 1220 Internal Storage 1225 IO interface 1235 Input / User Interface 1240 Output Device / Interface 1245 External Storage 1250 Network 1260 logical units 1265 API units 1270 input units 1275 output units
Claims
1. 1. A method for recognizing and unloading a plurality of objects, the method comprising: collecting, by a processor, object data through a visual sensor; performing, by the processor, object recognition by determining an object pose and an object position based on the object data; If an object of the plurality of objects is recognized and determined to be available for picking, picking and unloading the object using a robotic device; If no objects are determined to be available for picking, determining, by the processor, that an object recognition error has occurred; If the object recognition error is detected, executing, by the processor, a recovery process to address the object recognition error; if the object recognition error is not detected, recognizing, by the processor, completion of unloading of the plurality of objects; and The recovery process is sending a recovery request including the object data to a user; receiving a response to the recovery request, the response including at least one selected recognition error area and an error type indicating a type of the object recognition error corresponding to the recognition error area; updating data for object recognition based on the error type and re-performing object recognition; The method is configured to be performed by
2. The method of claim 1 , wherein the object pose comprises an object dimension.
3. The processor may configure the recovery process to respond to one error type by: receiving a corrective response from the user that determines the presence of an incorrectly segmented object or an incorrectly unsegmented object as a result of the object recognition; updating data for object recognition based on the corrected response and re-performing object recognition; The method of claim 1 , wherein the method is configured to be performed by:
4. The method of claim 3 , wherein object recognition is performed using at least one trainable object recognition function.
5. The method of claim 4 , wherein the training of the at least one learnable object recognition function is performed online or offline.
6. The method of claim 3 , wherein sending the recovery request including the object data to the user includes sending the recovery request to a graphic user interface (GUI) for review by the user.
7. the processor is configured to perform object recognition by performing an object belief operation on the plurality of objects to generate belief values, each belief value being associated with a corresponding object of the plurality of objects; Recognizing an object among the plurality of objects comparing the object confidence value to a confidence threshold; If the certainty value of the object is equal to or greater than the certainty threshold, the object is determined to have been recognized; by determining that the object has not been recognized if the confidence value of the object is less than the confidence threshold; the processor is configured to execute the recovery process to address the object recognition error by further performing a threshold adjustment on the confidence threshold. The method of claim 3.
8. The processor performs object recognition by: performing object edge detection using red, blue, and green (RGB) images of the plurality of objects to detect a first object boundary; performing object edge detection using the plurality of object depth images to detect a second object boundary; using the RGB image as an input to a first trained machine learning model to generate a first edge probability image having pixel values representing the likelihood that the pixels belong to the object edge of a first object boundary; using the depth image as an input to a second trained machine learning model to generate a second edge probability image having pixel values representing the likelihood that the pixels belong to the object edge of a second object boundary; assigning weights to the first edge probability image and the second edge probability image; combining the first edge probability image and the second edge probability image based on the weights; generating a binarized edge map for the plurality of objects, in which each pixel of the combined edge probability image generated by the combination is binarized to indicate whether it is an edge or not, and performing object recognition for the plurality of objects; is configured to run by The method of claim 3 , wherein the object data includes the RGB image and the depth image.
9. the processor executes the recovery process for addressing the object recognition error, adjusting the weights assigned to the first edge probability image and the second edge probability image to improve object recognition; The method of claim 8 , further configured to be performed by:
10. The processor calculates the weights assigned to the first edge probability image and the second edge probability image by: if a failed object segmentation is detected, increasing a first weight among the weights or decreasing a second weight among the weights, the first weight being associated with the first edge probability image and the second weight being associated with the second edge probability image; 10. The method of claim 9, configured to adjust by decreasing the first weight or increasing the second weight if an incorrectly segmented object is detected.
11. 1. A system for recognizing and unloading a plurality of objects, the system comprising: A robotic device; a processor in communication with the robotic device, the processor: Collect object data through visual sensors, performing object recognition by determining object pose and object position based on the object data; If an object of the plurality of objects is recognized and determined to be available for picking, using a robotic device to pick and unload the object; determining that an object recognition error has occurred if no object is determined to be available for picking; If the object recognition error is detected, executing a recovery process to address the object recognition error; If the object recognition error is not detected, repeating the process of recognizing completion of unloading of the plurality of objects; The recovery process is sending a recovery request including the object data to a user; receiving a response to the recovery request, the response including at least one selected recognition error area and an error type indicating a type of the object recognition error corresponding to the recognition error area; updating data for object recognition based on the error type and re-performing object recognition; The system is configured to run by
12. The system of claim 11 , wherein the object pose comprises an object dimension.
13. The processor may configure the recovery process to respond to one error type by: receiving a corrective response from the user that determines the presence of an incorrectly segmented object or an incorrectly unsegmented object as a result of the object recognition; updating data for object recognition based on the corrected response and re-performing object recognition; The system of claim 11 , configured to execute by:
14. The system of claim 13 , wherein object recognition is performed using at least one trainable object recognition function.
15. The system of claim 14 , wherein training of the at least one learnable object recognition function is performed online or offline.
16. The system of claim 13 , wherein sending the recovery request including the object data to the user includes sending the recovery request to a graphic user interface (GUI) for review by the user.
17. the processor is configured to perform object recognition by performing an object belief operation on the plurality of objects to generate belief values, each belief value being associated with a corresponding object of the plurality of objects; Recognizing an object among the plurality of objects comparing the object confidence value to a confidence threshold; If the certainty value of the object is equal to or greater than the certainty threshold, the object is determined to have been recognized; by determining that the object has not been recognized if the confidence value of the object is less than the confidence threshold; the processor is configured to execute the recovery process to address the object recognition error by further performing a threshold adjustment on the confidence threshold. The system of claim 13.
18. The processor performs object recognition by: performing object edge detection using red, blue, and green (RGB) images of the plurality of objects to detect a first object boundary; performing object edge detection using the plurality of object depth images to detect a second object boundary; using the RGB image as an input to a first trained machine learning model to generate a first edge probability image having pixel values representing the likelihood that the pixels belong to the object edge of a first object boundary; using the depth image as an input to a second trained machine learning model to generate a second edge probability image having pixel values representing the likelihood that the pixels belong to the object edge of a second object boundary; assigning weights to the first edge probability image and the second edge probability image; combining the first edge probability image and the second edge probability image based on the weights; The method is configured to perform object recognition by generating a binarized edge map for the plurality of objects, which binarizes each pixel of the combined edge probability image generated by the combination and indicates whether it is an edge or not, and then performing object recognition on the plurality of objects; the object data includes the RGB image and the depth image; The system of claim 13.
19. the processor executes the recovery process for addressing the object recognition error, adjusting the weights assigned to the first edge probability image and the second edge probability image to improve object recognition; 20. The system of claim 18, further configured to perform by:
20. The processor calculates the weights assigned to the first edge probability image and the second edge probability image by: if a failed object segmentation is detected, increasing a first weight among the weights or decreasing a second weight among the weights, the first weight being associated with the first edge probability image and the second weight being associated with the second edge probability image; and adjusting the first weight by decreasing the first weight or increasing the second weight if an incorrectly segmented object is detected.
20. The system of claim 19.
Citation Information
Patent Citations
Image recognition device
JP2005149074A
Mix-size depalletizing
JP2022045905A
ROBOTIC SYSTEM WITH OBJECT UPDATE MECHANISM AND METHOD FOR OPERATING A ROBOTIC SYSTEM - Patent application
JP2023052476A
Control device, control method, and storage medium
WO2022224447A1