System and method for picking item

The method automates remote recovery and improves object recognition accuracy by integrating human operator corrections, addressing the inefficiencies of current systems in handling mixed-load SKU recognition.

JP2025093890AActive Publication Date: 2025-06-24HITACHI LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024215128
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-12
Filing Date
2024-12-10
Publication Date
2025-06-24
Estimated Expiration
2044-12-10

Smart Images

  • Figure 2025093890000001_ABST
    Figure 2025093890000001_ABST
Patent Text Reader

Abstract

To provide a method and a system for recognizing and unloading a plurality of objects.SOLUTION: In a system including a robotic device, a method repeatedly executes until a plurality of objects are unloaded: collecting object data via a visual sensor; executing object recognition by determining an object pose and an object position based on the object data; picking up and unloading the object by using the robotic device if an object of the plurality of objects is determined to be available for picking; determining that an object recognition error has occurred if no object is determined to be available for picking; executing a recovery process to address the object recognition error if the object recognition error is detected; and recognizing completion of unloading of the plurality of objects if no object recognition error is detected.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to methods and systems for performing object recognition and unloading.

Background Art

[0002] As part of business operations, depalletizing of articles from pallets is performed by retailers, wholesalers, and other third-party logistics vendors. Some pallets may have only one type of parcel or a single Stock Keeping Unit (SKU), while other pallets may contain different types of parcels, also known as mixed-load SKUs. Robot devices such as robotic arms are used in performing both palletizing and depalletizing of single SKUs and mixed-load SKUs.

[0003] In related technologies, recognition software using computer vision techniques (e.g., conventional rule-based methods, machine learning-based methods, etc.) has been developed and used in recognizing objects and their locations. Once object recognition is performed, the robot device then proceeds to pick / grip and move the identified objects based on the recognition results. However, there remains a problem that the recognition software cannot accurately recognize articles / objects, leading to failures in palletizing and depalletizing. This is particularly problematic for depalletizing of mixed-load SKUs because many types of objects are involved.

[0004] In related technologies, methods for performing remote perception assistance and object identification correction have been disclosed. Remote assistance is desired to verify object identification and provide corrections. Based on the corrections, additional processing operations on the objects are then performed by the robot.

[0005] In the related art, a method for training a machine learning model to identify an object for picking is disclosed. The data used when training the machine learning model is collected from edge cases where the machine learning model fails to detect an object.

Summary of the Invention

Problems to be Solved by the Invention

[0006] Currently, remote recovery is required when misrecognition results or errors occur. Remote recovery involves having the robot device send an image of the objects on the pallet to a human operator, who then performs object area selection on the undetected objects or objects that have been incorrectly grouped with another object (e.g., generating a rectangular area for the object), and instructing the robot device to pick / grip the objects that have become identified. However, this remote recovery process is insufficient for two reasons. 1) Manual object selection when performed by a human can be time-consuming due to the complexity of object selection and cumbersome annotation, and 2) the recognition accuracy / recognition ability of the recognition software is not improved when useful information such as object boundaries is not provided or utilized.

[0007] Figures 1(A) to 1(C) are diagrams showing exemplary process flows of a conventional recovery process. As shown in Figure 1(A), recognition errors have occurred and been identified. Various recognition errors are associated with the objects shown in Figure 1(A). In particular, the upper left object and the upper right object are not recognized, the central object has a pose error, the lower right object does not fit correctly, and the lower left object is not segmented and is incorrect. Taking the central object on pallet 100 as an example, a pose error has been identified and remote recovery assistance is required. As shown in Figure 1(B), the central object is remotely selected by a human operator under a recovery approach. It requires expertise and may take time to determine which object to select and highlight that object. As shown in Figure 1(C), the information necessary to improve the object recognition function is manually annotated, which incurs additional cost and time.

Means for Solving the Problems

[0008] Aspects of the present disclosure involve an innovative method for recognizing and unloading a plurality of objects. The method includes, until the plurality of objects are unloaded, collecting object data through a vision sensor by a processor, performing object recognition by the processor by determining an object pose and an object position based on the object data, picking up and unloading an object using a robotic device when an object among the plurality of objects is recognized and determined to be available for picking, determining by the processor the occurrence of an object recognition error when there is no object determined to be available for picking, executing a recovery process by the processor to address the object recognition error when the object recognition error is detected, and recognizing by the processor the completion of unloading of the plurality of objects when the object recognition error is not detected, and may include repeatedly executing the above.

[0009] Aspects of the present disclosure involve an innovative non - transitory computer - readable medium that stores instructions for recognizing and unloading a plurality of objects. The instructions cause a processor to, until the plurality of objects are unloaded, collect object data through a vision sensor, perform object recognition by determining an object pose and an object position based on the object data, use a robotic device to pick up and unload an object if an object among the plurality of objects is recognized and determined to be available for picking, determine the occurrence of an object recognition error by the processor if there is no object determined to be available for picking, execute a recovery process by the processor to handle the object recognition error if the object recognition error is detected, recognize the completion of unloading of the plurality of objects by the processor if no object recognition error is detected, and repeat the execution of these operations.

[0010] Aspects of the present disclosure involve an innovative server system for recognizing and unloading a plurality of objects. The server system causes a processor to, until the plurality of objects are unloaded, collect object data through a vision sensor, perform object recognition by determining an object pose and an object position based on the object data, use a robotic device to pick up and unload an object if an object among the plurality of objects is recognized and determined to be available for picking, determine the occurrence of an object recognition error by the processor if there is no object determined to be available for picking, execute a recovery process by the processor to handle the object recognition error if the object recognition error is detected, recognize the completion of unloading of the plurality of objects by the processor if no object recognition error is detected, and repeat the execution of these operations.

[0011] Aspects of the present disclosure involve an innovative system for recognizing and unloading multiple objects. The system includes means for collecting object data through a visual sensor until the multiple objects are unloaded, means for performing object recognition by determining an object pose and an object position based on the object data, means for picking up and unloading an object using a robotic device when the object among the multiple objects is recognized and determined to be available for picking, means for determining the occurrence of an object recognition error when there is no object determined to be available for picking, means for executing a recovery process to address the object recognition error when the object recognition error is detected, and means for recognizing the completion of unloading of the multiple objects when no object recognition error is detected, and may include repeatedly executing the above.

[0012] Hereinafter, a general architecture for implementing various features of the present disclosure will be described with reference to the drawings. The drawings and the related description are provided to illustrate exemplary embodiments of the present disclosure and are not intended to limit the scope of the present disclosure. Throughout the drawings, reference numerals are reused to indicate the correspondence between the referenced elements.

Brief Description of the Drawings

[0013]

Figure 1(A)

Figure 1(B)

Figure 1(C)

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

[0014] The following detailed description provides details of the figures and exemplary embodiments of the present application. Reference numerals and descriptions of overlapping elements between the figures are omitted for clarity. The terms used throughout the description are provided as examples and are not intended to be limiting. For example, the use of the term "automated" may include fully automated or semi-automated implementations, including user or administrator control over specific aspects of the implementation, depending on the desired implementation by those skilled in the art of practicing the embodiments of the present application. The selection can be performed by the user via a user interface or other input means, or can be implemented via a desired algorithm. Exemplary embodiments as described herein are available either alone or in combination, and the functions of such exemplary embodiments can be implemented via any means according to the desired embodiment.

[0015] This exemplary embodiment relates to a method and system for recognizing and unloading a plurality of objects. The exemplary embodiment can utilize simple but informative input from a human operator during remote recovery to automatically generate useful information for improving the entire recognition function / process.

[0016] Figure 2 shows an exemplary system 200 according to an exemplary embodiment. The system 200 can be used for palletizing / depalletizing a single SKU or mixed-load SKUs. The system 200 can include components such as, but not limited to, a robotic device 202, a vision sensor 204, a processor 206, and a memory 208. The robotic device 202 can be a robotic arm that performs functions of picking / gripping and moving an object. The vision sensor 204 can be a camera directed at a pallet 212 that includes several objects 214. The vision sensor 204 captures an image or video of the objects 214 on the pallet 212, and that image or video is then used to measure the objects 214. In some exemplary embodiments, the vision sensor captures the objects 214 in a data format other than an image or video, such as a point cloud.

[0017] The processor 206 performs data processing on the data (such as images, videos, etc.) collected from the vision sensor 204, and issues commands for controlling the movement of the robot device 202 to the robot device 202. In particular, the processor 206 performs object recognition using the collected data. The memory 208 stores the data collected from the vision sensor 204, recognition data such as that generated by the processor 206, and the instructions / programs used by various components of the system 200. Requests for performing remote recovery can be sent to the user / operator through the graphical user interface (GUI) 210. The collected data and recognition results such as those generated from object recognition may be sent to the user / operator for consideration on the GUI 210, and the user response to the request is input through the GUI 210 and can be received by the processor 206.

[0018] Figure 3 shows an exemplary process flow 300 for performing object recognition and unloading according to an exemplary embodiment. The process flow 300 starts at step S302 where object data is collected / measured using the vision sensor 204. At step S304, object recognition is performed using the collected data to recognize the location and orientation (such as object dimensions, etc.) of the object and to identify the availability of the object for picking (for example, the object may be recognized but may not be operable due to object overlap or obstacles, etc.). The object recognition can be performed using at least one of a learning-based method and / or a rule-dependent method.

[0019] At step S306, the object is recognized and executed objectA determination is made as to whether it is available for picking based on the recognition. If the object is recognized as available for picking, the process proceeds to step S308, and the recognized object is picked up / grasped and moved. At step S318, a determination is made as to whether the pickup of the object was successful. If the answer at step S318 is "yes", the process then returns to step S302, and the collection of object data is performed again. If the answer at step S318 is "no", the process then proceeds to step S310, which is detailed below.

[0020] If there is no object recognized as available for picking at step S306, the process continues to step S310, and a determination is made as to whether a recognition error has occurred. In some exemplary embodiments, the recognition error detection is performed by comparing the current object recognition result with the previous object recognition result (past result) to check for no discrepancies. FIG. 5 shows an exemplary recognition error detection process according to an exemplary embodiment. As shown in FIG. 5, a recognition error is detected if the object 502 was correctly recognized in a previous object recognition cycle / step but is misrecognized as a single object in the current step. In an alternative exemplary embodiment, a recognition error may be detected if the depth data in one image area indicates the presence of an object when the object is not recognized. In an alternative exemplary embodiment, the failure of picking up or unloading a recognized object may indicate a recognition error. In an alternative exemplary embodiment, a recognition error may exist if none of the recognized objects are suitable for picking and unloading. Whether an object is suitable for picking can be determined in many ways, such as, for example, the minimum / maximum object size, the possibility of collision with other objects, etc. If the answer at step S310 is "no", the process ends.

[0021] If a recognition error is detected in step S310, a remote recovery request is sent / issued to the user / operator in S312, together with the data / objects collected for consideration. In step S314, the user / operator then generates a correction response that includes the selected area in the collected object data where the recognition error occurred and the specified error type of that recognition error. FIGS. 4(A) and 4(B) show an exemplary remote recovery process according to an exemplary embodiment. As shown in FIG. 4(A), a recognition error is detected and a remote recovery request is sent / issued to the user / operator. As shown in FIG. 4(B), the user / operator makes a selection on the provided object data and indicates the error type associated with the selected object (e.g., error type 1, error type 2, error type 3, etc.). In some exemplary embodiments, the area selection and error type display can be made using a GUI. FIG. 11 shows an exemplary GUI 1100 for performing area selection and error type display according to an exemplary embodiment. The remote recovery request is received on the user device 1102 and displayed through the GUI 1104. The information displayed on the GUI 1104 can include the current object placement and recognition results. As shown in FIG. 11, the user / operator can provide annotations such as boundary selection and error type identification.

[0022] When user input is received, the collected data is then updated based on the user input in step S316, and object recognition is performed in step S304 using the updated data from step S316. The process flow 300 is repeatedly executed until no objects for processing remain on the pallet.

[0023] Figure 6 shows an alternative process flow 600 for performing object recognition and unloading according to an exemplary embodiment. Process flow 600 is similar to process flow 300 of FIG. 3, except for some additional steps. Process flow 600 begins at step S602 where object data is collected / measured using vision sensor 204. At step S604, using the collected data, object recognition is performed to detect the location and orientation of the object, and an object confidence calculation of the object is performed based on the determined location and object orientation of the object to generate a confidence value. Alternatively, the object confidence can be calculated using other methods such as, but not limited to, the object probability output from the neural network during object recognition. Each confidence value is associated with the corresponding object. Object recognition can be performed using at least one of a learning-based method or a rule-dependent method of object recognition.

[0024] At step S606, a confidence threshold is applied and used when determining whether an object has been detected / recognized. At step S608, a determination is made as to whether the object has been recognized as available for picking based on a comparison of the confidence value with the applied confidence threshold. In particular, the confidence value of the object is compared against the confidence threshold. If the threshold of the object is greater than or equal to the confidence threshold, the object is determined to be detected / recognized. If the object is identified as available for picking, the process then proceeds to step S610 where the recognized object is picked up / gripped and moved. At step S624, a determination is made as to whether the pick-up of the object was successful. If the answer at step S624 is "yes", the process then returns to step S602 and the collection of object data is performed again. If the answer at step S624 is "no", the process proceeds to step S612, which is detailed below.

[0025] However, if no object available for picking is recognized at step S608 (e.g., all objects are below the confidence threshold) valueWhen there is such as having), the object is determined as non-detection / non-recognition, and the process then proceeds to step S612, where it is determined whether a recognition error has occurred. In some exemplary embodiments, recognition error detection is performed by comparing the current object recognition result with the previous object recognition result (past result) to check for no discrepancies. If the answer at step S612 is "no", the process ends.

[0026] If a recognition error is detected at step S612, a remote recovery request is collected for consideration at step S614 and data / sent / issued to the user / operator along with the object. At step S616, the user / operator then selects an area regarding the collected object data where a recognition error has occurred and that indicates the error type of the recognition error. In some exemplary embodiments, area selection and error type display can be done using a GUI. The process then continues to step S618, where the user / operator determines whether there are undetected objects caused by an inappropriate confidence threshold. If the answer to step S618 is "no", the process continues to step S620, where the collected data is then updated based on user input, and object recognition is run again at step S604 using the updated data from step S620.

[0027] If the answer to step S618 is "yes", the process continues to step S622, where an area associated with the undetected object is created, and accordingly, a confidence threshold adjustment is performed. When the confidence threshold is updated, the process then returns to step S606, where the updated confidence threshold is applied. Process flow 600 is repeatedly executed until there are no more objects left on the pallet for processing.

[0028] Figures 7(A) and 7(B) illustrate an exemplary confidence threshold update process according to an exemplary embodiment. As shown in FIG. 7(A), undetected objects and detected objects are identified. An undetected object is an object having a confidence value below the confidence threshold, and a detected object is an object having a confidence value equal to or greater than the confidence threshold. As shown in FIG. 7(B), the user / operator makes an adjustment or update to the confidence threshold, which leads to the detection of an object that was previously undetected.

[0029] FIG. 8 illustrates an alternative process flow 800 for performing object recognition and unloading according to an exemplary embodiment. Process flow 800 utilizes a neural network when performing object recognition. Process flow 800 begins at step S802 where object data is collected / measured using vision sensor 204. At step S804, edge / boundary detection of the object using a red, green, and blue (RGB) image is performed. At step S806, edge / boundary detection of the object using a depth image is performed.

[0030] Object edge / boundary detection through RGB and depth images is performed using a machine learning (ML) algorithm. FIG. 9 shows an exemplary ML model training and testing process according to an exemplary embodiment. During the model training phase, a sensor 902 such as a vision sensor 204 is used to capture an RGB image 904 and a depth image 910 of an object. In some exemplary embodiments, two or more sensors 902 may be utilized when capturing the RGB image 904 and the depth image 910 of the object. For example, a first sensor 902 may be used when capturing the RGB image 904, and a second sensor 902 may be used when capturing the depth image 910. Then, the RGB image 904 and the depth image 910 are used separately to train two ML models. In some exemplary embodiments, a deep neural network (DNN) is implemented as the ML model. As shown in FIG. 9, the RGB image 904 is used to train an RGB DNN 906 to generate a trained RGB DNN 908. The depth image 910 is used to train a depth DNN 912 to generate a trained depth DNN 914. The trained DNNs are utilized to detect the edges / boundaries of the object in the RGB image 904 and the depth image 910. The output of the trained DNNs may include a set of edge / boundary probability images having pixel values representing the likelihood that a pixel belongs to an edge / boundary.

[0031] During the implementation / testing phase, the sensor 902 captures an RGB image 916 and a depth image 920 of the object to be processed. In some exemplary embodiments, two or more sensors 902 may be utilized when capturing the RGB image 916 and the depth image 920 of the object. The RGB image 916 is sent to the trained RGB DNN 908 to generate a set of edge / boundary images indicating an edge probability 918. The depth image 920 is sent to the trained depth DNN 914 to generate a set of edge / boundary images indicating an edge probability 922. Then, weights are assigned to the edge probability 918 and the edge probability 922, and the two are combined to generate a combined edge probability 924.

[0032] Next, the pixel value of the combined edge probability 924 is compared with a threshold value to determine whether the pixel belongs to an edge / boundary. The result is a binary edge map (binary edge image 926) where each pixel is either an edge / boundary or not. Next, the object can be recognized (recognized object 930) using methods such as image segmentation 928 on the detected edge / boundary.

[0033] Referring back to FIG. 8, the edges / boundaries from the RGB image and the depth image are combined using the weighting assignment in step S808. In step S810, the object is recognized using the combined edge / boundary. The process then continues to step S812, where it is determined whether the recognized object is recognized as being available for picking based on the recognition. If the object is recognized as being available for picking, the process then proceeds to step S814, where the recognized object is picked up / gripped and moved. In step S826, it is determined whether the pickup of the object was successful. If the answer in step S826 is "yes", the process then returns to step S802 and the collection of object data is performed again. If the answer in step S826 is "no", the process proceeds to step S816, which is detailed below. object If no object is recognized as being available for picking in step S812, the process continues to step S816, where it is determined whether a recognition error has occurred. If the answer in step S816 is "no", the process ends. If a recognition error is detected in step S816, a remote recovery request is collected for consideration in step S818.

[0034] data / ​It is sent / issued to the user / operator together with the object. In step S820, the user / operator then selects an area regarding the collected object data that indicates that a recognition error has occurred and the error type of the recognition error. In some exemplary embodiments, the area selection and error type display can be made using a GUI.

[0035] In step S822, the user / operator determines the presence of an object that has been incorrectly segmented or a failure of object segmentation (an object that has not been incorrectly segmented). The process then continues to step S824, where an adjustment is made to the weights of the edge probabilities (edge probability 918 and edge probability 922). If there is an incorrect object segmentation, the weight associated with the trained RGB DNN908 (edge probability 918) is increased. Otherwise, the weight associated with the trained RGB DNN908 is decreased. Alternatively, the weight associated with the trained depth DNN914 (edge probability 922) is decreased if an incorrect object segmentation is detected and increased if there is a failure of object segmentation. When step S824 is completed, the process then returns to step S808, where the updated weights are applied. Process flow 800 is repeatedly executed until no objects remain on the pallet for processing.

[0036] FIG. 10 shows an exemplary process flow 1000 for performing object recognition and unloading using a learnable object recognition function according to an exemplary embodiment. Process flow 1000 begins at step S1002 where object data is collected / measured using vision sensor 204. In step S1004, object recognition is performed on the data collected using the learnable object recognition function. As part of the object recognition, the identification of the availability of the object for picking is also performed.

[0037] In step S1006, the object is recognized and executed objectA determination is made as to whether it is available for picking based on the recognition. If the object is recognized as available for picking, the process then proceeds to step S1008, and the recognized object is picked up / gripped and moved. At step S1020, a determination is made as to whether the picking of the object was successful. If the answer at step S1020 is "yes", the process then returns to step S1002, and the collection of object data is performed again. If the answer at step S1020 is "no", the process proceeds to step S1010, which is detailed below.

[0038] If there is no object recognized as available for picking at step S1006, the process continues to step S1010, and a determination is made as to whether a recognition error has occurred. In some exemplary embodiments, the recognition error detection is performed by comparing the current object recognition result with the previous object recognition result (past result) to see if there are no discrepancies. If the answer at step S1010 is "no", the process ends.

[0039] If a recognition error is detected in step S10I0, a remote recovery request is sent / issued to the user / operator in S1012 together with the data / objects collected for consideration. In step S1014, the user / operator then selects an area regarding the collected object data where a recognition error has occurred and which indicates the error type of the recognition error. When user input is received, the collected data is then updated based on the user input in step S1016, and object recognition is re-executed in step S1004 using the updated data. In addition to the re-execution of object recognition in step S1004, the learnable object recognition function can be trained online or offline using the updated data to improve the recognition accuracy of the function in step S1018. For example, the trained DNNs in FIG. 9 (the trained RGB DNN 908 and the trained depth DNN 914) can be further trained using the update results. The process flow 1000 is repeatedly executed until the objects for processing are no longer left on the pallet.

[0040] The above exemplary embodiments can have various benefits and advantages. For example, the exemplary embodiments enable a reduction in the costs and processing time associated with object recognition. In particular, the costs and processing time are reduced as a result of a reduction in the complexity of the annotation. Further, the utilization of the annotation and the updated recognition results for online / offline function training realizes an improved recognition accuracy in object recognition.

[0041] FIG. 12 illustrates an exemplary computing environment having an exemplary computer device suitable for use in some exemplary embodiments. The computer device 1205 in the computing environment 1200 can include one or more processing units, cores, or processors 1210, memory 1215 (e.g., RAM, ROM, and / or the like), internal storage 1220 (e.g., magnetic, optical, solid state storage, and / or organic), and / or an IO interface 1225, any of which can be coupled on a communication mechanism or bus 1230 for communicating information, or can be incorporated into the computer device 1205. The IO interface 1225 is further configured to receive images from a camera or provide images to a projector or display, depending on the desired embodiment.

[0042] Computer device 1205 may be communicatively coupled to input / user interface 1235 and output device / interface 1240. One or both of input / user interface 1235 and output device / interface 1240 may be a wired or wireless interface and may be removable. Input / user interface 1235 may include any physical or virtual device, component, sensor or interface (e.g., buttons, touch screen interface, keyboard, pointing / cursor control, microphone, camera, Braille, motion sensor, accelerometer, optical reader, and / or the like) that can be used to provide input. Output device / interface 1240 may include a display, television, monitor, printer, speaker, Braille, or the like. In some exemplary embodiments, input / user interface 1235 and output device / interface 1240 may be incorporated with or physically coupled to computer device 1205. In other embodiments, other computer devices may function as or provide the functionality of input / user interface 1235 and output device / interface 1240 for computer device 1205.

[0043] Examples of computer device 1205 may include, but are not limited to, highly mobile devices (e.g., smartphones, devices mounted on vehicles and other machines, devices held by a person or animal, and the like), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, and the like), and devices not designed for mobility (e.g., desktop computers, other computers, information kiosks, televisions with one or more processors incorporated therein and / or televisions with one or more processors coupled thereto, radios, and the like).

[0044] Computer device 1205 can be communicatively coupled to external storage 1245 and network 1250 (e.g., via IO interface 1225) for communication with any number of network-connected components, devices, and systems, including one or more computer devices of the same or different configurations. Computer device 1205 or any other connected computer device can function as, provide services as, or be referred to by names such as a server, client, thin server, general-purpose machine, dedicated machine, or others.

[0045] IO interface 1225 can include wired and / or wireless interfaces that use any communication or IO protocol or convention (e.g., Ethernet, 802.11x, Universal System Bus, WiMax, modem, cellular network protocol, and the like), but are not limited to, for information communication to and / or from at least all connected components, devices, and networks in computing environment 1200. Network 1250 can be any network or combination of networks (such as, for example, the Internet, local area network, wide area network, telephone network, cellular network, satellite network, and the like).

[0046] Computer device 1205 can use and / or communicate using computer-usable media or computer-readable media, including transient media and non-transient media. Transient media includes transmission media (e.g., metal cables, optical fibers), signals, carrier waves, and the like. Non-transient media includes magnetic media (e.g., disks and tapes), optical media (e.g., CD ROM, digital video disk, Blu-ray (registered trademark) disk), solid state media (e.g., RAM, ROM, flash memory, solid state storage), and other non-volatile storage or memory.

[0047] Computer device 1205 can be used to implement techniques, methods, applications, processes, or computer-executable instructions in some exemplary computing environments. The computer-executable instructions can be retrieved from a temporary medium, stored in a non-temporary medium, and retrieved therefrom. The executable instructions can be derived from one or more of any programming language, scripting language, and machine language (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, etc.).

[0048] Processor 1210 can execute under any operating system (OS) (not shown) in a native environment or a virtual environment. One or more applications can be deployed that include a logic unit 1260, an application programming interface (API) unit 1265, an input unit 1270, an output unit 1275, and an inter-unit communication mechanism 1295 for communicating with each other, with the OS, and with other applications (not shown). The units and elements described above can vary in design, function, configuration, or implementation, and are not limited to the above description. Processor 1210 can be in the form of a hardware processor such as a central processing unit (CPU), or can be a combination of hardware units and software units.

[0049] In some exemplary embodiments, when information or execution instructions are received by the API unit 1265, it can be communicated to one or more other units (e.g., the logic unit 1260, the input unit 1270, the output unit 1275). In some examples, the logic unit 1260 can be configured to control the information flow between units and, in some of the exemplary embodiments described above, to direct the services provided by the API unit 1265, the input unit 1270, and the output unit 1275. For example, the flow of one or more processes or embodiments can be controlled by the logic unit 1260 alone or in conjunction with the API unit 1265. The input unit 1270 may be configured to obtain inputs for the calculations described in the exemplary embodiments, and the output unit 1275 may be configured to provide outputs based on the calculations described in the exemplary embodiments.

[0050] As shown in FIG. 3, the processor 1210 can be configured to collect object data through a visual sensor. The processor 1210 can also be configured to perform object recognition by determining an object pose and an object position based on the object data, as shown in FIG. 3. The processor 1210 can also be configured to pick up and unload an object using a robotic device when the object among the plurality of objects is recognized, as shown in FIG. 3. The processor 1210 can also be configured to determine the occurrence of an object recognition error, as shown in FIG. 3. The processor 1210 can also be configured to execute a recovery process to handle the object recognition error when the object recognition error is detected, as shown in FIG. 3. The processor 1210 can also be configured to recognize the completion of unloading of a plurality of objects when no object recognition error is detected, as shown in FIG. 3.

[0051] Some portions of the detailed description are presented in terms of algorithms and symbolic representations of operations within a computer. These algorithmic descriptions and symbolic representations are the means used by those skilled in the data processing arts to convey the essence of their innovations to others skilled in the art. An algorithm is a defined sequence of steps leading to a desired end state or result. In an exemplary embodiment, the steps executed require physical manipulation of physical quantities to effect a tangible result.

[0052] Unless otherwise noted, as will be apparent from the description, throughout this specification, descriptions using terms such as "processing", "calculating", "computing", "determining", "displaying", or the like may include actions and processes of a computer system or other information processing device that manipulate data represented as physical (electronic) quantities within the registers and memories of the computer system, and transform them into other data similarly represented as physical quantities within the memories or registers or other information storage devices, transmission devices, or display devices of the computer system.

[0053] Exemplary embodiments may further relate to an apparatus for performing operations herein. The apparatus may be specially constructed for the required purposes, or may include one or more general purpose computers selectively activated or reconfigured by one or more computer programs. Such computer programs may be stored on a computer-readable medium such as a computer-readable storage medium or a computer-readable signal medium. The computer-readable storage medium may include tangible media such as, but not limited to, optical disks, magnetic disks, read-only memory, random access memory, solid state devices and drives, or any other type of tangible or non-transitory media suitable for storing electronic information. The computer-readable signal medium may include media such as carrier waves. The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. The computer program may include a software-only implementation including instructions to perform the operations of the desired embodiment.

[0054] Various general purpose systems may be used with the programs and modules according to the examples herein, or it may be convenient as a result to construct specialized apparatus by performing the steps of the desired method. Further, exemplary embodiments are not described with reference to any particular programming language. It will be understood that various programming languages may be used to implement the teachings of the exemplary embodiments as described herein. The instructions of the programming language may be executed by one or more processing devices such as, for example, a central processing unit (CPU), a processor, or a controller.

[0055] As is known in the art to which the present invention pertains, the operations described above can be executed by hardware, software, or some combination of software and hardware. While various aspects of the exemplary embodiments may be implemented using circuits and logic devices (hardware), other aspects may be implemented using instructions stored on a machine-readable medium (software) that, when executed by a processor, cause the processor to execute a method for implementing the present application. Further, some exemplary embodiments of the present application may be executed by hardware only, while other exemplary embodiments may be executed by software only. Further, the various functions described may be executed by a single unit or may be distributed across multiple components in any number of ways. When executed by software, the method may be executed by a processor, such as a general-purpose computer, based on instructions stored on a computer-readable medium. Optionally, the instructions may be stored on the medium in a compressed and / or encrypted format.

[0056] Furthermore, other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the teachings of the present application. The various aspects and / or components of the exemplary embodiments described may be used singly or in some combination. The specification and exemplary embodiments are intended to be considered as examples only, and the true scope and spirit of the present application are indicated by the following claims.

Description of Reference Numerals

[0057] 208 Memory 206 Processor 204 Visual Sensor 214 Object 212 Pallet 202 Robot Device 1205 Computer Device 1210 Processor 1215 Memory 1220 Internal Storage 1225 IO Interface 1235 Input / User Interface 1240 Output Device / Interface 1245 External Storage 1250 Network 1260 Logical Unit 1265 API Unit 1270 Input Unit 1275 Output Unit

Claims

1. 1. A method for recognizing and unloading a plurality of objects, comprising: collecting, by a processor, object data via a visual sensor; performing object recognition by determining, by the processor, an object pose and an object position based on the object data; If an object of the plurality of objects is recognized and determined to be available for picking, picking up and unloading the object using a robotic device; If no objects are determined to be available for picking, determining, by the processor, that an object recognition error has occurred; if the object recognition error is detected, executing, by the processor, a recovery process to address the object recognition error; and repeatedly performing, by the processor, recognizing completion of unloading of the plurality of objects and performing the steps if the object recognition error is not detected.

2. The method of claim 1 , wherein the object pose comprises an object dimension.

3. The processor performs the recovery process for addressing the object recognition error. sending a recovery request including the object data to a user; receiving a correction response from the user in response to the recovery request, the correction response including at least one selected recognition error area and at least one type of recognition error; updating data for object recognition based on the corrected response and re-performing object recognition; The method of claim 1 , wherein the method is configured to be performed by:

4. The method of claim 3 , wherein object recognition is performed using at least one trainable object recognition function.

5. The method of claim 4 , wherein the training of the at least one learnable object recognition function is performed online or offline.

6. The method of claim 3 , wherein sending the recovery request including the object data to the user includes sending the recovery request to a graphic user interface (GUI) for review by the user.

7. the processor is configured to perform object recognition by performing an object belief operation on the plurality of objects to generate belief values, each belief value being associated with a corresponding object of the plurality of objects; Recognizing an object among the plurality of objects comparing the object confidence value to a confidence threshold; If the certainty value of the object is equal to or greater than the certainty threshold, the object is determined to have been recognized; if the confidence value of the object is less than the confidence threshold, determining that the object has not been recognized; the processor is configured to execute the recovery process to address the object recognition error by further performing a threshold adjustment on the confidence threshold. The method according to claim 3.

8. The processor performs object recognition by: performing object edge detection using red, blue, and green (RGB) images of the plurality of objects to detect a first object boundary; performing object edge detection using the depth images of the plurality of objects to detect a second object boundary; generating a first probability image using the RGB image as an input to a first trained machine learning model; generating a second probability image using the depth image as an input to a second trained machine learning model; assigning weights to the first probability image and the second probability; combining the first probability image and the second probability image based on the weights to generate a binarized edge map of the plurality of objects; by performing object recognition of the plurality of objects using the binarized edge map; The method of claim 3 , wherein the object data includes the RGB image and the depth image.

9. The processor performs the recovery process for addressing the object recognition error. Detecting the presence of a falsely segmented or falsely unsegmented object as a result of the object recognition; adjusting the weights assigned to the first probability image and the second probability image to improve object recognition; The method of claim 8 , further configured to be performed by:

10. The processor calculates the weights assigned to the first probability image and the second probability image by: if a failed object segmentation is detected, increasing a first weight among the weights or decreasing a second weight among the weights, the first weight being associated with the first probability image and the second weight being associated with the second probability image; 10. The method of claim 9, configured to adjust by decreasing the first weight or increasing the second weight if an incorrectly segmented object is detected.

11. 1. A system for recognizing and unloading a plurality of objects, the system comprising: A robotic device; and a processor in communication with the robotic device, the processor: Collect object data through visual sensors, performing object recognition by determining an object pose and an object position based on the object data; If an object of the plurality of objects is recognized and determined to be available for picking, picking up and unloading the object using a robotic device; determining that an object recognition error has occurred if no object is determined to be available for picking; If the object recognition error is detected, executing a recovery process to address the object recognition error; The system is configured to repeat performing the step of recognizing completion of unloading of the plurality of objects if the object recognition error is not detected.

12. The system of claim 11 , wherein the object pose comprises an object dimension.

13. The processor performs the recovery process for addressing the object recognition error. sending a recovery request including the object data to a user; receiving a corrective response from the user in response to the recovery request; receiving the correction response, the correction response including at least one selected recognition error area and at least one type of recognition error; updating data for object recognition based on the corrected response and re-performing object recognition; The system of claim 11 , configured to execute by:

14. The system of claim 13 , wherein object recognition is performed using at least one trainable object recognition function.

15. The system of claim 14 , wherein the training of the at least one learnable object recognition function is performed online or offline.

16. The system of claim 13 , wherein sending the recovery request including the object data to the user includes sending the recovery request to a graphic user interface (GUI) for review by the user.

17. the processor is configured to perform object recognition by performing an object belief operation on the plurality of objects to generate belief values, each belief value being associated with a corresponding object of the plurality of objects; Recognizing an object among the plurality of objects comparing the object confidence value to a confidence threshold; If the certainty value of the object is equal to or greater than the certainty threshold, the object is determined to have been recognized; if the confidence value of the object is less than the confidence threshold, determining that the object has not been recognized; the processor is configured to execute the recovery process to address the object recognition error by further performing a threshold adjustment on the confidence threshold. The system of claim 13.

18. The processor performs object recognition by: performing object edge detection using red, blue, and green (RGB) images of the plurality of objects to detect a first object boundary; performing object edge detection using the depth images of the plurality of objects to detect a second object boundary; generating a first probability image using the RGB image as an input to a first trained machine learning model; generating a second probability image using the depth image as an input to a second trained machine learning model; assigning weights to the first probability image and the second probability; combining the first probability image and the second probability image based on the weights to generate a binarized edge map of the plurality of objects; by performing object recognition of the plurality of objects using the binarized edge map; the object data includes the RGB image and the depth image; The system of claim 13.

19. The processor performs the recovery process for addressing the object recognition error. Detecting the presence of a falsely segmented or falsely unsegmented object as a result of the object recognition; adjusting the weights assigned to the first probability image and the second probability image to improve object recognition; 20. The system of claim 18, further configured to perform by executing:

20. The processor calculates the weights assigned to the first probability image and the second probability image by: if a failed object segmentation is detected, increasing a first weight among the weights or decreasing a second weight among the weights, the first weight being associated with the first probability image and the second weight being associated with the second probability image; and adapted to adjust the first weight by decreasing the first weight or increasing the second weight when an incorrectly segmented object is detected.

20. The system of claim 19.

Citation Information

Patent Citations

  • Image recognition device

    JP2005149074A

  • Mix-size depalletizing

    JP2022045905A

  • ROBOTIC SYSTEM WITH OBJECT UPDATE MECHANISM AND METHOD FOR OPERATING A ROBOTIC SYSTEM - Patent application

    JP2023052476A

  • Control device, control method, and storage medium

    WO2022224447A1