Object Recognition System for Picking Up Items

The system addresses the inefficiencies in warehouse automation by using a sensor with a linear slider to enhance object recognition and picking, optimizing throughput and reducing costs by minimizing sensor movements and collisions.

JP7772883B2Active Publication Date: 2025-11-18HITACHI LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024120841
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-07-26
Filing Date
2024-07-26
Publication Date
2025-11-18
Estimated Expiration
2044-07-26

AI Technical Summary

Technical Problem

Existing warehouse automation systems face challenges in efficiently recognizing and picking objects without prior information, leading to increased costs, reduced throughput, and economic infeasibility due to the use of multiple vision sensors or vertical sliders, which require additional time and resources.

Method used

A system utilizing a sensor coupled with a linear slider to measure object surfaces, calculate confidence levels, distinguish indistinguishable objects, and adjust sensor position to improve recognition and minimize sensor movements, thereby reducing the need for expensive sliders and optimizing throughput.

Benefits of technology

The system effectively recognizes and picks multiple objects with high confidence levels while minimizing sensor movements and costs, enhancing system throughput and reducing the risk of collisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007772883000001
    Figure 0007772883000001
  • Figure 0007772883000002
    Figure 0007772883000002
  • Figure 0007772883000003
    Figure 0007772883000003
Patent Text Reader

Abstract

To provide a method and system for performing object arrangement recognition associated with a plurality of objects.SOLUTION: An object arrangement recognition system 100 includes: a vision sensor for measuring distances between the sensor and the plurality of objects; a linear slider to which the vision sensor is coupled, to linearly move the vision sensor; and a computer including a processor and a memory coupled to the processor. The memory stores instructions executable by the processor, the memory being configured to: in response to the instructions, measure surfaces of the plurality of objects using the sensor; recognize dimensions, positions, and orientations of the plurality of objects based on the measured surfaces to identify recognized objects; calculate a degree of confidence of each of the recognized objects; identify undistinguishable objects from the recognized objects based on the calculated degrees of confidence; calculate an approachable distance for each of the undistinguishable objects; and move the vision sensor towards the plurality of objects by a distance that corresponds to a minimum approachable distance from the calculated approachable distances.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure is generally directed to methods and systems for performing object placement recognition associated with multiple objects. [Background technology]

[0002] The automation of physical tasks through the use of warehouse automation systems is becoming mainstream due to an aging workforce and increasing labor market volatility. In performing warehouse operations, automation systems focus on tasks including depalletizing, debunking, and object picking, where warehouse workers pick products from multiple products arranged on pallets or roll box pallets, in truck bins, or in containers.

[0003] To automate these tasks, autonomous robots with one or more manipulators and one or more vision systems have been proposed and put to practical use. Generally, products and their placement / packaging can vary widely, and the automated system does not receive such prior information. The automated system / autonomous robot recognizes the size, position, and orientation of the products and, based on the recognized information, plans the manipulator's movements to pick up and move the recognized products.

[0004] Related art discloses depalletizer systems that utilize fixedly mounted vision sensors to receive vision data to generate an image and / or 3D model of the top object on the pallet. Figure 1 shows a conventional depalletizer system that utilizes a single vision sensor fixed on the equipment. Such systems require prior information, such as the dimensions of the target object, for the system to function properly, and generating such information requires additional time and resources.

[0005] Related art methods utilize a manipulator to grasp an area of ​​an object's detected top surface that is within a boundary with no confidence, slightly displacing the object to increase the boundary's confidence level. By measuring the resulting displacement, the object can be detected with an accurate size estimate. Figure 2 shows a conventional depalletizer system that utilizes a manipulator. While slight displacement of the object allows for better object size estimation, the process takes additional time to perform and tends to reduce system throughput.

[0006] In related art, multiple vision sensors are fixedly mounted to a depalletizer system to measure the top surface of an object from the perspective of the vision sensors. Figure 3 shows a conventional depalletizer system that utilizes multiple fixedly mounted vision sensors. However, using multiple vision sensors increases the cost of the depalletizer system, making it economically unfeasible.

[0007] Related art discloses a depalletizer system having one or more vision sensors mounted on a manipulator. FIG. 4 illustrates a conventional depalletizer system having a vision sensor mounted on a manipulator. FIG. 5 illustrates a conventional depalletizer system having a vision sensor mounted on a manipulator different from the object motion manipulator. As shown in FIG. 5, there are two manipulators: one with a vision sensor and one with a hand for object manipulation. If the vision sensor and the hand for grasping the object were mounted on the same manipulator as shown in FIG. 4, the manipulator would have to pause temporarily to measure the object. This requires measurements from a high position to recognize the overall object configuration and from a low position to accurately recognize the boundary of the target object, thereby reducing the system's throughput. While the throughput issue can be addressed using two manipulators as shown in FIG. 5, such a system would be too expensive and economically unfeasible.

[0008] In related art, a depalletizer system utilizes a vertical slider to enable vertical movement of a vision sensor. FIG. 6 shows a conventional depalletizer system equipped with a vertical slider. The system can measure the top surface of an object from above at a relatively short distance. However, there are several problems associated with using a vision system that utilizes a vertical slider. FIG. 7 shows the limitations of a depalletizer system that utilizes a vertical slider. As shown in FIG. 7, the approach distance of the vision sensor can be limited to include all objects within its field of view. Therefore, even when a vertical slider is applied, only small displacements of the object are frequently required. Summary of the Invention

[0009] Aspects of the present disclosure include an innovative method for performing object placement recognition associated with multiple objects. The method may include measuring distances between a sensor and the multiple objects using a sensor, moving a linear slider coupled to the sensor, measuring surfaces of the multiple objects using the sensor, recognizing dimensions, positions, and orientations of the multiple objects based on the measured surfaces to identify the recognized objects, calculating a confidence level for each of the recognized objects, distinguishing from the recognized objects indistinguishable objects having a confidence level lower than a preset confidence level threshold based on the calculated confidence levels, calculating an approachable distance for each of the indistinguishable objects, and moving the sensor toward the multiple objects by a distance corresponding to a minimum approachable distance from the calculated approachable distance.

[0010] Aspects of the present disclosure include an innovative system for object placement recognition associated with a plurality of objects. The system may include a sensor that measures distances between the sensor and the plurality of objects, a linear slider to which the sensor is coupled and that moves the sensor linearly, a processor, and a memory coupled to the processor, the memory storing instructions executable by the processor to measure surfaces of the plurality of objects using the sensor, recognize dimensions, positions, and orientations of the plurality of objects based on the measured surfaces, identify the recognized objects, calculate a confidence level for each of the recognized objects, distinguish from the recognized objects indistinguishable objects having a confidence level lower than a preset confidence level threshold based on the calculated confidence levels, calculate an approach distance for each of the indistinguishable objects, and move the sensor toward the plurality of objects by a distance corresponding to a minimum approach distance from the calculated approach distance.

[0011] Aspects of the present disclosure include an innovative system for performing object placement recognition associated with a plurality of objects. The system may include means for measuring distances between a measurement means and the plurality of objects, means for linearly moving the measurement means, means for measuring surfaces of the plurality of objects using the measurement means, means for recognizing dimensions, positions, and orientations of the plurality of objects based on the measured surfaces to identify the recognized objects, means for calculating a confidence level for each of the recognized objects, means for distinguishing from the recognized objects indistinguishable objects having a confidence level lower than a preset confidence level threshold based on the calculated confidence levels, means for calculating an approachable distance for each of the indistinguishable objects, and means for moving the measurement means toward the plurality of objects by a distance corresponding to a minimum approachable distance from the calculated approachable distance.

[0012] A general architecture for implementing various features of the present disclosure will now be described with reference to the drawings. The drawings and associated description are provided to illustrate example implementations of the present disclosure and are not intended to limit the scope of the disclosure. Reference numbers are re-used throughout the drawings to indicate correspondence between referenced elements. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 illustrates a conventional depalletizer system that utilizes a single vision sensor fixed on the equipment. [Figure 2] FIG. 1 illustrates a conventional depalletizer system that utilizes a manipulator. [Figure 3] FIG. 1 illustrates a conventional depalletizer system that utilizes multiple fixedly mounted vision sensors. [Figure 4] FIG. 1 illustrates a conventional depalletizer system having a vision sensor mounted on a manipulator. [Figure 5] FIG. 1 illustrates a conventional depalletizer system having a vision sensor mounted on a manipulator different from the object motion manipulator. [Figure 6] FIG. 1 illustrates a conventional depalletizer system with a vertical slider. [Figure 7] FIG. 1 illustrates limitations of a depalletizer system utilizing a vertical slider. [Figure 8] FIG. 1 illustrates an example of an object placement recognition system 100 according to one exemplary implementation. [Figure 9] FIG. 9 illustrates an example process flow for object recognition using the object placement recognition system 100 of FIG. 8, according to one example implementation. [Figure 10] FIG. 10 shows an illustrative example flow of the object recognition process flow of FIG. 9 according to one example implementation. [Figure 11] FIG. 2 illustrates an example of an object placement recognition system 200 with a manipulation system, according to one exemplary implementation. [Figure 12] FIG. 12 illustrates an example process flow for object recognition using the object placement recognition system 200 of FIG. 11, according to one example implementation. [Figure 13] FIG. 13 shows an illustrative example flow diagram of the object recognition process flow of FIG. 12 according to one example implementation. [Figure 14]FIG. 1 illustrates an example process flow for object recognition using an object placement recognition system involving objects of various loading heights, according to one illustrative implementation. [Figure 15] FIG. 15 shows an illustrative example flow diagram of the object recognition process flow of FIG. 14 according to one example implementation. [Figure 16] FIG. 16 illustrates an example of an object placement recognition system 1600 including a depth dimension sensor, according to one illustrative implementation. [Figure 17] FIG. 17 illustrates an example process flow for object recognition using the object placement recognition system 1600 of FIG. 16, according to one example implementation. [Figure 18] FIG. 18 shows an illustrative example flow diagram of the object recognition process flow of FIG. 17 according to one example implementation. [Figure 19] FIG. 1 illustrates an example computing environment having example computing devices suitable for use in some example implementations. DETAILED DESCRIPTION OF THE INVENTION

[0014] The following detailed description provides details of the drawings and exemplary implementations of the present application. Reference numbers and descriptions of elements that are duplicated between drawings are omitted for clarity. Terms used throughout the description are provided by way of example and are not intended to be limiting. For example, use of the term "automatic" can include fully automatic implementations or semi-automatic implementations that require user or administrator control over certain aspects of the implementation, depending on the desired implementation of those skilled in the art practicing the implementations of the present application. Selection can be performed by a user through a user interface or other input means, or can be achieved through a desired algorithm. The exemplary implementations as described herein can be utilized alone or in combination, and the functionality of the exemplary implementations can be achieved through any means depending on the desired implementation.

[0015] An example implementation reduces the frequency of small displacements of ambiguous objects and, by using a vertical slider, reduces the amount of vertical movement of sensors in the system, while allowing the system to pick up objects from a variety of object types placed on pallets or in containers.

[0016] FIG. 8 illustrates an example of an object placement recognition system 100 for object placement recognition associated with multiple objects, according to one exemplary implementation. As shown in FIG. 8, the system may include a computer 102, a vision sensor 104, and a linear slider 106. The computer 102 may include a processor 108 and a memory 110. The vision sensor 104 may be one of a time-of-flight (TOF) camera, a stereo camera, etc., that measures the distance to objects within a field of view. The linear slider 106 can move the vision sensor 104 vertically / linearly to move the vision sensor 104 toward or away from the objects on the pallet.

[0017] Figure 9 shows an example process flow for object recognition using the object placement recognition system 100 of Figure 8, according to one example implementation. Figure 10 shows an illustrative example flow for the object recognition process flow of Figure 9, according to one example implementation.

[0018] In S1001, the object placement recognition system 100 measures the surface of an object using the vision sensor 104. The second diagram in FIG. 10 shows the measurement results of S1001, where object heights are represented by various shadings / patterns. For example, an object measured to have a low height will have a lighter shading or dotted pattern in contrast to an object measured to have a high height. In S1002, the system uses the measurement results to recognize the size, position, and orientation of the object.

[0019] After performing the recognition process, in S1003, the system then calculates the confidence of each recognized result. If the change in curvature of the estimated surface boundary is relatively large and clear, the confidence is set to high; otherwise, the confidence is set to low. In addition, if there are several changes in curvature of the lines on the surface, the confidence is set to high; otherwise, the confidence is set to low. As shown in the third diagram of FIG. 10, recognized objects with a confidence above a preset confidence threshold are outlined with a black frame. Recognized objects with high confidence are called "distinguishable objects," and objects with low confidence are called "indistinguishable objects."

[0020] The system attempts to move the vision sensor 104 toward the object in order to measure changes in the curvature of the object's surface more clearly from a position closer to the object. In other words, the system attempts to increase the confidence of each recognized object by moving the vision sensor 104 closer to the object.

[0021] In S1004, the system also calculates an approachable distance for each indistinguishable object based on the angle of the field of view of the vision sensor 104 and its recognition result. The approachable distance of an object is the maximum movement / distance of the vision sensor 104 at which the field of view can still capture the entire object. A value is generated for each indistinguishable object as shown in the fourth diagram of FIG. 10. The minimum value of the approachable distance is 10 centimeters, as shown in the diagram.

[0022] In S1005, the system moves the vision sensor 104 by the minimum approachable distance using the linear slider 106. After moving the vision sensor 104 by the minimum approachable distance, all objects are still within the field of view of the vision sensor 104. The system can then attempt to recognize the indistinguishable objects again at a closer position than in the previous recognition process.

[0023] The system can move the vision sensor 104 as close as possible to each placed object by gradually moving it in small increments, thus eliminating the need for expensive sliders to rapidly slide the sensor.

[0024] 11 shows an example of an object placement recognition system 200 with a manipulation system according to one exemplary implementation. As with FIG. 8, a depalletizing operation is selected as an example. In addition to the components shown in FIG. 8, the object placement recognition system 200 includes a manipulator 202 that can grasp and move an object. The processor 108 can be used to control the manipulator 202.

[0025] Figure 12 shows an example process flow for object recognition using the object placement recognition system 200 of Figure 11, according to one example implementation. Figure 13 shows an illustrative example flow for the object recognition process flow of Figure 12, according to one example implementation.

[0026] As shown in Figure 12, in S1011, the system performs processes S1001 and S1002 of Figure 10 to measure the surface of the object and recognize the size, position, and orientation of the object. In S1012, it is determined whether there are any more recognized objects based on the result of S1002. If the answer is no, the process ends. Otherwise, the process continues to S1013, where S1003 of Figure 10 is performed to further input a confidence level for each recognized object. In S1014, the system calculates the proximity / distance from the vision sensor 104 to each recognized object.

[0027] In S1015, the system selects a manipulable object from the recognized objects. First, the system determines the object closest to the vision sensor based on the calculated proximity / distance. The system then selects an object with a proximity that can be included in the same level as the closest object. The manipulator cannot handle objects that are relatively far from the manipulator and the vision sensor because the manipulator may collide with closer objects. Therefore, the system considers the selected object as a manipulable object. For example, if the difference in proximity between one object and the closest object is less than a preset distance threshold, the system can consider the object's proximity to be included in the same level as the closest object. As shown in FIG. 13, the object with the darkest shadow is selected as a manipulable object.

[0028] At S1016, a determination is made as to whether any manipulable and distinguishable objects remain. If there are any manipulable (selected) and distinguishable objects remaining, the system picks / grabs and places these objects at S1017. After the picking operation, the system repeats the process from S1011 until no manipulable and distinguishable objects remain.

[0029] If there are no manipulable (selected) distinguishable objects, the system then proceeds to S1018, where S1004 is executed to calculate the approach distance for each manipulable and indistinguishable object. At S1019, a determination is made as to whether there are any manipulable and indistinguishable objects that the vision sensor 104 cannot approach any further. If the approach distances for all manipulable and indistinguishable objects are greater than zero, in other words, if the vision sensor 104 can approach all of the manipulable and indistinguishable objects, the process continues to S1021, where S1005 is executed to move the vision sensor 104 by the minimum approach distance. After completing step S1021, the process returns to S1011.

[0030] On the other hand, if there is at least one manipulable (selected) indistinguishable object that the vision sensor 104 cannot access, the system then slightly displaces at least one object in S1020. The manipulator 202 grasps an area near a corner of the target object, lifts it slightly, and displaces the hand in a direction where there are no other objects. This movement can help distinguish boundaries between objects. After the slight displacement movement, the process returns to S1011. In FIG. 13, the bottom left diagram shows step S1020 in progress. By grasping an area near the corner (x) and slightly displacing the hand, the system can then reliably recognize that there are two small objects.

[0031] As shown in Figure 13, the system picks up and places the two small objects as they become manipulable and distinguishable, and continues looping this sequence.

[0032] Next, we describe in detail the situation where the approachable distance of a low-height object (unselected or non-operable object) is shorter than that of a high-height object (selected or operable object). In this situation, it may be dangerous for the system to operate the low-height object first, since the manipulator may collide with the high-height object.

[0033] As described above, the system adjusts the height of the vision sensor 104 relative to the indistinguishable object at the minimum approach distance. If the vision sensor 104 cannot get any closer to the focused object, or if the system cannot operate the focused object because there are other objects at a higher height level than the focused object, the system cannot make the focused object distinguishable and therefore cannot complete its picking task.

[0034] Figure 14 shows an example process flow for object recognition using an object placement recognition system involving objects of various loading heights, according to one example implementation. Figure 15 shows an illustrative example flow for the object recognition process flow of Figure 14, according to one example implementation.

[0035] Initially, the system starts without any stored information and proceeds to steps S1101-S1105. Because there is no stored information, "No" is selected in both steps S1102 and S1105. In the upper left and upper center diagrams of FIG. 15, objects at high height levels remain in the center of the range image, while objects at medium or low height levels are located near the outer frame of the image. As shown in FIG. 15, the approachable distances of objects located near the outer frame of the range image tend to be shorter than those in the center of the image, so the minimum approachable distances of indistinguishable objects that are not operable (unselected, medium or low height levels) are shorter than the minimum approachable distances of indistinguishable objects that are operable (selected, high height levels). Returning to FIG. 14, in S1108, "Yes" is selected based on the comparison of the minimum approachable distances, and the process proceeds to step S1110.

[0036] In S1110, the system sets the minimum approach distance of the manipulable and indistinguishable objects as the next approach distance of the vision sensor 104. In S1111, the system then selects an inoperable (unselected) object that has an approach distance shorter than the set approach distance, and stores the recognition result (size, position, orientation, and confidence) of the selected object in S1112. The system then moves the vision sensor 104 toward the object using the set approach distance in S1113.

[0037] After moving the vision sensor 104, the system repeats this sequence, starting with measuring and recognizing the object in S1101. As shown in the top right diagram of FIG. 15, the captured range image focuses on the central area of ​​the object on the pallet. As a result of the sensor movement, all remaining manipulable (selected) objects at higher height levels are recognized with high confidence. The system can then pick up and place these manipulable and distinguishable objects. As a result of the picking operation, the heights of all manipulable (selected) objects fall within the intermediate height level. Because the heights of all stored objects, like the manipulable (selected) objects, fall within the intermediate height level, all stored objects are now manipulable. After receiving this confirmation in S1102, the manipulator 202 then picks up and places the stored manipulable and distinguishable objects in S1103. After the picking operation, the system removes the stored information for the picked and moved objects.

[0038] Now with the stored information from the first iteration, the system calculates the proximity from the vision sensor 104 to each stored indistinguishable object in S1104 and checks in S1105 whether there are any more stored indistinguishable objects that have the same or closer level of proximity to the vision sensor 104 than the closest operable (selected) object. If the condition is met in S1105, the system then moves the vision sensor 104 away from the object in S1106 to a position where the vision sensor 104 can measure the stored indistinguishable object. In S1107, the system then clears all information about the stored indistinguishable objects.

[0039] Next, a method for simultaneously picking up one of the objects and moving the vision sensor 104 closer to the object to improve the throughput of the picking task is described. As a result of the movement, a problem occurs in that the vision sensor 104 cannot measure the surface of an object behind the picked object. If the system can predict the size, position, and orientation of the hidden object, the system can determine whether it is possible to measure the surface of an object behind the picked object. To generate the prediction, the system would need to know the depth dimension of the picked object.

[0040] The system assumes that no prior object information has been received. Additionally, the vision sensor 104 generally has no way of obtaining any information to recognize the depth dimension of an object before it is picked up and moved away.

[0041] On the other hand, the depth dimension of the picked object is necessary for the manipulator 202 to safely place the object. FIG. 16 shows an example of an object placement recognition system 1600 including a depth dimension sensor according to one exemplary implementation. As shown in FIG. 16, the depth dimension sensor 1602 is included as part of the object placement recognition system 1600 to measure the depth dimension of the picked object. The manipulator 202 carries the picked object to a position where the depth dimension sensor 1602 can measure the picked object. In some exemplary implementations, the depth dimension sensor 1602 is fixedly attached to one of the surrounding devices. In some exemplary implementations, a time-of-flight camera can be implemented as the depth dimension sensor 1602. In some exemplary implementations, an obstacle sensor using a linear laser can be selected as the depth dimension sensor 1602, using information on the position of the hand of the manipulator 202. When the depth dimension sensor 1602 detects the underside of the side of the picked object, if the height of both the depth dimension sensor 1602 and the hand of the manipulator 202 are known, the system can calculate the depth dimension of the picked object.

[0042] Figure 17 shows an example process flow for object recognition using the object placement recognition system 1600 of Figure 16, according to one example implementation. Figure 18 shows an illustrative example flow for the object recognition process flow of Figure 17, according to one example implementation.

[0043] At S1201, the system initiates and executes processes S1011-S1015. If there are two or more manipulable (selected) and distinguishable objects, the system determines and selects one of the manipulable (selected) and distinguishable objects as the target object. At S1202, a determination is made as to whether there are two or more manipulable (selected) and distinguishable objects. If the answer is yes, the process continues at S1203. Otherwise, the process continues at S1204. At S1203, the manipulator 202 picks up and removes the other manipulable (selected) and distinguishable objects from the pallet, leaving the target object untouched on the pallet. As shown in FIG. 18, object (p) is determined as the target object.

[0044] After repeating steps S1201 and S1202, the system verifies in S1204 that the number of operable (selected) and distinguishable objects is less than two and calculates the approach distance of indistinguishable objects. Next, in S1205, the system determines whether there are any more operable and indistinguishable objects that the vision sensor 104 cannot approach any further. If the answer is yes, the process continues to S1206, where the system determines and executes a slight object displacement. As shown in FIG. 18, the system determines to slightly displace object (q) and executes the slight displacement (as described as S1206). On the other hand, if the answer is no in S1205, the process continues to S1207.

[0045] After repeating steps S1201, S1202, and S1204, the system determines in step S1205 not to perform a small displacement movement, and determines in S1207 whether only one operable (selected) and distinguishable object remains on the pallet. If it is determined that only one operable and distinguishable object (target object) remains on the pallet, the process continues to S1209. Otherwise, the process continues to S1208, where step S1021 is performed. As shown in the upper right diagram of FIG. 18, the target object (p) remains on the pallet. The minimum approachable distance is 6 centimeters, which corresponds to object (r). In S1209, the system simultaneously picks up object (p) and moves the vision sensor 104 by the minimum approachable distance (6 centimeters), as shown in the upper right diagram of FIG. 18.

[0046] After performing process S1209, the system also sequentially performs steps S1210, S1211, S1212, and S1213, which are described in more detail below. The process then returns to S1201, where the system recognizes two operable (selected) and distinguishable objects (s) and (r), as shown in the bottom right diagram of Figure 18. The system determines that only object (s) will remain and object (r) will be picked up and placed.

[0047] After repeating steps S1201, S1202, and S1204, the system determines in step S1205 that the minimum approach distance of the indistinguishable objects is not zero, so a slight object displacement movement will not be performed. As shown in FIG. 18, the minimum approach distance is 8 centimeters, which corresponds to object (t). The system also checks in step S1207 whether only one operable (selected) distinguishable object remains on the pallet. Next, as shown in the top center diagram of FIG. 18, the system simultaneously picks up object (s) and moves vision sensor 104 by the minimum approach distance (8 centimeters) corresponding to object (t) in step S1209.

[0048] As shown in the lower left diagram of FIG. 18 , the vertical (height) dimension of object(s) is shorter than that of other manipulable (selected) objects, so a portion of the object behind object(s) is outside the field of view of the vision sensor. In S1210, the system measures the depth dimension of the last-picked object through the depth dimension sensor 1602 before placing the last-picked object being held by the manipulator 202. In S1211, the last-picked object is placed. Then, in S1212, the system predicts / estimates the size, position, and orientation of the object behind the last-picked object based on the depth dimension of the last-picked object. For example, the system assumes / estimates that the size, 2D position, and orientation are the same between the last-picked object and the surface of the hidden object. The vertical position of the surface of the hidden object is the sum of the vertical position of the last-picked object and the measured depth dimension of the last-picked object.

[0049] After the prediction, the system also estimates in S1213 whether the vision sensor 104 can measure the object located behind the last-picked object based on the prediction result and the field of view of the vision sensor 104. If the answer is yes in S1213, the process returns to S1201 and repeats the steps. Otherwise, the process continues to step S1214. If the system determines in S1214 that the vision sensor 104 cannot measure the object located behind the last-picked object, the system moves the vision sensor 104 away from the object to a position where the vision sensor 104 can measure the object located behind the last-picked object.

[0050] Additionally, if an assumption can be made regarding the depth dimension of the processed object, the system avoids a situation in which the vision sensor 104 is unable to measure the surface of an object located behind the picked object as a result of moving the vision sensor 104 closer to the object. In some example implementations, a minimum depth dimension of the processed object can be preset. In this situation, after performing process S1207 of FIG. 17, the system uses the preset depth dimension to predict the approach distance at which the object will be located behind the last of the operative (selected) distinguishable objects. The system then selects the smaller of the predicted approach distance corresponding to the recognized indistinguishable object and the minimum of the calculated approach distances. After determining the selected approach distance (which is the minimum approach distance of both the recognized indistinguishable object and the predicted object), the system then proceeds to step S1209 of FIG. 17. Using the selected approach distance can bypass steps S1210, S1212, S1213, and S1214 of FIG. 17.

[0051] By moving the vision sensor 104 towards the object, the system attempts to better measure changes in the curvature of the object's surface and increase the confidence of each recognized object. If the object's confidence is equal to or higher than a preset threshold, the system can directly pick up the object without displacing it slightly.

[0052] There is a relationship between the confidence of a recognized object, the width of the gap between the recognized object and surrounding objects, and the distance between the vision sensor 104 and the recognized object. If the distance between the vision sensor 104 and the recognized object is relatively long and the gap is relatively small, the vision sensor 104 will therefore not be able to clearly measure the gap, which will lead to a decrease in confidence. If the vision sensor 104 gradually approaches the recognized object, the gap will become more visible on the measured distance image, which will lead to an increase in confidence. On the other hand, if the gap is relatively large, the vision sensor 104 will be able to measure the gap even at a distance, which will lead to an additional input of a high confidence value.

[0053] Therefore, by predetermining the associations of confidence, gap, and distance from many combinations of sample data, the amount of sensor movement to bring the confidence to equal or exceed a preset confidence threshold can be generated from the association, current distance, and current confidence corresponding to the recognized object.

[0054] The system predicts the minimum required movement of the vision sensor 104 corresponding to each indistinguishable object. The system then also determines the next movement amount, which is the smaller of the minimum approach distance and the minimum expected required movement amount, and moves the vision sensor 104 by the determined movement amount.

[0055] The above-described exemplary implementation may have various benefits and advantages. For example, the exemplary implementation can reliably recognize the placement of multiple placed objects and sequentially pick up each object while keeping system costs low and sensor movements to a minimum. Expensive sliders for high-speed sensor movement are not required. Object collisions are avoided by performing manipulator movements for slight displacement or object grasping while taking into account previously recognized objects that are out of the sensor's field of view. Additionally, unnecessary sensor movements can be skipped when performing sensor movement and picking movements, improving throughput.

[0056] 19 illustrates an example computing environment having example computer devices suitable for use in some illustrative implementations. The computing device 1905 of the computing environment 1900 can include one or more processing units, cores, or processors 1910, memory 1915 (e.g., RAM, ROM, and / or other), internal storage 1920 (e.g., magnetic, optical, solid-state storage, and / or organic), and / or I / O interface 1925, any of which can be coupled by a communication mechanism or bus 1930 for communicating information or embedded in the computing device 1905. The I / O interface 1925 is also configured to receive images from a camera or provide images to a projector or display, depending on the desired implementation.

[0057] The computing device 1905 may be communicatively coupled to an input / user interface 1935 and an output device / interface 1940. Either or both of the input / user interface 1935 and the output device / interface 1940 may be wired or wireless interfaces and may be detachable. The input / user interface 1935 may include any device, component, sensor, or interface, physical or virtual, that can be used to provide input (e.g., buttons, touchscreen interface, keyboard, pointing / cursor control, microphone, camera, Braille, motion sensor, accelerometer, optical reader, and / or the like). The output device / interface 1940 may include a display, television, monitor, printer, speaker, Braille, etc. In some example implementations, the input / user interface 1935 and the output device / interface 1940 may be embedded in or physically coupled to the computing device 1905. In other example implementations, other computing devices may function as or provide the functionality of the input / user interface 1935 and the output device / interface 1940 of the computing device 1905 .

[0058] Examples of computing devices 1905 may include, but are not limited to, highly mobile devices (e.g., smartphones, devices in automobiles or other machines, devices carried by people and animals, etc.), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, etc.), and devices not designed for mobility (e.g., desktop computers, other computers, information kiosks, televisions with one or more embedded processors and / or televisions, radios, etc.).

[0059] Computing device 1905 may be communicatively coupled (e.g., via I / O interface 1925) to external storage 1945 and a network 1950 for communicating with any number of networked components, devices, and systems, including one or more computing devices of the same or different configurations. Computing device 1905, or any connected computing device, may function as, provide services for, or be referred to as a server, client, thin server, general-purpose machine, special-purpose machine, or another level.

[0060] I / O interface 1925 can include, but is not limited to, wired and / or wireless interfaces using any communication or I / O protocol or standard (e.g., Ethernet, 802.11x, Universal System Bus, WiMax, modem, cellular network protocols, etc.) for communicating information to and from at least all connected components, devices, and networks of computing environment 1900. Network 1950 can be any network or combination of networks (e.g., the Internet, a local area network, a wide area network, a telephone network, a cellular network, a satellite network, etc.).

[0061] The computing device 1905 can use and / or communicate using computer-usable or computer-readable media, including transitory and non-transitory media. Transitory media include transmission media (e.g., metallic cables, fiber optics), signals, carried waves, etc. Non-transitory media include magnetic media (e.g., disks and tape), optical media (e.g., CD ROM, digital video disks, Blu-ray disks), solid media (e.g., RAM, ROM, flash memory, solid-state storage), and other non-volatile storage or memory.

[0062] The computing device 1905 can be used to implement techniques, methods, applications, processes, or computer-executable instructions in some example computing environments. The computer-executable instructions can be retrieved from transitory media and can be stored on and retrieved from non-transitory media. The executable instructions can be in one or more of any programming, scripting, and machine language (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, etc.).

[0063] The processor 1910 can run under any operating system (OS) (not shown) in a native or virtual environment. One or more applications can be deployed, including a logic unit 1960, an application programming interface (API) unit 1965, an input unit 1970, an output unit 1975, and an inter-unit communication mechanism 1995 through which different units communicate with each other, the OS, and other applications (not shown). The described units and elements may vary in design, function, configuration, or implementation and are not limited to the provided description. The processor 1910 can be in the form of a hardware processor, such as a central processing unit (CPU), or a combination of hardware and software units.

[0064] In some example implementations, once information or instructions to execute are received by API unit 1965, they may be communicated to one or more other units (e.g., logic unit 1960, input unit 1970, output unit 1975). In some examples, logic unit 1960 may be configured to control the flow of information between units and, in some example implementations described above, to direct the services provided by API unit 1965, input unit 1970, and output unit 1975. For example, the flow of one or more processes or implementations may be controlled solely by logic unit 1960 or in combination with API unit 1965. Input unit 1970 may be configured to obtain inputs for computations described in example implementations, and output unit 1975 may be configured to provide outputs based on computations described in example implementations.

[0065] The processor 1910 may be configured to measure the surfaces of multiple objects using a sensor as shown in FIG. 9 . The processor 1910 may also be configured to recognize dimensions, positions, and orientations of the multiple objects based on the measured surfaces to identify the recognized objects, as shown in FIG. 9 . The processor 1910 may also be configured to calculate a confidence level for each of the recognized objects, as shown in FIG. 9 . The processor 1910 may also be configured to distinguish indistinguishable objects from the recognized objects based on the calculated confidence levels, as shown in FIG. 9 , where the calculated confidence levels of the indistinguishable objects are lower than a preset confidence level threshold. The processor 1910 may also be configured to calculate an approachable distance for each of the indistinguishable objects, as shown in FIG. 9 . The processor 1910 may also be configured to move the sensor toward the multiple objects by a distance corresponding to a minimum approachable distance from the calculated approachable distance, as shown in FIG. 9 .

[0066] The processor 1910 may also be configured to calculate a distance from the sensor to each of the recognized objects, as shown in FIG. 12 . The processor 1910 may also be configured to select, from the recognized objects, objects for which the distance between the selected object and the object closest to the sensor is respectively less than a preset distance threshold, as shown in FIG. 12 . The processor 1910 may also be configured to grasp, by the manipulator, one of the selected distinguishable objects, which is a subset of the selected objects, and the calculated confidence of the selected distinguishable object is equal to or higher than a preset confidence threshold, as shown in FIG. 12 . The processor 1910 may also be configured to move, by the manipulator, the grasped object to a destination area, as shown in FIG. 12 . The processor 1910 may also be configured to displace, by the manipulator, one of the selected indistinguishable objects, which is a subset of the selected objects, and which is the indistinguishable objects, as shown in FIG. 12 .

[0067] The processor 1910 may also be configured to move the sensor toward the multiple objects by a distance corresponding to the minimum of all the approach distances of the selected indistinguishable objects if the minimum of all the approach distances of the selected indistinguishable objects is longer than the minimum of all the approach distances of unselected indistinguishable objects that are not part of the selected object but belong to the indistinguishable object, as shown in FIG. 14 .

[0068] The processor 1910 may also be configured to select an object that belongs to the unselected objects and has an approachable distance that is less than the minimum of all the approachable distances of the selected objects, as shown in Figure 14. The processor 1910 may also be configured to store the dimensions, position, and orientation of the selected object, as shown in Figure 14.

[0069] The processor 1910 may also be configured to calculate a distance from the sensor to each stored distinguishable object that is part of the selected object and belongs to the distinguishable object, as shown in Figure 14. The processor 1910 may also be configured to use the distance of each stored distinguishable object from the sensor to determine whether at least one of the stored distinguishable objects is manipulable, as shown in Figure 14. The processor 1910 may also be configured to grasp and move at least one of the stored distinguishable objects that is manipulable, as shown in Figure 14.

[0070] The processor 1910 may also be configured to calculate a distance from the sensor to each stored indistinguishable object that is part of the selected object and belongs to the indistinguishable object, as shown in Figure 14. The processor 1910 may also be configured to move the sensor away from the objects to a position where the sensor can measure at least one of the stored indistinguishable objects if the difference in distance between the nearest selected object to the sensor and the nearest stored indistinguishable object to the sensor is less than a second preset distance threshold, or if the distance from the nearest selected object to the sensor is greater than the distance from the nearest stored indistinguishable object to the sensor, as shown in Figure 14.

[0071] The processor 1910 may also be configured to move the sensor away from the plurality of objects to a position where the sensor can measure all of the selected objects or all of the stored indistinguishable objects, as shown in Figure 14. The processor 1910 may also be configured to clear the stored indistinguishable object information after moving the sensor away from the plurality of objects, as shown in Figure 14.

[0072] The processor 1910 may also be configured to continue grasping and moving the selected distinguishable objects to the destination area until all of the selected distinguishable objects have been moved, for a number of selected distinguishable objects greater than or equal to two, as shown in Figure 17. The processor 1910 may also be configured to simultaneously perform the movement of the sensor and the grasping of the last of the selected distinguishable objects, as shown in Figure 17.

[0073] The processor 1910 may also be configured to move the last-grasped selected distinguishable object to enable the depth dimension sensor to measure the depth dimension of the last-grasped selected distinguishable object, as shown in FIG. 17. The processor 1910 may also be configured to measure the depth dimension of the last-grasped selected distinguishable object, as shown in FIG. 17. The processor 1910 may also be configured to use the measured depth dimension to estimate the size, position, and orientation of the top surface of an object located behind the last-grasped selected distinguishable object before moving the last-grasped selected distinguishable object, as shown in FIG. 17. The processor 1910 may also be configured to determine whether the sensor can measure the top surface of the estimated object after moving the sensor and grasping the last of the selected distinguishable objects to move the last of the selected distinguishable objects away from the field of view of the sensor, as shown in FIG. The processor 1910 may also be configured to move the sensor away from the plurality of objects to a position where the sensor can measure the estimated top surfaces of the objects, as shown in FIG. 17 .

[0074] The processor 1910 may also be configured to grasp and continue moving the selected distinguishable objects to the destination area until all of the selected distinguishable objects have been moved, for a number of selected distinguishable objects greater than or equal to two, as shown in FIG. 17. The processor 1910 may also be configured to estimate an approachable distance of an object located behind the last of the selected distinguishable objects using a preset depth dimension, as shown in FIG. 17. The processor 1910 may also be configured to select the smaller of the estimated approachable distance and a minimum approachable distance of the calculated approachable distances as the selected distance, as shown in FIG. 17. The processor 1910 may also be configured to simultaneously move the sensor at the selected distance and grasp the last of the selected distinguishable objects, as shown in FIG.

[0075] The processor 1910 may also be configured to estimate a required movement of the sensor to each of the indistinguishable objects, where the required movement is a distance the sensor can move that increases the confidence of the indistinguishable object to or above a preset confidence threshold, as shown in Figure 17. The processor 1910 may also be configured to move the sensor toward the objects by a distance corresponding to the smaller of the calculated minimum approach distance and a minimum required movement, as shown in Figure 17.

[0076] Some portions of the detailed descriptions are presented in terms of algorithms and symbolic representations of operations within a computer. These algorithmic descriptions and symbolic representations are the means used by those skilled in the data processing arts to convey the essence of their innovations to others skilled in the art. An algorithm is a prescribed sequence of steps leading to a desired end state or result. In exemplary implementations, the performed steps require physical manipulations of tangible quantities to achieve a tangible result.

[0077] Unless otherwise specifically indicated, as will be apparent from the discussion, it will be recognized that throughout the description, discussion utilizing terms such as "processing," "computing," "calculating," "determining," "displaying," and the like can include operations and processes of a computer system or other information processing device that manipulate and convert data represented as physical (electronic) quantities in the computer system's registers and memory into other data that is similarly represented as physical quantities in the computer system's memory or registers or other information storage, transmission, or display device.

[0078] Example implementations may also relate to apparatuses for performing the operations herein. This apparatus may be specially constructed for the required purposes, or may include one or more general-purpose computers selectively activated or reconfigured by one or more computer programs. Such computer programs may be stored on computer-readable media, such as computer-readable storage media or computer-readable signal media. Computer-readable storage media may include tangible media, such as, but not limited to, optical disks, magnetic disks, read-only memory, random-access memory, solid-state devices and drives, or any other type of tangible or non-transitory medium suitable for storing electronic information. Computer-readable signal media may include media such as carried waves. The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. A computer program may include a pure software implementation containing instructions that perform the operations of a desired implementation.

[0079] Various general-purpose systems may be used with the programs and modules according to the examples herein, or it may prove convenient to construct more specialized apparatus to perform the desired method steps. Additionally, the example implementations are not described with reference to any particular programming language. It will be appreciated that various programming languages ​​may be used to implement the teachings of the implementations as described herein. Instructions in the programming language may be executed by one or more processing devices, such as a central processing unit (CPU), processor, or controller.

[0080] As is known in the art, the operations described above can be implemented by hardware, software, or some combination of software and hardware. Various aspects of the implementations may be implemented using circuits and logic devices (hardware), while other aspects may be implemented using instructions stored on a machine-readable medium (software) that, when executed by a processor, cause the processor to perform methods that implement the examples of the present application. Furthermore, some implementations of the present application may be implemented solely by hardware, while other implementations may be implemented solely by software. Furthermore, the various functions described may be implemented in a single unit or may be spread across multiple components in various ways. When implemented by software, the methods may be executed by a processor, such as a general-purpose computer, based on instructions stored on a computer-readable medium. If desired, the instructions may be stored on the medium in compressed and / or encrypted format.

[0081] Additionally, other implementations of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the teachings of the present application. Various aspects and / or components of the described exemplary implementations may be used alone or in any combination. It is intended that the specification and exemplary implementations be considered as examples only, with the true scope and spirit of the present application being indicated by the following claims.

Claims

1. 1. A system for object placement recognition associated with a plurality of objects, comprising: a sensor for measuring distances between the sensor and the plurality of objects; a linear slider to which the sensor is coupled and which moves the sensor linearly; a processor; a memory coupled to the processor, the memory storing instructions executable by the processor, the instructions causing measuring the surfaces of the plurality of objects using the sensor; Recognizing the dimensions, positions, and orientations of the plurality of objects based on the measured surfaces, and identifying the recognized objects; calculating a confidence score for each of the recognized objects; Identifying indistinguishable objects from the recognized objects based on the calculated confidence, wherein the calculated confidence of the indistinguishable objects is lower than a preset confidence threshold; Calculating an approachable distance for each of the indistinguishable objects, which is a maximum movement amount of the sensor that can capture the entire object; moving the sensor toward the plurality of objects by a distance corresponding to the smallest approachable distance among the calculated approachable distances; system.

2. further comprising a manipulator for grasping and moving the plurality of objects; The memory further stores instructions executable by the processor, the instructions causing: calculating a distance from the sensor to each of the recognized objects; From the recognized objects, an object is selected in which the distance between the selected object and the object closest to the sensor is shorter than a preset distance threshold. grasping, with the manipulator, one of the selected subset of objects, the selected distinguishable object having a calculated confidence level equal to or higher than the preset confidence level threshold; The manipulator moves the grasped object to a destination area; displacing, by the manipulator, one of the selected indistinguishable objects, the selected indistinguishable object being a subset of the selected objects; The system of claim 1 .

3. The memory further stores instructions executable by the processor, the instructions causing: if the minimum value of all approach distances of the selected indistinguishable objects is longer than the minimum value of all approach distances of unselected indistinguishable objects belonging to the indistinguishable objects, moving the sensor toward the plurality of objects by a distance corresponding to the minimum value of all approach distances of the selected indistinguishable objects; The system of claim 2 .

4. The memory further stores instructions executable by the processor, the instructions causing: Selecting an object belonging to the unselected objects and having an approachable distance that is shorter than the minimum value of all the approachable distances of the selected objects; storing the dimensions, position, and orientation of the selected object; The system of claim 3 .

5. The memory further stores instructions executable by the processor, the instructions causing: calculating a distance from said sensor to each stored distinguishable object that is part of said selected object and belongs to said distinguishable object; using the distance of each of the stored distinguishable objects from the sensor to determine whether at least one of the stored distinguishable objects becomes operable; grasping and moving the at least one of the stored distinguishable objects that is manipulable; The system of claim 4.

6. The memory further stores instructions executable by the processor, the instructions causing: calculating a distance from said sensor to each stored indistinguishable object that is part of said selected object and belongs to said distinguishable object; if a difference between a distance from a nearest selected object to the sensor and a distance from a nearest stored indistinguishable object to the sensor is less than a second preset distance threshold, or if the distance from the nearest selected object to the sensor is greater than the distance from the nearest stored indistinguishable object to the sensor, moving the sensor away from the plurality of objects to a position where the sensor can measure at least one of the stored indistinguishable objects; The system of claim 4.

7. The memory further stores instructions executable by the processor, the instructions causing: moving the sensor away from the plurality of objects to a position where the sensor can measure all of the selected objects or all of the stored indistinguishable objects; clearing the stored indistinguishable object information after moving the sensor away from the plurality of objects; The system of claim 6.

8. The memory further stores instructions, the instructions causing for a number of the selected distinguishable objects greater than or equal to two, continuing to grasp and move the selected distinguishable objects to the destination area until all of the selected distinguishable objects have been moved; Simultaneously moving the sensor and grasping the last of the selected distinguishable objects. The system of claim 2 .

9. a depth dimension sensor for measuring a depth dimension of the last-grasped selected distinguishable object; The memory further stores instructions, the instructions causing moving the last-grasped selected distinguishable object to enable the depth dimension sensor to measure a depth dimension of the last-grasped selected distinguishable object; measuring the depth dimension of the last-grasped selected distinguishable object; using the measured depth dimension to estimate a size, position, and orientation of a top surface of an object located behind the last-grasped selected distinguishable object after moving the last-grasped selected distinguishable object; determining whether the sensor can measure the estimated top surface of the object after moving the sensor to grasp the last of the selected distinguishable objects and thereby moving the last of the selected distinguishable objects away from the field of view of the sensor; when it is determined that the sensor cannot measure the top surface of the estimated object, moving the sensor away from the plurality of objects to a position where the sensor can measure the top surface of the estimated object; The system of claim 8.

10. The memory further stores instructions, the instructions causing for a number of the selected distinguishable objects greater than or equal to two, continuing to grasp and move the selected distinguishable objects to the destination area until all of the selected distinguishable objects have been moved; using a preset depth dimension to estimate the approach distance of an object located behind the last of the selected distinguishable objects; selecting the smaller of the estimated approachable distance and the minimum approachable distance of the calculated approachable distance as a selected distance; simultaneously moving the sensor the selected distance and grasping the last of the selected distinguishable objects. The system of claim 2 .

11. The memory further stores instructions, the instructions causing estimating a required movement of the sensor to each indistinguishable object, where the required movement is the distance the sensor can move that increases the confidence of the indistinguishable object to or above the preset confidence threshold; moving the sensor toward the plurality of objects by a distance corresponding to the smaller of the minimum accessible distance of the calculated accessible distances and the minimum required movement; The system of claim 1 .

Citation Information

Patent Citations

  • Variable type article holding device, transferring device, robot handling system, and method of controlling transferring device

    JP2018122945A

  • Robotic system with automatic package registration mechanism and method of operation

    JP2021507857A