Picking device
The picking device addresses the cost issue of conventional bulk picking by using a low-cost 3D camera and AI-driven image processing, achieving high-precision and cost-effective bulk picking operations.
Patent Information
- Application Number
- PCT/JP2024/042044
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-11-27
- Publication Date
- 2025-06-05
AI Technical Summary
Conventional bulk picking devices are costly due to the use of expensive 3D sensors, making them inefficient for high-precision and low-cost bulk picking operations.
A picking device equipped with a camera unit comprising a 2D camera and a low-cost 3D camera, along with a processing unit that performs stereo calibration and AI-driven image processing to recognize and pick pre-registered objects with high precision.
The device achieves low-cost and high-precision bulk picking by reducing the overall system cost through the use of affordable 3D cameras and enhancing recognition capabilities with AI, allowing it to handle loose stacking and environmental variations effectively.
Smart Images

Figure JP2024042044_05062025_PF_FP_ABST
Abstract
Description
Picking Device
[0001] The present invention relates to a picking device, and more particularly to a picking device that can achieve high-precision bulk picking at low cost.
[0002] Conventional bulk picking generally uses 3D sensors.
[0003] A technology has been proposed that provides a processing system for controlling a three-dimensional distance sensor to project a light pattern onto a target surface (see, for example, Patent Document 1).
[0004] Patent document 1 discloses a method including the steps of controlling a three-dimensional distance sensor to project a light pattern onto a target surface, controlling a two-dimensional camera to acquire a first image of the light pattern on the target surface, controlling a light receiving system of the three-dimensional distance sensor to acquire a second image of the light pattern on the target surface, associating a set of two-dimensional coordinates of a first light point among the plurality of light points with distance coordinates of the set of three-dimensional coordinates of the first light point to form a single set of coordinates for the first light point, and deriving a relational equation from the single set of coordinates.
[0005] Japanese Patent Application Laid-Open No. 2022-532725
[0006] However, the technology of Patent Document 1 uses an expensive 3D sensor, and therefore the cost of the entire unit including the 2D camera is too high.
[0007] The present invention has been made in consideration of the above points, and provides a picking device that can achieve high-precision bulk picking at a low cost.
[0008] In other words, the picking device of the embodiment is a picking device that performs bulk picking, and is equipped with a picking unit that picks objects present within a specified area, a camera unit equipped with a 2D camera and a 3D camera that photographs the objects, and a processing unit that recognizes the objects from the images acquired by the camera unit, and the processing unit is characterized by being equipped with an acquisition unit that separates an image of an area corresponding to the object from the image and acquires it as a separated image, a determination unit that determines whether the object in the separated image is a pre-registered object, and an operation control unit that executes operation control on the picking unit to pick the object if the object is a pre-registered object.
[0009] Furthermore, in the picking device, the processing unit may further include a parameter setting unit that performs stereo calibration between the 2D camera and the 3D camera and sets conversion parameters for converting a 3D point cloud generated by image capture with the 3D camera to a corresponding position in a planar image generated by image capture with the 2D camera, a coordinate conversion unit that uses the parameters to convert the coordinates of the 3D point cloud into a planar image, and a coordinate estimation unit that projects the 3D point cloud that has been coordinate-transformed into the planar image onto the planar image and estimates the spatial coordinates of the 3D point cloud projected onto the planar image generated by image capture with the 2D camera.
[0010] Furthermore, in the picking device, the objects may be loosely packed small bags.
[0011] Furthermore, in the picking device, the parameter setting unit may perform stereo calibration between the 2D camera and the 3D camera, and set conversion parameters using AI image processing to convert the 3D point cloud generated by the 3D camera into corresponding positions in the planar image generated by the 2D camera.
[0012] Furthermore, the picking device may include a base, an arm extending from the base, and an adsorption unit connected to the tip of the arm for adsorbing the target object, and the adsorption unit may include an adsorption pad at its tip and a suction unit for creating negative pressure on the adsorption pad.
[0013] Furthermore, in the picking device, the processing unit may further include a priority assignment unit that assigns picking priorities in descending order of area of the separated images; an area extraction unit that extracts an image area in the separated image where the suction pad fits and sets the center of gravity of the image area as the plane coordinate position when the suction pad comes into contact; an angle detection unit that compares the orientation of the object with the orientation of the object in a predetermined image by feature point matching and detects the angle centered on the height direction of the object; and an suction position and orientation determination unit that estimates the spatial coordinates of the suction position where the suction pad will suction the object and determines the position and orientation where the suction pad will suction the object by applying the angle of the object.
[0014] The picking device of the present invention includes a picking unit that picks objects present within a predetermined area, a camera unit equipped with a 2D camera and a 3D camera that photographs the objects, and a processing unit that recognizes the objects from the images acquired by the camera unit. The processing unit includes an acquisition unit that separates an image of an area corresponding to the object from the image and acquires it as a separated image, a determination unit that determines whether the object in the separated image is a pre-registered object, and an operation control unit that controls the picking unit to pick the object if the object is a pre-registered object. Therefore, by using an in-house camera unit equipped with a 2D camera and a low-cost 3D camera, the overall system can be made inexpensive. In addition, because the image of the object is separated and extracted using AI, bulk recognition that is resistant to environmental changes is possible.
[0015] FIG. 1 is a schematic diagram showing the configuration of a picking device of an embodiment; FIG. 2 is a perspective view showing the configuration of a picking device of an embodiment; FIG. 3 is a block diagram showing the functional units of a picking device of an embodiment; FIG. 4 is a top view showing the process of detecting the area of a pouch in a container in an embodiment; FIG. 5 is a top view showing the process of extracting candidate pickup points for a pouch in a container in an embodiment; FIG. 6 is a top view showing the process of detecting the orientation of a pouch in a container in an embodiment; FIG. 7 is a top view showing the process of determining the pickup position and orientation in a container in an embodiment; FIG. 8 is a diagram showing the process of converting a point cloud of 3D coordinates into 2D coordinates in an embodiment; FIG. 9 is a diagram showing processing by an angle detection unit in an embodiment; and FIG. 10 is a flowchart showing the picking process in an embodiment.
[0016] The picking device of the embodiment is a picking device that can achieve high-precision bulk picking at low cost.
[0017] Conventional picking devices use expensive 3D sensors, which makes the entire unit with the 2D camera too costly.
[0018] Therefore, the picking device (method) of the embodiment is positioned as a technology that enables low-cost, high-precision bulk picking by using a camera unit consisting of a 2D camera and an inexpensive 3D camera. The 2D camera captures two-dimensional information in horizontal and vertical directions (X, Y), while the 3D camera captures not only vertical and horizontal information but also depth information. The 3D camera of the embodiment complements the 2D camera by being able to measure heights that the 2D camera cannot measure. The 2D camera is highly accurate, complementing the low-precision 3D camera. Bulk picking refers to the process of picking up multiple small bags 3 piled loosely in a container 2, removing them from the container 2, and placing them in a small box.
[0019] The configuration of picking device 1 of this embodiment is shown as a schematic diagram in Figure 1. Picking device 1 includes camera unit 10 that captures images of an object. Camera unit 10 includes 2D camera 11 and 3D camera 12. 2D camera 11 and 3D camera 12 capture images of multiple pouches 3 stacked in container 2.
[0020] 2 is a perspective view showing the entire picking device 1. A plurality of pouches 3 are stacked in a container 2. Each pouch 3 contains, for example, snacks, jam, sauce, screws, or electronic components, and measures, for example, 10 cm in length and 5 cm in width. A 2D camera 11 is installed directly above the container 2 at a distance of 50 cm to 1 m, and captures two-dimensional images of the plurality of pouches 3 stacked in the container 2. A 3D camera 12 is installed immediately behind the 2D camera 11, and captures three-dimensional images of the plurality of pouches 3 stacked in the container 2.
[0021] Picking unit 20 is positioned closer to container 2 than camera unit 10, and picks a plurality of pouches 3 piled up in container 2. Picking unit 20 has a robotic structure and includes a base 21, an arm 22 extending from base 21, and a suction unit 23 connected to the tip of arm 22 and adapted to suck an object. Suction unit 23 includes a suction pad 25 at its tip, and a pump-type suction unit 24 that creates negative pressure on suction pad 25. Suction pad 25 is made of a suction cup or rubber.
[0022] The processing unit 30 is disposed behind the picking unit 20 and is connected to the camera unit 10 and the picking unit 20. The processing unit 30 is configured as a computer in terms of hardware, and is internally equipped with processing elements such as a CPU, a GPU, and storage elements such as a ROM, a RAM, a HDD, and an SSD. The processing unit 30 may be implemented as various electronic computers (computing resources), such as a personal computer (PC), a mainframe, a workstation, a cloud computing system, a tablet terminal, a smartphone, or the like. In the embodiment, a tablet terminal (computer) is used as the processing unit 30. The picking method is realized in terms of software by a picking program loaded into main memory, etc.
[0023] 3 is a block diagram showing the functional units of the picking device 1. The picking device 1 is made up of a picking unit 20, a camera unit 10, and a processing unit 30. The processing unit 30 includes an acquisition unit 40, a determination unit 50, an operation control unit 60, a parameter setting unit 70, a coordinate conversion unit 80, a coordinate estimation unit 90, a priority assignment unit 100, an area extraction unit 110, an angle detection unit 120, and a pickup position and attitude determination unit 130. The picking unit 20 and the camera unit 10 are as described above.
[0024] The acquisition unit 40 separates an image of a region corresponding to an object from the image acquired by the camera unit 10 and acquires it as a separated image. A separated image is an image separated by segmentation. In other words, the object to be picked is enclosed in a rectangular frame to form a section and distinguish it from other objects. This makes it possible to separate only the object to be picked.
[0025] The determination unit 50 determines whether the object in the separated image is a pre-registered object. For example, if the object to be picked is a "bag of snacks" and "bag of snacks" is pre-registered, the determination unit 50 determines that the object is a pre-registered object.
[0026] If the target object is a pre-registered target object, the operation control unit 60 controls the picking unit 20 to pick up the target object.
[0027] The parameter setting unit 70 performs stereo calibration between the 2D camera 11 and the 3D camera 12 and sets transformation parameters for converting the 3D point cloud generated by the 3D camera 12 into corresponding positions in the planar image generated by the 2D camera 11. Stereo calibration is a type of stereo matching, and stereo matching is a technique for estimating the depth of a scene captured in an image using two images of the same still scene captured from different viewpoints. The transformation parameters are a predetermined matrix, and are used to calculate the distance between a certain point (x) in space and a certain point (x) in space. 1 , y 1 , z 1 ) to another position (x 2 , y 2 , z 2)
[0028] The coordinate transformation unit 80 transforms the 3D point cloud into a planar image using the parameters, as will be described later with reference to FIG.
[0029] The coordinate estimation unit 90 projects the 3D point cloud, which has been coordinate-transformed into a planar image, onto the planar image, and estimates the spatial coordinates of the 3D point cloud projected onto the planar image captured by the 2D camera 11. The spatial coordinates are Cartesian coordinates expressed as (x, y, z).
[0030] The priority assigning unit 100 assigns picking priorities in descending order of the area of the separated image. For example, the object with the largest area of the separated image is picked first, the object with the second largest area of the separated image is picked second, and so on. As a result, the objects with the largest area are picked first, making it easier to see the multiple pouches 3 piled up in the container 2. This will be described later with reference to Figure 4.
[0031] The region extraction unit 110 extracts an image region in the separated image where the suction pad 25 fits, and sets the center of gravity of the image region as the position of the plane coordinates when the suction pad 25 comes into contact with the image region. This will be described later with reference to FIG.
[0032] Angle detection unit 120 compares the orientation of the object with the orientation of the object in a predetermined image to detect the angle centered in the height direction of the object. The comparison is performed by feature point matching. Feature point matching involves comparing geometrically characterized positions, such as the four corners of the pouch or the center of gravity of the pouch, between three-dimensional and two-dimensional coordinates to measure the angular deviation from the perpendicular.
[0033] The suction position and orientation determination unit 130 estimates the spatial coordinates of the suction position where the suction pad 25 suctions the object, and determines the position and orientation where the suction pad 25 suctions the object by applying the angle θ of the object, thereby enabling the suction pad 25 to accurately suction the object.
[0034] Figure 4 is a top view showing the process of detecting the area of a pouch 3 in a container in this embodiment. The process of detecting the area of this pouch 3 is performed by the 2D camera 11 using instance segmentation by AI (artificial intelligence). Instance segmentation is a technique for distinguishing and detecting the foreground mask of an "instance of an object class" that appears in an image or RGB-D image, for each instance. The area enclosed by a square in Figure 4 is the area of one pouch 3.
[0035] 5 is a top view showing the process of extracting candidate suction points for pouch 3 in container 2. This process is also performed by 2D camera 11, and points where suction pad 25 fits are extracted as candidate suction points P from the area of pouch 3 detected by the area detection process in FIG.
[0036] Figure 6 is a top view showing the process of detecting the orientation of pouch 3 in container 2. This process is also performed by 2D camera 11, and involves comparing the orientation of the object extracted by region extraction unit 120 with the orientation of the object in a pre-registered image of the object, and detecting the orientation deviation as the angle θ of deviation from the perpendicular line in the height direction of the object. This orientation comparison is performed by feature point matching. In Figure 6, the vertical orientation of pouch 3 is indicated by arrow A, and the horizontal orientation is indicated by arrow B.
[0037] Figure 7 is a top view showing the process of determining the suction position and orientation of pouch 3 inside the container. This process is performed by 3D camera 12, and the position and orientation at which suction pad 25 will suction the object are determined using the coordinates of the 3D point cloud of the three-dimensional image. In Figure 7, the vertical orientation of pouch 3 is indicated by arrow A, and the horizontal orientation is indicated by arrow B.
[0038] FIG. 8 is a diagram showing the process of converting a point cloud of 3D coordinates into 2D coordinates. As shown in FIG. 8, the upper side shows the three-dimensional coordinates onto which the 3D point cloud is mapped. P 1 , P 2 , P 3Each of these is a point in the 3D point cloud. When a point on a 3D coordinate system is transferred to a 2D coordinate system, some distortion occurs. Therefore, depending on the location, an algorithm created by AI is used to optimize, correct, and rectify the matrix. The original point P is multiplied by a certain matrix A to obtain a new point P'. For example, it can be expressed as a 3-row, 3-column equation as follows:
[0039]
[0040] 9 is a diagram showing the processing performed by angle detection unit 120. The orientation of the object extracted by area extraction unit 110 is compared with the orientation of the object in a pre-registered image of the object, and the deviation in orientation is detected as an angle θ of deviation with respect to a perpendicular line H in the height direction of the object. 3A is the small pouch that is the object to be grasped, and 3B and 3C are other objects. A perpendicular line H is dropped from the center of gravity of 3A, and the rotation angle θ about perpendicular line H is detected as the angle of deviation between the orientation of the object extracted by area extraction unit 110 and the orientation of the object in the pre-registered image of the object.
[0041] The picking method and picking program of the embodiment will now be described with reference to the flowchart of Figure 10. The picking method of the embodiment is executed by the computer (processing unit 30) of the picking device 1 of the embodiment based on the picking program (see Figures 2 and 3). The picking program of the embodiment causes the computer of the picking device 1 to realize a parameter setting function, a coordinate conversion function, a coordinate estimation function, an acquisition function, a determination function, a priority assignment function, an area extraction function, an angle detection function, a pickup position and attitude determination function, a pickup function, and an operation control function. Each function overlaps with the description of the picking device 1 of the embodiment described above, so details will be omitted.
[0042] 10 shows the flow of an information processing method according to an embodiment, and includes various steps such as a parameter setting step (S1), an acquisition step (S2), a coordinate conversion step (S3), a coordinate estimation step (S4), an area extraction step (S5), a determination step (S6), a priority assignment step (S7), an angle detection step (S8), a pickup position and orientation determination step (S9), a pickup step (S10), and an operation control step (S11). In addition, the picking method also includes various other steps not shown in the drawings as needed.
[0043] The parameter setting function performs stereo calibration between the 2D camera 11 and the 3D camera 12 and sets transformation parameters for transforming the 3D point cloud generated by the 3D camera 12 into corresponding positions in the planar image generated by the 2D camera 11 (S1: parameter setting step). The acquisition function separates an image of a region corresponding to the object from the image acquired by the camera unit 10 and acquires it as a separated image (S2: acquisition function).
[0044] The coordinate conversion function converts the coordinates of the 3D point cloud into a planar image using the parameters (S3: coordinate conversion step). The coordinate estimation function projects the 3D point cloud, whose coordinates have been converted into a planar image, onto the planar image, and estimates the spatial coordinates of the 3D point cloud projected onto the planar image captured by the 2D camera 11 (S4: coordinate estimation step).
[0045] The area extraction function extracts an image area in the separated image where the suction pad 25 fits, and sets the center of gravity of the image area as the position of the plane coordinates when the suction pad 25 comes into contact (S5: area extraction function). The determination function determines whether the object in the separated image is a pre-registered object (S6: determination function).
[0046] The priority assignment function assigns picking priorities to separated images in descending order of area (S7: priority assignment function). The angle detection function compares the orientation of the object extracted by the area extraction function with the orientation of the object in a pre-registered image of the object, and detects the deviation in orientation as the angle of deviation from the perpendicular line in the height direction of the object (S8: angle detection function). The suction position and orientation determination function estimates the spatial coordinates of the suction position where the suction pad 25 will suction the object, and determines the position and orientation where the suction pad 25 will suction the object by applying the angle of the object (S9: suction position and orientation determination function).
[0047] The suction function sucks the target object (S10: suction function). The operation control function executes operation control for the picking unit to pick up the target object if the target object is a pre-registered target object (S11: operation control function).
[0048] The information processing program of the embodiment can be implemented using, for example, scripting languages such as ActionScript, JavaScript (registered trademark), Python, and Ruby, or compiler languages such as C, C++, C#, Objective-C, Swift, and Java (registered trademark).
[0049] REFERENCE SIGNS LIST 1 Picking device 2 Container 3 Pouch 10 Camera unit 11 2D camera 12 3D camera 20 Picking section 30 Processing section 40 Acquisition section 50 Determination section 60 Operation control section 70 Parameter setting section 80 Coordinate conversion section 90 Coordinate estimation section 100 Priority assignment section 110 Area extraction section 120 Angle detection section 130 Pickup position and orientation determination section
Claims
1. A picking device comprising: a picking unit which picks an object present within a specified area; a camera unit equipped with a 2D camera and a 3D camera which photographs the object; and a processing unit which recognizes the object from the image acquired by the camera unit, wherein the processing unit comprises: an acquisition unit which separates an image of an area corresponding to the object from the image and acquires it as a separated image; a determination unit which determines whether the object in the separated image is a pre-registered object; and an operation control unit which, if the object is a pre-registered object, executes operation control on the picking unit to pick the object.
2. The picking device described in claim 1, further comprising: a parameter setting unit that performs stereo calibration between the 2D camera and the 3D camera and sets conversion parameters for converting a 3D point cloud generated by the 3D camera's image capture to corresponding positions within a planar image generated by the 2D camera's image capture; a coordinate conversion unit that uses the parameters to perform coordinate conversion of the 3D point cloud onto the planar image; and a coordinate estimation unit that projects the 3D point cloud whose coordinates have been converted into the planar image onto the planar image and estimates spatial coordinates of the 3D point cloud projected onto the planar image generated by the 2D camera's image capture.
3. The picking device according to claim 1, wherein the object is a bulk bag.
4. The picking device according to claim 1, characterized in that the recognition unit performs image processing using AI. The picking device according to claim 1, characterized in that the parameter setting unit performs stereo calibration between the 2D camera and the 3D camera and sets, by image processing using AI, conversion parameters for converting a 3D point cloud generated by shooting with the 3D camera into corresponding positions within a planar image generated by shooting with the 2D camera.
5. The picking device described in claim 1, characterized in that the picking unit comprises a base, an arm extending from the base, and an suction unit connected to the tip of the arm for suctioning an object, and the suction unit comprises a suction pad at its tip and a suction unit for creating negative pressure on the suction pad.
6. The picking device according to claim 4, further comprising: a priority assignment unit which assigns priorities for picking to the separated images in the order of size, an area extraction unit which extracts an image area in which the suction pad fits in the separated image and sets the position of the center of gravity of the image area as the position of plane coordinates when the suction pad comes into contact, an angle detection unit which compares the orientation of the object extracted by the area extraction unit with the orientation of the object in a pre-registered image of the object and detects a deviation in orientation as the angle of deviation relative to a perpendicular line in the height direction of the object, and a suction position and attitude determination unit which estimates the spatial coordinates of the suction position where the suction pad will suction the object, and applies the angle of the object to determine the position and attitude where the suction pad will suction the object.
Citation Information
Patent Citations
Picking system, picking method, and program
JP2021088011A