Moving object identification system, moving object identification device, object identification method, and program

JPWO2024105855A5Active Publication Date: 2025-07-18NEC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024558603
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-18
Estimated Expiration
2042-11-17

AI Technical Summary

Technical Problem

Existing methods for determining the position and orientation of a pallet, such as those described in Patent Documents 1 and 2, face challenges in accuracy and efficiency, particularly when the pallet is not directly facing the forklift or when learning is insufficient, leading to potential inaccuracies in estimating the 6D Pose of the pallet.

Method used

A moving object identification system that includes a holding means, height acquisition means, recognition means, search means, and estimation means to identify the position and orientation of a pallet by recognizing corners on the front surface and estimating corners on the rear surface using similar data, enabling accurate 6D Pose estimation even when the rear of the pallet is not visible or learning is difficult, employing an RGB-D camera and machine learning for corner recognition and pose estimation.

Benefits of technology

The system efficiently and accurately estimates the position and orientation of the pallet, assisting operators and enabling automatic operation by solving the PnP problem using corner positions, thereby improving accuracy and stability in pallet identification.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Provided is a system for identifying object to be moved, said system comprising: a holding unit (102) that holds an object to be moved; a height acquisition unit (202) that acquires the height of an imaging device (103) attached to the holding unit; a recognition unit (203) that recognizes the position of a corner existing at a front surface of the object to be moved which has been imaged using the imaging device; a search unit (204) that searches for similar data, which is similar to an image of the object to be moved, in accordance with the acquired height of the imaging device and the recognized position of the corner existing at the front surface of the object to be moved; an estimation unit (205) that estimates the position of a corner existing at a rear surface of the object to be moved, in accordance with similar data which has been retrieved; and an object identification unit (206) that identifies a state of the object to be moved in accordance with the recognized position of the corner existing at the front surface of the object to be moved and the estimated position of the corner existing at the rear surface of the object to be moved.
Need to check novelty before this filing date? Find Prior Art

Description

Mobile object identification system, mobile object identification device, object identification method, and computer-readable medium

[0001] The present disclosure relates to a mobile object identification system, a mobile object identification device, an object identification method, and a computer-readable medium.

[0002] There is a technology for determining the position and orientation of a pallet in a system in which a mobile object transports a pallet. Patent Document 1 describes a method for estimating the position and orientation of a pallet based on the line segments of a bounding box (hereinafter referred to as BB) that surrounds either the front of the pallet or one of its two holes, as captured by a camera. It also describes that the position and orientation of the pallet may be estimated based on BB data and reference data.

[0003] Patent Document 2 describes a method of determining whether a forklift is facing a pallet directly based on whether the shape of the pallet included in the image is symmetrical.

[0004] JP 2021-24718 A JP 2020-109030 A

[0005] In the invention described in Patent Document 1, the position and orientation of the pallet are calculated by comparing the length of the line segment of BB and the BB data with reference data, so there is a possibility that the position and orientation of the pallet cannot be estimated with high accuracy.

[0006] The invention described in Patent Document 2 is based on whether the shape of the pallet included in the image is bilaterally symmetrical, and therefore cannot estimate the position and orientation of a pallet that is not directly facing the forklift.

[0007] Therefore, the inventions described in Patent Documents 1 and 2 may not be able to efficiently estimate the position and orientation of the pallet.

[0008] The moving target identification system of the present disclosure is a moving target identification system comprising: a holding means for holding a moving target; a height acquisition means for acquiring the height of an imaging device attached to the holding means; a recognition means for recognizing the position of a corner present on the front side of the moving target imaged using the imaging device; a search means for searching for similar data that resembles the image of the moving target according to the acquired height of the imaging device and the recognized position of a corner present on the front side of the moving target; an estimation means for estimating the position of a corner present on the rear side of the moving target according to the searched similar data; and an identification means for identifying the state of the moving target according to the recognized position of a corner present on the front side of the moving target and the estimated position of a corner present on the rear side of the moving target.

[0009] The moving target identification device of the present disclosure is a moving target identification device comprising: a holding means for holding a moving target; a height acquisition means for acquiring the height of an imaging device attached to the holding means; a recognition means for recognizing the position of a corner present on the front side of the moving target imaged using the imaging device; a search means for searching for similar data that resembles the image of the moving target according to the acquired height of the imaging device and the recognized position of a corner present on the front side of the moving target; an estimation means for estimating the position of a corner present on the rear side of the moving target according to the searched similar data; and an identification means for identifying the state of the moving target according to the recognized position of a corner present on the front side of the moving target and the estimated position of a corner present on the rear side of the moving target.

[0010] The target identification method disclosed herein is a target identification method that captures an image of a target using an imaging device, acquires the height of the imaging device, recognizes the positions of corners present in front of the target imaged using the imaging device, searches for similar data that resembles the image of the target based on the acquired height of the imaging device and the recognized positions of corners present in front of the target, estimates the positions of corners present on the rear of the target based on the searched similar data, and identifies the state of the target based on the recognized positions of corners present in front of the target and the estimated positions of corners present on the rear of the target.

[0011] The computer-readable medium of the present disclosure is a non-transitory computer-readable medium that stores a program that causes an information processing device to perform the following steps: capture an image of an object using an imaging device; acquire the height of the imaging device; recognize the positions of corners present in front of the object imaged using the imaging device; search for similar data that resembles the image of the object based on the acquired height of the imaging device and the recognized positions of corners present in front of the object; estimate the positions of corners present on the rear of the object based on the searched similar data; and identify the state of the object based on the recognized positions of corners present in front of the object and the estimated positions of corners present on the rear of the object.

[0012] According to the present disclosure, a moving object identification system that can efficiently estimate the position and orientation of a pallet can be provided.

[0013] 1 is a schematic diagram of a moving target identification system according to an embodiment; FIG. 2 is a block diagram showing the configuration of a moving target identification system according to an embodiment; FIG. 3 is a flowchart of a target identification method according to an embodiment; FIG. 4 is a block diagram of a moving target identification system according to a first embodiment; FIG. 5 is a diagram showing the hierarchy of a dictionary data section according to a first embodiment; FIG. 6 is a flowchart of a target identification method according to a first embodiment; FIG. 7 is a diagram showing recognition of the position of a front corner, searching dictionary data, and estimation of the position of a rear corner according to a first embodiment; and FIG. 8 is a diagram showing determination of the position and orientation of a moving target from the positions of the front corner and the rear corner by solving a PnP problem according to a first embodiment.

[0014] Embodiments Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the invention according to the claims is not limited to the following embodiments. Furthermore, not all of the configurations described in the embodiments are necessarily essential as means for solving the problems. For clarity of explanation, the following description and drawings have been omitted and simplified as appropriate. In each drawing, the same elements are given the same reference numerals, and repeated explanations are omitted as necessary.

[0015] (Description of a moving target identification system according to an embodiment) Fig. 1 is a schematic diagram of a moving target identification system according to an embodiment. Fig. 2 is a block diagram showing the configuration of a moving target identification system according to an embodiment. The moving target identification system according to an embodiment will be described with reference to Figs. 1 and 2.

[0016] As shown in FIG. 1, a moving object identification system 100 according to the embodiment includes a moving object 101, a holding unit 102, an imaging device 103, a sensor 104, and an information processing device 105.

[0017] The moving body 101 is, for example, a forklift. The moving body 101 can use the holding unit 102 to transport and move a moving object 701 (see FIG. 7 ) of a fixed shape. The moving body 101 itself does not need to move. The moving body 101 only needs to be able to change the height of the moving object 701 using the holding unit 102. The moving object 701 is, for example, a pallet and has a fixed size. The moving object 701 is a loading platform for carrying luggage. The moving object 701 has a rectangular parallelepiped front surface. That is, there are four corner positions on the front surface of the moving object 701. Similarly, there are four corner positions on the rear surface of the moving object 701. The moving object 701 has insertion holes at the corners of its front surface, and can be lifted by inserting the holding unit 102 into them.

[0018] The holding unit 102 is, for example, a fork attached to a forklift. The holding unit 102 is L-shaped in side view, and holds the object 701 to be moved by inserting the bottom part into the object 701 to be moved. The holding unit 102 can be moved up and down. Therefore, the height of the holding unit 102 can be changed.

[0019] The imaging device 103 is, for example, an RGB-D camera. An RGB-D camera is a camera that outputs depth data and color data. Alternatively, for example, an RGB camera and a depth sensor may be used as the imaging device 103. There may be multiple imaging devices 103. The imaging device 103 is attached to the holding unit 102 and captures images of the surroundings of the imaging device 103 and the moving object 701. The imaging device 103 may be configured to enable driving assistance or automatic driving by capturing images of the surroundings of the imaging device 103. Furthermore, by capturing an image of the moving object 701, the imaging device 103 can identify the position and orientation of the moving object 701, as will be described later.

[0020] The sensor 104 is a variety of sensors that sense the state of the mobile body 101 or the holding unit 102. Because the holding unit 102 is attached to the mobile body 101, the sensor 104 particularly acquires the height of the holding unit 102 relative to the mobile body 101. The sensor 104 may acquire the height of the holding unit 102 based on operation information for a lift cylinder of the mobile body 101. The sensor 104 may also measure the height of the holding unit 102 from the ground. The sensor 104 may measure the height from the ground by attaching a distance sensor such as a LiDAR, laser sensor, radar sensor, or Time of Flight sensor to the holding unit 102. Because the imaging device 103 is attached to the holding unit 102, acquiring the height of the holding unit 102 is synonymous with acquiring the height of the imaging device 103. By acquiring the height of the holding unit 102 when an image of the moving object 701 is captured using the imaging device 103, the height of the imaging device 103 at the time of capturing the image can be acquired. The sensor 104 may sense the position, speed, distance from the moving object 701 to the holding unit 102, etc. of the moving body 101 for driving assistance or automatic driving of the moving body 101.

[0021] The information processing device 105 processes data collected from various devices and sensors attached to the mobile object 101. The information processing device 105 is network-connected to the mobile object 101 via, for example, Wi-Fi (registered trademark), Bluetooth (registered trademark), or the like. The information processing device includes at least one processor that executes instructions and at least one memory that stores the instructions. For example, the information processing device 105 acquires images from the imaging device 103. The information processing device 105 also acquires sensor information from the sensor 104. Furthermore, the information processing device 105 issues control commands to the mobile object 101 and controls the mobile object 101 to provide driving assistance or automatic driving. The information processing device 105 may include, for example, a machine learning machine. Furthermore, some or all of the functions of the information processing device 105 may be distributed to the cloud. Furthermore, the information processing device 105 may be a single device or multiple devices. Here, the information processing device 105 remotely controls the moving object 101, but the information processing device 105 may be installed in the moving object 101 and the moving object 101 may operate independently. In this case, the information processing device 105 can be regarded as one moving object identification device.

[0022] If the position and orientation of the moving object 701 can be identified, it is possible to assist the operator of the moving body 101 in operating the moving body 101. Furthermore, it is possible to contribute to realizing automatic driving of the moving body 101.

[0023] The processing of the information processing device 105 according to the embodiment will be described with reference to Fig. 2. As shown in Fig. 2, the information processing device 105 includes an image acquisition unit 201, a height acquisition unit 202, a recognition unit 203, a search unit 204, an estimation unit 205, and a target identification unit 206.

[0024] The image acquisition unit 201 acquires an image from the imaging device 103 attached to the holding unit 102. The image from the imaging device 103 may be a normal RGB image without depth information. The image also includes a moving object 701.

[0025] The height acquisition unit 202 acquires the height of the holding unit 102 when the image is captured. The height of the holding unit 102 when the image is captured is, in other words, the height of the imaging device 103 when the image is captured.

[0026] The recognition unit 203 recognizes the positions of the corners on the front surface of the moving object 701 from the captured image. As described above, the positions of the four corners on the front surface of the moving object 701, which is a rectangular parallelepiped, are recognized. Here, "recognizing" the positions of the four corners means specifying the positions where the four corners are located. A known method can be used for the recognition. For example, the front surface of the pallet is cut out using the recognition result of the pallet holes, and the positions P of the corners on the front surface of the pallet are detected using edge detection. k Edge detection is a method for recognizing discontinuous changes in an image according to the amount of change in the features in the image. Also, feature point information on the corner positions is machine-learned in advance, and the image of the moving object 701 is input using feature point matching to detect the position P of the front corner. k The image of the moving object may be input to a machine learning machine that has learned images of multiple moving objects to recognize the position of the corners in front of the moving object. Alternatively, the image of the moving object 701 may be input using a 6D pose estimation technique using a convolutional neural network to recognize the position P of the corners in front of the moving object. k 6D Pose is information that represents the position and orientation of an object using three-axis rotation vectors and three-axis translation vectors.

[0027] The search unit 204 is a part that has a function of searching for similar data that resembles the image of the moving object 701 according to the height of the holding unit 102 when the image was captured and the positions of the corners on the front surface of the recognized moving object 701. Similar data is pre-stored images of the moving object 701. The search unit 204 stores similar data captured under many conditions. The search unit 204 searches for similar data based on the height of the holding unit 102 when the image was captured. The search unit 204 is a trained machine learning machine that stores and learns from multiple combinations of the positions of the corners on the front surface of the captured moving object 701 and the positions of the corners present on the rear surface as training data sets.

[0028] The estimation unit 205 has a function of estimating the position of a corner on the rear surface of the moving object based on the searched similar data. That is, when the position of a corner on the front surface of the moving object 701 is input, the estimation unit 205 estimates the position of the corner on the rear surface using a trained machine learning device that outputs the position of a corner on the rear surface of the moving object 701.

[0029] The target identification unit 206 is a part that has the function of identifying the state of the moving target 701 according to the positions of corners present on the front surface of the recognized moving target 701 and the positions of corners present on the rear surface of the estimated moving target 701. The target identification unit 206 identifies the three-dimensional posture and position (6D Pose) of the moving target 701 by solving a PnP problem based on, for example, the positions of corners present on the front surface of the recognized moving target 701 and the positions of corners present on the rear surface of the estimated moving target 701. In order to solve the PnP problem, the internal parameters of the imaging device 103 are known. The size of the moving target 701 is also known. The PnP problem can be solved using a known method.

[0030] In the description of the embodiment, the holding unit 102, image acquisition unit 201, height acquisition unit 202, recognition unit 203, search unit 204, estimation unit 205, and target identification unit 206 are referred to, but these may also be read as holding means, acquisition means, recognition means, search means, estimation means, and target identification means.

[0031] The method of solving the PnP problem to determine the 6D Pose from nine points, including the positions of the front corners, the rear corners, and the center position, may not be able to accurately estimate the 6D Pose when a large load is loaded and the rear of the pallet is not visible in the image, or when learning is insufficient. By using the moving object identification system disclosed herein, the position and orientation of the object can be estimated with higher accuracy than a technology that estimates the 6D Pose of the object based on a convolutional neural network when the rear of the moving object is not visible in the image or when learning of the moving object is insufficient. This can therefore assist the operator of the moving object 101. Furthermore, this can contribute to the realization of autonomous driving of the moving object 101.

[0032] When estimating the position and orientation of a pallet using only BB data, some method must be used to estimate the corner positions after acquiring the BB data. This can result in unstable calculations. However, the moving object identification system according to the embodiment can accurately identify the positions of corners on the rear surface in a short time, allowing for stable identification of the position and orientation.

[0033] (Description of Target Identification Method According to the Embodiment) Fig. 3 is a flowchart of the target identification method according to the embodiment. The target identification method according to the embodiment will be described with reference to Fig. 3 .

[0034] As shown in FIG. 3 , first, an image is captured (step S301). An image of the target is captured using the imaging device 103. Next, the height of the imaging device is acquired (step S302). The information processing device 105 acquires the height of the imaging device 103 at the time of capturing the image. Next, the positions of corners on the front surface are recognized (step S303). The information processing device 105 recognizes the positions of corners present in front of the captured target. Next, similar data is searched for (step S304). The information processing device 105 searches for similar data that is similar to the image of the target based on the height of the imaging device at the time of capturing the image and the positions of corners present in front of the recognized target.

[0035] Next, the positions of the corners on the rear surface are estimated (step S305). The information processing device 105 estimates the positions of the corners on the rear surface of the object based on the searched similar data. Next, the object is identified using the positions of the corners on the front surface and the corners on the rear surface (step S306). The information processing device 105 identifies the state of the object based on the positions of the corners on the front surface of the recognized object and the estimated positions of the corners on the rear surface of the object.

[0036] By using such an object identification method, the position and orientation of the object can be determined more efficiently than related techniques and with higher accuracy than a method of calculating a 6D pose from the positions of corners present in the front.

[0037] (Description of a moving target identification system according to a first embodiment) The first embodiment is a form showing an example of an embodiment, and includes examples of non-essential configurations and operations. FIG. 4 is a block diagram of a moving target identification system according to the first embodiment. FIG. 5 is a diagram showing the hierarchy of a dictionary data section according to the first embodiment. The moving target identification system according to the first embodiment will be described with reference to FIGS. 4 and 5.

[0038] As shown in FIG. 4, the moving target identification system 400 according to the first embodiment differs from the moving target identification system 100 according to the embodiment in that it further includes a moving object 101 and a dictionary data unit 401 .

[0039] The dictionary data section 401 registers a plurality of similar data. The similar data is a captured image of the moving object 701. The dictionary data section 401 registers a large number of similar data captured from various heights and angles.

[0040] 5, the dictionary data section 401 has a hierarchical structure. The dictionary data section 401 stores the dictionary data of the plurality of cameras C attached to the holding section 102. id (id=1, 2, . . .), the height H of the holding portion 102 i The similar data captured at (i=1, 2, . . .) is saved. The similar data is stored at the positions P 1 (u, v), P 2 (u, v), P 5 (u, v), P 6 (u, v) and the position of the corner on the back surface P 3 (u, v), P 4 (u, v), P 7 (u, v), P 8 Point cloud data D including (u, v) i (i=1, 2, . . . ) are registered.

[0041] The search unit 204 identifies, from the dictionary data unit 401, similar data that has the same ID as the camera that captured the moving object 701, has a small error in the height of the holding unit 102 when captured, and has a small error in the position of the corners present in front of the moving object 701. In this way, the search unit 204 searches for similar data that resembles the image of the moving object 701. The similar data searched for by the search unit 204 is preferably similar data that has the smallest error in the height of the holding unit 102 when captured and the smallest error in the position of the corners present in front of the moving object 701.

[0042] If similar data is found and completely matches the image of the moving object, the position of the corner on the rear surface of the similar data is used. However, there is little similar data that completely matches the image of the moving object, in which case the position of the corner on the rear surface is estimated using the following calculation method.

[0043] First, for example, P 1 From the reference point such as k The conversion formula λ k Calculate λ k is expressed by the following formula: Here, D m is the point cloud data with the smallest error in the position of the corner present in the front. 1 From the above conversion formula λ k Using the virtual point P k Calculate.

[0044] In this way, the position of the corner on the rear surface, which is a virtual point, is estimated from the position of the corner on the front surface. Then, by solving the PnP problem using the positions of the corners on the front surface and the corners on the rear surface, it is possible to estimate the 6D Pose with high accuracy.

[0045] In this way, the position and orientation of the pallet can be estimated efficiently.

[0046] (Description of Object Identification Method According to First Embodiment) Fig. 6 is a flowchart of the object identification method according to the first embodiment. Fig. 7 is a diagram showing recognition of the positions of the front corners, searching dictionary data, and estimation of the positions of the rear corners according to the first embodiment. Fig. 8 is a diagram showing how the position and orientation of a moving object are determined from the positions of the front corners and the rear corners by solving a PnP problem according to the first embodiment. The object identification method according to the first embodiment will be described with reference to Figs. 6 to 8.

[0047] As shown in Fig. 6, first, an image is acquired (step S601). The imaging device 103 captures an image of a moving object 701. Next, the information processing device 105 acquires an imaging device ID (Identification), imaging device internal parameters, and imaging device height (step S602). The internal parameters differ for each imaging device 103. The internal parameters of the imaging device 103 are required when solving the PnP problem. Therefore, the information processing device 105 needs to acquire the imaging device ID and its internal parameters.

[0048] Next, the information processing device 105 recognizes the positions of the corners on the front surface of the moving object (step S603). Next, the information processing device 105 searches for matching data from the dictionary data unit 401 (step S604). Next, the information processing device 105 estimates the positions of the corners on the rear surface of the moving object 701 (step S605). These three processes will be described with reference to FIG. 7.

[0049] As shown in the upper left of FIG. 7, the position P of the corner in front of the moving object 701 1 , P 2 , P 5 , P 6 These four are recognized using machine learning or the like. Next, similar data is searched for in the dictionary data section 401 as shown in the bottom of Fig. 7. Similar data is found that matches the image capture device 103, has the smallest error in the height of the image capture device, and has the position of the corner at the front that most closely matches the captured image, and the position P of the corner present on the rear surface is then identified as shown in the top left of Fig. 7. 3 , P 4 , P 7 , P 8 The estimation method is as described above,1 The virtual point P is the position of the corner on the rear surface. k Conversion formula to λ k Calculate the reference point P 1 Transform from virtual point P k This is a method for calculating

[0050] Finally, the information processing apparatus 105 solves the PnP problem using the positions of the corners on the front surface and the corners on the rear surface to identify the position and orientation (step S606). k (u k , v k ) (k=0, 1...) to solve the PnP problem. 0 (u 0 , v 0 ) is the center point of the moving object 701. By solving the PnP problem, R|t of the 6D Pose, which is the three-dimensional posture of the moving object 701, is obtained.

[0051] By estimating the 6D Pose, it becomes possible to calculate the turning radius and the number of degrees of turning of the moving body 101 so that the moving body 101 can face the moving target 701 head-on.

[0052] It is also possible to geometrically estimate the 6D pose of the pallet by using a depth sensor to reconstruct the pallet's front surface in 3D without solving the PnP problem. However, in this case, a highly accurate depth sensor is required, and stable distance information such as LiDAR must be acquired. Therefore, highly accurate estimation is difficult with an inexpensive RGB-D camera. As described above, the position and orientation of a moving object can be estimated using an inexpensive RGB-D camera by solving the PnP problem using the positions of the front corners and the rear corners to identify the position and orientation.

[0053] As disclosed in the embodiments and embodiment 1, the position and orientation of a moving object can be estimated with high accuracy by estimating the position of a corner on the rear surface of the moving object based on the position of a corner on the front surface of the moving object.

[0054] Furthermore, part or all of the processing in the information processing device 105 described above can be realized as a computer program. Such a program can be stored using various types of non-transitory computer-readable media and supplied to a computer. Non-transitory computer-readable media include various types of tangible recording media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory)). The program may also be supplied to a computer by various types of temporary computer-readable media. Examples of temporary computer-readable media include electrical signals, optical signals, and electromagnetic waves. The temporary computer-readable medium can supply the program to the computer via a wired communication path such as an electric wire or an optical fiber, or via a wireless communication path.

[0055] The present invention is not limited to the above-described embodiment, and can be modified as appropriate within the scope of the invention.

[0056] Some or all of the above embodiments may be described as, but are not limited to, the following supplementary notes. (Supplementary Note 1) A moving object identification system comprising: a holding means for holding a moving object; a height acquisition means for acquiring the height of an imaging device attached to the holding means; a recognition means for recognizing the position of a corner present on the front side of the moving object imaged using the imaging device; a search means for searching for similar data similar to an image of the moving object based on the acquired height of the imaging device and the recognized position of a corner present on the front side of the moving object; an estimation means for estimating the position of a corner present on the rear side of the moving object based on the searched similar data; and an identification means for identifying a state of the moving object based on the recognized position of a corner present on the front side of the moving object and the estimated position of a corner present on the rear side of the moving object. (Supplementary Note 2) The moving object identification system according to Supplementary Note 1, wherein the recognition means recognizes the position of a corner present on the front side of the moving object based on an amount of change in a feature within the image. (Supplementary Note 3) The moving object identification system according to Supplementary Note 1, wherein the recognition means inputs an image of the moving object into a machine learning machine that has trained on images of multiple moving objects to recognize the position of a corner present on the front side of the moving object. (Supplementary Note 4) The moving target identification system according to Supplementary Note 1, wherein a plurality of pieces of the similar data are registered in dictionary data, and the search means searches for the similar data similar to the image of the moving target by identifying similar data in which the height of the imaging device acquired from the dictionary data is similar and the error in the positions of the corners present on the front surface of the moving target is small. (Supplementary Note 5) The moving target identification system according to Supplementary Note 1, wherein the state of the moving target is the three-dimensional posture and position of the moving target. (Supplementary Note 6) The moving target identification system according to Supplementary Note 1, wherein the moving target has a rectangular parallelepiped shape, the number of positions of the corners present on the front surface of the moving target is four, and the number of positions of the corners present on the rear surface of the moving target is four. (Supplementary Note 7) The moving target identification system according to Supplementary Note 1, wherein the holding means is a fork, the moving target is a pallet of a uniform size, and the imaging device is an RGB-D camera.(Supplementary Note 8) The moving target identification system according to Supplementary Note 5, wherein the identification means identifies the three-dimensional posture and position of the moving target by solving a PnP problem using the recognized position of a corner present on the front of the moving target and the estimated position of a corner present on the rear of the moving target. (Supplementary Note 9) A moving target identification device comprising: a holding means for holding a moving target; a height acquisition means for acquiring the height of an imaging device attached to the holding means; a recognition means for recognizing the position of a corner present on the front of the moving target imaged using the imaging device; a search means for searching for similar data similar to an image of the moving target according to the acquired height of the imaging device and the recognized position of the corner present on the front of the moving target; an estimation means for estimating the position of a corner present on the rear of the moving target according to the searched similar data; and an identification means for identifying the state of the moving target according to the recognized position of the corner present on the front of the moving target and the estimated position of the corner present on the rear of the moving target. (Supplementary Note 10) The moving target identification device according to Supplementary Note 9, wherein the recognition means recognizes the position of a corner present on the front of the moving target according to an amount of change in a feature in an image. (Supplementary Note 11) The moving target identification device according to Supplementary Note 9, wherein the recognition means inputs an image of the moving target into a machine learning machine that has learned images of a plurality of moving targets and recognizes the positions of corners present on the front surface of the moving target. (Supplementary Note 12) The moving target identification device according to Supplementary Note 9, wherein a plurality of pieces of the similar data are registered in dictionary data, and the search means searches for the similar data similar to the image of the moving target by identifying similar data that has a similar height of the imaging device acquired from the dictionary data and has little error in the positions of corners present on the front surface of the moving target. (Supplementary Note 13) The moving target identification device according to Supplementary Note 9, wherein the state of the moving target is the three-dimensional posture and position of the moving target. (Supplementary Note 14) The moving target identification device according to Supplementary Note 9, wherein the moving target has a rectangular parallelepiped shape, the number of positions of corners present on the front surface of the moving target is four, and the number of positions of corners present on the rear surface of the moving target is four.(Supplementary Note 15) The moving object identification device according to Supplementary Note 9, wherein the holding means is a fork, the moving object is a pallet of a fixed size, and the imaging device is an RGB-D camera. (Supplementary Note 16) The moving object identification device according to Supplementary Note 13, wherein the identification means identifies the three-dimensional posture and position of the moving object by solving a PnP problem using the positions of corners present on the front of the recognized moving object and the estimated positions of corners present on the rear of the moving object. (Supplementary Note 17) A target identification method comprising: capturing an image of the object using an imaging device; acquiring a height of the imaging device; recognizing the positions of corners present on the front of the object imaged using the imaging device; searching for similar data similar to the image of the object according to the acquired height of the imaging device and the recognized positions of the corners present on the front of the object; estimating the positions of corners present on the rear of the object according to the searched similar data; and identifying the state of the object according to the positions of corners present on the front of the recognized object and the estimated positions of corners present on the rear of the object. (Supplementary Note 18) The object identification method according to Supplementary Note 17, wherein the recognition recognizes the positions of corners present on the front side of the object according to a change in features within the image. (Supplementary Note 19) The object identification method according to Supplementary Note 17, wherein the recognition recognizes the positions of corners present on the front side of the object by inputting an image of the object into a machine learning machine that has learned images of multiple objects. (Supplementary Note 20) The object identification method according to Supplementary Note 17, wherein a plurality of pieces of similar data are registered in dictionary data, and the search searches for similar data similar to the image of the object by identifying similar data obtained from the dictionary data that has a similar height of the imaging device and has small errors in the positions of corners present on the front side of the object. (Supplementary Note 21) The object identification method according to Supplementary Note 17, wherein the state of the object is the three-dimensional posture and position of the object. (Supplementary Note 22) The object identification method according to Supplementary Note 17, wherein the object has a rectangular parallelepiped shape, the number of positions of corners present on the front side of the object is four, and the number of positions of corners present on the rear side of the object is four. (Supplementary Note 23) The object identification method according to Supplementary Note 17, wherein the object is a palette having a constant size, and the imaging device is an RGB-D camera.(Supplementary Note 24) The object identification method according to Supplementary Note 21, wherein the identification identifies the three-dimensional posture and position of the object by solving a PnP problem using the positions of corners present on the front side of the recognized object and the estimated positions of corners present on the rear side of the object. (Supplementary Note 25) A non-transitory computer-readable medium storing a program that causes an information processing device to perform the following steps: capture an image of the object using an imaging device; acquire the height of the imaging device; recognize the positions of corners present on the front side of the object imaged using the imaging device; search for similar data that is similar to the image of the object according to the acquired height of the imaging device and the positions of the corners present on the front side of the recognized object; estimate the positions of corners present on the rear side of the object according to the searched similar data; and identify the state of the object according to the positions of the corners present on the front side of the recognized object and the estimated positions of corners present on the rear side of the object. (Supplementary Note 26) A non-transitory computer-readable medium storing the program according to Supplementary Note 25, wherein the recognition recognizes the positions of corners present on the front side of the object according to an amount of change in a feature within the image. (Supplementary Note 27) A non-transitory computer-readable medium storing the program of Supplementary Note 25, wherein the recognition involves inputting an image of the target into a machine learning machine that has learned images of multiple targets and recognizing the positions of corners present on the front surface of the target. (Supplementary Note 28) A non-transitory computer-readable medium storing the program of Supplementary Note 25, wherein a plurality of pieces of the similar data are registered in dictionary data, and the search involves searching for similar data similar to the image of the target by identifying similar data that has a similar height of the imaging device acquired from the dictionary data and has small errors in the positions of corners present on the front surface of the target. (Supplementary Note 29) A non-transitory computer-readable medium storing the program of Supplementary Note 25, wherein the state of the target is the three-dimensional posture and position of the target. (Supplementary Note 30) A non-transitory computer-readable medium storing the program of Supplementary Note 25, wherein the target has a rectangular parallelepiped shape, and the number of positions of corners present on the front surface of the target is four, and the number of positions of corners present on the rear surface of the target is four.(Supplementary Note 31) A non-transitory computer-readable medium storing the program according to claim 25, wherein the object is a palette of a fixed size, and the imaging device is an RGB-D camera. (Supplementary Note 32) A non-transitory computer-readable medium storing the program according to Supplementary Note 29, wherein the identification identifies the three-dimensional posture and position of the object by solving a PnP problem using the positions of recognized corners present on the front surface of the object and estimated positions of corners present on the rear surface of the object.

[0057] 100 Moving object identification system, 101 Moving object, 102 Holding unit, 103 Imaging device, 104 Sensor, 105 Information processing device, 201 Image acquisition unit, 202 Height acquisition unit, 203 Recognition unit, 204 Search unit, 205 Estimation unit, 206 Object identification unit, 400 Moving object identification system, 401 Dictionary data unit, 701 Moving object

Claims

1. holding means for holding the object to be moved; height acquisition means for acquiring the height of the imaging device attached to the holding means; recognition means for recognizing the position of the corner existing on the front surface of the object to be moved imaged by the imaging device; search means for searching for similar data similar to the image of the object to be moved according to the acquired height of the imaging device and the position of the corner existing on the front surface of the recognized object to be moved; estimation means for estimating the position of the corner existing on the rear surface of the object to be moved according to the searched similar data; specification means for specifying the state of the object to be moved according to the position of the corner existing on the front surface of the recognized object to be moved and the position of the corner existing on the rear surface of the estimated object to be moved; An object to be moved specifying system comprising:

2. The object to be moved specifying system according to claim 1, wherein the recognition means recognizes the position of the corner existing on the front surface of the object to be moved according to the amount of change of features in the image.

3. The object to be moved specifying system according to claim 1, wherein the recognition means inputs the image of the object to be moved to a machine learning device that has learned images of a plurality of objects to be moved to recognize the position of the corner existing on the front surface of the object to be moved.

4. A plurality of the similar data are registered in the dictionary data, The object to be moved specifying system according to claim 1, wherein the search means searches for the similar data similar to the image of the object to be moved by specifying similar data in the dictionary data, in which the acquired height of the imaging device is similar and the error in the position of the corner existing on the front surface of the object to be moved is small.

5. The object to be moved specifying system according to claim 1, wherein the state of the object to be moved is the three-dimensional posture and position of the object to be moved.

6. The object to be moved has a rectangular parallelepiped shape, The number of positions of the corners existing on the front surface of the object to be moved is four, The object to be moved specifying system according to claim 1, wherein the number of positions of the corners existing on the rear surface of the object to be moved is four.

7. The object to be moved specifying system according to claim 1, wherein the holding means is a fork, the object to be moved is a pallet having a constant size, and the imaging device is an RGB-D camera.

8. holding means for holding the object to be moved; height acquisition means for acquiring the height of the imaging device attached to the holding means; recognition means for recognizing the position of the corner existing on the front surface of the object to be moved imaged by the imaging device; Search means for searching for similar data similar to the image of the moving object according to the height of the obtained imaging device and the position of the corner existing on the front surface of the recognized moving object; Estimation means for estimating the position of the corner existing on the rear surface of the moving object according to the searched similar data; Specification means for specifying the state of the moving object according to the position of the corner existing on the front surface of the recognized moving object and the position of the corner existing on the rear surface of the estimated moving object; A moving object specifying device comprising:

9. An image of an object is captured using an imaging device, The height of the imaging device is obtained, The position of the corner existing on the front surface of the object imaged using the imaging device is recognized, According to the obtained height of the imaging device and the position of the corner existing on the front surface of the recognized object, similar data similar to the image of the object is searched, The position of the corner existing on the rear surface of the object is estimated according to the searched similar data, A target specifying method for specifying the state of the target according to the position of the corner existing on the front surface of the recognized target and the position of the corner existing on the rear surface of the estimated target.

10. An image of an object is captured using an imaging device, The height of the imaging device is obtained, The position of the corner existing on the front surface of the object imaged using the imaging device is recognized, According to the obtained height of the imaging device and the position of the corner existing on the front surface of the recognized object, similar data similar to the image of the object is searched, The position of the corner existing on the rear surface of the object is estimated according to the searched similar data, A program for causing an information processing device to execute specifying the state of the object according to the position of the corner existing on the front surface of the recognized object and the position of the corner existing on the rear surface of the estimated object.