Image processing device

The image processing device enhances three-dimensional position information acquisition by using bottom contour information to reduce computational load and maintain accuracy over a wide range, addressing the limitations of existing technologies in collaborative environments.

JP2025179461APending Publication Date: 2025-12-10SOKEN CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024086222
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-28
Publication Date
2025-12-10

AI Technical Summary

Technical Problem

Existing technologies for acquiring three-dimensional position information of dynamic objects in collaborative environments require significant computational resources and struggle to maintain accuracy over a wide acquisition range, especially when the target object is far from the camera.

Method used

An image processing device that utilizes predetermined first information about the bottom contour of a target object, acquired from a captured image, to output three-dimensional position information of the target object's center, reducing the computational load by using two-dimensional information matching processes.

Benefits of technology

The device achieves high accuracy and wide-range three-dimensional position information with reduced computational requirements, suitable for real-time processing in collaborative environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025179461000001_ABST
    Figure 2025179461000001_ABST
Patent Text Reader

Abstract

To obtain, with a small amount of computation. three-dimensional position information of a target object.SOLUTION: There is provided an image processing device comprising: a storage unit for storing predetermined first information regarding a bottom surface outline of a target object; and a computation unit configured to acquire second information regarding the bottom surface outline of the target object from an image of the target object captured by a camera and output three-dimensional position information on a center of the bottom surface of the target object obtained based on the first information and the second information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device. [Background technology]

[0002] For example, Patent Document 1 discloses a technology that uses a simulation model of a target object to extract image features that distinguish between multiple types of postures, and acquires information about the position and posture of the target object from an image of the target object captured using a camera. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2023-121899 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the technology of Patent Document 1 leaves room for improvement in that the amount of calculation required to obtain three-dimensional position information of a target object is large. [Means for solving the problem]

[0005] In order to solve the above-mentioned problems, an image processing device according to one embodiment has a memory unit that stores predetermined first information regarding the bottom contour of a target object, and a calculation unit that acquires second information regarding the bottom contour of the target object from an image of the target object captured by a camera, and outputs three-dimensional position information of the center of the bottom of the target object obtained based on the first information and the second information. [Effects of the Invention]

[0006] According to an image processing device according to an embodiment, three-dimensional position information of a target object can be obtained with a small amount of calculation. [Brief explanation of the drawings]

[0007] [Figure 1] 1 is a block diagram illustrating a configuration of an image processing device according to an embodiment. [Figure 2] 10A and 10B are schematic diagrams illustrating a process of acquiring three-dimensional position information of the center of the bottom surface of a target object by an image processing device according to an embodiment. [Figure 3] 10A and 10B are schematic diagrams illustrating a process of comparing first information and second information related to a bottom surface contour by an image processing device according to an embodiment. [Figure 4] 4 is a schematic diagram showing a point cloud representing first information and second information relating to the bottom surface contour. FIG. [Figure 5] FIG. 10 is a schematic diagram showing an extended region having an area larger than the bottom surface region of a target object. [Figure 6] 10 is a flowchart illustrating a process performed by an image processing device according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0008] An embodiment of the present invention will now be described with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configuration are designated by the same reference numerals, and redundant description will be omitted.

[0009] <Configuration of an image processing device according to an embodiment> Fig. 1 is a block diagram showing the configuration of an image processing device 100 according to an embodiment. Fig. 2 is a schematic diagram showing the process of acquiring three-dimensional position information Ct of the bottom center 32 of a target object 3 by the image processing device 100. Fig. 3 is a schematic diagram showing the process of matching first information C1 and second information C2 related to the bottom contour 33 by the image processing device 100. Fig. 4 is a schematic diagram showing point cloud information representing the first information C1 and second information C2 related to the bottom contour 33. Fig. 5 is a schematic diagram showing an extended region 34 having an area larger than the bottom region 31 of the target object 3.

[0010] The image processing device 100 is a device that outputs three-dimensional position information of a target object 3 acquired using an image Im1 captured by a camera. The target object 3 is a mobile object such as an AGV (Automatic Guided Vehicle), an AMR (Autonomous Mobile Robot), a transport vehicle, or a forklift that moves on the floor of a factory, warehouse, or the like.

[0011] In recent years, collaborative environments have been attracting attention in factories, where they are expected to improve work efficiency by allowing the coexistence of robot workspaces and workspaces of pedestrians, manned guided vehicles, etc., without separating them. For a robot to move autonomously in a collaborative environment without compromising productivity, it is necessary to grasp the positions of dynamic objects that could become obstacles over a wide area. To meet this requirement, technology that uses cameras installed on pillars or ceilings in factories to acquire three-dimensional position information of dynamic objects is expected.

[0012] A known technology for acquiring three-dimensional position information of a dynamic object (hereinafter referred to as a target object) from which three-dimensional position information is to be acquired using images captured by a camera is to convert the captured images into three-dimensional point cloud information using a stereo camera and compare it with a model of the target object. However, this technology has room for improvement in that it is difficult to achieve both high accuracy and a wide acquisition range for three-dimensional position information because the accuracy of acquiring the three-dimensional position information decreases when the target object is located far away from the camera.

[0013] Furthermore, a technique for acquiring three-dimensional position information of a target object using an image captured by a camera has been disclosed in which a simulation model of the target object, the appearance of which changes due to a change in posture, is compared with the target object shown in the image captured by the camera. However, this technique requires time and effort to create an elaborate model that includes color, pattern, etc., and there is room for improvement in that the amount of calculation required to acquire the three-dimensional position information of the target object is large.

[0014] The image processing device 100 according to this embodiment has a memory unit 1 that stores predetermined first information C1 related to the bottom contour of the target object 3. The image processing device 100 also has a calculation unit 2 that acquires second information C2 related to the bottom contour of the target object 3 from an image Im1 captured by a camera of the target object 3, and outputs three-dimensional position information Ct of the bottom center 32 of the target object 3 obtained based on the first information C1 and the second information C2. The three-dimensional position information Ct of the bottom center 32 of the target object 3 corresponds to the three-dimensional position information of the target object 3.

[0015] When a target object 3 moves on the floor of a factory, warehouse, or the like, the bottom surface of the target object 3 is approximately parallel to the floor surface. The orientation of the bottom surface of the target object 3 remains almost unchanged even when the target object 3 moves on the floor surface. Therefore, the image processing device 100 can acquire three-dimensional position information Ct of the bottom center 32 of the target object 3 with high accuracy by performing a process such as matching the first information C1 with the second information C2. Furthermore, even when the target object 3 is located far from the camera, the image processing device 100 can suppress a decrease in the accuracy of acquiring three-dimensional position information compared to a technology using a stereo camera. As described above, the image processing device 100 can achieve both high accuracy and a wide acquisition range for three-dimensional position information.

[0016] Furthermore, the image processing device 100 performs the matching process using information about the bottom surface contour 33, which is two-dimensional information. This reduces the search space and the amount of calculations required compared to when performing the matching process using an elaborate three-dimensional model of the target object 3, including color, pattern, etc. In addition, because the first information C1 is information about the bottom surface contour 33, creating the information does not require time or effort compared to three-dimensional shape information of the entire target object 3, and the capacity of the storage unit 1 required to store the information is reduced.

[0017] As described above, the image processing device 100 can acquire three-dimensional position information of the target object 3 with high accuracy and over a wide range with a small amount of calculation. For this reason, the image processing device 100 is particularly suitable for acquiring three-dimensional position information of the target object 3 with high accuracy and in real time using a captured image Im1 obtained by a fixed camera installed in a high place such as the ceiling of a factory or warehouse. Note that the image processing device 100 may perform a process other than the process of comparing the first information C1 with the second information C2 as a process for obtaining three-dimensional position information Ct of the bottom center 32 of the target object 3, as long as it is based on the first information C1 and the second information C2.

[0018] 1, the storage unit 1 stores information such as annotation information An, camera parameters Pm, area extension information Mg, past information Pd, and first information C1. The calculation unit 2 has functions such as a calibration unit 21, a correction unit 22, a rectangular area extraction unit 23, a second information acquisition unit 24, a bottom center acquisition unit 25, an extension area acquisition unit 26, and an output unit 27.

[0019] The image processing device 100 realizes the functions of the storage unit 1 using a memory such as a read-only memory (ROM), a hard disk drive (HDD), or a solid state drive (SSD). The image processing device 100 also performs various processes by executing instruction codes stored in the memory using an electronic circuit, or by using an electronic circuit designed for a specific application, to realize the functions of the calculation unit 2. Examples of the electronic circuit include a central processing unit (CPU), a field programmable gate array (FPGA), and an application specific integrated circuit (ASIC). However, some of the functions of the image processing device 100 may be realized by an external device other than the image processing device 100, or may be realized by distributed processing between the image processing device 100 and the external device. Examples of the external device include a personal computer (PC), a server, etc.

[0020] The annotation information An is information indicating the correspondence between the coordinates of each pixel constituting the image Im1 captured by the camera and the position in the space 200 (see FIG. 2) where the image is actually captured.

[0021] The camera parameters Pm are information including extrinsic parameters, intrinsic parameters, distortion coefficients, etc. The extrinsic parameters are information about the position and orientation of the camera used to convert the camera coordinate system into the world coordinate system. The intrinsic parameters are information about the focal length, optical center, shear coefficient, etc. of the lens of the camera used to convert the world coordinate system into the camera coordinate system. The distortion coefficients are information used to correct distortion of the image Im1 captured by the camera.

[0022] The region expansion information Mg is information about the expansion region 34 having an area larger than the bottom region 31 of the target object 3. The region expansion information Mg is, for example, predetermined information such as an additional area or an expansion rate relative to the area of ​​the bottom region 31. Alternatively, the region expansion information Mg may be information for determining the expansion region 34 depending on the time it takes from when the camera starts capturing the target object 3 until the computation unit 2 inputs the captured image Im1. Such information for determining the expansion region 34 includes information such as the time it takes from when the camera starts capturing the target object 3 until the computation unit 2 inputs the captured image Im1, the moving direction 35 of the target object 3, and the moving speed of the target object 3.

[0023] The past information Pd is past information relating to at least one of the position, posture, and movement speed of the target object 3. This past information is, for example, information relating to at least one of the position, posture, and movement speed of the target object 3 stored in the storage unit 1 when the image processing device 100 most recently finished acquiring three-dimensional position information. Alternatively, the past information may be information relating to at least one of the position, posture, and movement speed of the target object 3 in the immediately preceding image capture performed by the camera at a frame cycle.

[0024] The first information C1 is information that is determined in advance based on CAD (Computer Aided Design) information of the target object 3 or measurement results of the bottom surface of the target object 3. For example, the first information C1 is expressed by point cloud information having two-dimensional coordinates. The two-dimensional coordinates are coordinate information that indicates the two-dimensional position of each of a plurality of pixels that form the bottom contour 33a in the bottom region 31a of the target object 3a in a virtual image corresponding to the captured image Im1.

[0025] The calibration unit 21 has a functional configuration that executes processing to acquire camera parameters Pm before the image processing device 100 performs processing to acquire three-dimensional position information of the target object 3. For example, the calibration unit 21 inputs an image Im1 captured by a camera of a calibration board. The calibration board is a flat plate on which a grid pattern with known dimensions is provided. Based on the captured image Im1 of the calibration board, the calibration unit 21 calculates a rotation matrix and a translation vector as external parameters, a focal length and an optical center as internal parameters, and a distortion coefficient. The calibration unit 21 passes the acquired results of the camera parameters Pm to the storage unit 1 for storage.

[0026] The correction unit 22 is a functional component that executes image processing to correct distortion in an image Im1 captured by a camera using camera parameters Pm stored in the storage unit 1 during processing by the image processing device 100 to acquire three-dimensional position information of the target object 3. The correction by the correction unit 22 results in a corrected image Im2 in which distortion aberrations of lenses such as barrel-shaped or pincushion-shaped lenses are substantially corrected. The image processing device 100 can acquire three-dimensional position information Ct of the bottom center 32 of the target object 3 with high accuracy by performing various processes on the corrected image Im2 using the rectangular area extraction unit 23, second information acquisition unit 24, and bottom center acquisition unit 25. The correction unit 22 passes the corrected image Im2 after correction to the rectangular area extraction unit 23.

[0027] The rectangular area extraction unit 23 is a functional component that executes image processing to extract, from the corrected image Im2, a rectangular area 30 that includes the target object 3 that appears in the corrected image Im2. The rectangular area 30 is a partial image area in the corrected image Im2, and is a rectangular image area that shows the entire target object 3. The rectangular area extraction unit 23 can use, as the image processing to extract the rectangular area 30, AI (Artificial Intelligence) image processing or the like that recognizes the target object 3 that appears in the corrected image Im2.

[0028] There is no limit to the size of the rectangular region 30 as long as it is large enough to capture the entire target object 3. However, from the perspective of reducing the amount of calculation, it is preferable that the size of the rectangular region 30 be as small as possible. The rectangular region extraction unit 23 passes information Sq about the extracted rectangular region 30 to the second information acquisition unit 24. The information Sq about the rectangular region 30 is, for example, coordinate information indicating two or more of the four corners of the rectangular region 30.

[0029] Corrected image Im2 shown in FIG. 2 shows target objects 3, which include target objects 3-1, 3-2, and 3-3, and an obstacle 4, which is an object other than target object 3. A first processing area IP1 in FIG. 2 is an image area obtained by enlarging area A in corrected image Im2. A rectangular area 30 includes target object 3-1. Note that the following description will be given taking as an example the process of acquiring three-dimensional position information Ct of the bottom center 32 of target object 3-1, but similar acquisition processes can also be applied to target objects 3-2 and 3-3.

[0030] The second information acquisition unit 24 has a functional configuration that executes image processing to acquire second information C2 related to a bottom contour 33 that indicates the contour of the bottom region 31 of the target object 3-1 reflected in the rectangular region 30. However, the second information acquisition unit 24 may execute image processing to simultaneously extract the bottom contours 33 of all the target objects 3 reflected in the corrected image Im2, not limited to the rectangular region 30. For example, the second information C2 is expressed by point cloud information having two-dimensional coordinates. The two-dimensional coordinates are coordinate information that indicates the two-dimensional position of each of multiple pixels that constitute the bottom contour 33 in the bottom region 31 of the target object 3-1 in the corrected image Im2.

[0031] The second processing area IP2 shown in Figures 2 and 3 shows the result of image processing performed on the first processing area IP1 by the second information acquisition unit 24 to acquire second information C2. The second processing area IP2 in Figure 3 is an enlarged version of the second processing area IP2 in Figure 2. A bottom area 31 is extracted inside a rectangular area 30. In the second processing area IP2 in Figure 3, a bottom contour 33 is further extracted. The second information acquisition unit 24 can use, for example, the following methods (1) and (2) as a method for extracting the bottom contour 33.

[0032] (1) A method of adding features to the target object 3 to distinguish the bottom surface In this method, a mark of a predetermined color that makes it easy to identify the bottom surface is applied in advance to the lower end of the side of the target object 3, and the bottom contour 33 is extracted using the mark as a guide to obtain the second information C2. The mark may be written on the lower end of the side of the target object 3 with a pen or the like, or tape may be attached to the lower end of the side of the target object 3. The second information acquisition unit 24 can easily extract the bottom contour 33 by performing image processing such as color discrimination. It is preferable to select a color with high brightness and saturation, or a color complementary to the color of the target object 3, as the predetermined color used for the mark.

[0033] (2) Using segmentation AI In this method, information such as rectangular areas, points, and text is provided to the second information acquisition unit 24 as hints for extracting the bottom contour 33, and image processing using segmentation AI is performed to extract the bottom contour 33 and acquire the second information C2. This method makes it easier for the second information acquisition unit 24 to separate the target object 3 from the background image area other than the target object 3. The second information acquisition unit 24 can use MobileSAM (see, for example, https: / / github.com / ChaoningZhang / MobileSAM) as the segmentation AI. If the lighting conditions of the environment in which the camera captures images do not change easily over time, the second information acquisition unit 24 may calculate the difference from a background image captured in advance, thereby making it easier to separate the target object 3 from the background image area other than the target object 3.

[0034] The bottom center acquisition unit 25 is functionally configured to acquire three-dimensional position information Ct of the bottom center 32 of the target object 3 by calculation based on the first information C1 and the second information C2. The bottom center acquisition unit 25 executes a process of matching the geometric shape of the second information C2 projected onto the world coordinate system with the geometric shape corresponding to the first information C1. As a result, the bottom center acquisition unit 25 can calculate the three-dimensional coordinates of the bottom center 32 of the target object 3-1 reflected in the corrected image Im2 and acquire the three-dimensional position information Ct of the bottom center 32. The world coordinates Wc shown in FIG. 3 represent the second information C2 projected onto the world coordinate system. The world coordinates Wc are expressed in terms of directions using mutually orthogonal x-, y-, and z-axes. The simulation model Me represents the target object 3a, the bottom area 31a, and the bottom contour 33a in the virtual space of the simulation. The bottom center acquisition unit 25 performs a process of comparing the bottom contour 33 with the bottom contour 33a in the world coordinate Wc. The bottom center acquisition unit 25 passes the acquired three-dimensional position information Ct of the bottom center 32 to the output unit 27.

[0035] The matching process performed by the bottom center acquisition unit 25 will now be described in more detail. As shown in the first process Ps1 in FIG. 4, the bottom center acquisition unit 25 first searches for a pair of points from one of the point groups representing the first information C1 and the point group representing the second information C2 that are closest to the other point group. This search process can be performed using a so-called nearest point search process. Next, as shown in the second process Ps2 in FIG. 4, the bottom center acquisition unit 25 performs rigid body transformation so that the pairs selected as a result of the search best match each other. This transformation process can be performed using singular value decomposition (SVD). The second information C2t shown in the second process Ps2 in FIG. 4 represents a point group resulting from rigid body transformation of the second information C2 using the singular value decomposition process. The bottom center acquisition unit 25 repeatedly performs the first process Ps1 and the second process Ps2. The bottom center acquisition unit 25 determines whether the first process Ps1 and the second process Ps2 have been executed a predetermined number of times, or whether the distance between the entire point group representing the first information C1 and the entire point group representing the second information C2 has converged to a minimum value. If it is determined that the process has been executed a predetermined number of times or has converged to a minimum value, the bottom center acquisition unit 25 ends the matching process and ends the process of acquiring the three-dimensional position information Ct of the bottom center 32.

[0036] The image processing device 100 can execute a matching process between the first information C1 and the second information C2 by solving a general-purpose two-dimensional point cloud matching problem using the first information C1 and the second information C2 expressed by point cloud information having two-dimensional coordinates via the bottom center acquisition unit 25. This reduces the search space and the amount of calculation compared to when performing a matching process by solving a three-dimensional point cloud matching problem. An iterative closest point (ICP) algorithm or the like can be used as an algorithm for solving the two-dimensional point cloud matching problem.

[0037] The bottom center acquisition unit 25 can acquire three-dimensional position information Ct of the bottom center 32 of the target object 3-1 using past information Pd related to at least one of the position, orientation, and movement speed of the target object 3-1. For example, in the iterative closest point algorithm, the number of iterative calculations until a convergent solution is obtained depends on the initial values ​​of the position, orientation, and movement speed of the input point group. If past information Pd related to at least one of the position, orientation, and movement speed of the target object 3-1 is available, the number of iterative calculations in the matching process can be reduced by setting this past information Pd as the initial value, thereby shortening the calculation time.

[0038] The processing result RS shown in Fig. 2 displays the three-dimensional position of the bottom center 32 of the target object 3-1 in the space 200 that is actually photographed. A section 201 represents an area that partitions the floor surface of the space 200. For example, an external device that receives the three-dimensional position information Ct of the bottom center 32 output from the image processing device 100 displays the processing result RS. For ease of understanding, the processing result RS displays the bottom center 32 as a cylindrical object.

[0039] The extended area acquisition unit 26 has a functional configuration that acquires information Ea about the extended area 34 based on the area expansion information Mg stored in the storage unit 1. The information Ea about the extended area 34 is, for example, coordinate information indicating two or more of the four corners of the extended area 34. The extended area acquisition unit 26 passes the acquired information Ea about the extended area 34 to the output unit 27.

[0040] For example, consider a case where the three-dimensional position information Ct of the bottom center 32 acquired by the image processing device 100 is used to prevent the target object 3-1 from coming into contact with an obstacle 4, a person, or the like present in a factory. When an external device that has received the three-dimensional position information Ct estimates the position of the outer edge of the target object 3-1 from the three-dimensional position information Ct and information related to the size of the target object 3-1, the probability of contact with the target object 3-1 may increase when the obstacle 4, a person, or the like is moving faster than expected.

[0041] 5, the extended region acquisition unit 26 of the image processing device 100 further outputs information Ea about the extended region 34, which is larger in area than the bottom region 31 of the target object 3. This allows the external device that has received the information Ea about the extended region 34 to estimate the position of the outer edge of the target object 3-1 with a higher safety margin than when using information about the size of the target object 3-1. As a result, the probability of contact between the target object 3-1 and an obstacle 4, a person, etc. is reduced, improving safety.

[0042] Furthermore, for example, depending on the data communication environment between the camera and the image processing device, or processing conditions such as data encoding or decoding, the time it takes from when the camera starts photographing the target object 3 until the calculation unit 2 inputs the photographed image Im1 may be long.

[0043] The extended region 34 in the image processing device 100 is a region obtained by extending the bottom region 31 of the target object 3 in the movement direction 35, depending on the time it takes from when the camera starts capturing the target object 3 until the calculation unit 2 inputs the captured image Im1, and the movement speed of the target object 3 in the movement direction 35. This allows the external device that has input the information Ea related to the extended region 34 to estimate the position of the outer edge of the target object 3-1 with an even higher safety margin. As a result, the probability that the target object 3-1 will come into contact with an obstacle 4, a person, etc. is reduced, further improving safety.

[0044] The method for determining the extended region 34 will be described in more detail. As shown in FIG. 5, an area determined based on predetermined region extension information Mg, such as an additional area or an extension rate relative to the area of ​​the bottom region 31 of the target object 3, is defined as extended region 34-1. An area determined based on the time it takes from when the camera starts capturing the target object 3 until the calculation unit 2 inputs the captured image Im1, and the moving speed of the target object 3 in the moving direction 35, is defined as extended region 34-2. The extended region 34-2 includes the target object 3' after it has moved. The image processing device 100 can determine the union of the extended region 34-1 and the extended region 34-2 as the extended region 34.

[0045] In order to predict the future position of the target object 3, the image processing device 100 may obtain the movement trajectory of the target object 3 using image tracking, and then calculate the speed from the rate of change (differential) of the position to estimate the movement amount of the target object 3. If the position estimate value contains noise, the image processing device 100 may combine a statistical method such as a Kalman filter or RANSAC (Random Sample Consensus) in the process of estimating the movement amount of the target object 3.

[0046] <Processing by an image processing device according to an embodiment> Fig. 6 is a flowchart showing the processing by the image processing device 100. The image processing device 100 starts the processing in Fig. 6 when, for example, a captured image Im1 is input from a camera.

[0047] First, in step 1, the image processing device 100 executes image processing by the correction unit 22 to correct distortion of the image Im1 captured by the camera, using the camera parameters Pm stored in the storage unit 1. The correction unit 22 passes the corrected image Im2 to the rectangular area extraction unit 23.

[0048] Next, in step 2, the image processing device 100 acquires, via the bottom center acquisition unit 25, first information C1 relating to the bottom contour 33 of the target object 3 stored in the storage unit 1. Note that the order of steps 1 and 2 may be reversed. Also, steps 1 and 2 may be performed in parallel.

[0049] Next, in step 3, the image processing device 100 executes image processing to extract, from the corrected image Im2, a rectangular area 30 including the target object 3 shown in the corrected image Im2, using the rectangular area extraction unit 23. The rectangular area extraction unit 23 passes information Sq related to the extracted rectangular area 30 to the second information acquisition unit 24.

[0050] Next, in step 4, the image processing device 100 executes image processing to acquire, by the second information acquisition unit 24, second information C2 relating to the bottom surface contour 33 indicating the contour of the bottom surface region 31 of the target object 3 reflected in the rectangular region 30. The second information acquisition unit 24 passes the acquired second information C2 to the bottom surface center acquisition unit 25.

[0051] Next, in step 5, the image processing device 100 projects the second information C2 onto the world coordinate system using the bottom center acquisition unit 25. Note that step 5 may be performed immediately before step 8.

[0052] Next, in step 6, the image processing device 100 determines, by the bottom center acquisition unit 25, whether or not past information Pd relating to at least one of the position, orientation, and moving speed of the target object 3 is stored in the storage unit 1.

[0053] If it is determined in step 6 that the past information Pd of the target object 3 is not stored (step 6, NO), the image processing device 100 proceeds to step 8. On the other hand, if it is determined in step 6 that the past information Pd of the target object 3 is stored (step 6, YES), in step 7, the image processing device 100 sets the past information Pd of the target object 3 to the initial value in the matching process between the first information C1 and the second information C2 by the bottom surface center acquisition unit 25.

[0054] Next, in step 8, the image processing device 100 causes the bottom center acquisition unit 25 to execute a process of matching the geometric shape of the second information C2 with the geometric shape corresponding to the first information C1. As a result, the bottom center acquisition unit 25 calculates the three-dimensional coordinates of the bottom center 32 of the target object 3-1 reflected in the corrected image Im2 and acquires three-dimensional position information Ct of the bottom center 32. The bottom center acquisition unit 25 passes the acquired three-dimensional position information Ct of the bottom center 32 to the output unit 27.

[0055] Next, in step 9, the image processing device 100 causes the extended area acquisition unit 26 to acquire information Ea about the extended area 34 based on the area extension information Mg stored in the storage unit 1. The extended area acquisition unit 26 passes the acquired information Ea about the extended area 34 to the output unit 27.

[0056] Subsequently, in step 10, the image processing device 100 outputs the three-dimensional position information Ct of the bottom center 32 and the information Ea about the extended area 34 to an external device via the output unit 27.

[0057] Next, in step 11, the image processing device 100 determines whether or not the three-dimensional position information Ct of the bottom center 32 and the information Ea about the extended area 34 of all the target objects 3 captured in the captured image Im1 have been output to an external device.

[0058] If it is determined in step 11 that the information has been output (step 11, YES), the image processing device 100 ends the processing. On the other hand, if it is determined that the information has not been output (step 11, NO), the image processing device 100 executes the processing from step 3 onwards for the target objects 3 for which the three-dimensional position information Ct of the bottom center 32 and the information Ea regarding the extended region 34 have not been output. The image processing device 100 repeats the processing from step 3 onwards until the three-dimensional position information Ct of the bottom center 32 and the information Ea regarding the extended region 34 for all target objects 3 captured in the captured image Im1 have been output to the external device.

[0059] In this way, the image processing device 100 can execute processing to output the three-dimensional position information Ct of the bottom center 32 and the information Ea about the extended area 34 of all target objects 3 captured in the captured image Im1 to an external device.

[0060] Although a preferred embodiment has been described in detail above, the present invention is not limited to the above-described embodiment of the present invention, and various modifications and substitutions can be made to the above-described embodiment of the present invention without departing from the scope of the claims.

[0061] All ordinal numbers, quantitative numbers, and other figures used in the description of one embodiment of the present invention are provided as examples to specifically explain the technology of the present invention, and the present invention is not limited to the illustrated figures. Furthermore, the connection relationships between components are provided as examples to specifically explain the technology of the present invention, and do not limit the connection relationships that realize the functions of the present invention. [Explanation of symbols]

[0062] 1 Storage section 2 Arithmetic section 21 Calibration section 22 Correction unit 23 Rectangular area extraction part 24 2nd Information Acquisition Department 25 Bottom center acquisition part 26 Extended area acquisition unit 27 Output section 3, 3-1, 3-2, 3-3, 3a, 3' Target object 30 rectangular area 31, 31a bottom area 32 Bottom center 33, 33a Bottom contour 34, 34-1, 34-2 Expansion Area 35 Direction of movement 4. Obstacles 100 Image processing device 200 space Section 201 Annotation information C1 1st information C2 Second Information Ea Extension Area Information Im1 image Im2 corrected image IP1 First processing area IP2 Second processing area Me Simulation Model Mg Area Extension Information Pd past information Pm camera parameters Ps1 First treatment Ps2 Second Processing RS processing result Sq Information about a rectangular area Wc world coordinates

Claims

1. a storage unit that stores predetermined first information related to a bottom surface contour of the target object; An image processing device having a calculation unit that acquires second information regarding the bottom contour of the target object from an image of the target object captured by a camera, and outputs three-dimensional position information of the bottom center of the target object obtained based on the first information and the second information.

2. The image processing device according to claim 1 , wherein the first information and the second information are expressed by point cloud information having two-dimensional coordinates.

3. The image processing device according to claim 1 , wherein the calculation unit acquires three-dimensional position information of the center of the bottom surface of the target object using past information on at least one of the position, orientation, and moving speed of the target object.

4. The image processing device according to claim 1 , wherein the calculation unit further outputs information about an extended region having an area larger than the bottom region of the target object.

5. 5. The image processing device according to claim 4, wherein the extended area is an area in which the bottom area of ​​the target object is extended in the movement direction, depending on the time it takes from when the camera starts photographing the target object until the calculation unit inputs the photographed image, and the movement speed of the target object in the movement direction.

Citation Information

Patent Citations

  • Method, system, and computer program for recognizing position and posture of object photographed by camera

    JP2023121899A