Determination device, determination system, determination method, and program

The determination device enhances object detection accuracy by using an image acquisition, detection, and estimation process, ensuring precise control and transportation of objects by autonomously traveling mobile bodies.

JP7798545B2Active Publication Date: 2026-01-14MAXELL LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021192022
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-26
Publication Date
2026-01-14
Estimated Expiration
2041-11-26

AI Technical Summary

Technical Problem

Conventional techniques for detecting the shape of an object using machine learning in autonomously traveling mobile objects often result in inaccurate inference calculations, leading to improper control of the mobile body.

Method used

A determination device that includes an image information acquisition unit, a detection unit to identify key points, a calculation unit to determine depths based on actual object sizes and depth information, and an estimation unit to estimate the object's posture, utilizing a trained model for accurate detection.

Benefits of technology

Enables precise detection of the object's orientation and posture, allowing for accurate control and transportation of the object by the mobile body.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007798545000002
    Figure 0007798545000002
  • Figure 0007798545000003
    Figure 0007798545000003
  • Figure 0007798545000004
    Figure 0007798545000004
Patent Text Reader

Abstract

To provide a determination device, a determination system, a determination method, and a program, which accurately detect the orientation of an object.SOLUTION: A determination device 12 includes: an image information acquisition unit 121 configured to acquire image information including an image I1 of an object and depth information I2 of a position corresponding to the image; a detection unit 122 configured to detect a key point, which is a specific position of the object, based on the image included in the image information; a calculation unit 123 configured to calculate the depth of the position corresponding to the detected key point based on the depth information included in the image information; and an estimating unit 124 configured to estimate the orientation of the object based on the calculated depth of the position corresponding to the key point.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a determination device, a determination system, a determination method, and a program. [Background technology]

[0002] BACKGROUND ART Conventionally, in the technical field of autonomously traveling mobile objects, a technique for detecting the shape of an object using machine learning is known (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-099383 Summary of the Invention [Problem to be solved by the invention]

[0004] Such conventional technology may be able to estimate the position and orientation of an object for autonomous driving of a mobile body, or update map information for calculating the position and orientation of a sensor. However, when detecting the shape of an object using machine learning, errors contained in the inference calculation results of the machine learning cannot be ignored. Therefore, there is a problem that the mobile body may not be controlled appropriately.

[0005] The present invention has been made in view of the above circumstances, and aims to provide a technique that can accurately detect the orientation of an object. [Means for solving the problem]

[0006] A determination device according to one aspect of the present invention includes an image information acquisition unit that acquires image information including an image of an object and depth information of a position corresponding to the image; a detection unit that detects key points that are specific positions of the object based on the image included in the image information; a calculation unit that calculates depths of positions corresponding to the detected key points based on the depth information included in the image information; and an estimation unit that estimates a posture of the object based on the calculated depths of the positions corresponding to the key points. The calculation unit calculates the depth of the position corresponding to the detected key point based on the actual size of the object and the depth information included in the image information. do.

[0007] In the determination device according to one aspect of the present invention, the estimation unit estimates, as the attitude of the object, an inclination of the object with respect to an imaging unit that captured the image.

[0008] In the determination device according to one aspect of the present invention, the object is a basket cart having wheels, and the key points include at least one of the wheels.

[0009] In addition, in a determination device according to one embodiment of the present invention, the basket cart has a bottom surface capable of carrying luggage and four of the wheels, and the detection unit detects four points that identify the bottom surface and four points that identify the positions of the four wheels, respectively, as the key points.

[0010] In addition, in the determination device according to one aspect of the present invention, the estimation unit estimates the coordinates and angle of the center point of one of the multiple sides of the object that is closest to the imaging unit that captured the image.

[0012] In addition, in the determination device according to one aspect of the present invention, the detection unit includes a trained model and detects the key points through machine learning.

[0013] In addition, a judgment system according to one aspect of the present invention includes a storage device that stores the image of the object and the position coordinates of the key points of the object, a model generation means that uses the object and the key points contained in the image stored in the storage device as training data to generate the trained model, and the above-mentioned judgment device that detects the key points using the trained model.

[0014] Furthermore, a determination method according to one aspect of the present invention includes: A determination method executed by a processor, the processor comprising: an image information acquisition step of acquiring image information including an image of an object and depth information corresponding to the image; the processor: a detection step of detecting key points, which are specific positions of the object, based on the image included in the image information; the processor: a calculation step of calculating a depth of a coordinate corresponding to the detected key point based on the depth information included in the image information; the processor: and an estimation step of estimating the posture of the object. The calculation step calculates the depth of the position corresponding to the detected key point based on the actual size of the object and the depth information included in the image information. .

[0015] A program according to one aspect of the present invention causes a computer to execute an image information acquisition step of acquiring image information including an image of an object and depth information corresponding to the image; a detection step of detecting key points that are specific positions of the object based on the image included in the image information; a calculation step of calculating depths of coordinates corresponding to the detected key points based on the depth information included in the image information; and an estimation step of estimating a posture of the object. The calculation step calculates the depth of the position corresponding to the detected key point based on the actual size of the object and the depth information included in the image information. do. [Effects of the Invention]

[0016] According to the present invention, the orientation of an object can be detected with high accuracy. [Brief explanation of the drawings]

[0017] [Figure 1] FIG. 1 is a diagram for explaining an overview of a mobile object control system according to an embodiment. [Figure 2] 1 is a diagram illustrating a moving object and an imaging device according to an embodiment. [Figure 3] FIG. 10 is a diagram illustrating an example of an angle of an object relative to an imaging unit according to the embodiment. [Figure 4] FIG. 2 is an example of a functional configuration diagram for explaining functions of a moving body according to an embodiment. [Figure 5] FIG. 2 is an example of a functional configuration diagram for explaining functions of a determination device according to an embodiment. [Figure 6] FIG. 2 is a diagram for explaining learning of the determination device according to the embodiment. [Figure 7] 1A and 1B are diagrams for explaining an example of an image of an object captured according to an embodiment and an example of a depth image. [Figure 8] FIG. 2 is a diagram illustrating an example of a function of a determination device according to an embodiment. [Figure 9] FIG. 2 is a diagram for explaining key points of an object according to the embodiment. [Figure 10] 10A and 10B are diagrams for explaining a method for calculating the depth of a key point according to an embodiment. [Figure 11] FIG. 10 is a diagram illustrating an example of key points calculated by the determination device according to the embodiment. [Figure 12] 10A to 10C are diagrams for explaining an example of a series of operations of a moving object according to an embodiment. [Figure 13] FIG. 10 is a diagram for explaining a modified example of the mobile object control system according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0018] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The embodiments described below are merely examples, and the embodiments to which the present invention is applied are not limited to the following embodiments.

[0019] [Mobile Control System] FIG. 1 is a diagram illustrating an overview of a mobile object control system according to an embodiment. An overview of the mobile object control system 1 will be described with reference to the figure. The mobile object control system 1 includes at least one mobile object 10. The mobile object 10 includes a mechanism for moving an object 20 to a predetermined position.

[0020] The mobile body 10 moves the object 20 to a destination by grasping, towing, or lifting and transporting the object 20. Lifting and transporting may mean inserting forks (claws) (not shown) provided on the mobile body 10 into the bottom of the object 20 and lifting the object 20 by raising the forks, and transporting it to a destination. Alternatively, the mobile body 10 itself may get under the object 20, lift the object 20 by a predetermined lifting mechanism, and transport it to a destination. The moving body 10 may be, for example, an AGV (Automatic Guided Vehicle) such as an automatic guided vehicle or an automatic guided robot.

[0021] The object 20 may be any object that can be transported by the moving body 10. In the following description, an example will be described in which the object 20 has a platform on which luggage can be placed and wheels. The object 20 may have wheels, and may be pulled or pushed by a person to transport the luggage placed on it. The object 20 may be, for example, a basket cart (basket cart or basket dolly), a cart, a dolly, or the like, or may be one of a plurality of moving bodies 10. In the following explanation, an example in which the object 20 is a basket cart will be explained.

[0022] If the position where the moving body 10 is located is position P1 and the position where the target object 20 is located is position P3, we will explain an example of when the moving body 10 transports the target object 20 located at position P3 to the destination position, position P4. First, the mobile object 10 moves to a position P2 based on map information. The map information may be location information estimated by a communication unit (not shown) included in the mobile object 10 based on radio waves received from an artificial satellite such as a GPS (Global Positioning System) or a predetermined base station. The communication unit may also acquire information about the position P2 by performing short-range communication with a central control device (not shown) via a wireless LAN or the like.

[0023] Position P2 is a position where the posture of object 20 can be confirmed. Specifically, position P2 is a position where the object 20 can be imaged in order to determine the posture of object 20. More specifically, position P2 may be at a distance of about 5 m (meters) from position P3 where object 20 is present.

[0024] The moving body 10 may image the object 20 from one position, or may image the object 20 from multiple positions. When the moving body 10 images the object 20 from multiple positions, the multiple positions at which the moving body 10 images the object 20 are collectively referred to as position P2. At position P2, the moving body 10 captures an image of the object 20 using the imaging device 11 provided therein. The moving body 10 determines the attitude of the object 20 based on the image capture result. The attitude of the object 20 may be the distance from the moving body 10 to the object 20, the angle of the object 20 relative to the moving body 10, the angle of the position where the object 20 is located relative to the position where the moving body 10 is located, etc.

[0025] The mobile body 10 moves to a position where it can contact the object 20, and puts the object 20 into a transportable state based on the determined posture of the object 20. The transportable state may be a state in which the object 20 is grasped, pulled, lifted up (raised), etc. After putting the object 20 into a transportable state, the mobile body 10 transports the object 20 to the destination position P4.

[0026] [Moving object] 2 is a diagram for explaining a moving object and an imaging device according to an embodiment, and a moving object 10 and an imaging device 11 will be explained with reference to the drawing. The moving body 10 is equipped with an imaging device 11. The attitude of the moving body 10 is described using a three-dimensional Cartesian coordinate system in which the imaging direction of the imaging device 11 is the y-axis, the direction directly above the imaging device 11 is the z-axis, and the direction perpendicular to the y-axis and z-axis is the x-axis. The position coordinates where the imaging device 11 exists are set as the origin (0, 0, 0).

[0027] The imaging device 11 captures an image in the y-axis direction at a predetermined angle of view. The moving body 10 moves to position P2, which is an imaging position where the object 20 is located within the predetermined angle of view, and captures an image of the object 20. Based on the captured image, the moving body 10 determines the distance to the object 20, the angle (tilt) of the object 20 relative to its own angle, and the angle of the object 20 relative to its own position, etc.

[0028] The imaging device 11 may be attached to the top of the moving body 10 as shown in the figure, or may be provided in another location such as inside the moving body 10.

[0029] 3 is a diagram showing an example of the angle of an object relative to the imaging unit according to the embodiment. The angle (tilt) of the object 20 relative to the angle of the moving body 10 will be described with reference to the diagram. The diagram is a plan view showing the positional relationship between the moving body 10 and the object 20 when the moving body 10 moves to position P2, where the object 20 is imaged. In explaining the diagram, the attitudes of the moving body 10 and the object 20 may be explained using a coordinate system consisting of an x-axis and a y-axis.

[0030] The moving body 10 moves to position P2 based on the map information. Since the map information does not include information about the posture of the moving body 10, information such as the inclination of the target object 20 relative to the moving body 10 is unknown at the time of moving to position P2.

[0031] Fig. 3(A) shows an example where the moving body 10 and the object 20 are positioned parallel to each other. Fig. 3(B) shows an example where the object 20 is positioned with a positive inclination. Fig. 3(C) shows an example where the object 20 is positioned with a negative inclination. 3A, when the moving body 10 and the object 20 are positioned parallel to each other, the angle θ of the object 20 with respect to the imaging unit 11 is approximately 0 degrees. The range of approximately 0 degrees may be from about -1 to +1 degrees, or from about -3 to +3 degrees. As shown in Fig. 3(B), when the distance between the moving body 10 and the object 20 increases toward the positive x-axis direction, the angle θ of the object 20 with respect to the imaging unit 11 is positive. As shown in Fig. 3(C), when the distance between the moving body 10 and the object 20 decreases toward the positive x-axis direction, the angle θ of the object 20 with respect to the imaging unit 11 is negative.

[0032] Based on the image captured at position P2, the mobile body 10 determines the distance to the object 20, the angle of the object 20 relative to itself, and the angle between its own position and the position where the object 20 is located. Based on the determined distance and angle, the mobile body 10 moves to the vicinity of the object 20, puts the object 20 in a transportable state, and then transports it to position P4, which is the destination position.

[0033] 4 is an example of a functional configuration diagram for explaining the functions of a moving body according to an embodiment. An example of the functional configuration of the moving body 10 will be described with reference to the same figure. The moving body 10 includes an imaging device 11, a determination device 12, a moving body control device 13, and a moving body drive device 14 as components.

[0034] The imaging device 11 acquires an image by capturing an image of the object 20. The imaging device 11 also acquires information about the depth of the object 20. The depth of the object 20 is information about the distance between the imaging device 11 and the object 20. The imaging device 11 may acquire an image by including a camera, and may acquire information about the depth by including an infrared laser, a compound eye camera, or the like. The imaging device 11 outputs acquired information about the object 20, such as an image and depth information, to the determination device 12 as image information II.

[0035] The determination device 12 determines the posture of the object 20 based on the image, depth information, etc. acquired from the imaging device 11. The determination device 12 outputs information related to the determined posture of the object 20 to the mobile object control device 13 as determination information JI.

[0036] The mobile body control device 13 controls the mobile body 10 to move its position based on information about the attitude of the object 20 determined by the determination device 12. The mobile body control device 13 includes a storage device such as a central processing unit (CPU), a graphics processing unit (GPU), a read only memory (ROM), or a random access memory (RAM), all of which are not shown. The mobile body control device 13 outputs control information CI to the mobile body drive device 14, thereby causing the mobile body 10 to move to a position where the mobile body 10 can grasp, tow, lift up, etc. the object 20. After the mobile body 10 moves to a position where the mobile body 10 can grasp, tow, lift up, etc. the object 20, the mobile body control device 13 controls the mobile body 10 to transport the object 20 to a destination position P4 by grasping, towing, lifting up, etc. the object 20.

[0037] The mobile body drive device 14 includes a transport drive unit such as a grip drive unit, a towing drive unit, or an elevation drive unit (not shown), and a movement drive unit such as a wheel drive unit or a belt drive unit. Based on a control signal CI acquired from the mobile body control device 13, the mobile body drive device 14 grips, tows, lifts up, etc. the mobile body 10, and transports the target object 20 to the destination.

[0038] [Judgment device] 5 is an example of a functional configuration diagram for explaining the functions of a determination device according to an embodiment. An example of the functional configuration of the determination device 12 will be described with reference to the same diagram. The determination device 12 includes an image information acquisition unit 121, a detection unit 122, a calculation unit 123, and an estimation unit 124. The determination device 12 includes a CPU, a GPU, a storage device such as a ROM or a RAM, etc., which are connected via a bus, and functions as a device including the image information acquisition unit 121, the detection unit 122, the calculation unit 123, and the estimation unit 124 by executing a determination program.

[0039] All or part of the functions of the determination device 12 may be realized using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field-Programmable Gate Array). In particular, it is preferable that the detection unit 122 is configured using hardware, and the calculation unit 123 and the estimation unit 124 are configured using software.

[0040] The image information acquisition unit 121 acquires image information II from the imaging device 11. The image information II includes at least an image I1 and depth information I2. The image I1 is a two-dimensional image of an object 20. The depth information I2 is information including depth information of a position corresponding to the image I1. The depth information I2 may be information such as a depth image acquired by an infrared laser, a compound eye camera, or the like.

[0041] The detection unit 122 detects the position coordinates of keypoints KP, which are specific positions of the object 20, based on the image I1 included in the image information acquired by the image information acquisition unit 121. Specifically, the detection unit 122 is provided with a pre-trained model and detects keypoints KP through machine learning. The pre-trained model provided by the detection unit 122 includes a neural network consisting of multiple layers. The multiple layers include a convolution operation unit that performs a convolution operation between pre-trained weight parameters and input data based on the image I1, and a quantization operation unit that performs a quantization operation on the result of the convolution operation. In this way, by providing a quantization operation unit in at least one layer of the neural network included in the detection unit 122, it is possible to reduce the computational load and perform detection based on the trained model on an edge device such as the mobile object 10.

[0042] Here, the key points KP are points used to identify the posture of the object 20. The key points KP may be determined in advance according to the object 20. For example, if the object 20 is a basket cart with wheels, the key points KP may include at least one wheel or a vertex of the basket cart. In the determination device 12, the detection unit 122 detects the position coordinates of at least two key points KP in order to determine the posture of the object 20. The detection unit 122 outputs information about the position coordinates of the detected key points KP to the calculation unit 123 as detection information DI.

[0043] The calculation unit 123 calculates the depth of the position corresponding to the detected keypoint KP based on the depth information I2 included in the image information. Specifically, the calculation unit 123 may acquire the depth of the position coordinates corresponding to the keypoint KP in the image I1 by referring to the depth information I2, and perform calculation based on the acquired depth. The calculation unit 123 outputs information about the calculated depth of the position corresponding to the keypoint KP to the estimation unit 124 as keypoint depth information KDI.

[0044] The estimation unit 124 estimates the orientation of the object 20 based on the depth of the position corresponding to the position coordinates of the key point KP calculated by the calculation unit 123. The orientation of the object 20 may be, for example, the inclination of the object 20 with respect to the imaging unit 11 that captured the image I1. That is, the estimation unit 124 may estimate the inclination of the object 20 with respect to the imaging unit 11 as the orientation of the object 20.

[0045] 6 is a diagram for explaining learning of the determination device according to the embodiment. An example of learning will be described with reference to the same figure. The configuration of the determination device 12 during learning will be referred to as a determination system 2. The determination system 2 is an example of a computer connected to a server equipped with a GPU, and the determination system 2 may be a system independent of the mobile object control system 1. The determination system 2 includes a storage device 30 and a model generation means 31 .

[0046] The storage device 30 stores an image I1 and position coordinate information I3 corresponding to the image I1. The position coordinate information I3 is information about the position coordinates of the key points KP of the object 20 captured in the image I1. That is, the storage device 30 stores the image I1 in which the object 20 is captured and the position coordinates of the key points KP of the object 20.

[0047] The model generation means 31 generates a trained model using information about the object 20 included in the image I1 stored in the storage device 30 and the position coordinates of the keypoints KP as training data. More specifically, the storage device 30 preferably stores images captured in various scenes including the object 20 used in the mobile object control system 1, but may also include images that do not include the object 20 to improve the generalization performance of the trained model. Furthermore, multiple learning results may be stored as checkpoints so that parameters can be switched depending on the type of object 20. The detection unit 122 included in the determination device 12 uses the trained model generated by the model generation means 31 to detect key points KP.

[0048] 7A and 7B are diagrams illustrating an example of an image of an object captured according to an embodiment and an example of a depth image. An example of an image I1 and depth information I2 will be described with reference to the figure. Fig. 7A shows an example of the image I1, and Figs. 7B and 7C show examples of the depth information I2 corresponding to the image I1.

[0049] Image I1 is an image captured by the imaging device 11 at position P2, where the moving body 10 captures an image of the object 20. Image I1 captures the object 20. The object 20 is a basket cart loaded with luggage. Depth information I2 is a depth image captured at approximately the same angle of view and approximately the same imaging angle at the point where image I1 was captured. As shown in the figure, the depth is shallower at the position where object 20 is present. In this embodiment, an example has been shown in which images are captured at approximately the same angle of view and approximately the same imaging angle, but depending on the imaging position and the configuration of imaging device 11, it may not always be possible to capture images at the same angle of view and angle. In this case, the detection unit 122 is required to detect the position coordinates of key points KP with higher accuracy.

[0050] 8 is a diagram for explaining an example of the functions of the determination device according to the embodiment. The functions of the determination device 12 will be explained with reference to the diagram. First, the determination device 12 detects the position coordinates of the keypoints KP based on the image I1 acquired by the imaging device 11. Specifically, the determination device 12 detects the position coordinates of the keypoints KP by inputting the image I1 into the trained model 32. When there are multiple keypoints KP, the determination device 12 detects the position coordinates of each of the multiple keypoints KP.

[0051] Here, the position coordinates of the keypoints KP detected by the trained model 32 are position coordinates on a two-dimensional image and do not include information about depth. The process of detecting the position coordinates of the key points KP is preferably performed using hardware such as an ASIC, a PLD, or an FPGA.

[0052] The calculation unit 123 calculates the depth at the key point KP based on the position coordinates of the key point KP detected by the detection unit 122 and the depth information 12. In the following description, information containing information about the depth at the key point KP will be referred to as key point depth information 14. It is preferable that the depth calculation process for the key points KP is performed using software.

[0053] The estimation unit 124 estimates the orientation of the object 20 based on the key point depth information I4. Specifically, the estimation unit 124 estimates the coordinates (X, Y) of the midpoint of the multiple key points KP and the inclination θ of the object 20 with respect to the image capture device 11.

[0054] Next, an example of key points KP of the object 20 and an example of calculating the depth of the key points KP will be described with reference to FIGS. In the explanation of FIGS. 9 and 10, the positional relationship between the moving body 10 and the target object 20 may be explained using a three-dimensional Cartesian coordinate system of x-axis, y-axis, and z-axis.

[0055] 9 is a diagram for explaining key points of an object according to the embodiment. With reference to the drawing, an example of key points KP in the case where the object 20 is a basket cart will be described. The object 20, which is a basket cart, has four wheels W and a bottom surface B. The bottom surface B is a surface on which a predetermined load can be loaded. The bottom surface B may be a rectangle with at least two sides of the same length. Alternatively, the bottom surface B may be a square with four sides of the same length. The wheels W may be attached to the bottom surface B as part of a caster unit.

[0056] The key points KP may be four points for specifying the bottom surface B and the rotation centers of the four wheels W (shaft portions of the wheels W). That is, the target object 20, which is a basket cart, has a bottom surface B on which luggage can be loaded and four wheels W. The detection unit 122 detects the four points for specifying the bottom surface B and four points for specifying the positions of the four wheels W, respectively, as the key points KP.

[0057] Specifically, the basket cart has a bottom surface B located 200 mm (millimeters) above the ground G. The basket cart is provided with four caster units to support the bottom surface B. Each caster unit has a mounting seat attached to the bottom surface B. Each caster unit is provided with a wheel W. In the following description, the four key points KP for identifying the bottom surface B will be referred to as key points KP1 to KP4, and the four key points KP for identifying the wheel W will be referred to as key points KP5 to KP8.

[0058] 9 is a diagram illustrating the two-dimensional case, with image I1 simplified for the sake of explanation. In the figure, key points KP1 and KP4 are shown as key points KP for identifying the bottom surface B, and key points KP5 and KP8 are shown as key points KP for identifying the wheels W. For simplicity, the explanation in the same figure will be given for the two-dimensional case, but since the image actually captured by the imaging device 11 is three-dimensional, image I1 may capture all four key points KP for identifying the bottom surface B and all four key points KP for identifying the wheels W.

[0059] Here, an example of the case where the determination device 12 estimates the distance between the object 20 and the moving body 10 will be described. The detection unit 122 detects, based on the image I1 captured by the imaging device 11, the position coordinates of key points KP1 and KP4 for identifying the bottom surface B, and the position coordinates of key points KP5 and KP8 for identifying the wheels W. The detection unit 122 outputs information about the position coordinates of the detected key points KP to the calculation unit 123. If an abnormal value is detected in the distance between the position coordinates of the key points KP, a result indicating that detection was not possible may be output. Specifically, if it is determined that the length of the long side or short side of the basket cart estimated from the position coordinates is too short, information indicating that detection was not possible may be output.

[0060] The detection unit 122 may detect a class corresponding to the key point KP in addition to the position coordinates of the key point KP, and output the position coordinates of the key point KP and the class in association with each other to the calculation unit 123. The class is information indicating whether the key point KP identifies the bottom surface B or the wheel W. Furthermore, the detection unit 122 may perform a threshold determination based on the confidence level, and output to the calculation unit 123 information on whether or not the position coordinates of the key point KP have been detected.

[0061] The detection unit 122 may detect the position coordinates of the key point KP based on a plurality of images I1 captured at a plurality of locations. For example, the imaging device 11 may detect the position coordinates of the key point KP based on a plurality of images I1 captured at a first position near the object 20 and a second position that is closer to the object 20 than the first position. In other words, the detection unit 122 may detect the position coordinates of the key point KP with higher reliability by averaging the results of detection based on a plurality of images captured at a distance and a close distance.

[0062] The calculation unit 123 calculates the depth of each key point KP based on the position coordinates of the key point KP detected by the detection unit 122 and the depth information I2. Here, the depth is the distance in the y direction from the imaging device 11 to the target point. The depth information included in the depth information I2 may not be accurate depending on the angle of view and the imaging angle. Therefore, the calculation unit 123 converts the value indicated in the depth information I2 into an actual depth based on the actual size of the object 20. In other words, the calculation unit 123 calculates the depth of a position corresponding to the position coordinates of the detected key point KP based on the actual size of the object 20 and the depth information I2 included in the image information II. Here, a trained model using machine learning can achieve high detection accuracy in various scenes. The detection unit 122 in this embodiment detects keypoints KP based on a trained model using machine learning, but detection accuracy may decrease in scenes not anticipated during learning (for example, the placement of a basket cart or the background). Specifically, deviations may occur in the coordinates of the detected keypoints KP. In particular, in order to improve the accuracy of detecting the position of the basket cart in the calculation unit 123, the keypoints KP to be detected are the four corners of the basket cart, so deviations may occur in the detected coordinates of the keypoints KP, which may result in the background being detected. In this embodiment, the calculation unit 123 calculates not only the coordinates of the keypoint KP detected by the detection unit 122, but also the depth of the surrounding area using those coordinates. The calculation unit 123 then determines the final depth corresponding to the position coordinates of the keypoint KP by combining multiple depths of the surrounding area other than the depth corresponding to the calculated position coordinates of the keypoint KP. In other words, the calculation unit 123 determines the depth of the position coordinates of the keypoint KP by selectively combining all or part of the depths of the surrounding area of ​​the keypoint KP. Examples of selective combinations include using a short distance from multiple ranging points, calculating the average or median value of multiple ranging points, or using ranging points in an area located on a side of the multiple ranging points. Furthermore, when ranging is performed repeatedly, ranging points may be selected taking into account the previous result, or may be calculated based on multiple depth change points of the surrounding area. The calculation of the depth of the surrounding area may also be based on the reliability output by the detection unit 122 when detecting the keypoint KP.

[0063] The calculation unit 123 first calculates the lengths (number of pixels) of line segments LS extending perpendicularly from the center points of the wheels W to the bottom surface B. Specifically, the calculation unit 123 calculates the length of a line segment LS1 extending perpendicularly from the key point KP5 to the bottom surface B, and the length of a line segment LS2 extending perpendicularly from the key point KP8 to the bottom surface B. The calculation unit 123 calculates the depth of the two wheels W based on the length of the line segment LS and the depth information I2.

[0064] 10 is a diagram for explaining a method for calculating the depth of a key point according to an embodiment, and the details of the calculation of the depth will be described with reference to the same figure. Fig. 10 shows the positional relationship on the yz plane in the three-dimensional orthogonal coordinate system in Fig. 9. "Focal length: f" is information obtained based on the depth information I2. The calculation unit 123 calculates the depth based on "focal length: f".

[0065] The calculation unit 123 stores the actual size of the object 20 (for example, the distance between multiple key points KP) in advance, and calculates the depth from the length (number of pixels) of a line segment LS extended perpendicularly from the center point of the wheel of the basket to the bottom surface B and the actual size of the object 20.

[0066] 9, the estimation unit 124 calculates the distance between the moving body 10 and the target object 20 based on the depths of the key points KP calculated by the calculation unit 123. Specifically, the estimation unit 124 calculates the depth at the midpoint of the line segment connecting the key points KP. More specifically, the estimation unit 124 determines the average of the depths at the two wheels as the depth at the midpoint.

[0067] In the example shown in the figure, the average of the depths of keypoints KP5 and KP8 is set to the depth at midpoint C1 of line segment LS4 connecting keypoints KP5 and KP8, and the average of the depths of keypoints KP1 and KP4 is set to the depth at midpoint C2 of line segment LS4 connecting keypoints KP1 and KP4. The estimation unit 124 calculates actual coordinates using the depths of the midpoints C1 and C2. Based on the calculated actual coordinates, the estimation unit 124 estimates the distance from the imaging unit 11 that captured the image I1 to the basket car. The estimation unit 124 may also estimate the angle of the position where the basket car is located relative to the direction of the imaging unit 11 that captured the image I1. The midpoint C1 is the side of the plurality of sides of the basket car that is closest to the imaging unit 11 that captured the image. That is, the estimation unit 124 estimates the coordinates and angle of the center point of the side of the plurality of sides of the basket car that is closest to the imaging unit 11 that captured the image.

[0068] Furthermore, the estimation unit 124 estimates the inclination (angle θ) of the basket with respect to the imaging unit 11 that captured the image I1, based on the actual coordinates of the two key points KP. Specifically, the angle θ is calculated by the following equation (1).

[0069]

number

[0070] More specifically, when the actual coordinates of two key points KP are (x1, y1, z1) and (x2, y2, z2), and y1 < y2, if the difference between x1 and x2 is within the range of a predetermined error ε, the angle θ is set to 90 degrees. Also, if the difference between x1 and x2 is greater than the range of the predetermined error ε, the angle θ is set to 90 degrees. The angle θ is calculated by the arctangent of the value obtained by dividing y2 - y1 by x2 - x1. The predetermined error ε may be about 1° to 3°.

[0071] Note that the estimation unit 124 may detect outliers based on the previous measurement results or the like. When the estimated result is abnormal, the imaging device 11 may be configured to acquire the normal value by acquiring the image I1 and the depth information I2 again. Also, based on the distance between the object 20 and the moving body 10 when an outlier is detected and the distance between the object 20 and the moving body 10 when a normal value is detected, the imaging device 11 may determine the position P2 where the object 20 is imaged.

[0072] FIG. 11 is a diagram showing an example of key points calculated by the determination device according to the embodiment. In the description made while referring to FIGS. 9 and 10, for simplicity, the two-dimensional case was described. However, since the object imaged by the imaging device 11 is three-dimensional, the detection unit 122 detects all eight key points KP.

[0073] Specifically, the detection unit 122 detects key points KP1 to KP4 as key points KP for specifying the bottom surface B, and detects key points KP5 to KP8 as key points KP for specifying the wheels W.

[0074] [Series of operations of the moving body] FIG. 12 is a diagram for explaining an example of a series of operations of the moving body according to the embodiment. With reference to this figure, a series of operations of the moving body 10 will be described. (Step S110) The moving body 10 moves to a position P2 near the target object 20 based on the map information. (Step S120) The moving object 10 acquires a camera stream of the surroundings of the moving object at position P2. Specifically, the camera stream may be a stream of a depth camera. In other words, the moving object 10 acquires image information II by the imaging device 11. The image information II includes an image I1 and depth information I2.

[0075] (Step S130) The detection unit 122 detects the position coordinates of the key points KP from the acquired image I1 by machine learning using a trained model trained in advance by the determination system 2. (Step S140) The calculation unit 123 calculates the depth at the position coordinates of the key point KP based on information about the detected position coordinates and the depth information I2.

[0076] (Step S150) The estimation unit 124 estimates the posture of the object 20 based on the depth at the calculated position coordinates of the key points KP. (Step S160) The mobile object control device 13 moves the object 20 to the destination position based on the estimated posture of the object 20.

[0077] [Modification of the mobile control system] 13 is a diagram for explaining a modified example of the mobile object control system according to the embodiment. With reference to the same figure, a mobile object control system 1A, which is a modified example of the mobile object control system 1, will be explained. The mobile object control system 1 differs from the mobile object control system 1 in that it includes a central control device 50 that controls multiple mobile objects. In explaining the mobile object control system 1A, components similar to those in the mobile object control system 1 will be assigned the same reference numerals as those in the mobile object control system 1, and explanations thereof may be omitted.

[0078] The central control device 50 has a function equivalent to the determination device 12 possessed by the mobile body 10. The central control device 50 is connected to a plurality of mobile bodies 10A via a predetermined communication network NW. The mobile body 10A is a modified version of the mobile body 10. The mobile body 10A transmits image information about the object 20 to the central control device 50 via the communication network NW. The central control device 50 estimates the posture of the object 20 based on the acquired image information and transmits the estimated result to the mobile body 10A. The mobile body 10A transports the object 20 based on the information about the posture of the object 20 acquired from the central control device 50.

[0079] That is, the mobile object control system 1A differs from the mobile object control system 1 in that the function equivalent to the determination device 12 possessed by the mobile object 10 is performed by the central control device 50 rather than the mobile object 10. The mobile object control system 1A can reduce the resources of the mobile object 10A by having the central control device 50 have the function equivalent to the determination device 12 instead of the multiple mobile objects 10A.

[0080] [Summary of the embodiment] According to the embodiment described above, the determination device 12 is equipped with an image information acquisition unit 121 to acquire image information II including an image I1 and depth information I2, a detection unit 122 to detect the position coordinates of key points KP of the object 20 captured in the image I1, a calculation unit 123 to calculate the depth of the key points KP, and an estimation unit 124 to estimate the posture of the object 20. That is, according to this embodiment, after detecting the positions of the key points KP, the depth is calculated based on the positions of the detected key points KP. In other words, the depth of the key points KP is not directly calculated from the image I1 by machine learning, but after detecting the positions of the key points KP, the depth is detected based on the depth information I2. Therefore, according to this embodiment, the orientation of the object 20 can be detected with high accuracy.

[0081] Furthermore, according to the embodiment described above, the estimation unit 124 of the determination device 12 estimates the inclination of the object 20 relative to the imaging unit 11 that captured the image I1 as the attitude of the object 20. Therefore, according to this embodiment, the moving body 10 equipped with the determination device 12 can recognize the angle of the object 20 relative to itself. Since the moving body 10 can recognize the angle of the object 20 relative to itself, it can accurately transport the object 20 to a destination position by, for example, gripping the object 20.

[0082] Furthermore, according to the embodiment described above, the determination device 12 detects the position of at least one wheel W as a key point KP in the image I1 in which the basket cart, which is the object 20, is captured. Therefore, according to this embodiment, the determination device 12 can determine the attitude of the basket cart.

[0083] Furthermore, according to the embodiment described above, the basket cart whose attitude is determined by the determination device 12 has a bottom surface B on which luggage can be loaded, and four wheels W. The determination device 12 detects, by the detection unit 122, four points that identify the bottom surface B and four points that identify the positions of the four wheels W, as key points KP. Therefore, according to this embodiment, the determination device 12 can easily determine the attitude of the basket cart.

[0084] Furthermore, according to the embodiment described above, the determination device 12 includes the estimation unit 124, and thereby estimates the distance from the imaging unit 11 that captured the image I1 to the basket cart and the inclination of the basket cart with respect to the imaging unit 11 that captured the image I1. Therefore, according to this embodiment, the moving body 10 including the determination device 12 can recognize the distance between itself and the object 20 and the angle of the object 20 relative to itself. Because the moving body 10 can recognize the distance between itself and the object 20 and the angle of the object 20 relative to itself, it can easily grasp the object 20 and transport it to a destination position with high accuracy.

[0085] Furthermore, according to the embodiment described above, in the determination device 12, the detection unit 122 is provided with a trained model and detects key points KP by machine learning. Therefore, according to this embodiment, the detection of key points KP in a two-dimensional image is performed by machine learning, and the subsequent estimation of the posture of the three-dimensional object 20 is performed by software processing. Therefore, according to this embodiment, the posture of the object 20 can be determined with high accuracy. Furthermore, according to this embodiment, the time required for learning can be shortened.

[0086] Furthermore, according to the embodiment described above, the determination system 2 includes the storage device 30, which stores the image I1 of the object 20 and the position coordinates of the key points of the object 20, and the model generation means 31, which generates a trained model using the information stored in the storage device 30 as training data. Therefore, according to this embodiment, a model required for detecting the key points KP can be trained.

[0087] Note that all or part of the functions of each unit of the mobile object control system 1 in the above-described embodiment may be realized by recording a program for realizing these functions on a computer-readable recording medium, and reading and executing the program recorded on the recording medium into a computer system. Note that the term "computer system" here includes hardware such as an OS and peripheral devices.

[0088] Furthermore, "computer-readable recording media" refers to portable media such as optical magnetic disks, ROMs, and CD-ROMs, as well as storage units such as hard disks built into computer systems. Furthermore, "computer-readable recording media" may also include devices that dynamically store programs for a short period of time, such as communication lines used when transmitting programs over a network such as the Internet, and devices that store programs for a fixed period of time, such as volatile memory within a computer system that serves as a server or client in such a case. Furthermore, the program may be one that realizes part of the aforementioned functions, or may be one that can realize the aforementioned functions in combination with a program already stored in the computer system.

[0089] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the spirit of the present invention. [Explanation of symbols]

[0090] 1...mobile body control system, 10...mobile body, 11...imaging device, 12...determination device, 13...mobile body control device, 14...mobile body drive device, 20...target object, 121...image information acquisition unit, 122...detection unit, 123...calculation unit, 124...estimation unit, 2...determination system, 30...storage device, 31...model generation means, 32...trained model, 50...central control device, I1...image, I2...depth information, I3...position coordinate information, I4...keypoint depth information, II...image information, JI...determination information, CI...control information, DI...detection information, KDI...keypoint depth information, W...wheel, B...bottom, KP...keypoint

Claims

1. an image information acquisition unit that acquires image information including an image of an object and depth information of a position corresponding to the image; a detection unit that detects key points, which are specific positions of the object, based on the image included in the image information; a calculation unit that calculates a depth at a position corresponding to the detected key point based on the depth information included in the image information; an estimation unit that estimates the posture of the object based on the calculated depth of the position corresponding to the key point; Equipped with The calculation unit calculates the depth of the position corresponding to the detected key point based on the actual size of the object and the depth information included in the image information. Judgment device.

2. The estimation unit estimates the tilt of the object with respect to the imaging unit that captured the image as the posture of the object. The determination device according to claim 1 .

3. The object is a basket cart having wheels, The key points include at least one of the wheels. The determination device according to claim 1 or 2.

4. The basket cart has a bottom surface on which luggage can be loaded and four of the wheels, The detection unit detects, as the key points, four points that identify the bottom surface and four points that identify the positions of the four wheels. The determination device according to claim 3 .

5. The estimation unit estimates the coordinates and angle of the center point of a side of the object that is closest to an imaging unit that captured the image. The determination device according to claim 4 .

6. The detection unit includes a trained model and detects the keypoints through machine learning. The determination device according to any one of claims 1 to 5.

7. a storage device that stores the image of the object and the position coordinates of the key points of the object; a model generation means for generating the trained model by using the object and the key points included in the image stored in the storage device as training data; The determination device according to claim 6, wherein the key points are detected using the trained model. A determination system comprising:

8. A determination method executed by a processor, comprising: an image information acquisition step in which the processor acquires image information including an image of an object and depth information corresponding to the image; a detection step in which the processor detects key points, which are specific positions of the object, based on the image included in the image information; a calculation step in which the processor calculates depths of coordinates corresponding to the detected key points based on the depth information included in the image information; an estimation step in which the processor estimates the pose of the object; and The calculation step calculates the depth of the position corresponding to the detected key point based on the actual size of the object and the depth information included in the image information. Judgment method.

9. On the computer, an image information acquisition step of acquiring image information including an image of an object and depth information corresponding to the image; a detection step of detecting key points, which are specific positions of the object, based on the image included in the image information; a calculation step of calculating a depth of a coordinate corresponding to the detected key point based on the depth information included in the image information; an estimation step of estimating the posture of the object; Execute The calculation step calculates the depth of the position corresponding to the detected key point based on the actual size of the object and the depth information included in the image information. program.

Citation Information

Patent Citations

  • Human body key point detection method and device, electronic device and storage medium

    CN110348524A

  • Three-dimensional tracking apparatus

    JP1995151844A

  • Information processing apparatus, information processing method, and program

    JP2021099383A

  • Method and System for Stereo Based Vehicle Pose Estimation

    US20190205670A1