Estimation program, method, and device

By using a moving body with an ego camera and fixed cameras to estimate trajectories and optimize camera positions through a factor graph, the method addresses inefficiencies in existing camera positioning methods, enhancing object tracking and spatial mapping accuracy.

WO2025154178A1PCT designated stage expired Publication Date: 2025-07-24FUJITSU LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/001003
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing methods for estimating the position and orientation of multiple cameras in indoor spaces are inefficient and prone to human error, especially when the number of cameras is large, leading to inaccurate tracking of objects and tasks such as customer flow management and security monitoring.

Method used

A method that utilizes a moving body with an ego camera to capture images of the space, combined with fixed cameras, to estimate the first and second trajectories of a moving object, allowing for the accurate grouping and relative positioning of fixed cameras based on keypoint associations and SLAM, followed by optimization using a factor graph to correct and construct a spatial three-dimensional map.

Benefits of technology

This approach enables efficient and accurate estimation of camera positions and orientations, improving the accuracy of object tracking and enabling precise spatial mapping for tasks like traffic management and safety monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024001003_24072025_PF_FP_ABST
    Figure JP2024001003_24072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention estimates a first path of a moving body moving through a designated space and the relative positions and orientations of a plurality of cameras disposed in the space on the basis of a first image of the moving body taken by each of the cameras, estimates a second path of the moving body on the basis of a second image of the space taken from the viewpoint of the moving body as the moving body moves through the space, and estimates the position and orientation, with respect to the space, of each of the plurality of cameras on the basis of an association between the first and the second paths, and the relative positions and orientations.
Need to check novelty before this filing date? Find Prior Art

Description

Estimation program, method, and device

[0001] The disclosed technology relates to an estimation program, an estimation method, and an estimation device.

[0002] Conventionally, in indoor spaces such as shopping stores and underground parking lots, there are systems that install multiple cameras on walls, ceilings, etc., and perform traffic management, safety monitoring, behavior guidance, etc. of people, vehicles, etc. in the indoor space based on images captured by the cameras. In order to accurately execute tasks in such systems, it is necessary to accurately recognize the positions of objects such as people and vehicles in the space. To achieve this, it is necessary to appropriately grasp the position and orientation of the camera relative to the space.

[0003] As a technology for setting up such a system, for example, an image processing system has been proposed in which multiple cameras are installed so that a field of view overlap region is formed where the fields of view of adjacent cameras overlap, and the system acquires internal parameters of each camera. This system tracks a moving object, outputs the time and coordinates of the moving object, acquires a trajectory of the moving object, and connects the trajectories of the moving object based on image features of the moving object. This system then calculates external parameters related to the mounting state of the camera as setting values ​​from the acquired coordinates, and stores the internal parameters in association with the external parameters.

[0004] Also, a technology has been proposed in which, for example, in a scene layout, four points, such as the four corners of a room, are dropped, and then four points corresponding to the camera's field of view are dropped, and the camera is calibrated by matching the dropped points.

[0005] Japanese Patent Application Laid-Open No. 2019-41261

[0006] Pushpak Pujari, "Optimize the Use of Physical Spaces With People Heatmaps," Feb 14, 2023.

[0007] However, the method of setting the external parameters of a camera based on the trajectory of a moving object calculates the relative external parameters between cameras, and has the problem that it is not possible to accurately estimate the position and orientation of the camera relative to the space in which the camera is installed.

[0008] Furthermore, calibration based on point correspondences dropped onto the scene layout and camera fields of view is inefficient and prone to human error, especially when there are a large number of cameras in the space.

[0009] In one aspect, the disclosed technology aims to efficiently and accurately estimate the positions and orientations of multiple cameras arranged in a space.

[0010] In one aspect, the disclosed technology estimates a first trajectory of a moving object and the relative positions and orientations of the multiple cameras based on first images of the moving object moving through a predetermined space taken by each of multiple cameras arranged in the space. The disclosed technology also estimates a second trajectory of the moving object based on second images of the space taken from a viewpoint of the moving object while the moving object is moving through the space. The disclosed technology then estimates the position and orientation of each of the multiple cameras with respect to the space based on the correspondence between the first trajectory and the second trajectory and the relative positions and orientations.

[0011] As one aspect, the present invention has an effect of being able to efficiently and accurately estimate the positions and orientations of a plurality of cameras arranged in a space.

[0012] 1 is a diagram for explaining an example system. FIG. 2 is a diagram for explaining problems in estimating the position and orientation of a fixed camera. FIG. 3 is a diagram for explaining problems in estimating the position and orientation of a fixed camera. FIG. 4 is a schematic configuration diagram of an estimation system according to the present embodiment. FIG. 5 is a functional block diagram of an estimation device according to the present embodiment. FIG. 6 is a diagram showing an example of trajectories of key points extracted from a first image. FIG. 7 is a diagram for explaining grouping of fixed cameras. FIG. 8 is a diagram showing an example of grouping of fixed cameras. FIG. 9 is a diagram for explaining estimation of the relative position and orientation of fixed cameras. FIG. 10 is a diagram for explaining estimation of a real-scale footprint and 3D point cloud information of a space. FIG. 11 is a diagram for explaining merging of footprint points. FIG. 12 is a diagram for explaining correspondence between a footprint of unknown scale and a footprint of real scale. FIG. 13 is a diagram for explaining estimation of the position and orientation of a fixed camera relative to a space. FIG. 14 is a diagram for explaining estimation of the position and orientation of a single fixed camera relative to a space. FIG. 15 is a diagram showing an example of a factor graph of the prior art. FIG. 16 is a diagram showing an example of a factor graph of the present embodiment. FIG. 17 is a diagram showing an example of optimization of a real-scale footprint using a factor graph. FIG. 18 is a diagram for explaining construction of a spatial 3D map. FIG. 19 is a block diagram showing a schematic configuration of a computer functioning as an estimation device. FIG. 19 is a flowchart showing an example of estimation processing according to the present embodiment. FIG. 19 is a flowchart showing an example of a first estimation processing. It is a flowchart showing an example of a second estimation process It is a schematic diagram showing an example of an application result of this embodiment It is a diagram for explaining another example of construction of a spatial three-dimensional map.

[0013] Hereinafter, an example of an embodiment of the disclosed technology will be described with reference to the drawings.

[0014] In this embodiment, the positions and orientations of multiple cameras arranged in a space are estimated relative to the space. Here, the reason why it is necessary to accurately estimate the positions and orientations of multiple cameras arranged in a space relative to the space will be described.

[0015] For example, as shown in FIG. 1 , assume that multiple fixed cameras are placed on the walls, ceiling, etc. of a space such as a shopping store (indicated by the dashed lines in FIG. 1 ) to perform tasks such as managing shopper traffic, monitoring safety, and guiding behavior within the space. To accurately perform these tasks, it is necessary to accurately recognize the position of people within the space. To achieve this, it is necessary to appropriately grasp the positions and orientations (six degrees of freedom) of the fixed cameras relative to the space. For example, if only the relative positions and orientations of the multiple fixed cameras, which are not correlated with the spatial layout, are grasped, it is impossible to accurately observe shoppers entering or leaving the store through the door, shoppers moving to cash registers, specific shelves, etc., or the duration of shopper stays.

[0016] A simple method for estimating the position and orientation of each of multiple fixed cameras relative to space is to associate static keypoints in the images captured by each fixed camera. However, this method may result in an erroneous estimation of the position of the fixed camera when the structures and textures of the images captured by fixed cameras Cam_0 and Cam_1 are similar, as shown in FIG.

[0017] Another simple method for estimating the position and orientation of each of multiple fixed cameras relative to a space is to use images captured by an egoistic camera (hereinafter referred to as an "ego-camera") that captures the space from the viewpoint of a moving object moving through the space. This method associates static keypoints in the space between images captured by the fixed camera and images captured by the ego-camera. However, this method makes it difficult to accurately estimate the orientation of the fixed camera Cam_0 when the appearance of the same static keypoint is significantly different between images captured by the fixed camera Cam_0 and images captured by the ego-camera Cam_ego, as shown in FIG. 3 , for example.

[0018] In this embodiment, a method is proposed for estimating the position and orientation of each of a plurality of fixed cameras relative to space, which is robust against the visual ambiguity of static keypoints as described above.

[0019] 4 shows a schematic configuration of an estimation system 100 according to this embodiment. The estimation system 100 includes an estimation device 10, a plurality of fixed cameras 20, a moving object 22, and an ego-camera 24.

[0020] The fixed cameras 20 are arranged in a predetermined space 30 such as a room. Specifically, each of the multiple fixed cameras 20 is arranged on a wall, ceiling, or the like of the space 30 so that the shooting range of the multiple fixed cameras 20 covers the entire space 30. In FIG. 4 , each fixed camera 20 is represented by a quadrangular pyramid, with the apex of the pyramid representing the position of the fixed camera 20 and the square at the base of the pyramid representing the viewing angle, i.e., the attitude, of the fixed camera 20. In addition, in FIG. 4 , "cami (i = 0, 1, 2, 3, 4, 5 in the example of FIG. 4 )" written next to each quadrangular pyramid representing each fixed camera 20 is an identification number of each fixed camera 20. In the following, when the fixed cameras 20 are described without distinction, they will be referred to as "fixed camera 20," and when the fixed cameras 20 are described with distinction, the fixed camera 20 with the identification number cami will be referred to as "fixed camera cami."

[0021] Each of the fixed cameras 20 outputs an image of its respective imaging range in the space 30 to the estimation device 10. Note that in Fig. 4, some of the connecting lines between the fixed cameras 20 and the estimation device 10 are omitted. In this embodiment, the fixed cameras 20 output a first image of a moving object 22 moving in the space 30 to the estimation device 10. The first image includes multiple frames, and each frame is associated with a timestamp indicating the time of capture.

[0022] The mobile object 22 is a person, a self-propelled robot, or the like, moving through the space 30. The mobile object 22 moves so as to cover the entire space 30 while capturing images of the space 30 with the ego-camera 24. If the mobile object 22 is a self-propelled robot, a movement path may be programmed to realize the movement described above.

[0023] The ego camera 24 is a camera that captures images of the space 30 from the viewpoint of the moving body 22. For example, if the moving body 22 is a person, the ego camera 24 is attached to the chest, head, or the like of the person. Also, for example, if the moving body 22 is a robot, the ego camera 24 is mounted on the robot. Specifically, the ego camera 24 is attached to or mounted on the moving body 22 so that the vertical direction of the ego camera 24 coincides with the vertical direction of the moving body 22 and the shooting direction of the ego camera 24 coincides with the direction of travel of the moving body 22.

[0024] The ego-camera 24 captures images, such as visible light images and depth images, necessary for estimating the self-position of the mobile object 22 and generating an environmental map of the area around the mobile object 22, and outputs the captured images to the estimation device 10. Hereinafter, the image captured by the ego-camera 24 will be referred to as a second image. The second image includes multiple frames, and each frame is associated with a timestamp indicating the time of capture.

[0025] As shown in FIG. 5 , the estimation device 10 functionally includes a first estimating unit 12 , a second estimating unit 14 , a third estimating unit 16 , an optimizing unit 18 , and a constructing unit 20 .

[0026] The first estimation unit 12 estimates the footprint of a moving body 22 of unknown scale (hereinafter referred to as the "footprint of unknown scale") and the relative positions and attitudes of the multiple fixed cameras 20 based on the first images input from each of the multiple fixed cameras 20.

[0027] Specifically, the first estimator 12 detects key points on the moving object 22 from each frame of the first image. The key points may be, for example, the top of the head and the feet of the moving object 22 (black circles in FIG. 4 ), as shown in FIG. 4 . The first estimator 12 tracks the key points detected from each frame between frames, and extracts a trajectory connecting the key points in the first image (dashed line in FIG. 6 ), as shown in FIG. 6 . The first estimator 12 may apply a filter to the trajectory connecting the key points to smooth the trajectory.

[0028] 4, when the space 30 is divided into multiple rooms, the image capturing ranges covered by the fixed cameras 20 may correspond to different rooms. In such a case, if the position and orientation of each fixed camera 20 are estimated by associating the fixed cameras 20 capturing images of different rooms with each other, appropriate estimation cannot be performed.

[0029] Therefore, the first estimation unit 12 groups the multiple fixed cameras 20 based on the degree of overlap of the shooting ranges of each of the multiple fixed cameras 20. Specifically, as shown in the upper diagram of FIG. 7 , for each pair of two fixed cameras 20 among the multiple fixed cameras 20, the first estimation unit 12 calculates the frequency with which common key points are detected as key points of the moving object 22 detected from each frame of the first image. For example, as shown in the lower diagram of FIG. 7 , the first estimation unit 12 prepares a matrix in which each row and each column corresponds to each fixed camera 20. The first estimation unit 12 counts the number of key points with the same timestamp detected by both of the pair of fixed cameras 20 corresponding to each element of the matrix. In this embodiment, because two key points are set on the moving object 22, the maximum value of the matrix elements is 2× the number of frames, and the minimum value is 0. The value of this matrix element corresponds to the degree of overlap of the shooting ranges between the pair of fixed cameras 20.

[0030] The first estimation unit 12 clusters each element of the matrix based on the value of that element, thereby grouping the multiple fixed cameras 20. For example, as shown in Fig. 8 , fixed cameras cam0, cam1, and cam2 are classified into group 0, and fixed cameras cam3, cam4, and cam5 are classified into group 1.

[0031] The first estimator 12 estimates, for each group, a footprint of the moving object 22 at an unknown scale and the relative positions and orientations of the multiple fixed cameras 20. Specifically, as shown in A of FIG. 9 , the first estimator 12 estimates the relative positions and orientations of two fixed cameras 20 included in a group from the correspondence of key points, and also estimates the trajectories of the key points in three dimensions. For example, SfM (Structure from Motion) may be applied to this estimation. Furthermore, as shown in B of FIG. 9 , the first estimator 12 estimates the relative positions and orientations of other fixed cameras 20 included in the group based on first images captured by the other fixed cameras 20 and the estimated three-dimensional trajectories. Furthermore, as shown in C of FIG. 9 , the first estimator 12 bundle adjusts the relative positions and orientations of all fixed cameras 20 included in the group.

[0032] Furthermore, as shown in FIG. 9D, the first estimation unit 12 merges a three-dimensional trajectory based on key points at the top of the head with a three-dimensional trajectory based on key points at the feet, and estimates the resulting footprint. This makes it possible to estimate a footprint with the undetected portion interpolated, even when key points at the feet of the moving object 22 are not detected in the first image due to obstruction by an obstacle. This footprint is based on the relative position and orientation of the fixed camera 20, and is a footprint of unknown scale, in which the scaling factor and rotation relative to the actual-scale footprint are unknown. The unknown-scale footprint is an example of the "first trajectory" of the disclosed technology.

[0033] The second estimation unit 14 estimates a real-scale footprint of the moving object 22 based on the second image input from the ego camera 24. The real-scale footprint is an example of a "second trajectory" of the disclosed technology.

[0034] Specifically, the second estimation unit 14 applies, for example, SLAM (Simultaneous Localization and Mapping) based on the second image including the visible light image (RGB image) and the depth image, to simultaneously estimate the self-position of the moving object 22 and generate an environmental map of the area around the moving object 22. As a result, as shown in A of Fig. 10 , the second estimation unit 14 estimates a real-scale footprint of the moving object 22 by connecting the results of the self-position estimation at each time, and also estimates three-dimensional point cloud information of the space 30.

[0035] 10B, the second estimation unit 14 extracts a 3D point cloud corresponding to the floor surface of the space 30 from the 3D point cloud information of the space 30, for example, by applying an existing segmentation method. Then, as shown in FIG. 10C, the second estimation unit 14 transforms the coordinates of the 3D point cloud information so that the normal vector of the plane indicated by the extracted 3D point cloud is parallel to the vertical axis of the world coordinate system. The second estimation unit 14 also transforms the coordinates of the real-scale footprint to correspond to the transformed coordinates. This makes it possible to suppress bias in the association between the real-scale footprint and the unknown-scale footprint (described in detail below).

[0036] The third estimation unit 16 estimates the position and orientation of each of the multiple fixed cameras 20 relative to the space 30 for each group based on the correspondence between the real-scale footprint and the unknown-scale footprint and the relative positions and orientations of the multiple fixed cameras 20.

[0037] Specifically, the third estimation unit 16 merges adjacent points in the spatiotemporal domain for each footprint. As described above, the unknown-scale footprint is a series of key points of the moving object 22 detected from the first image captured at each time, and the real-scale footprint is a series of the moving object's own position at each time estimated based on the second image. Each of the key points at each time (timestamp T) of the unknown-scale footprint and the own position at each time (timestamp T) of the real-scale footprint is also referred to as a "point."

[0038] More specifically, when the positions of points corresponding to multiple consecutive timestamps in the actual-scale footprint are within a predetermined range, the third estimating unit 16 merges the multiple points by associating the points with a single timestamp.Then, the third estimating unit 16 merges points in the unknown-scale footprint that correspond to the same timestamp as the points merged in the actual-scale footprint.

[0039] For example, the third estimator 16 clusters the points of the actual-scale footprint based on the distance between them, thereby detecting points included in the same cluster as neighboring points. For example, as shown in the upper left diagram of FIG. 11 , assume that the points corresponding to timestamps T32-35 of the actual-scale footprint (white circles in FIG. 11 ) are included in one cluster (dashed circles in the upper diagram of FIG. 11 ). In this case, as shown in the upper right diagram of FIG. 11 , the third estimator 16 determines, for example, the average position of the points corresponding to timestamps T32-35 as the new point with timestamp T32 (dotted circle in FIG. 11 ). Furthermore, the third estimator 16 associates points from timestamp T36 onwards with timestamps 33 onwards. Furthermore, as shown in the lower left and right diagrams of FIG. 11 , the third estimator 16 merges the points with timestamps T32-35 in the unknown-scale footprint and adjusts the timestamps of points from timestamp T36 onwards.

[0040] In this way, by merging points that are close to each other in the spatiotemporal domain for each footprint, it is possible to prevent points from concentrating too much at the same location, which would result in that location being over-emphasized when matching footprints as described below.

[0041] 12, the third estimator 16 associates each of the unknown-scale footprints of each group after the points are merged with a real-scale footprint. Based on this association, the third estimator 16 then calculates parameters for converting the relative position and orientation of the fixed camera 20 into a position and orientation with respect to the space 30. For example, the third estimator 16 uses a point pattern matching algorithm (Reference 1) to calculate the parameters [c^, R^, t^] according to the following equation (1):

[0042] Reference 1: Shinji Umeyama, "Point Pattern Matching Algorithm," Transactions of the Institute of Electronics, Information and Communication Engineers, Vol. J72-D2, No. 2, pp. 218-228, published February 25, 1989

[0043]

[0044] X i are the x and y coordinate values ​​of the point at timestamp T=i on an unknown scale, and Y i are the x and y coordinate values ​​of the point at timestamp T=i in real scale. c is a scale coefficient indicating the scaling ratio of the footprint at unknown scale relative to the footprint at real scale. R is the rotation matrix of the footprint at unknown scale relative to the footprint at real scale. t is the translation matrix of the footprint at unknown scale relative to the footprint at real scale. c^, R^, and t^ are the optimal values ​​of c, R, and t, respectively. Note that "c^" is represented by a "^ (hat)" above "c" in formulas and drawings. The same applies to the other parameters.

[0045] 12, the unknown-scale footprints of group 0 are associated with the actual-scale footprints for timestamps T0-134, and the unknown-scale footprints of group 1 are associated with the actual-scale footprints for timestamps T135-220.

[0046] The third estimator 16 converts the unknown-scale footprint into a real-scale footprint for each group by applying the calculated parameters [c^, R^, t^], as shown in the upper diagram of Fig. 13. Furthermore, the third estimator 16 converts the relative positions and orientations of the multiple fixed cameras 20 in the group into positions and orientations with respect to the space 30 in accordance with the conversion from the unknown-scale footprint to the real-scale footprint, as shown in the lower diagram of Fig. 13.

[0047] Note that when multiple fixed cameras 20 are grouped, if a fixed camera 20 is not classified into any group, i.e., if only a single fixed camera 20 is included in a group, the relative position and orientation of that fixed camera 20 are not estimated. Therefore, the position and orientation of that fixed camera 20 relative to the space 30 cannot be estimated using the parameters calculated by the above formula (1). Therefore, for such a single fixed camera 20, the third estimation unit 16 associates the real-scale footprint in the world coordinate system estimated from the second image with the trajectory of key points on the first image captured by the single fixed camera 20, as shown in FIG. 14 . Then, based on this association, the third estimation unit 16 estimates the position and orientation of the single fixed camera 20 in the world coordinate system, i.e., its position and orientation relative to the space 30.

[0048] The optimization unit 18 optimizes the estimation result by the third estimation unit 16. The estimation result by the third estimation unit 16 is the position and orientation of each of the multiple fixed cameras 20 relative to the space 30, estimated based on the correspondence between the footprint at an unknown scale and the footprint at a real scale, and the relative positions and orientations of the multiple fixed cameras 20. In this embodiment, optimizing the estimation result refers to processing that further improves the consistency of the estimation result by the third estimation unit 16 by using information on the first image and the second image, etc. For example, the optimization unit 18 optimizes the estimation result by the third estimation unit 16 by using a factor graph.

[0049] Here, the upper diagram of FIG. 15 shows, for example, a factor graph for estimating the position and orientation of a fixed camera by associating keypoints captured by the fixed camera. In this example, the graph includes a first node representing the position and orientation of each fixed camera and a second node representing the trajectory (position at each time) of the keypoints of the moving object. Furthermore, between the first node and the second node, graph factors of the keypoints of the fixed camera and the moving object are set as constraints. Furthermore, between the second nodes, graph factors of the trajectory of the keypoints of the moving object are set as constraints.

[0050] A factor graph for estimating the self-position of a moving object based on images captured by an ego camera is shown in the lower diagram of Fig. 15. In this example, the graph includes a third node representing the position and orientation of the ego camera at each time and a fourth node representing static keypoints in a scene in which the moving object is moving. Between the third nodes, the graph factor of the trajectory of the moving object's keypoints is set as a constraint, and between the third node and the fourth node, the graph factor of the ego camera and the 3D point cloud is set as a constraint.

[0051] On the other hand, the factor graph of this embodiment represents, as nodes, the position and orientation of each of the multiple fixed cameras 20 relative to the space 30, the position and orientation of the ego camera 24, and each of the landmarks in the space 30. Furthermore, the factor graph of this embodiment uses the footprint at an unknown scale, the footprint at a real scale, information about the landmarks in the space 30 obtained from the first image, and information about the landmarks in the space 30 obtained from the second image as constraints between the nodes.

[0052] 16 , the factor graph of this embodiment includes a first node representing the position and orientation of the fixed camera 20 and a second node representing the position and orientation of the ego camera 24 at each time. The factor graph of this embodiment also includes a third node representing a landmark in the space 30 (scene) in which the moving body 22 moves. Between the first node and the second node, a graph factor of the key points of the moving body 22 captured by the fixed camera 20 is set as a constraint. Between the second nodes, a graph factor of the key points of the moving body 22 is set as a constraint. Between the second node and the third node, a graph factor of the 3D point cloud information generated from the second image captured by the ego camera 24 is set as a constraint. Between the first node and the third node, a graph factor of the 3D point cloud information generated from the first image captured by the fixed camera 20 and the second image captured by the ego camera 24 is set as a constraint.

[0053] The optimization unit 18 simultaneously optimizes the position and orientation of the fixed camera 20 relative to the space 30 and the real-scale footprint, for example, by using a factor graph such as that shown in Fig. 16. Fig. 17 shows how the real-scale footprint is optimized using a factor graph. The optimization unit 18 outputs information on the optimized position and orientation of the fixed camera 20 relative to the space 30 (hereinafter referred to as "fixed camera position and orientation information").

[0054] The constructor 20 modifies the three-dimensional point cloud information of the space 30 based on the optimized estimation result, and constructs a layout map (hereinafter referred to as a "three-dimensional space map") representing the three-dimensional shape of the space 30 based on the modified three-dimensional point cloud information. Specifically, the constructor 20 calculates a transformation matrix for converting the pre-optimization real-scale footprint into the optimized real-scale footprint. Then, as shown in A of Fig. 18, the constructor 20 applies the calculated transformation matrix to the three-dimensional point cloud information of the space 30 estimated by the second estimator 14 to modify the three-dimensional point cloud information.

[0055] The construction unit 20 also detects a plane by trimming a portion corresponding to the floor surface from the corrected 3D point cloud information using, for example, Open3D Planar Patch Detection (http: / / www.open3d.org / docs / latest / tutorial / geometry / pointcloud.html#Planar-patch-detection). Then, as shown in B of FIG. 18, the construction unit 20 adds a wall surface to the periphery of the detected plane to construct a spatial 3D map. The construction unit 20 outputs the constructed spatial 3D map.

[0056] The output spatial 3D map and fixed camera position and orientation information are registered in association with a predetermined area of ​​a system that performs tasks such as managing the flow of people, vehicles, and other objects, safety monitoring, and behavior guidance using multiple fixed cameras 20. In this system, the positions of objects can be accurately recognized from images captured by the fixed cameras 20 based on this registered information.

[0057] The estimation device 10 may be realized by, for example, a computer 40 shown in FIG. 19 . The computer 40 includes a CPU (Central Processing Unit) 41, a GPU (Graphics Processing Unit) 42, a memory 43 serving as a temporary storage area, and a non-volatile storage device 44. The computer 40 also includes an input / output device 45 such as an input device and a display device, and an R / W (Read / Write) device 46 that controls reading and writing of data from and to a storage medium 49. The computer 40 also includes a communication I / F (Interface) 47 that is connected to a network such as the Internet. The CPU 41, GPU 42, memory 43, storage device 44, input / output device 45, R / W device 46, and communication I / F 47 are connected to one another via a bus 48.

[0058] The storage device 44 is, for example, a hard disk drive (HDD), a solid state drive (SSD), a flash memory, etc. The storage device 44 serving as a storage medium stores an estimation program 50 for causing the computer 40 to function as the estimation device 10. The estimation program 50 includes a first estimation process control command 52, a second estimation process control command 54, a third estimation process control command 56, an optimization process control command 58, and a construction process control command 60.

[0059] The CPU 41 reads the estimation program 50 from the storage device 44, loads it into the memory 43, and sequentially executes the control instructions included in the estimation program 50. The CPU 41 operates as the first estimation unit 12 shown in FIG. 5 by executing the first estimation process control instruction 52. The CPU 41 operates as the second estimation unit 14 shown in FIG. 5 by executing the second estimation process control instruction 54. The CPU 41 operates as the third estimation unit 16 shown in FIG. 5 by executing the third estimation process control instruction 56. The CPU 41 operates as the optimization unit 18 shown in FIG. 5 by executing the optimization process control instruction 58. The CPU 41 operates as the construction unit 20 shown in FIG. 5 by executing the construction process control instruction 60. As a result, the computer 40 that executes the estimation program 50 functions as the estimation device 10. The CPU 41 that executes the program is hardware. A portion of the program may be executed by the GPU 42.

[0060] The functions realized by the estimation program 50 may be realized by, for example, a semiconductor integrated circuit, more specifically, an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or the like.

[0061] Next, the operation of the estimation system 100 according to this embodiment will be described.

[0062] First, the moving object 22 starts moving so as to cover the entire space 30 while capturing images of the space 30 with the ego-camera 24. At the same time, the multiple fixed cameras 20 also start capturing images. Then, first images captured by each of the multiple fixed cameras 20 during a period including the start and end of the movement of the moving object 22, and second images including a visible light (RGB) image and a depth image captured by the ego-camera 24 are input to the estimation device 10. When the first and second images are input to the estimation device 10, the estimation device 10 executes the estimation process shown in FIG. 20 . Note that the estimation process is an example of an estimation method of the disclosed technology.

[0063] In step S10, the first estimation unit 12 executes a first estimation process. The first estimation process will now be described with reference to FIG.

[0064] In step S12, the first estimation unit 12 detects key points on the moving object 22 from each frame of the first image, tracks the detected key points from each frame between frames, and extracts a trajectory connecting the key points in the first image. Here, it is assumed that two points, the position of the top of the head and the position of the feet of the moving object 22, are detected as key points.

[0065] Next, in step S14, the first estimator 12 calculates, as the degree of overlap of the shooting range, the frequency at which common key points are detected as key points of the moving object 22 detected from each frame for each pair of two fixed cameras 20 among the multiple fixed cameras 20. Next, in step S16, the first estimator 12 groups the multiple fixed cameras 20 based on the calculated degree of overlap.

[0066] Next, in step S18, the first estimation unit 12 estimates the relative positions and orientations of the two fixed cameras 20 included in the group from the correspondence of the key points, and also estimates the trajectories of the key points in three dimensions. Then, the first estimation unit 12 estimates the relative positions and orientations of the other fixed cameras 20 included in the group based on the first images captured by the other fixed cameras 20 and the estimated three-dimensional trajectories.

[0067] Next, in step S20, the first estimator 12 bundle adjusts the relative positions and orientations of all of the fixed cameras 20 included in the group. Next, in step S22, the first estimator 12 merges the three-dimensional trajectory based on the keypoints at the head position of the moving object 22 estimated in step S18 with the three-dimensional trajectory based on the keypoints at the feet position, and estimates it as a footprint of unknown scale.

[0068] When the processing of steps S18 to S22 is completed for each group divided in step S16, the process returns to the estimation processing (FIG. 20).

[0069] Next, in step S30, the second estimation unit 14 executes the second estimation process. The second estimation process will now be described with reference to FIG.

[0070] In step S32, the second estimation unit 14 simultaneously performs self-location estimation of the moving object 22 and generates an environmental map of the area around the moving object 22, for example, by applying SLAM, based on the second image. As a result, the second estimation unit 14 estimates a real-scale footprint of the moving object 22 by connecting the results of the self-location estimation at each time, and also estimates three-dimensional point cloud information of the space 30.

[0071] Next, in step S34, the second estimator 14 extracts a plane corresponding to the floor of the space 30 from the three-dimensional point cloud information of the space 30, and transforms the coordinates of the three-dimensional point cloud information so that the normal vector of the plane is parallel to the vertical axis of the world coordinate system. Next, in step S36, the second estimator 14 applies the coordinate transformation of step S34 to also transform the coordinates of the actual-scale footprint, and returns to the estimation process ( FIG. 20 ).

[0072] The first estimation process in step S10 and the second estimation process in step S30 may be performed in either order, or may be performed in parallel.

[0073] Next, in step S40, the third estimator 16 merges adjacent points in the actual-scale footprint by associating them with a single timestamp, and also merges points in the unknown-scale footprint that correspond to the same timestamp as the merged points in the actual-scale footprint.

[0074] Next, in step S42, the third estimator 16 associates each of the unknown-scale footprints of each group after the points have been merged with a real-scale footprint. Based on this association, the third estimator 16 then calculates parameters for converting the relative positions and orientations of the fixed cameras 20 into positions and orientations with respect to the space 30. The third estimator 16 also applies the calculated parameters to the unknown-scale footprints for each group to convert them into real-scale footprints. The third estimator 16 then converts the relative positions and orientations of the multiple fixed cameras 20 in the group in accordance with this conversion, thereby estimating their positions and orientations with respect to the space 30.

[0075] For a group that includes only a single fixed camera 20, the third estimation unit 16 associates the real-scale print in the world coordinate system estimated from the second image with the trajectory of key points on the first image captured by the single fixed camera 20. Then, based on this association, the third estimation unit 16 estimates the position and orientation of the single fixed camera 20 in the world coordinate system, i.e., its position and orientation relative to the space 30.

[0076] Next, in step S44, the optimization unit 18 simultaneously optimizes the position and orientation of the fixed camera 20 relative to the space 30 and the real-scale footprint using a factor graph such as that shown in FIG.

[0077] Next, in step S46, the constructor 20 calculates a transformation matrix for converting the pre-optimization real-scale footprint into the optimized real-scale footprint. The constructor 20 also applies the calculated transformation matrix to the 3D point cloud information of the space 30 estimated in step S32 above to correct the 3D point cloud information. The constructor 20 then constructs a 3D map of the space 30 based on the corrected 3D point cloud information.

[0078] Next, in step S48, the optimization unit 18 outputs the fixed camera position and orientation information estimated in step S44, and the construction unit 20 outputs the three-dimensional map of the space 30 constructed in step S46, thereby completing the estimation process.

[0079] As described above, according to the estimation system of this embodiment, the estimation device acquires first images captured by each of multiple fixed cameras arranged in a predetermined space of a moving object moving through the space. The estimation device estimates the unknown-scale footprint of the moving object and the relative positions and orientations of the multiple fixed cameras based on the first images. The estimation device also estimates the real-scale footprint of the moving object based on second images captured of the space from the viewpoint of the moving object while the moving object is moving through the space. The estimation device then estimates the position and orientation of the fixed camera relative to the space based on the correspondence between the unknown-scale footprint and the real-scale footprint and the relative positions and orientations of the fixed camera. This makes it possible to efficiently and accurately estimate the position and orientation of the multiple cameras arranged in the space.

[0080] 23 shows an example of a spatial 3D map and fixed camera position and orientation information estimated by applying this embodiment to an actual indoor scene. The fixed camera position and orientation information is expressed by a quadrangular pyramid representing the fixed camera.

[0081] In the above embodiment, a case where a visible light image and a depth image are acquired as the second image captured by the ego camera has been described, but this is not limiting. It is sufficient if information necessary for estimating the self-position of a moving object and generating an environmental map is acquired. For example, information from an IMU (Inertial Measurement Unit) sensor may be acquired instead of or in addition to the depth image.

[0082] Furthermore, the method of constructing the spatial 3D map in the above embodiment is merely an example, and the spatial 3D map may be constructed by other methods, such as the method shown in FIG. 24 . In this method, the construction unit extracts a plane corresponding to the floor surface from the optimized 3D point cloud information, as shown in FIG. 24A, and projects the 3D point cloud information onto the plane. The construction unit also extracts line segments from the projected point cloud using a Hough transform, as shown in FIG. 24B, and merges the shapes of the line segments to identify the positions of walls, as shown in FIG. 24C. The construction unit then adds walls to the identified positions to construct the spatial 3D map, as shown in FIG. 24D.

[0083] In the above embodiment, the estimation program is stored (installed) in advance in a storage device, but this is not limiting. The program according to the disclosed technology may be provided in a form stored in a storage medium such as a CD-ROM, a DVD-ROM, or a USB memory.

[0084] REFERENCE SIGNS LIST 10 Estimation device 12 First estimation unit 14 Second estimation unit 16 Third estimation unit 18 Optimization unit 20 Construction unit 20 Fixed camera 22 Moving body 24 Ego-camera 30 Space 40 Computer 41 CPU 42 GPU 43 Memory 44 Storage device 45 Input / output device 46 R / W device 47 Communication I / F 48 Bus 49 Storage medium 50 Estimation program 52 First estimation process control command 54 Second estimation process control command 56 Third estimation process control command 58 Optimization process control command 60 Construction process control command 100 Estimation system

Claims

1. A estimation program for causing a computer to execute a process including: estimating a first trajectory of a moving object, relative positions and postures of a plurality of cameras, based on a first image obtained by photographing the moving object moving in the space by each of the plurality of cameras arranged in a predetermined space; estimating a second trajectory of the moving object based on a second image obtained by photographing the space from the viewpoint of the moving object while the moving object moves in the space; and estimating positions and postures of each of the plurality of cameras with respect to the space based on association between the first trajectory and the second trajectory, and the relative positions and postures.

2. The estimation program according to claim 1, further optimizing an estimation result of positions and postures of each of the plurality of cameras with respect to the space, the estimation result being estimated based on association between the first trajectory and the second trajectory, and the relative positions and postures.

3. The estimation program according to claim 2, further optimizing the estimation result by using a factor graph that represents positions and postures of each of the plurality of cameras with respect to the space, positions and postures of the camera that captured the second image, and each of landmarks in the space as nodes, and using the first trajectory, the second trajectory, information of landmarks in the space obtained from the first image, and information of landmarks in the space obtained from the second image as constraints between the nodes.

4. The second image includes a visible light image and a depth image. Based on the second image, by performing self-position estimation of the moving object and generation of an environmental map around the moving object, the second trajectory of the moving object is estimated, and three-dimensional point cloud information of the space is estimated. Based on the optimized estimation result, the three-dimensional point cloud information of the space is corrected, and based on the corrected three-dimensional point cloud information of the space, a three-dimensional map of the space is constructed. The estimation program according to claim 3.

5. The estimation program according to any one of claims 1 to 4, further grouping the plurality of cameras based on a degree of overlap of photographing ranges of each of the plurality of cameras, and estimating the first trajectory of the moving object and relative positions and postures of the plurality of cameras for each group.

6. The estimation program according to claim 5, wherein for each pair of two cameras among the plurality of cameras, the frequency at which a common keypoint is detected as the keypoint of the moving object detected from the first image at each time is calculated as the degree of overlap of the imaging ranges.

7. For two cameras included in the group, based on the common keypoint detected from the first image, the relative position and orientation between the two cameras are estimated, the first trajectory is estimated in three dimensions, and based on the first image taken by other cameras other than the two cameras included in the group and the three-dimensional first trajectory, the relative position and orientation of the other cameras with respect to the two cameras are estimated, and the relative positions and orientations of all the cameras included in the group are adjusted to estimate the relative positions and orientations of the plurality of cameras for each group. The estimation program according to claim 6.

8. The estimation program according to claim 4, wherein a plane corresponding to the floor surface of the space is extracted from the three-dimensional point cloud information of the space, the coordinates of the three-dimensional point cloud information of the space are transformed so that the normal vector of the plane is parallel to the vertical axis of the world coordinate system, and the coordinates of the second trajectory are transformed corresponding to the transformed coordinates.

9. The first trajectory is a series of keypoints of the moving object detected from the first image taken at each time, the second trajectory is a series of self-positions of the moving object estimated at each time based on the second image, in the second trajectory, when a plurality of self-positions estimated at consecutive times are included within a predetermined range, the plurality of self-positions are merged into one time, and in the first trajectory, the keypoints corresponding to the consecutive times are merged. The estimation program according to any one of claims 1 to 4.

10. The estimation program according to claim 5, wherein when there is one camera included in the group, based on the association between the first trajectory corresponding to the one camera and the second trajectory in three dimensions, the position and orientation of the one camera with respect to the space are estimated.

11. Based on a first image captured by each of a plurality of cameras arranged in a predetermined space of a moving body moving in the space, estimating a first trajectory of the moving body and relative positions and postures of the plurality of cameras, and estimating a second trajectory of the moving body based on a second image captured of the space from the viewpoint of the moving body while the moving body moves in the space, and estimating positions and postures of each of the plurality of cameras with respect to the space based on association between the first trajectory and the second trajectory and the relative positions and postures, a estimation method in which a computer executes a process including this.

12. The estimation method according to claim 11, optimizing an estimation result of positions and postures of each of the plurality of cameras with respect to the space estimated based on association between the first trajectory and the second trajectory and the relative positions and postures.

13. Representing positions and postures of each of the plurality of cameras with respect to the space, positions and postures of the camera that captured the second image, and each of landmarks in the space as nodes, and using a factor graph in which the first trajectory, the second trajectory, information of landmarks in the space obtained from the first image, and information of landmarks in the space obtained from the second image are constraints between the nodes to optimize the estimation result. The estimation method according to claim 12.

14. The second image includes a visible light image and a depth image. Based on the second image, estimating the second trajectory of the moving body by performing self-position estimation of the moving body and generation of an environmental map around the moving body, and estimating three-dimensional point cloud information of the space, correcting the three-dimensional point cloud information of the space based on the optimized estimation result, and constructing a three-dimensional map of the space based on the corrected three-dimensional point cloud information of the space. The estimation method according to claim 13.

15. Grouping the plurality of cameras based on the degree of overlap of the shooting ranges of each of the plurality of cameras, and estimating the first trajectory of the moving body and relative positions and postures of the plurality of cameras for each group. The estimation method according to any one of claims 11 to 14.

16. For each pair of two cameras among the plurality of cameras, as the degree of overlap of the imaging ranges, calculate the frequency at which a common keypoint is detected as a keypoint of the moving object detected from the first image at each time, according to the estimation method of claim 15.

17. For two cameras included in the group, based on the common keypoints detected from the first image, estimate the relative position and orientation between the two cameras, estimate the first trajectory in three dimensions, and based on the first image taken by another camera other than the two cameras included in the group and the three-dimensional first trajectory, estimate the relative position and orientation of the other camera with respect to the two cameras, and adjust the relative positions and orientations of all the cameras included in the group to estimate the relative positions and orientations of the plurality of cameras for each group, according to the estimation method of claim 16.

18. Extract a plane corresponding to the floor surface of the space from the three-dimensional point cloud information of the space, transform the coordinates of the three-dimensional point cloud information of the space so that the normal vector of the plane is parallel to the vertical axis of the world coordinate system, and transform the coordinates of the second trajectory corresponding to the transformed coordinates, according to the estimation method of claim 14.

19. The first trajectory is a series of keypoints of the moving object detected from the first image taken at each time, and the second trajectory is a series of self-positions of the moving object estimated at each time based on the second image. In the second trajectory, when a plurality of self-positions estimated at consecutive times are included within a predetermined range, merge the plurality of self-positions into one time. In the first trajectory, merge the keypoints corresponding to the consecutive times, according to the estimation method according to any one of claims 11 to 14. A first estimation unit that estimates a first trajectory of the moving object, and relative positions and postures of the plurality of cameras, based on a first image obtained by each of the plurality of cameras arranged in a predetermined space and capturing the moving object moving in the space; a second estimation unit that estimates a second trajectory of the moving object based on a second image obtained by capturing the space from the viewpoint of the moving object while the moving object moves in the space; and a third estimation unit that estimates positions and postures of each of the plurality of cameras with respect to the space based on association between the first trajectory and the second trajectory, and the relative positions and postures. An estimation device including the above components.

Citation Information

Patent Citations

  • Information processor, system, method, and program

    JP2019125354A

  • Mobile location estimation system and mobile location method

    JP2020153956A

  • Method for estimating connection relation among wide-area distributed camera and program for estimating connection relation

    WO2007026744A1