Position and orientation estimation system, position and orientation estimation device, position and orientation estimation method, and program

The system addresses accuracy issues in imaging device positioning by using multiple point cloud data sets with varying stability to estimate and project images accurately, enhancing structural inspection efficiency.

JP7861907B2Active Publication Date: 2026-05-19NEC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NEC CORP
Filing Date
2023-02-28
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing methods for estimating the position and orientation of an imaging device using three-dimensional point cloud data and imaging images suffer from reduced accuracy due to differences in timing between data acquisition and imaging, leading to poor feature point correspondence.

Method used

A system that stores multiple derived point cloud data sets with varying static stability, allowing for sequential estimation based on these data sets and imaging images to determine the position and orientation of the imaging device, using techniques like DeepI2P, Direct Regression, Monodepth2+USIP, and 2D3D-MatchNet, and projecting the image onto a virtual environment model.

Benefits of technology

Enables accurate and efficient estimation of the imaging device's position and orientation, facilitating rapid and precise mapping of structural deformations by reducing the impact of temporal discrepancies in data acquisition and imaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007861907000001
    Figure 0007861907000001
  • Figure 0007861907000002
    Figure 0007861907000002
  • Figure 0007861907000003
    Figure 0007861907000003
Patent Text Reader

Abstract

An information processing device (1) estimates the position and orientation of a camera (2) at the time of imaging, on the basis of three-dimensional point cloud data of an environment and a captured image obtained by imaging, with the camera (2), a structure included in the environment. The information processing device (1) comprises a storage unit (3), an acquisition unit (4), and an estimation unit (5). The storage unit (3) stores a plurality of pieces of derived point cloud data (DPC), which are generated from the three-dimensional point cloud data and which have static stabilities that are different from each other. The acquisition unit (4) acquires the captured image. The estimation unit (5) estimates the position and the orientation on the basis of at least one of the plurality of pieces of the derived point cloud data (DPC) and the captured image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] ,

[0006] , , , , , ,

[0005] , , , , , ,

[0001] The present disclosure relates to a position and orientation estimation system, a position and orientation estimation device, and a position and orientation estimation method.

Background Art

[0002] Patent Document 1 discloses a technique of mapping an image obtained by imaging a structure onto the surface of the structure in a virtual space constituted by the three-dimensional design data of the structure.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] By the way, the inventors of the present application have developed a technique for estimating the position and orientation of an imaging device at the time of imaging based on three-dimensional point cloud data and an imaging image.

[0005] An object of the present disclosure is to provide a technique for estimating the position and orientation of an imaging device at the time of imaging with high accuracy based on three-dimensional point cloud data and an imaging image.

Means for Solving the Problems

[0006] According to a first aspect of this disclosure, a position and orientation estimation system is provided for estimating the position and orientation of an imaging device at the time of imaging, based on three-dimensional point cloud data of an environment and an image obtained by imaging an object included in the environment with an imaging device, the system comprising: storage means for storing a plurality of derived point cloud data generated from the three-dimensional point cloud data and having different static stabilities; acquisition means for acquiring the image; and estimation means for estimating the position and orientation based on at least one of the plurality of derived point cloud data and the image.

[0007] A second aspect of this disclosure provides a position and orientation estimation system for estimating the position and orientation of an imaging device at the time of imaging, based on three-dimensional point cloud data of an environment and an image obtained by imaging an object included in the environment with an imaging device, the system comprising: storage means for storing a plurality of derived point cloud data generated from the three-dimensional point cloud data and having different static stabilities; acquisition means for acquiring the image; and estimation means for estimating the position and orientation based on at least one of the plurality of derived point cloud data and the image.

[0008] A third aspect of this disclosure provides a position and orientation estimation method for estimating the position and orientation of an imaging device at the time of imaging, based on three-dimensional point cloud data of an environment and an image obtained by imaging an object included in the environment with the imaging device, the method comprising: an acquisition step of acquiring the image; and an estimation step of estimating the position and orientation based on the image and at least one derived point cloud data from a plurality of derived point cloud data generated from the three-dimensional point cloud data and having different static stabilities. [Effects of the Invention]

[0009] According to this disclosure, the position and orientation can be estimated with high accuracy. [Brief explanation of the drawing]

[0010] [Figure 1] This is a functional block diagram of the position and orientation estimation system. (Summary of this disclosure) [Figure 2] This is a functional block diagram of an information processing device. (First Embodiment) [Figure 3] This is the data structure of the derived point cloud database. (First Embodiment) [Figure 4] This is an explanatory diagram of the static stability in each derived point cloud data. (First Embodiment) [Figure 5] This is an explanatory diagram of the method for generating each derived point cloud data. (First Embodiment) [Figure 6] This is the control flow of an information processing device. (First Embodiment) [Figure 7] This is the control flow of an information processing device. (Second Embodiment) [Figure 8] This is the control flow of an information processing device. (Third Embodiment) [Figure 9] This is the control flow of an information processing device. (Fourth embodiment) [Figure 10] This is the control flow of an information processing device. (Fifth embodiment) [Modes for carrying out the invention]

[0011] (Summary of this disclosure) First, an overview of this disclosure will be given with reference to Figure 1. Figure 1 shows a functional block diagram of the position and attitude estimation system 100.

[0012] The position and orientation estimation system 100 estimates the position and orientation of the imaging device at the time of imaging based on three-dimensional point cloud data of the environment and an image obtained by imaging an object included in the environment with the imaging device. The position and orientation estimation system 100 includes a storage means 101, an acquisition means 102, and an estimation means 103.

[0013] The storage means 101 stores multiple derived point cloud data generated from three-dimensional point cloud data, each having a different static stability.

[0014] The acquisition means 102 acquires the captured image.

[0015] The estimation means 103 estimates the position and orientation based on at least any one of the plurality of derived point cloud data and the captured image.

[0016] According to the above configuration, the position and orientation of the imaging device during imaging can be estimated with high accuracy.

[0017] (First Embodiment) Next, referring to FIGS. 2 to 6, the first embodiment will be described.

[0018] FIG. 2 shows a functional block diagram of the information processing apparatus 1. The information processing apparatus 1 is a specific example of a position and orientation system. The information processing apparatus 1 is a specific example of a position and orientation device.

[0019] The information processing apparatus 1 shown in FIG. 2 is typically used for managing structures such as bridges, dams, tunnels, towers, houses, etc. That is, since structures deteriorate over time, it is necessary to regularly inspect for the presence of deformation, and when deformation is detected, appropriate measures are required. Deformation typically refers to the floating, peeling, and cracking of the concrete constituting the structure.

[0020] When an inspector discovers a deformation of a structure, the inspector captures the deformation with the camera 2 (imaging device) and acquires a captured image. Here, when three-dimensional point cloud data of the structure is available, it is conceivable to map the above captured image onto the structure in the virtual space constituted by the three-dimensional point cloud data. If the structure in the virtual space onto which the captured image is mapped can be displayed on the display, it becomes possible to easily grasp where the deformation is located on the structure and how large it is. As a result, when the inspector captures the deformation with the camera 2, the inspector does not need to record the position of the deformation in detail or measure the size of the deformation on-site, so that the inspection of the structure can be performed in a short time with a small number of people.

[0021] Incidentally, the position and orientation of camera 2 at the time of imaging are essential for mapping captured images to three-dimensional point cloud data. The position and orientation of camera 2 at the time of imaging typically refers to the transformation parameters between the point cloud coordinate system of the three-dimensional point cloud data and the camera coordinate system of camera 2. In other words, by transforming the coordinates of the three-dimensional point cloud data using the transformation parameters, the three-dimensional point cloud data can be represented in the camera coordinate system. Conversely, by transforming the coordinates of the captured image using the transformation parameters, the captured image can be represented in the point cloud coordinate system. The transformation parameters typically include a rotation matrix and a translation matrix.

[0022] Here, registration techniques such as DeepI2P (Image-to-Point Cloud Registration via Deep Classification), Direct Regression, Monodepth2+USIP, Monodepth2+GT-ICP, and 2D3D-MatchNet are known as methods for estimating the position and orientation of Camera 2 during image acquisition. Regardless of the method used, the following unresolved problems remain that significantly reduce the estimation accuracy (matching score) when calculating the above position and orientation. In short, these unresolved problems are that when measuring the distance of a structure to generate three-dimensional point cloud data, objects other than the structure are also measured simultaneously, and when imaging the structure with Camera 2, objects other than the structure are also captured simultaneously. For example, it is quite conceivable that a car is parked near a structure when measuring its distance to generate three-dimensional point cloud data, and that the car is not in the same location when imaging the structure with Camera 2. This is because, in actual operation, the distance of a structure to generate three-dimensional point cloud data is measured only once, and inspectors then image the structure with Camera 2 every six months thereafter. In other words, the main cause of the above problem is the large difference between the time when the structure is measured to generate three-dimensional point cloud data and the time when the inspector images the structure with camera 2. In this case, it is not possible to establish a good correspondence between feature points in the three-dimensional point cloud data and feature points in the captured image, and as a result the estimation accuracy of the conversion parameters plateaus.

[0023] The information processing device 1 shown in Figure 2 was devised to solve the above technical problems, and the information processing device 1 will be described in detail below.

[0024] As shown in Figure 2, the information processing device 1 comprises a CPU 1a (Central Processing Unit), memory 1b, LCD 1c (Liquid Crystal Display), communication interface 1d, and input means 1e.

[0025] Memory 1b consists of RAM (Random Access Memory), ROM (Read Only Memory), HDD (Hard Disk Drive), etc. The control program is stored in Memory 1b.

[0026] The input means 1e is typically a keyboard.

[0027] CPU 1a reads and executes the control program stored in memory 1b. This causes the control program to make the CPU 1a and other hardware function as a storage unit 3, an acquisition unit 4, an estimation unit 5, a mapping unit 6, and an output unit 7.

[0028] The memory unit 3 is a specific example of a storage means. The memory unit 3 stores multiple derived point cloud data generated from three-dimensional point cloud data of the environment, each having a different static stability. Here, "environment" includes the structure being inspected and the objects surrounding the structure. Specifically, the memory unit 3 stores the derived point cloud DB8 shown in Figure 3. The derived point cloud DB8 will be described below with reference to Figure 3.

[0029] As shown in Figure 3, the derived point cloud DB8, as an example, holds three derived point cloud data DPCs (Derived Point Clouds). The three derived point cloud data DPCs consist of derived point cloud data DPC1, derived point cloud data DPC2, and derived point cloud data DPC3. In this embodiment, the number of derived point cloud data DPCs held by the derived point cloud DB8 is set to three, but it is not limited to this; it may be two, four or more, or any number.

[0030] Each derived point cloud data (DPC) is a three-dimensional point cloud data set, generated from the three-dimensional point cloud data of the environment.

[0031] In the derived point cloud DB8, each derived point cloud data DPC is associated with a static stability (SS). As shown in Figure 3, for example, the static stability SS of derived point cloud data DPC1 is "60 years", the static stability SS of derived point cloud data DPC2 is "3 years", and the static stability SS of derived point cloud data DPC3 is "1 second". Thus, static stability SS is expressed, for example, by the length of the period on the time axis. It can be said that the longer the period, the relatively higher the static stability SS, and the shorter the period, the relatively lower the static stability SS. Therefore, the static stability SS of derived point cloud data DPC1 is higher than the static stability SS of derived point cloud data DPC2. The static stability SS of derived point cloud data DPC2 is higher than the static stability SS of derived point cloud data DPC3. In other words, the static stability SS of derived point cloud data DPC1 is the highest among the multiple derived point cloud data DPCs held by derived point cloud DB8. Furthermore, derived point cloud DB8 can be said to hold multiple derived point cloud data DPCs, each with a different static stability SS. Note that static stability SS may be expressed indirectly using level notation, such as Level 1, Level 2, Level 3, rather than directly expressed by the length of the period on the time axis.

[0032] Here, we will explain the differences between derived point cloud data DPC1, derived point cloud data DPC2, and derived point cloud data DPC3. Simply put, derived point cloud data DPC2 is derived point cloud data DPC3 with some point clouds removed, and derived point cloud data DPC1 is derived point cloud data DPC2 with some point clouds removed. As shown in Figure 3, derived point cloud data DPC3 includes building point cloud PPC1 (Partial Point Cloud) corresponding to buildings, tree point cloud PPC2 corresponding to trees, automobile point cloud PPC3 corresponding to automobiles, and pedestrian point cloud PPC4 corresponding to pedestrians. In contrast, derived point cloud data DPC2 includes building point cloud PPC1 and tree point cloud PPC2, but does not include automobile point cloud PPC3 and pedestrian point cloud PPC4. Also, derived point cloud data DPC1 includes building point cloud PPC1, but does not include tree point cloud PPC2, automobile point cloud PPC3, or pedestrian point cloud PPC4.

[0033] Next, we will explain the static stability SS in detail with reference to Figure 4. Figure 4 shows the static stability SS for each derived point cloud data DPC. The horizontal axis in Figure 4 is the time axis.

[0034] As shown in Figure 4, the derived point cloud data DPC1 contains only the building point cloud PPC1. Buildings corresponding to the building point cloud PPC1 generally remain stationary on the time axis for about 60 years. Therefore, the maintenance period of the buildings is 60 years. Consequently, the static stability SS of the derived point cloud data DPC1 is 60 years, which is the maintenance period of the buildings.

[0035] In contrast, the derived point cloud data DPC2 includes the building point cloud PPC1 and the tree point cloud PPC2. Trees corresponding to the tree point cloud PPC2 generally remain stationary on the time axis for about three years. That is, trees are cut down or transplanted after a few years. Therefore, the maintenance period for trees is three years. Consequently, the static stability SS of the derived point cloud data DPC2 is three years, which is the shortest maintenance period among the maintenance periods of buildings and trees. In other words, the static stability SS of the derived point cloud data DPC2 is the length of the maintenance period of trees, which are the objects that maintain a stationary state on the time axis for the shortest period among buildings and trees.

[0036] Furthermore, the derived point cloud data DPC3 includes the building point cloud PPC1, the tree point cloud PPC2, the automobile point cloud PPC3, and the pedestrian point cloud PPC4. When an automobile is parked, it generally remains stationary on the time axis for about 9 hours. Therefore, the maintenance period for an automobile is 9 hours. Pedestrians, on the other hand, hardly remain stationary on the time axis, and even when stationary, it is only for about 1 second. Therefore, the maintenance period for a pedestrian is 1 second. As a result, the static stability SS of the derived point cloud data DPC3 is 1 second, which is the shortest maintenance period among the maintenance periods of buildings, trees, automobiles, and pedestrians. In other words, the static stability SS of the derived point cloud data DPC3 is the length of the maintenance period of a pedestrian, which is the shortest maintenance period among buildings, trees, automobiles, and pedestrians in that it remains stationary on the time axis.

[0037] Next, with reference to Figure 5, we will illustrate the method for generating each derived point cloud data DPC.

[0038] (Step 1) First, we acquire three-dimensional point cloud data of the environment. Methods for acquiring three-dimensional point cloud data of the environment include using Lidar (Light Detection and Ranging) and using photogrammetry.

[0039] In the LiDAR-based method, the environment is measured from various angles using LiDAR, and multiple three-dimensional point cloud data output from the LiDAR are combined using registration techniques such as ICP (Iterative Closest Point) to generate three-dimensional point cloud data of the environment.

[0040] In photogrammetry, the three-dimensional structure of the environment is reconstructed by solving a geometric inverse problem using multiple images obtained by capturing the environment from various angles, thereby generating three-dimensional point cloud data of the environment. A typical method for reconstructing the three-dimensional structure of an environment from multiple images is Structure from Motion (SfM). When MVS (Multi-View Stereo) is used in conjunction with this, more precise three-dimensional point cloud data of the environment can be generated.

[0041] Alternatively, three-dimensional point cloud data of the environment may be generated using both Lidar and photogrammetry. That is, three-dimensional point cloud data of the environment may be generated by combining three-dimensional point cloud data of the environment generated using Lidar and three-dimensional point cloud data of the environment generated by photogrammetry using the registration technique described above.

[0042] (Step 2) Next, the three-dimensional point cloud data of the environment obtained in Step 1 is classified according to the maintenance period of each object included in the environment. Specifically, the building point cloud PPC1, which corresponds to buildings with a maintenance period longer than 10 years, is classified into Layer 1; the tree point cloud PPC2, which corresponds to trees with a maintenance period of 1 year or less but longer than 10 years, is classified into Layer 2; and the automobile point cloud PPC3, which corresponds to automobiles, and the pedestrian point cloud PPC4, which corresponds to pedestrians, are classified into Layer 3.

[0043] First, to detect objects in an environment based on three-dimensional point cloud data of the environment, known Deep Neural Networks (DNNs) such as PointNet, PointNet++, and VoteNet can be used. If an image of the environment is obtained by simultaneously measuring distance, then known DNNs such as R-CNN (Regions with Convolutional Neural Networks) or YOLO (You Only Look Once) can be used to detect objects in the environment. Then, a table is created in advance showing the correspondence between objects and their duration, and by referring to this table, the three-dimensional point cloud data of the environment obtained in step 1 is classified according to the duration of each object in that environment.

[0044] The task of classifying the three-dimensional point cloud data of the environment obtained in Step 1 according to the maintenance period of each object included in that environment may be performed manually by the operator.

[0045] In Figure 5, the objects classified as Layer 1 are not limited to the exemplified structures. For example, objects such as road surfaces and terrain have a longer maintenance period than structures. Therefore, these objects are also classified as Layer 1.

[0046] Objects classified as Layer 2 are not limited to the example of trees. For example, objects such as chairs, desks, and doors may also be classified as Layer 2.

[0047] Objects classified as Layer 3 are not limited to the example of automobiles and pedestrians. For example, animals and drones may also be classified as Layer 3.

[0048] (Step 3) Next, Layer 1 is stored in the derived point cloud DB8 as derived point cloud data DPC1. Additionally, the point cloud data obtained by combining Layer 1 and Layer 2 is stored in the derived point cloud DB8 as derived point cloud data DPC2. Furthermore, the point cloud data obtained by combining Layer 1, Layer 2, and Layer 3 is stored in the derived point cloud DB8 as derived point cloud data DPC3.

[0049] The generation of each of the derived point cloud data DPCs described above is typically performed on the same day or within a few days of generating the three-dimensional point cloud data of the environment. This generation should ideally be completed at least before imaging using camera 2. This reduces the time required from the start of inspection to the completion of mapping. However, the above derived point cloud data DPCs may also be generated after imaging using camera 2.

[0050] Returning to Figure 2, the acquisition unit 4 is a specific example of an acquisition means. The acquisition unit 4 acquires the captured image stored in the memory of the camera 2 via the communication interface 1d. Alternatively, the acquisition unit 4 may acquire the image by reading it from a storage medium.

[0051] The estimation unit 5 is a specific example of an estimation means. The estimation unit 5 estimates the position and orientation of camera 2 at the time of imaging based on at least one of the multiple derived point cloud data DPCs and the image acquired by the acquisition unit 4. As a method for estimating the position and orientation of camera 2 at the time of imaging, registration technologies such as DeepI2P (Image-to-Point Cloud Registration via Deep Classification), Direct Regression, Monodepth2+USIP, Monodepth2+GT-ICP, and 2D3D-MatchNet can be employed, as described above. Any of these registration technologies output the estimated position and orientation of camera 2 at the time of imaging along with the estimation accuracy (matching score).

[0052] In this embodiment, the estimation unit 5 repeatedly estimates the position and orientation based on the derived point cloud data DPCs and the captured image, sequentially from the derived point cloud data DPCs with relatively high static stability SS to the derived point cloud data DPCs with relatively low static stability SS. Then, depending on whether the accuracy of the position and orientation estimation exceeds a predetermined value, the estimation unit 5 determines the position and orientation of the camera 2 at the time of imaging to be the position and orientation that was last estimated.

[0053] The mapping unit 6 projects the captured image onto a virtual environment model in a virtual space composed of three-dimensional point cloud data of the environment, based on the position and orientation of the camera 2 at the time of imaging, as estimated by the estimation unit 5.

[0054] The output unit 7 displays a virtual environment model on the LCD 1c with the captured image projected onto it. When an inspector selects a deformation visible in the captured image displayed on the LCD 1c using the input means 1e, the output unit 7 highlights the deformation by enclosing it in a rectangle and displays the lengths of the long and short sides of the rectangle near it. This allows the inspector to easily understand the location and magnitude of the deformation in the structure.

[0055] Next, the control flow of the information processing device 1 will be explained with reference to Figure 6.

[0056] First, the acquisition unit 4 acquires the captured image (S100).

[0057] Next, the estimation unit 5 estimates the position and orientation of the camera 2 at the time of imaging based on the derived point cloud data DPC1 and the image acquired by the acquisition unit 4 (S110).

[0058] Next, the estimation unit 5 determines whether the estimation accuracy in step S110 exceeds a predetermined value (S120). If the estimation unit 5 determines that the estimation accuracy in step S110 exceeds a predetermined value, the estimation unit 5 proceeds to step S160. On the other hand, if the estimation unit 5 determines that the estimation accuracy in step S110 does not exceed a predetermined value, the estimation unit 5 proceeds to step S130.

[0059] Next, the estimation unit 5 estimates the position and orientation of the camera 2 at the time of imaging based on the derived point cloud data DPC2 and the image acquired by the acquisition unit 4 (S130).

[0060] Next, the estimation unit 5 determines whether the estimation accuracy in step S130 exceeds a predetermined value (S140). If the estimation unit 5 determines that the estimation accuracy in step S130 exceeds a predetermined value, the estimation unit 5 proceeds to step S160. On the other hand, if the estimation unit 5 determines that the estimation accuracy in step S130 does not exceed a predetermined value, the estimation unit 5 proceeds to step S150.

[0061] Next, the estimation unit 5 estimates the position and orientation of the camera 2 at the time of imaging based on the derived point cloud data DPC3 and the image acquired by the acquisition unit 4 (S150).

[0062] Next, the mapping unit 6 projects the captured image onto a virtual environment model in a virtual space composed of three-dimensional point cloud data of the environment, based on the position and orientation of the camera 2 at the time of imaging, which was last estimated by the estimation unit 5 (S160).

[0063] Next, the output unit 7 displays the virtual environment model on the LCD 1c with the captured image projected onto it (S170).

[0064] The second embodiment has been described above, and the second embodiment has the following features.

[0065] For example, as shown in Figures 2 and 6, the information processing device 1 (position and orientation estimation system, position and orientation estimation device) estimates the position and orientation of camera 2 at the time of imaging based on three-dimensional point cloud data of the environment and an image obtained by imaging a structure (imaging target) included in the environment with camera 2 (imaging device). The information processing device 1 includes a storage unit 3 (storage means), an acquisition unit 4 (acquisition means), and an estimation unit 5 (estimation means). The storage unit 3 stores a plurality of derived point cloud data DPCs generated from the three-dimensional point cloud data and having different static stabilities. The acquisition unit 4 acquires the image. The estimation unit 5 estimates the position and orientation based on at least one of the plurality of derived point cloud data DPCs and the image. With the above configuration, the position and orientation can be estimated with high accuracy.

[0066] Furthermore, as shown in Figure 6, the estimation unit 5 sequentially estimates the position and orientation based on the derived point cloud data DPC and the captured image, starting from derived point cloud data DPC1, which has a relatively high static stability SS, to derived point cloud data DPC3, which has a relatively low static stability SS (S110, S130, S150). Then, depending on whether the accuracy of the position and orientation estimation exceeds a predetermined value, the estimation unit 5 determines the position and orientation of camera 2 at the time of imaging to be the position and orientation estimated last. With this configuration, the position and orientation can be estimated with high accuracy in a short time.

[0067] The following explains the mechanism by which the estimation accuracy changes for each derived point cloud data (DPC), as described above. Specifically, as shown in Figure 3, when a derived point cloud data (DPC) includes objects other than structures, these objects may either contribute to or hinder the estimation of position and orientation. For example, the tree point cloud PPC2 included in derived point cloud data DPC2 can be a valuable feature point for estimating the position and orientation of DPC2. Similarly, the car point cloud PPC3 and pedestrian point cloud PPC4 included in derived point cloud data DPC3 can be valuable feature points for estimating the position and orientation of DPC3. Thus, even if a derived point cloud data (DPC) includes objects other than structures, these objects often contribute to the estimation of position and orientation. Therefore, the estimation accuracy based on derived point cloud data DPC2 and DPC3 may be higher or lower than the estimation accuracy based on derived point cloud data DPC1. As described above, different estimation accuracy can be obtained for each derived point cloud data DPC. Therefore, by providing a storage unit 3 that stores multiple derived point cloud data DPCs having different static stability values ​​SS, the information processing device 1 can estimate the position and orientation with high accuracy.

[0068] The second embodiment described above can be modified as follows, for example.

[0069] That is, in the second embodiment described above, the estimation unit 5 repeatedly estimates the position and orientation based on the derived point cloud data DPCs and the captured image, sequentially from derived point cloud data DPCs with relatively high static stability SS to derived point cloud data DPCs with relatively low static stability SS. However, instead, the estimation unit 5 may repeatedly estimate the position and orientation based on the derived point cloud data DPCs and the captured image, sequentially from derived point cloud data DPCs with relatively low static stability SS to derived point cloud data DPCs with relatively high static stability SS. Alternatively, the estimation unit 5 may randomly rearrange the multiple derived point cloud data DPCs and repeatedly estimate the position and orientation based on the derived point cloud data DPCs and the captured image in the rearranged order.

[0070] Furthermore, in this embodiment, the position and orientation estimation system is implemented by an information processing device 1, which is a single device. However, instead, the position and orientation estimation system may be implemented by distributed processing across multiple devices. That is, the position and orientation estimation system may be implemented by an external server equipped with a derived point cloud DB 8 and an information processing device 1 that can access the derived point cloud DB 8 on the external server.

[0071] (Third embodiment) The third embodiment will be described below with reference to Figure 7. The following description will focus on the differences between this embodiment and the second embodiment, omitting any redundant explanations. Figure 7 shows the processing flow of the information processing device 1.

[0072] First, the acquisition unit 4 acquires the captured image (S200).

[0073] Next, the estimation unit 5 estimates multiple positional orientations based on all the derived point cloud data DPCs held by the derived point cloud DB8 and the captured image (S210). In this embodiment, as shown in Figure 3, the derived point cloud DB8 holds three derived point cloud data DPCs, so the estimation unit 5 estimates three different positional orientations.

[0074] Next, the estimation unit 5 determines the position and orientation of the camera 2 at the time of imaging to be the position and orientation with the highest estimation accuracy among multiple position and orientations (S220).

[0075] Next, the mapping unit 6 projects the captured image onto a virtual environment model in a virtual space composed of three-dimensional point cloud data of the environment, based on the position and orientation of the camera 2 at the time of imaging, as estimated by the estimation unit 5 (S230).

[0076] Then, the output unit 7 displays the virtual environment model on the LCD 1c with the captured image projected onto it (S240).

[0077] With the above configuration, the position and orientation of camera 2 during image capture can be estimated with the highest estimation accuracy.

[0078] (Fourth Embodiment) The fourth embodiment will be described below with reference to Figure 8. The following description will focus on the differences between this embodiment and the second embodiment, omitting any redundant explanations. Figure 8 shows the processing flow of the information processing device 1.

[0079] First, the acquisition unit 4 acquires the captured image (S300).

[0080] Next, the estimation unit 5 receives user input to select one derived point cloud data DPC from among the multiple derived point cloud data DPCs held by the derived point cloud DB8 (S310).

[0081] Next, the estimation unit 5 estimates the position and orientation based on the derived point cloud data DPC specified by the user input and the captured image (S320).

[0082] Next, the mapping unit 6 projects the captured image onto a virtual environment model in a virtual space composed of three-dimensional point cloud data of the environment, based on the position and orientation of the camera 2 at the time of imaging, as estimated by the estimation unit 5 (S330).

[0083] Then, the output unit 7 displays the virtual environment model on the LCD 1c with the captured image projected onto it (S340).

[0084] With the above configuration, the processing time required for the estimation process of the estimation unit 5 can be reduced compared to the second embodiment described above.

[0085] For example, an inspector can select a derived point cloud data DPC (Digital Point Cloud) based on the difference between the time the structure was distanced to generate the three-dimensional point cloud data and the time the structure was imaged with camera 2. If the difference is relatively small, the inspector may select derived point cloud data DPC2 or DPC3, and if the difference is relatively large, they may select derived point cloud data DPC1.

[0086] Furthermore, if the inspector can identify a derived point cloud data DPC that consistently yields high estimation accuracy as a result of repeated mapping using the information processing device 1, they can select that derived point cloud data DPC thereafter. This allows for obtaining a position and orientation with high estimation accuracy while shortening the processing time of the information processing device 1.

[0087] (Fifth embodiment) The fifth embodiment will be described below with reference to Figure 9. The following description will focus on the differences between this embodiment and the second embodiment, omitting any redundant explanations. Figure 9 shows the processing flow of the information processing device 1.

[0088] First, the acquisition unit 4 acquires the captured image (S400).

[0089] Next, the estimation unit 5 estimates the position and orientation based on the previously selected derived point cloud data DPC and the captured image (S410).

[0090] Next, the mapping unit 6 projects the captured image onto a virtual environment model in a virtual space composed of three-dimensional point cloud data of the environment, based on the position and orientation of the camera 2 at the time of imaging, as estimated by the estimation unit 5 (S430).

[0091] Then, the output unit 7 displays the virtual environment model on the LCD 1c with the captured image projected onto it (S440).

[0092] With the above configuration, the processing time required for the estimation process of the estimation unit 5 can be reduced compared to the second embodiment described above.

[0093] For example, if an inspector identifies a derived point cloud data DPC that consistently yields high estimation accuracy as a result of repeated mapping using the information processing device 1, they no longer need to select that derived point cloud data DPC each time they perform an inspection. Instead, they can pre-select and fix the derived point cloud data DPC to be used in subsequent estimation processes. This allows for a reduction in the processing time of the information processing device 1 while obtaining a position and orientation with high estimation accuracy, and also reduces the effort required for user input.

[0094] (Sixth Embodiment) The sixth embodiment will be described below with reference to Figure 10. The following description will focus on the differences between this embodiment and the second embodiment described above, omitting any redundant explanations. Figure 10 shows the processing flow of the information processing device 1.

[0095] First, the acquisition unit 4 acquires the captured image (S500).

[0096] Next, the estimation unit 5 calculates the difference between the distance measurement time when the structure was measured to generate three-dimensional point cloud data and the imaging time when the structure was imaged by the camera 2 (S510).

[0097] Next, the estimation unit 5 selects one derived point cloud data DPC from among the multiple derived point cloud data DPCs held by the derived point cloud DB8 based on the difference calculated in step S510 (S520). Specifically, if the difference is relatively small, the estimation unit 5 selects derived point cloud data DPC2 or derived point cloud data DPC3, and if the difference is relatively large, it selects derived point cloud data DPC1.

[0098] Next, the estimation unit 5 estimates the position and orientation based on the derived point cloud data DPC selected in step S520 and the captured image (S530).

[0099] Next, the mapping unit 6 projects the captured image onto a virtual environment model in a virtual space composed of three-dimensional point cloud data of the environment, based on the position and orientation of the camera 2 at the time of imaging, as estimated by the estimation unit 5 (S540).

[0100] Then, the output unit 7 displays the virtual environment model on the LCD 1c with the captured image projected onto it (S550).

[0101] With the above configuration, the processing time required for the estimation process of the estimation unit 5 can be shortened compared to the second embodiment described above, and since the optimal derived point cloud data DPC is used for the estimation process from among multiple derived point cloud data DPCs, the position and orientation can be estimated with a high estimation accuracy.

[0102] In the above example, the program can be stored and supplied to the computer using various types of non-transitory computer-readable medium. Non-transitory computer-readable medium includes various types of tangible storage medium. Examples of non-transitory computer-readable medium include magnetic storage media (e.g., flexible disks, magnetic tapes, hard disk drives) and magneto-optical storage media (e.g., magneto-optical disks). Examples of non-transitory computer-readable medium further include CD-ROM (Read Only Memory), CD-R, CD-R / W, and semiconductor memory (e.g., mask ROM; examples of non-transitory computer-readable medium further include PROM (Programmable ROM), EPROM (Erasable PROM), flash ROM, and RAM (random access memory)). Alternatively, the program may be supplied to the computer by various types of transient computer-readable medium. Examples of transient computer-readable medium include electrical signals, optical signals, and electromagnetic waves. Temporary computer-readable media can supply programs to a computer via wired communication channels such as electric wires and optical fibers, or via wireless communication channels.

[0103] Some or all of the above embodiments may also be described as follows, but are not limited to the following:

[0104] (Note 1) A position and orientation estimation system that estimates the position and orientation of an imaging device at the time of imaging, based on three-dimensional point cloud data of the environment and an image obtained by imaging an object included in the environment with an imaging device, A storage means for storing multiple derived point cloud data generated from the aforementioned three-dimensional point cloud data, each having a different static stability; The acquisition means for acquiring the aforementioned captured image, An estimation means for estimating the position and orientation based on at least one derived point cloud data from the plurality of derived point cloud data and the captured image, including, Position and orientation estimation system. (Note 2) The estimation means is, From the plurality of derived point cloud data, starting with those with relatively high static stability and moving toward those with relatively low static stability, the estimation of the position and orientation based on the derived point cloud data and the captured image is repeated. In response to the estimation accuracy of the position and orientation exceeding a predetermined value, the position and orientation of the imaging device during imaging is determined to be the last estimated position and orientation. The position and orientation estimation system described in Appendix 1. (Note 3) The estimation means is, Multiple positional orientations are estimated based on all derived point cloud data stored in the storage means and the captured image. The position and orientation of the imaging device during imaging is determined to be the position and orientation with the highest estimation accuracy among the plurality of position and orientations. The position and orientation estimation system described in Appendix 1. (Note 4) The estimation means is, The difference between the time the captured image is captured and the time the three-dimensional point cloud data is measured is calculated, and based on the calculated difference, one of the multiple derived point cloud data is selected, and the position and orientation are estimated based on the selected derived point cloud data and the captured image. The position and orientation estimation system described in Appendix 1. (Note 5) The estimation means is, The position and orientation are estimated based on the derived point cloud data selected in advance from the plurality of derived point cloud data and the captured image. The position and orientation estimation system described in Appendix 1. (Note 6) The estimation means is, The position and orientation are estimated based on the derived point cloud data specified by user input and the captured image, among the plurality of derived point cloud data. The position and orientation estimation system described in Appendix 1. (Note 7) The static stability of the derived point cloud data corresponds to the length of the maintenance period of the object that maintains a stationary state on the time axis for the shortest period among the multiple objects represented by the derived point cloud data. A position and orientation estimation system described in any one of the items 1 through 6 of the appendix. (Note 8) A position and orientation estimation system that estimates the position and orientation of an imaging device at the time of imaging, based on three-dimensional point cloud data of the environment and an image obtained by imaging an object included in the environment with an imaging device, A storage means for storing multiple derived point cloud data generated from the aforementioned three-dimensional point cloud data, each having a different static stability; The acquisition means for acquiring the aforementioned captured image, An estimation means for estimating the position and orientation based on at least one derived point cloud data from the plurality of derived point cloud data and the captured image, including, Position and orientation estimation device. (Note 9) The estimation means is, From the plurality of derived point cloud data, starting with those with relatively high static stability and moving toward those with relatively low static stability, the estimation of the position and orientation based on the derived point cloud data and the captured image is repeated. In response to the estimation accuracy of the position and orientation exceeding a predetermined value, the position and orientation of the imaging device during imaging is determined to be the last estimated position and orientation. The position and orientation estimation device described in Appendix 8. (Note 10) The estimation means is, Multiple positional orientations are estimated based on all derived point cloud data stored in the storage means and the captured image. The position and orientation of the imaging device during imaging is determined to be the position and orientation with the highest estimation accuracy among the plurality of position and orientations. The position and orientation estimation device described in Appendix 8. (Note 11) The estimation means is, The difference between the time the captured image is captured and the time the three-dimensional point cloud data is measured is calculated, and based on the calculated difference, one of the multiple derived point cloud data is selected, and the position and orientation are estimated based on the selected derived point cloud data and the captured image. The position and orientation estimation device described in Appendix 8. (Note 12) The estimation means is, The position and orientation are estimated based on the derived point cloud data selected in advance from the plurality of derived point cloud data and the captured image. The position and orientation estimation device described in Appendix 8. (Note 13) The estimation means is, The position and orientation are estimated based on the derived point cloud data specified by user input and the captured image, among the plurality of derived point cloud data. The position and orientation estimation device described in Appendix 8. (Note 14) A position and orientation estimation method for estimating the position and orientation of an imaging device at the time of imaging, based on three-dimensional point cloud data of the environment and an image obtained by imaging an object included in the environment with an imaging device, The acquisition step of acquiring the aforementioned captured image, An estimation step of estimating the position and orientation based on the captured image and at least one derived point cloud data from a plurality of derived point cloud data generated from the three-dimensional point cloud data, each having a different static stability; including, Position and orientation estimation method. (Note 15) In the estimation step described above, From the plurality of derived point cloud data, starting with those with relatively high static stability and moving toward those with relatively low static stability, the estimation of the position and orientation based on the derived point cloud data and the captured image is repeated. In response to the estimation accuracy of the position and orientation exceeding a predetermined value, the position and orientation of the imaging device during imaging is determined to be the last estimated position and orientation. The position and orientation estimation method described in Appendix 14. (Note 16) In the estimation step described above, Multiple positional orientations are estimated based on all of the derived point cloud data and the captured image. The position and orientation of the imaging device during imaging is determined to be the position and orientation with the highest estimation accuracy among the plurality of position and orientations. The position and orientation estimation method described in Appendix 14. (Note 17) In the estimation step described above, The difference between the time the captured image is captured and the time the three-dimensional point cloud data is measured is calculated, and based on the calculated difference, one of the multiple derived point cloud data is selected, and the position and orientation are estimated based on the selected derived point cloud data and the captured image. The position and orientation estimation method described in Appendix 14. (Note 18) In the estimation step described above, The position and orientation are estimated based on the derived point cloud data selected in advance from the plurality of derived point cloud data and the captured image. The position and orientation estimation method described in Appendix 14. (Note 19) In the estimation step described above, The position and orientation are estimated based on the derived point cloud data specified by user input and the captured image, among the plurality of derived point cloud data. The position and orientation estimation method described in Appendix 14. (Note 20) A position and orientation estimation program for estimating the position and orientation of an imaging device at the time of imaging, based on three-dimensional point cloud data of the environment and an image obtained by imaging an object included in the environment with the imaging device, On the computer, The acquisition step of acquiring the aforementioned captured image, An estimation step of estimating the position and orientation based on the captured image and at least one derived point cloud data from a plurality of derived point cloud data generated from the three-dimensional point cloud data, each having a different static stability; To execute Position and orientation estimation program. (Note 21) In the estimation step described above, From the plurality of derived point cloud data, starting with those with relatively high static stability and moving toward those with relatively low static stability, the estimation of the position and orientation based on the derived point cloud data and the captured image is repeated. In response to the estimation accuracy of the position and orientation exceeding a predetermined value, the position and orientation of the imaging device during imaging is determined to be the last estimated position and orientation. The position and orientation estimation program described in Appendix 20. (Note 22) In the estimation step described above, Multiple positional orientations are estimated based on all of the derived point cloud data and the captured image. The position and orientation of the imaging device during imaging is determined to be the position and orientation with the highest estimation accuracy among the plurality of position and orientations. The position and orientation estimation program described in Appendix 20. (Note 23) In the estimation step described above, The difference between the time the captured image is captured and the time the three-dimensional point cloud data is measured is calculated, and based on the calculated difference, one of the multiple derived point cloud data is selected, and the position and orientation are estimated based on the selected derived point cloud data and the captured image. The position and orientation estimation program described in Appendix 20. (Note 24) In the estimation step described above, The position and orientation are estimated based on the derived point cloud data selected in advance from the plurality of derived point cloud data and the captured image. The position and orientation estimation program described in Appendix 20. (Note 25) In the estimation step described above, The position and orientation are estimated based on the derived point cloud data specified by user input and the captured image, among the plurality of derived point cloud data. The position and orientation estimation program described in Appendix 20. [Industrial applicability]

[0105] This disclosure can be applied to techniques for estimating the position and orientation of an imaging device during imaging. [Explanation of symbols]

[0106] 1. Information Processing Device 1b Memory 1d communication interface 1e Input means 2 cameras 3 Storage section 4 Acquisition part 5 Estimation part 6. Mapping section 7 Output section 8 Derived point cloud DB DPC-derived point cloud data DPC1 Derived Point Cloud Data DPC2 Derived Point Cloud Data DPC3 Derived Point Cloud Data SS static stability PPC1 Building point cloud PPC2 Tree point cloud PPC3 Automobile Point Cloud PPC4 Pedestrian Point Cloud

Claims

1. A position and orientation estimation system that estimates the position and orientation of an imaging device at the time of imaging, based on three-dimensional point cloud data of an environment and an image obtained by imaging an object included in the environment with an imaging device, A storage means for storing multiple derived point cloud data generated from the aforementioned three-dimensional point cloud data, each having a different static stability; The acquisition means for acquiring the aforementioned captured image, An estimation means for estimating the position and orientation based on at least one derived point cloud data from the plurality of derived point cloud data and the captured image, Includes, The estimation means repeatedly estimates the position and orientation based on the derived point cloud data and the captured image, sequentially from the derived point cloud data with relatively high static stability to the derived point cloud data with relatively low static stability, and determines the position and orientation of the imaging device at the time of imaging to be the last estimated position and orientation, depending on whether the estimation accuracy of the position and orientation exceeds a predetermined value. Position and orientation estimation system.

2. A position and orientation estimation system that estimates the position and orientation of an imaging device at the time of imaging, based on three-dimensional point cloud data of an environment and an image obtained by imaging an object included in the environment with an imaging device, A storage means for storing multiple derived point cloud data generated from the aforementioned three-dimensional point cloud data, each having a different static stability; The acquisition means for acquiring the aforementioned captured image, An estimation means for estimating the position and orientation based on at least one derived point cloud data from the plurality of derived point cloud data and the captured image, Includes, The estimation means estimates a plurality of positional orientations based on all derived point cloud data stored in the storage means and the captured image, and determines the positional orientation of the imaging device at the time of imaging to be the positional orientation with the highest estimation accuracy among the plurality of positional orientations. Position and orientation estimation system.

3. A position and orientation estimation system that estimates the position and orientation of an imaging device at the time of imaging, based on three-dimensional point cloud data of an environment and an image obtained by imaging an object included in the environment with an imaging device, A storage means for storing multiple derived point cloud data generated from the aforementioned three-dimensional point cloud data, each having a different static stability; The acquisition means for acquiring the aforementioned captured image, An estimation means for estimating the position and orientation based on at least one derived point cloud data from the plurality of derived point cloud data and the captured image, Includes, The estimation means calculates the difference between the time the captured image is captured and the time the three-dimensional point cloud data is measured, selects one of the plurality of derived point cloud data based on the calculated difference, and estimates the position and orientation based on the selected derived point cloud data and the captured image. Position and orientation estimation system.

4. A position and orientation estimation system that estimates the position and orientation of an imaging device at the time of imaging, based on three-dimensional point cloud data of an environment and an image obtained by imaging an object included in the environment with an imaging device, A storage means for storing multiple derived point cloud data generated from the aforementioned three-dimensional point cloud data, each having a different static stability; The acquisition means for acquiring the aforementioned captured image, An estimation means for estimating the position and orientation based on at least one derived point cloud data from the plurality of derived point cloud data and the captured image, Includes, The static stability of the derived point cloud data corresponds to the length of the maintenance period of the object that maintains a stationary state on the time axis for the shortest period among the multiple objects represented by the derived point cloud data. Position and orientation estimation system.

5. A position and orientation estimation system that estimates the position and orientation of an imaging device at the time of imaging, based on three-dimensional point cloud data of an environment and an image obtained by imaging an object included in the environment with an imaging device, A storage means for storing multiple derived point cloud data generated from the aforementioned three-dimensional point cloud data, each having a different static stability; The acquisition means for acquiring the aforementioned captured image, An estimation means for estimating the position and orientation based on at least one derived point cloud data from the plurality of derived point cloud data and the captured image, Includes, The estimation means repeatedly estimates the position and orientation based on the derived point cloud data and the captured image, sequentially from the derived point cloud data with relatively high static stability to the derived point cloud data with relatively low static stability, and determines the position and orientation of the imaging device at the time of imaging to be the last estimated position and orientation, depending on whether the estimation accuracy of the position and orientation exceeds a predetermined value. Position and orientation estimation device.

6. A position and orientation estimation method for estimating the position and orientation of an imaging device at the time of imaging, based on three-dimensional point cloud data of the environment and an image obtained by imaging an object included in the environment with an imaging device, Computers Multiple derived point cloud data, generated from the aforementioned three-dimensional point cloud data and having different static stabilities, are stored. The aforementioned captured image is acquired, The position and orientation are estimated based on at least one of the multiple derived point cloud data and the captured image. The estimation described above involves repeatedly estimating the position and orientation based on the derived point cloud data and the captured image, starting from the derived point cloud data with relatively high static stability to the derived point cloud data with relatively low static stability, and determining the position and orientation of the imaging device at the time of imaging to be the last estimated position and orientation, depending on whether the estimation accuracy of the position and orientation exceeds a predetermined value. Position and orientation estimation method.

7. A position and orientation estimation method for estimating the position and orientation of an imaging device at the time of imaging, based on three-dimensional point cloud data of the environment and an image obtained by imaging an object included in the environment with an imaging device, Computers Multiple derived point cloud data, generated from the aforementioned three-dimensional point cloud data and having different static stabilities, are stored. The aforementioned captured image is acquired, The position and orientation are estimated based on at least one of the multiple derived point cloud data and the captured image. The estimation involves estimating multiple positional orientations based on all stored derived point cloud data and the captured image, and determining the positional orientation of the imaging device at the time of imaging to be the positional orientation with the highest estimation accuracy among the multiple positional orientations. Position and orientation estimation method.

8. A position and orientation estimation method for estimating the position and orientation of an imaging device at the time of imaging, based on three-dimensional point cloud data of the environment and an image obtained by imaging an object included in the environment with an imaging device, Computers Multiple derived point cloud data, generated from the aforementioned three-dimensional point cloud data and having different static stabilities, are stored. The aforementioned captured image is acquired, The position and orientation are estimated based on at least one of the multiple derived point cloud data and the captured image. The estimation involves calculating the difference between the acquisition time of the captured image and the distance measurement time of the three-dimensional point cloud data, selecting one of the multiple derived point cloud data based on the calculated difference, and estimating the position and orientation based on the selected derived point cloud data and the captured image. Position and orientation estimation method.

9. A position and orientation estimation method for estimating the position and orientation of an imaging device at the time of imaging, based on three-dimensional point cloud data of the environment and an image obtained by imaging an object included in the environment with an imaging device, Computers Multiple derived point cloud data, generated from the aforementioned three-dimensional point cloud data and having different static stabilities, are stored. The aforementioned captured image is acquired, The position and orientation are estimated based on at least one of the multiple derived point cloud data and the captured image. The static stability of the derived point cloud data corresponds to the length of the maintenance period of the object that maintains a stationary state on the time axis for the shortest period among the multiple objects represented by the derived point cloud data. Position and orientation estimation method.

10. Computers, The position and orientation estimation method described in any one of claims 6 to 9 is performed. program.