Coordinate estimation system, coordinate estimation device, and coordinate estimation method
The coordinate estimation system aligns and estimates pixel coordinates in point clouds, addressing the resolution gap between point clouds and images to efficiently manage deformations in large concrete structures.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2023-03-27
- Publication Date
- 2026-05-15
AI Technical Summary
Existing methods struggle to accurately estimate the positional relationship between point clouds and captured images of concrete structures with significantly different resolutions, which is crucial for efficient maintenance management of large infrastructure like bridges and tunnels.
A coordinate estimation system and method that aligns close-up and distant images with a point cloud, estimating the imaging position and orientation, and calculating specific pixel coordinates in the point cloud coordinate system using alignment and position estimation techniques.
Enables accurate estimation of the positional relationship between point clouds and images with different resolutions, facilitating efficient superimposition and management of deformations in large concrete structures.
Smart Images

Figure 0007859594000001 
Figure 0007859594000002 
Figure 0007859594000003
Abstract
Description
Technical Field
[0001] The present invention relates to a coordinate estimation system, a coordinate estimation device, and a coordinate estimation method.
Background Art
[0002] Patent Document 1 discloses a technique for generating a superimposed image obtained by superimposing an image captured using an infrared camera on a three-dimensional image showing a concrete structure part.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] By the way, some of the existing infrastructures have already advanced considerably in aging. In particular, in the case of concrete structures, when deformations such as cracks, floating, and peeling are discovered during inspection, it is an important issue to appropriately record and manage the position and size of the deformation.
[0005] Therefore, the inventors of the present application have developed a digital twin of a concrete structure in order to realize efficient maintenance management of the concrete structure. Specifically, they are considering superimposing an imaging image of a deformation discovered by an inspector during inspection on the point cloud of the concrete structure. By referring to the point cloud on which the imaging image is superimposed, it will be possible to efficiently grasp the position and size of the deformation.
[0006] However, when the aforementioned concrete structures are large, such as bridges, dams, or tunnels, three-dimensional distance measurements of the concrete structures are performed from a distance of several tens of meters. Considering the practical resolution of such three-dimensional distance measurements, the resolution of the point cloud obtained by these measurements will be at most one point per square centimeter. In other words, there will be approximately only one point corresponding to a one-square-centimeter area of the surface of the concrete structure.
[0007] In contrast, when inspectors take images of deformations during inspections, the images are taken from several meters away from the deformation. Considering the practical resolution of such images, the resolution of the images obtained is approximately 2500 pixels per square centimeter. That is, there are roughly 2500 pixels corresponding to a one-square-centimeter area on the surface of a concrete structure.
[0008] Because of the large resolution gap between the point cloud and the captured image, it is considered difficult to estimate where in the point cloud the captured image corresponds.
[0009] The purpose of this disclosure is to provide a technique for estimating the positional relationship between point clouds and captured images that have significantly different resolutions. [Means for solving the problem]
[0010] According to a first aspect of this disclosure, a coordinate estimation system is provided, which includes: alignment means for aligning a close-up image obtained by imaging a first imaging area of an object and at least one distant image obtained by imaging a second imaging area of the object that is larger than the first imaging area and includes the first imaging area; position and orientation estimation means for estimating the imaging position and orientation, including the imaging position and orientation in the point cloud coordinate system of an imaging device that captured the at least one distant image, based on a point cloud of the object and the at least one distant image; and coordinate estimation means for estimating the coordinates of a specific pixel in the point cloud coordinate system of at least one specific pixel, which is an arbitrary pixel of the close-up image, based on the point cloud, the alignment result by the alignment means, and the imaging position and orientation. According to a second aspect of this disclosure, a coordinate estimation device is provided, which includes: alignment means for aligning a close-up image obtained by imaging a first imaging area of an object and at least one distant image obtained by imaging a second imaging area of the object that is larger than the first imaging area and includes the first imaging area; position and orientation estimation means for estimating the imaging position and orientation, including the imaging position and orientation in the point cloud coordinate system of an imaging device that captured the at least one distant image, based on a point cloud of the object and the at least one distant image; and coordinate estimation means for estimating the coordinates of a specific pixel in the point cloud coordinate system of at least one specific pixel that is an arbitrary pixel of the close-up image, based on the point cloud, the alignment result by the alignment means, and the imaging position and orientation. A third aspect of this disclosure provides a coordinate estimation method comprising: an alignment step in which a computer aligns a foreground image obtained by imaging a first imaging area of an object with at least one background image obtained by imaging a second imaging area of the object that is larger than the first imaging area and includes the first imaging area; a position and orientation estimation step in which the computer estimates an imaging position and orientation, including the imaging position and orientation in the point cloud coordinate system of an imaging device that captured the at least one background image, based on a point cloud of the object and the at least one background image; and a coordinate estimation step in which the computer estimates the coordinates of a specific pixel in the point cloud coordinate system of at least one specific pixel that is any pixel of the foreground image, based on the point cloud, the alignment result from the alignment step, and the imaging position and orientation. [Effects of the Invention]
[0011] According to this disclosure, it is possible to estimate the positional relationship between point clouds and captured images that have significantly different resolutions. [Brief explanation of the drawing]
[0012] [Figure 1] This is a functional block diagram of the coordinate estimation system. (Summary of this disclosure) [Figure 2] This is a functional block diagram of the coordinate estimation device. (First Embodiment) [Figure 3] This is the workflow for the preparation phase. (First Embodiment) [Figure 4] This is a photograph of a bridge. (First embodiment) [Figure 5] This is a three-dimensional image of a bridge composed of point clouds. (First embodiment) [Figure 6] This figure shows the imaging areas for the foreground and background images. (First Embodiment) [Figure 7] This is an example of a close-up image. (First embodiment) [Figure 8] This is an example of a distant view image. (First embodiment) [Figure 9] This is a data structure diagram of the image storage unit. (First Embodiment) [Figure 10] It is an explanatory diagram of a homography matrix. (First Embodiment) [Figure 11] It is an explanatory diagram of the operation of the coordinate estimation unit. (First Embodiment) [Figure 12] It is an explanatory diagram of the operation of the coordinate estimation unit. (First Embodiment) [Figure 13] It is an example of a superimposed image. (First Embodiment) [Figure 14] It is an output example of a superimposed image. (First Embodiment) [Figure 15] It is a data structure diagram of the history DB. (First Embodiment) [Figure 16] It is a control flow of the coordinate estimation device. (First Embodiment) [Figure 17] It is an explanatory diagram of the operation of the coordinate estimation unit. (Second Embodiment) [Figure 18] It is an explanatory diagram of the operation of the coordinate estimation unit. (Third Embodiment) [Figure 19] It is a plan view showing imaging conditions of a close-up image and a long-distance image. (Fourth Embodiment) [Figure 20] It is a control flow of the coordinate estimation device. (Fourth Embodiment) [Figure 21] It is a control flow of the coordinate estimation device. (Fourth Embodiment) [Figure 22] It is a plan view showing the positional relationship between an imaging position, a bridge, and the position of specific pixel coordinates. (Fifth Embodiment) [Figure 23] It is a diagram exemplifying the angle for each extracted long-distance image. (Fifth Embodiment) [Figure 24] It is a control flow of the coordinate estimation device. (Fifth Embodiment) [Figure 25] It is a control flow of the coordinate estimation device. (Fifth Embodiment)
Modes for Carrying Out the Invention
[0013] (Overview of the Present Disclosure) The outline of this disclosure will be described below with reference to Figure 1. As shown in Figure 1, the coordinate estimation system 100 includes alignment means 101, position and orientation estimation means 102, and coordinate estimation means 103.
[0014] The alignment means 101 aligns a close-up image obtained by imaging a first imaging area of the object with at least one distant image obtained by imaging a second imaging area that is larger than the first imaging area of the object and includes the first imaging area.
[0015] The position and orientation estimation means 102 estimates the imaging position and orientation, including the imaging position and orientation of at least one distant image, based on the point cloud of the object and at least one distant image.
[0016] The coordinate estimation means 103 estimates the coordinates of at least one specific pixel, which is any pixel in the foreground image, in the point cloud coordinate system, based on the point cloud, the alignment result by the alignment means, and the imaging position and orientation.
[0017] With the above configuration, it is possible to estimate the positional relationship between point clouds and foreground images (captured images) that have significantly different resolutions.
[0018] (First Embodiment) The first embodiment of this disclosure will be described below with reference to Figures 2 to 16.
[0019] Figure 2 shows a functional block diagram of the coordinate estimation device 1. The coordinate estimation device 1 shown in Figure 2 is used to achieve efficient maintenance of concrete structures, for example, by superimposing images of deformations captured during inspection of concrete structures onto a point cloud of the concrete structure, such as bridges, dams, and tunnels. Concrete structures are one specific example of objects that are subject to maintenance. Deformations in concrete structures typically include cracks, delamination, and spalling. Hereafter, bridges will be used as an example of a concrete structure that is subject to maintenance.
[0020] (Preparation Flow) Here, referring to Figure 3, we will explain the preparatory flow to be performed before actually using the coordinate estimation device 1. As shown in Figure 3, first, a point cloud of the bridge is prepared by measuring the distance of the bridge prior to inspecting the bridge (S100). Figure 4 shows a photograph of bridge 2. Figure 5 shows the point cloud of bridge 2. Methods for generating the point cloud of bridge 2 shown in Figure 5 include using LiDAR (Light Detection And Ranging) and using photogrammetry.
[0021] In the LiDAR-based method, the distance to bridge 2 is measured from various angles using LiDAR, and the point clouds output from the LiDAR are combined using registration techniques such as ICP (Iterative Closest Point) to generate a point cloud of bridge 2.
[0022] In the photogrammetry method, the three-dimensional structure of bridge 2 is reconstructed by solving a geometric inverse problem from multiple images obtained by photographing bridge 2 from various angles, thereby generating a point cloud of bridge 2. A typical method for reconstructing the three-dimensional structure of bridge 2 from multiple images is Structure from Motion (SfM). In this case, using Multi-View Stereo (MVS) in conjunction can generate a more precise point cloud of bridge 2.
[0023] Alternatively, a point cloud of Bridge 2 may be generated using both LiDAR and photogrammetry. That is, a point cloud of Bridge 2 may be generated by combining the point cloud of Bridge 2 generated using LiDAR and the point cloud of Bridge 2 generated by photogrammetry using the registration technique described above.
[0024] Returning to Figure 3, when the inspection timing set, for example, once every five years, arrives (S110: YES), the inspector visually inspects Bridge 2 to check for any deformation (S120). If there is deformation in Bridge 2, the inspector takes a close-up image of the deformation with the imaging device (S130). Subsequently, the inspector takes a distant image of the deformation with the imaging device (S140). Now, referring to Figures 6 to 8, the close-up and distant images will be explained. In Figure 6, the close-up imaging area R1 (first imaging area), which is the imaging area for the close-up image, and the distant imaging area R2 (second imaging area), which is the imaging area for the distant image, are shown by solid rectangles. Figures 7 and 8 show the close-up image 3 and the distant image 4, respectively. As shown in Figure 6, both the close-up imaging area R1 and the distant imaging area R2 are imaging areas that include deformation 5. The distant imaging area R2 is a larger imaging area than the near imaging area R1. The distant imaging area R2 is an imaging area that includes at least the near imaging area R1. In this embodiment, when an inspector discovers a deformation 5 on the bridge 2, the inspector first captures a near-view image 3 of the deformation 5 with the telephoto side of the imaging device, and then captures a distant image 4 of the deformation 5 with the wide-angle side of the imaging device. By utilizing the telephoto and wide-angle sides of the imaging device in this way, the near-view image 3 and the distant image 4 can be captured in a short time. However, instead, the inspector may move closer to the deformation 5 to capture a near-view image 3 of the deformation 5, and then move far away from the deformation 5 to capture a distant image 4 of the deformation 5.
[0025] In this disclosure, "foreground" and "background" merely define relative characteristics, not absolute characteristics. When interpreting the technical scope of this disclosure, no interpretation that deviates from these definitions should be made.
[0026] Returning to Figure 3, once close-up images 3 and distant images 4 have been captured for all deformations 5 (S150: YES), the inspector waits until the next inspection timing (S160), and when the next inspection timing arrives, steps S120 to S150 are executed again.
[0027] (Coordinate estimation device 1) Returning to Figure 2, the coordinate estimation device 1 comprises a CPU 1a (Central Processing Unit), memory 1b, LCD 1c (Liquid Crystal Display), and media R / W 1d.
[0028] Memory 1b consists of RAM (Random Access Memory), ROM (Read Only Memory), HDD (Hard Disk Drive), etc. The control program is stored in Memory 1b.
[0029] The CPU 1a reads and executes the control program stored in memory 1b. This control program then causes the CPU 1a and other hardware to function as various functional units. These functional units include a data reception unit 10, a point cloud storage unit 11, an image storage unit 12, a positioning unit 13, a position and orientation estimation unit 14, and a coordinate estimation unit 15. The functional units also include a superimposed image generation unit 20, a superimposed image output unit 21, a history database 22, a history database update unit 23, a history database extraction unit 24, and a history image output unit 25.
[0030] In this embodiment, the coordinate estimation device 1 is implemented by a single device. However, instead, the coordinate estimation device 1 may be implemented by distributed processing using multiple devices.
[0031] The data receiving unit 10 receives the point cloud of the bridge 2, multiple close-up images 3, and multiple distant images 4 via the media R / W 1d. The data receiving unit 10 stores the point cloud of the bridge 2 in the point cloud storage unit 11, and stores the multiple close-up images 3 and multiple distant images 4 in the image storage unit 12.
[0032] Figure 9 is a data structure diagram of the image storage unit 12. Multiple images are stored in the image storage unit 12, associated with the date and time of acquisition. In this embodiment, when an inspector discovers a deformation 5 on the bridge 2, the rule is to first capture a close-up image 3 of the deformation 5 with the telephoto side of the imaging device, and then capture a distant image 4 of the deformation 5 with the wide-angle side of the imaging device. In this case, since the close-up image 3 and distant image 4 corresponding to the same deformation 5 can be captured within a few minutes, according to Figure 9, it can be seen that image 1 is the close-up image 3, and the distant image 4 corresponding to the close-up image 3 is image 2. Similarly, image 3 is the close-up image 3, and the distant image 4 corresponding to the close-up image 3 is image 4. Similarly, image 5 is the close-up image 3, and the distant image 4 corresponding to the close-up image 3 is image 6. In the image storage unit 12 shown in Figure 9, multiple images may be stored in association with imaging conditions such as focal length, in addition to the date and time of acquisition.
[0033] The alignment unit 13 aligns the foreground image 3 and background image 4 corresponding to the same deformation 5. Specifically, the alignment unit 13 calculates a homography matrix established between the foreground image 3 and background image 4 corresponding to the same deformation 5. Figure 10 shows an explanatory diagram of the homography matrix H established between the foreground image 3 and background image 4 corresponding to the same deformation 5. The alignment unit 13 detects feature points in both the foreground image 3 and background image 4 corresponding to the same deformation 5, and associates similar feature points between the foreground image 3 and background image 4 corresponding to the same deformation 5. The alignment unit 13 calculates the above homography matrix based on the correspondence between the feature points. At this time, the alignment unit 13 can ensure the reliability of the calculation result by calculating the homography matrix using RANSAC (Random Sample Consensus). As shown in Figure 10, the homography matrix H is a matrix used to transform any coordinate (u0,v0) in the foreground coordinate system uv, which is the coordinate system of the foreground image 3, to coordinate (u'0,v'0) in the background coordinate system u'-v', which is the coordinate system of the background image 4.
[0034] The position and orientation estimation unit 14 estimates the imaging position and orientation, including the imaging position and orientation of the imaging device that captured the distant image 4, based on the point cloud of the bridge 2 and the distant image 4. Here, the point cloud coordinate system is the coordinate system that defines the coordinates of the point cloud of the bridge 2. The imaging position and orientation can be estimated by known techniques. An example of a known technique is "C. Jaramillo, et al., "6-DoF pose localization in 3d point-cloud dense maps using a monocular camera," 2013." In short, the estimation of the imaging position and orientation can be performed by repeatedly comparing the projected image, which is formed by projecting the point cloud of the bridge 2 at an arbitrary imaging position and orientation, with the distant image 4 while changing the imaging position and orientation, and searching for an imaging position and orientation that matches the two as closely as possible. At this time, the position and orientation estimation unit 14 can also simultaneously determine the imaging conditions of the distant image 4. The imaging condition is typically the focal length.
[0035] The coordinate estimation unit 15 estimates the coordinates of at least one specific pixel Q, which is any pixel in the foreground image 3, in the point cloud coordinate system, based on the point cloud of the bridge 2, the alignment result by the alignment unit 13, and the imaging position and orientation estimated by the position and orientation estimation unit 14. That is, as shown in Figure 10, the coordinate estimation unit 15 calculates the coordinates of at least one specific pixel Q in the distant coordinate system u'-v' based on the homography matrix H as the alignment result by the alignment unit 13. The coordinate estimation unit 15 estimates the coordinates of the specific pixel based on the point cloud of the bridge 2, the calculation result, and the imaging position and orientation estimated by the position and orientation estimation unit 14.
[0036] Figure 11 shows the positional relationship between the imaging position, the distant image 4, and the point cloud. Figure 12 is a diagram illustrating collision detection between the projection line L extending from the imaging position and the point cloud. As shown in Figures 11 and 12, the coordinate estimation unit 15 sets the distant image 4 in the point cloud coordinate system at a position separated from the imaging position by the focal length at the time of imaging, in the direction of imaging. At this time, the line segment M that passes through the central point of the distant image 4 and is perpendicular to the distant image 4 passes through the imaging position. In this state, the coordinate estimation unit 15 calculates the projection line L that radiates from the imaging position toward a specific pixel Q converted to the distant coordinate system u'-v', and performs collision detection between the projection line L and the bridge 2. The projection line L is calculated as the equation of the line segment in the point cloud coordinate system. Then, based on the result of this detection, the coordinates of the specific pixel Q in the point cloud coordinate system are estimated.
[0037] However, as shown in Figure 12, since bridge 2 is represented by a point cloud, it is practically impossible for the projection line L to collide with the point cloud of bridge 2. Therefore, the coordinate estimation unit 15 extracts a sub-point cloud from the point cloud of bridge 2, from points p1 to p25 of bridge 2, where the distance from the projection line L is less than or equal to a predetermined value. In the example in Figure 12, the shortest distance d9 between point p9 and the projection line L, and the shortest distance d10 between point p10 and the projection line L are both less than or equal to the predetermined value. Therefore, the coordinate estimation unit 15 extracts points p9 and p10 as sub-point clouds from the point cloud of bridge 2 (points p1 to p25). Then, the coordinate estimation unit 15 selects the point p9 and p10 that are closest to the imaging position. In the example in Figure 12, point p10 is slightly closer to the imaging position than point p9. Therefore, the coordinate estimation unit 15 selects point p10 as the point closest to the imaging position among points p9 and p10, and estimates the coordinates of a specific pixel Q in the foreground image 3 as coordinates in the point cloud coordinate system based on the coordinates of point p10. In short, the coordinates of a specific pixel Q in the foreground image 3 in the point cloud coordinate system coincide with the coordinates of point p10.
[0038] Here, as shown in Figure 10, the deformation 5 in the foreground image 3 generally spans multiple pixels. Therefore, the coordinate estimation unit 15 estimates multiple specific pixel coordinates by performing collision detection on each of the multiple pixels that constitute the deformation 5 in the foreground image 3. Alternatively, in order to reduce computational cost, the coordinate estimation unit 15 may sample some pixels from the multiple pixels that constitute the deformation 5 in the foreground image 3 and perform collision detection on each of the sampled pixels. In this case, the coordinates of pixels that were not sampled may be estimated from the specific pixel coordinates of multiple pixels adjacent to those pixels.
[0039] Furthermore, the coordinate estimation unit 15 can detect deformation 5 using, for example, a well-known DNN (Deep Neural Network) such as R-CNN (Regions with Convolutional Neural Networks) or YOLO (You Only Look Once).
[0040] As shown in Figure 13, the superimposed image generation unit 20 superimposes the foreground image 3 onto the point cloud of the bridge 2 based on the specific pixel coordinates estimated by the coordinate estimation unit 15. Specifically, the superimposed image generation unit 20 generates a superimposed image 20a by superimposing the portion of the foreground image 3 occupied by the deformation 5 onto the three-dimensional image consisting of the point cloud of the bridge 2.
[0041] The superimposed image output unit 21 outputs the superimposed image 20a to the LCD 1c.
[0042] As shown in Figure 15, the history DB22 stores the close-up image 3, the date it was captured, and the specific pixel coordinates of the portion of the close-up image 3 occupied by the deformation 5, associating them with each other.
[0043] The history database update unit 23 stores the estimation results from the coordinate estimation unit 15 in the history database 22. Specifically, the history database update unit 23 stores in the history database 22 the foreground image 3 processed by the coordinate estimation unit 15, the date the foreground image 3 was captured, and the specific pixel coordinates of the portion of the foreground image 3 occupied by the deformation 5, associating them with each other.
[0044] The history database extraction unit 24 uses the representative specific pixel coordinates of the deformation 5 of the foreground image 3 processed by the coordinate estimation unit 15 as a key to perform a search within the history database 22 and extracts images associated with the same specific pixel coordinates.
[0045] As shown in Figure 14, the history image output unit 25 outputs the image extracted by the history DB extraction unit 24, along with its imaging date, to the LCD 1c displaying the superimposed image 20a. This makes it easy to visually grasp the changes in deformation 5 over time.
[0046] Next, with reference to Figure 16, the control flow of the coordinate estimation device 1 will be briefly explained.
[0047] First, the data receiving unit 10 receives the point cloud of the bridge 2, multiple close-up images 3, and multiple distant images 4 via the medium R / W 1d (S200).
[0048] Next, the alignment unit 13 aligns the near view image 3 and the far view image 4 that correspond to the same deformation 5 (S210).
[0049] Next, the position and orientation estimation unit 14 estimates the imaging position and orientation, including the imaging position and orientation of the imaging device that captured the distant image 4, based on the point cloud of the bridge 2 and the distant image 4 (S220).
[0050] Next, the coordinate estimation unit 15 estimates the coordinates of at least one specific pixel Q, which is any pixel in the foreground image 3, in the point cloud coordinate system based on the point cloud of the bridge 2, the alignment result by the alignment unit 13, and the imaging position and orientation estimated by the position and orientation estimation unit 14 (S230).
[0051] Next, the superimposed image generation unit 20 superimposes the foreground image 3 onto the point cloud of the bridge 2 based on the specific pixel coordinates estimated by the coordinate estimation unit 15 (S240).
[0052] Next, the superimposed image output unit 21 outputs the superimposed image 20a to the LCD 1c (S250).
[0053] Next, the history DB update unit 23 stores the estimation results from the coordinate estimation unit 15 in the history DB 22 (S260).
[0054] Next, the history DB extraction unit 24 performs a search in the history DB 22 using the representative specific pixel coordinates of the deformation 5 of the foreground image 3 processed by the coordinate estimation unit 15 as a key, and extracts images associated with the same specific pixel coordinates (S270).
[0055] Next, the history image output unit 25 outputs the image extracted by the history DB extraction unit 24 to the LCD 1c along with its imaging date (S280).
[0056] The first embodiment of this disclosure has been described above, and the above embodiment has the following features.
[0057] As shown in Figure 2, the coordinate estimation device 1 (coordinate estimation system) comprises a positioning unit 13, a position and orientation estimation unit 14, and a coordinate estimation unit 15. The positioning unit 13 aligns the foreground image 3 with at least one background image 4. The foreground image 3 is an image obtained by imaging the foreground imaging area R1 (first imaging area) of the bridge 2 (object). The at least one background image 4 is an image obtained by imaging the background imaging area R2 (second imaging area) of the bridge 2, which is larger than the foreground imaging area R1 and includes the foreground imaging area R1. The position and orientation estimation unit 14 estimates the imaging position and orientation, including the imaging position and orientation of the imaging device that captured at least one background image 4, based on the point cloud of the bridge 2 and at least one background image 4. The coordinate estimation unit 15 estimates the coordinates of at least one specific pixel Q, which is any pixel in the foreground image 3, in the point cloud coordinate system, based on the point cloud of the bridge 2, the alignment result by the alignment unit 13, and the imaging position and orientation. With this configuration, when estimating the positional relationship between the point cloud and the foreground image 3 (imported image), which have significantly different resolutions, the estimation can be performed without problems by using at least one distant image 4.
[0058] Furthermore, as shown in Figures 10 and 12, for example, the coordinate estimation unit 15 calculates the coordinates of at least one specific pixel Q in at least one distant image 4 based on the alignment result, and estimates the specific pixel coordinates based on the point cloud, the calculation result, and the imaging position and orientation. With the above configuration, the specific pixel coordinates can be efficiently estimated using the alignment result from the alignment unit 13.
[0059] Furthermore, as shown in Figure 12, for example, the coordinate estimation unit 15 calculates a projection line L that radiates from the imaging position toward at least one specific pixel Q based on the imaging position and orientation, performs collision detection between the projection line L and the bridge 2, and estimates the coordinates of the specific pixel based on the result of the detection. With this configuration, the coordinates of the specific pixel can be estimated with low computational cost.
[0060] Furthermore, as shown in Figure 12, for example, the coordinate estimation unit 15 extracts a sub-point cloud (points p9 and p10) from the point cloud (points p1 to p25) whose distance from the projection line L is less than or equal to a predetermined value. The coordinate estimation unit 15 estimates a specific pixel coordinate based on the coordinate of point p10, which is the closest point in the sub-point cloud (points p9 and p10) to the imaging position. With the above configuration, a pseudo-collision detection between the projection line L and the point cloud can be realized.
[0061] (Second Embodiment) A second embodiment of this disclosure will be described below with reference to Figure 17. The following description will focus on the differences between this embodiment and the first embodiment, omitting any redundant explanations.
[0062] In the first embodiment described above, as shown in Figure 12, the coordinate estimation unit 15 calculates a projection line L that radiates from the imaging position toward a specific pixel Q based on the imaging position and orientation estimated by the position and orientation estimation unit 14. The coordinate estimation unit 15 performs a collision determination between the projection line L and the bridge 2 and estimates the coordinates of the specific pixel based on the determination result. However, since the bridge 2 is represented by a point cloud, there is a problem that it is practically impossible for the projection line L to collide with the point cloud of the bridge 2. Therefore, the coordinate estimation unit 15 extracts a sub-point cloud (points p9 and p10) from the point cloud whose distance from the projection line L is less than or equal to a predetermined value, and estimates the coordinates of the specific pixel based on the coordinates of point p10, which is the closest of the sub-point cloud (points p9 and p10) to the imaging position.
[0063] In contrast, in this embodiment, as shown in Figure 17, the coordinate estimation unit 15 converts the point cloud of the bridge 2 into mesh data and performs collision detection between the projection line L and the bridge 2 based on the mesh data. In this case, the projection line L can always collide with the mesh represented by the mesh data of the bridge 2, so collision detection between the projection line L and the bridge 2 can be performed without any problems.
[0064] (Third embodiment) A third embodiment of this disclosure will be described below with reference to Figure 18. The following description will focus on the differences between this embodiment and the first embodiment, omitting any redundant explanations.
[0065] In the first embodiment described above, as shown in Figure 12, the coordinate estimation unit 15 calculates a projection line L that radiates from the imaging position toward a specific pixel Q based on the imaging position and orientation estimated by the position and orientation estimation unit 14. The coordinate estimation unit 15 performs a collision determination between the projection line L and the bridge 2 and estimates the coordinates of the specific pixel based on the determination result. However, since the bridge 2 is represented by a point cloud, there is a problem that it is practically impossible for the projection line L to collide with the point cloud of the bridge 2. Therefore, the coordinate estimation unit 15 extracts a sub-point cloud (points p9 and p10) from the point cloud whose distance from the projection line L is less than or equal to a predetermined value, and estimates the coordinates of the specific pixel based on the coordinates of point p10, which is the closest of the sub-point cloud (points p9 and p10) to the imaging position.
[0066] However, the following problems may arise in the first embodiment described above. Please refer to Figure 18. Figure 18 shows a conceptual diagram of calculating specific pixel coordinates corresponding to multiple specific pixels Q. In Figure 18, when specific pixel coordinates are determined for each specific pixel Q, as in the first embodiment described above, the subgroup of points where the distance from the projection line L1 corresponding to specific pixel Q1 is less than or equal to a predetermined value is points p41 and p42. Similarly, the subgroup of points where the distance from the projection line L2 corresponding to specific pixel Q2 is less than or equal to a predetermined value is point p63. Similarly, the subgroup of points where the distance from the projection line L3 corresponding to specific pixel Q3 is less than or equal to a predetermined value is points p46 and p47. Here, points p40 to p48 correspond to the point group of the front of the bridge 2 as seen from the imaging position estimated by the position and orientation estimation unit 14, and points p60 to p65 correspond to the point group of the back of the bridge 2 as seen from the imaging position estimated by the position and orientation estimation unit 14. In this case, the specific pixel coordinates corresponding to specific pixels Q1, Q2, and Q3 would be the coordinates of points p41, p63, and p47, respectively. Therefore, it is conceivable that in the superimposed image generated by the superimposed image generation unit 20, a portion of the deformation 5 will be scattered to coordinates far removed from the coordinates where the deformation 5 originally existed. As a result, a portion of the deformation 5 will be effectively missing from the superimposed image generated by the superimposed image generation unit 20. This missing portion could become a major problem when measuring the size of the deformation 5 on the digital twin.
[0067] Therefore, in this embodiment, the coordinate estimation unit 15 sets a predetermined value used for collision determination to be larger than that in the first embodiment, and then performs collision determination in the same manner as in the first embodiment. In this case, the subgroup of points whose distance from the projection line L1 corresponding to a specific pixel Q1 is less than or equal to the predetermined value is points p41, p42, p60, and p61. Similarly, the subgroup of points whose distance from the projection line L2 corresponding to a specific pixel Q2 is less than or equal to the predetermined value is points p43, p44, and p63. Similarly, the subgroup of points whose distance from the projection line L3 corresponding to a specific pixel Q3 is less than or equal to the predetermined value is points p46 and p47.
[0068] Next, the coordinate estimation unit 15 performs clustering on all the subpoints extracted by the collision detection process, according to the distance from the imaging position estimated by the position and orientation estimation unit 14. All the points represent points p41, p42, p43, p44, p46, 47, p60, p61, and p63. As a result, cluster C1 is obtained, to which points p41, p42, p43, p44, p46, and 47 belong, and cluster C2 is obtained, to which points p60, p61, and p63 belong.
[0069] Next, the coordinate estimation unit 15 selects either cluster C1 or cluster C2 obtained by the clustering and estimates the coordinates of a specific pixel Q for each specific pixel Q based on the selected cluster. Specifically, the coordinate estimation unit 15 may select the cluster that is closest to the imaging position estimated by the position and orientation estimation unit 14 from among the multiple clusters obtained by the clustering. Alternatively, the coordinate estimation unit 15 may select the cluster with the largest number of points from among the multiple clusters obtained by the clustering. In either selection criterion, in the example of Figure 18, the coordinate estimation unit 15 will select cluster C1. Then, when determining the coordinates of a specific pixel Q for each specific pixel Q, the coordinate estimation unit 15 selects the point p that is closest to the imaging position estimated by the position and orientation estimation unit 14 from among the multiple points corresponding to the specific pixel Q and belonging to cluster C1. For example, with respect to the specific pixel Q1, the coordinate estimation unit 15 selects point p41 from among points p41 and p42 that is closest to the imaging position estimated by the position and orientation estimation unit 14. Furthermore, with respect to a specific pixel Q2, the coordinate estimation unit 15 selects point p44, which is closest to the imaging position estimated by the position and orientation estimation unit 14, from among points p43 and p44. Similarly, with respect to a specific pixel Q3, the coordinate estimation unit 15 selects point p47, which is closest to the imaging position estimated by the position and orientation estimation unit 14, from among points p46 and p47. Then, for each specific pixel Q, the coordinate estimation unit 15 estimates the specific pixel coordinates corresponding to that specific pixel Q based on the selected point. With this configuration, it is possible to suppress the substantial loss of part of the deformation 5 in the superimposed image generated by the superimposed image generation unit 20.
[0070] In short, this embodiment has the following features.
[0071] That is, at least one specific pixel Q contains multiple specific pixels Q. The coordinate estimation unit 15 extracts from the point cloud subgroups for each specific pixel Q whose distance from the projection line L is less than or equal to a predetermined value. The coordinate estimation unit 15 clusters all of the subgroups. Then, the coordinate estimation unit 15 estimates the specific pixel coordinates for each specific pixel Q based on one of the clusters obtained by clustering. With the above configuration, it is possible to suppress the substantial loss of part of the deformation 5 in the superimposed image generated by the superimposed image generation unit 20.
[0072] (Fourth Embodiment) A fourth embodiment of this disclosure will be described below with reference to Figures 19 to 23. The following description will focus on the differences between this embodiment and the first embodiment, omitting any redundant explanations.
[0073] In the first embodiment described above, when an inspector discovers a deformation 5 on the bridge 2, the rule is to first capture a close-up image 3 of the deformation 5 with the telephoto side of the imaging device, and then capture a distant image 4 of the deformation 5 with the wide-angle side of the imaging device. In this case, since the close-up image 3 and distant image 4 corresponding to the same deformation 5 can be captured within a few minutes, as shown in Figure 9, it can be seen that Image 1 is the close-up image 3, and Image 2 is the distant image 4 corresponding to the close-up image 3.
[0074] However, even if the imaging rules for the foreground image 3 and background image 4 are defined as described above, these rules may not always be followed during actual inspections. For example, if the foreground image 3 fails to be captured due to camera shake, the foreground image 3 will need to be recaptured for the same deformation 5. Also, the background image 4 corresponding to the foreground image 3 may be forgotten. In this case, as shown in Figure 9, the correspondence between foreground image 3 and background image 4 may not be accurately determined solely from the date and time of capture.
[0075] As a result, as shown in Figure 19, the near-view image 3a, which captures the deformation 5 of bridge 2 from the front, and the far-view image 4a, which captures the deformation 5 of bridge 2 at a shallow angle, will be treated as corresponding to each other, and the estimation of specific pixel coordinates for specific pixels in the near-view image 3a will be performed. In this case, the accuracy of the estimation of specific pixel coordinates may deteriorate for the following reasons. Firstly, the area of the near-view image 3a in the far-view image 4a becomes smaller, which deteriorates the calculation accuracy of the homography matrix H. Secondly, since the far-view image 4a is an image of the deformation 5 of bridge 2 captured at a shallow angle, the projection line L used in collision determination by the coordinate estimation unit 15 will also be generated at a shallow angle to the deformation 5 of bridge 2. Consequently, it is possible that the subgroup of points whose distance from the projection line L is less than or equal to a predetermined value will be shifted overall to approach the position captured in the far-view image 4a.
[0076] Therefore, in this embodiment, the coordinate estimation unit 15 determines whether the distant view image 4 capturing the deformation 5 was taken from a shallow angle, and if it was taken from a shallow angle, it re-estimates the specific pixel coordinates using another distant view image 4 capturing the deformation 5. The operation of the coordinate estimation unit 15 will be described below with reference to Figures 20 to 23.
[0077] Figures 20 and 21 show the operation flow of the coordinate estimation device 1. In the operation flow shown in Figures 20 and 21, steps S200 to S230 and steps S240 to S280 are the same as in the first embodiment, so their explanation will be omitted as appropriate.
[0078] Referring to Figure 20, the coordinate estimation unit 15 estimates the coordinates of a specific pixel based on the foreground image 3a and background image 4a, which are estimated to be in a corresponding relationship (S230). Next, as shown in Figure 22, the coordinate estimation unit 15 calculates the angle θ formed by the normal S of the surface of the bridge 2 at the specific pixel coordinates and the line segment T connecting the imaging position of the background image 4a estimated by the position and orientation estimation unit 14 to the specific pixel coordinates (S231). Then, the coordinate estimation unit 15 determines whether the calculated angle θ is greater than or equal to a threshold (S232). If the coordinate estimation unit 15 determines that the calculated angle θ is not greater than or equal to a threshold, the coordinate estimation unit 15 proceeds to step S240. That is, if the calculated angle θ is not greater than or equal to a threshold, the estimation accuracy of the specific pixel coordinates is guaranteed for the reasons described above. On the other hand, if the coordinate estimation unit 15 determines that the calculated angle θ is greater than or equal to a threshold, the coordinate estimation unit 15 proceeds to step S233 shown in Figure 21.
[0079] In step S233, the coordinate estimation unit 15 extracts multiple distant view images 4 from among the multiple distant view images 4 stored in the image storage unit 12 that include the specific pixel coordinates estimated in S230 within the imaging range (S233). Next, the coordinate estimation unit 15 calculates the specific pixel coordinates for each extracted distant view image 4 and also calculates the angle θ shown in Figure 22 for each distant view image 4 (S234). Then, as shown in Figure 23, the coordinate estimation unit 15 stores the multiple distant view images 4 extracted in step S233 and the angle θ calculated in step S234 in the memory 1b, associating them.
[0080] Next, the coordinate estimation unit 15 selects the distant view image 4 with the smallest angle θ (distant view image number No. 5) from among the multiple distant view images 4 extracted in step S233 (S235). Then, based on the selected distant view image 4, the coordinate estimation unit 15 re-estimates the specific pixel coordinate of the specific pixel Q (S236), and proceeds to processing in S240. This makes it possible to estimate the specific pixel coordinate with high accuracy even when the corresponding near view image 3 and distant view image 4 cannot be accurately obtained from the image storage unit 12 shown in Figure 9.
[0081] In short, the above embodiment has the following features.
[0082] That is, at least one distant view image 4 includes multiple distant view images 4 with different imaging positions and orientations. The coordinate estimation unit 15 estimates a specific pixel coordinate for each distant view image 4 and calculates the angle θ formed by the normal S of the bridge 2 at the specific pixel coordinate and the line segment T connecting the imaging position and the specific pixel coordinate for each distant view image 4 (S234). Then, the coordinate estimation unit 15 estimates the specific pixel coordinate based on the distant view image 4 from among the multiple distant view images 4 in which the angle θ is smallest (S236). With the above configuration, even if the correspondence between the foreground image 3 and the distant view image 4 cannot be guaranteed, the specific pixel coordinate can be estimated with high accuracy.
[0083] (Fifth embodiment) A fifth embodiment of this disclosure will be described below with reference to Figures 24 and 25. The following description will focus on the differences between this embodiment and the fourth embodiment described above, omitting any redundant explanations.
[0084] In the fourth embodiment described above, as shown in Figure 22, the angle θ is determined for each distant view image 4, and the optimal distant view image 4 is selected based on this angle θ, thereby ensuring the accuracy of the estimation of specific pixel coordinates.
[0085] In contrast, in this embodiment, the proportion occupied by the foreground image 3 is determined for each background image 4, and the optimal background image 4 is selected based on this proportion, thereby ensuring the accuracy of the estimation of specific pixel coordinates. Specifically, this is done as follows.
[0086] Figures 24 and 25 show the operation flow of the coordinate estimation device 1. In the operation flow shown in Figures 24 and 25, steps S200 to S230 and steps S240 to S280 are the same as in the first embodiment, so their explanation will be omitted as appropriate.
[0087] Referring to Figure 24, the coordinate estimation unit 15 estimates specific pixel coordinates based on the foreground image 3a and background image 4a, which are estimated to be in a corresponding relationship (S230). Next, the coordinate estimation unit 15 calculates the proportion of the background image 3a that occupies the background image 4a (S300). The area occupied by the foreground image 3 in the background image 4 can typically be easily determined by converting the coordinates of the four corners of the foreground image 3a to the background coordinate system u'-v' of the background image 4a, based on the homography matrix H calculated by the alignment unit 13. Therefore, the coordinate estimation unit 15 can calculate this proportion by dividing the area occupied by the foreground image 3 in the background image 4 by the area of the background image 4. Then, the coordinate estimation unit 15 determines whether the calculated proportion is greater than or equal to a threshold (S301). If the coordinate estimation unit 15 determines that the calculated proportion is greater than or equal to a threshold, the coordinate estimation unit 15 proceeds to step S240. In other words, if the calculated ratio is greater than or equal to the threshold, the estimation accuracy of the specific pixel coordinates is guaranteed for the reasons described above. On the other hand, if the coordinate estimation unit 15 determines that the calculated angle θ is not greater than or equal to the threshold, the coordinate estimation unit 15 proceeds to step S302 shown in Figure 25.
[0088] In step S2302, the coordinate estimation unit 15 extracts multiple distant view images 4 from among the multiple distant view images 4 stored in the image storage unit 12 that include the specific pixel coordinates estimated in S230 within the imaging range (S302). Next, the coordinate estimation unit 15 calculates the homography matrix H and the above-mentioned ratio for each extracted distant view image 4 (S303). Then, the coordinate estimation unit 15 stores the multiple distant view images 4 extracted in step S302 and the ratio calculated by the coordinate estimation unit 15 in step S303 in the memory 1b, associating them.
[0089] Next, the coordinate estimation unit 15 selects the distant image 4 with the largest proportion from among the multiple distant images 4 extracted in step S303 (S304). Then, based on the selected distant image 4, the coordinate estimation unit 15 re-estimates the specific pixel coordinate of the specific pixel Q (S305), and proceeds to processing in S240. With this configuration, even if the correspondence between the foreground image 3 and the distant image 4 cannot be guaranteed, the specific pixel coordinate can be estimated with high accuracy.
[0090] In short, the above embodiment has the following features.
[0091] In other words, at least one distant image 4 includes multiple distant images 4 whose imaging positions and orientations differ from each other. The distant imaging region R2 of the multiple distant images 4 is larger than the near imaging region R1 of the near image 3 and includes at least the near imaging region R1 of the near image 3. The coordinate estimation unit 15 then estimates the specific pixel coordinates based on the distant image 4 in which the near image 3 occupies the largest proportion among the multiple distant images 4. With this configuration, even if the correspondence between the near image 3 and the distant images 4 cannot be guaranteed, the specific pixel coordinates can be estimated with high accuracy.
[0092] In the above example, the program can be stored and supplied to the computer using various types of non-transitory computer-readable medium. Non-transitory computer-readable medium includes various types of tangible storage medium. Examples of non-transitory computer-readable medium include magnetic storage media (e.g., flexible disks, magnetic tapes, hard disk drives) and magneto-optical storage media (e.g., magneto-optical disks). Examples of non-transitory computer-readable medium further include CD-ROM (Read Only Memory), CD-R, CD-R / W, and semiconductor memory (e.g., mask ROM; examples of non-transitory computer-readable medium further include PROM (Programmable ROM), EPROM (Erasable PROM), flash ROM, and RAM (random access memory)). Alternatively, the program may be supplied to the computer by various types of transient computer-readable medium. Examples of transient computer-readable medium include electrical signals, optical signals, and electromagnetic waves. Temporary computer-readable media can supply programs to a computer via wired communication channels such as electric wires and optical fibers, or via wireless communication channels.
[0093] Some or all of the above embodiments may also be described as follows, but are not limited to the following:
[0094] (Note 1) Alignment means for aligning a close-up image obtained by imaging a first imaging area of an object and at least one distant image obtained by imaging a second imaging area of the object that is larger than the first imaging area and includes the first imaging area, A position and orientation estimation means for estimating the imaging position and orientation of an imaging device that captured the at least one distant image, including the imaging position and orientation in the point cloud coordinate system, based on the point cloud of the object and the at least one distant image. A coordinate estimation means for estimating the coordinates of a specific pixel in the point cloud coordinate system of at least one specific pixel, which is any pixel in the foreground image, based on the point cloud, the alignment result by the alignment means, and the imaging position and orientation. including, Coordinate estimation system. (Note 2) The coordinate estimation means calculates the coordinates of the at least one specific pixel in the at least one distant image based on the alignment result, and estimates the coordinates of the specific pixel based on the point cloud, the calculation result, and the imaging position and orientation. The coordinate estimation system described in Appendix 1. (Note 3) The at least one distant view image includes a plurality of distant view images whose imaging positions and orientations differ from each other. The coordinate estimation means is For each of the aforementioned distant view images, the specific pixel coordinates are estimated. For each of the aforementioned distant view images, the angle formed by the normal of the object at the specific pixel coordinates and the line segment connecting the imaging position and the specific pixel coordinates is calculated. The specific pixel coordinates are estimated based on the distant view image in which the angle is smallest among the multiple distant view images. The coordinate estimation system described in Appendix 1. (Note 4) The at least one distant view image includes a plurality of distant view images whose imaging positions and orientations differ from each other. The coordinate estimation means estimates the specific pixel coordinates based on the distant image in which the foreground image occupies the largest proportion among the plurality of distant images. The coordinate estimation system described in Appendix 1. (Note 5) The coordinate estimation means calculates a projection line emanating from the imaging position toward at least one specific pixel based on the imaging position orientation, performs collision determination between the projection line and the object, and estimates the coordinates of the specific pixel based on the determination result. The coordinate estimation system described in Appendix 1. (Note 6) The coordinate estimation means extracts a subgroup of points from the point group whose distance from the projection line is less than or equal to a predetermined value, and estimates the specific pixel coordinates based on the coordinates of the point in the subgroup that is closest to the imaging position. The coordinate estimation system described in Appendix 5. (Note 7) The coordinate estimation means converts the point cloud into mesh data, Based on the mesh data, a collision determination is performed between the projection line and the object. The coordinate estimation system described in Appendix 5. (Note 8) The aforementioned at least one specific pixel includes a plurality of specific pixels, The coordinate estimation means is A subgroup of points is extracted from the point group in which the distance from the projection line to each specific pixel is less than or equal to a predetermined value. Cluster all subgroups of points, Based on one of the clusters obtained by the clustering, the coordinates of the specific pixel are estimated for each specific pixel. The coordinate estimation system described in Appendix 5. (Note 9) Alignment means for aligning a close-up image obtained by imaging a first imaging area of an object and at least one distant image obtained by imaging a second imaging area of the object that is larger than the first imaging area and includes the first imaging area, A position and orientation estimation means for estimating the imaging position and orientation of an imaging device that captured the at least one distant image, including the imaging position and orientation in the point cloud coordinate system, based on the point cloud of the object and the at least one distant image. A coordinate estimation means for estimating the coordinates of a specific pixel in the point cloud coordinate system of at least one specific pixel, which is any pixel in the foreground image, based on the point cloud, the alignment result by the alignment means, and the imaging position and orientation. including, Coordinate estimation device. (Note 10) The coordinate estimation means calculates the coordinates of the at least one specific pixel in the at least one distant image based on the alignment result, and estimates the coordinates of the specific pixel based on the point cloud, the calculation result, and the imaging position and orientation. The coordinate estimation device described in Appendix 9. (Note 11) The at least one distant view image includes a plurality of distant view images whose imaging positions and orientations differ from each other. The coordinate estimation means is For each of the aforementioned distant view images, the specific pixel coordinates are estimated. For each of the aforementioned distant view images, the angle formed by the normal of the object at the specific pixel coordinates and the line segment connecting the imaging position and the specific pixel coordinates is calculated. The specific pixel coordinates are estimated based on the distant view image in which the angle is smallest among the multiple distant view images. The coordinate estimation device described in Appendix 9. (Note 12) The at least one distant view image includes a plurality of distant view images whose imaging positions and orientations differ from each other. The coordinate estimation means estimates the specific pixel coordinates based on the distant image in which the foreground image occupies the largest proportion among the plurality of distant images. The coordinate estimation device described in Appendix 9. (Note 13) The coordinate estimation means calculates a projection line emanating from the imaging position toward at least one specific pixel based on the imaging position orientation, performs collision determination between the projection line and the object, and estimates the coordinates of the specific pixel based on the determination result. The coordinate estimation device described in Appendix 9. (Note 14) The coordinate estimation means extracts a subgroup of points from the point group whose distance from the projection line is less than or equal to a predetermined value, and estimates the specific pixel coordinates based on the coordinates of the point in the subgroup that is closest to the imaging position. The coordinate estimation device described in Appendix 13. (Note 15) The coordinate estimation means converts the point cloud into mesh data, Based on the mesh data, a collision determination is performed between the projection line and the object. The coordinate estimation device described in Appendix 13. (Note 16) The aforementioned at least one specific pixel includes a plurality of specific pixels, The coordinate estimation means is A subgroup of points is extracted from the point group in which the distance from the projection line to each specific pixel is less than or equal to a predetermined value. Cluster all subgroups of points, Based on one of the clusters obtained by the clustering, the coordinates of the specific pixel are estimated for each specific pixel. The coordinate estimation device described in Appendix 13. (Note 17) A computer performs a positioning step of aligning a close-up image obtained by imaging a first imaging area of an object with at least one distant image obtained by imaging a second imaging area of the object that is larger than the first imaging area and includes the first imaging area. The computer performs a position and orientation estimation step in which it estimates the imaging position and orientation of an imaging device that captured the at least one distant image, including the imaging position and orientation in the point cloud coordinate system, based on the point cloud of the object and the at least one distant image. The computer performs a coordinate estimation step in which it estimates the coordinates of a specific pixel in the point cloud coordinate system of at least one specific pixel, which is any pixel of the foreground image, based on the point cloud, the alignment result obtained by the alignment step, and the imaging position and orientation. including, Coordinate estimation method. (Note 18) In the coordinate estimation step, the coordinates of the at least one specific pixel in the at least one distant image are calculated based on the alignment result, and the coordinates of the specific pixel are estimated based on the point cloud, the calculation result, and the imaging position and orientation. The coordinate estimation method described in Appendix 17. (Note 19) The at least one distant view image includes a plurality of distant view images whose imaging positions and orientations differ from each other. In the aforementioned coordinate estimation step, For each of the aforementioned distant view images, the specific pixel coordinates are estimated. For each of the aforementioned distant view images, the angle formed by the normal of the object at the specific pixel coordinates and the line segment connecting the imaging position and the specific pixel coordinates is calculated. The specific pixel coordinates are estimated based on the distant view image in which the angle is smallest among the multiple distant view images. The coordinate estimation method described in Appendix 17. (Note 20) The at least one distant view image includes a plurality of distant view images whose imaging positions and orientations differ from each other. In the coordinate estimation step, the coordinates of the specific pixel are estimated based on the distant image in which the foreground image occupies the largest proportion among the multiple distant images. The coordinate estimation method described in Appendix 17. (Note 21) In the coordinate estimation step, a projection line is calculated from the imaging position toward at least one specific pixel based on the imaging position and orientation, a collision determination is performed between the projection line and the object, and the coordinates of the specific pixel are estimated based on the determination result. The coordinate estimation method described in Appendix 17. (Note 22) In the coordinate estimation step, a subgroup of points whose distance from the projection line is less than or equal to a predetermined value is extracted from the point group, and the coordinates of the specific pixel are estimated based on the coordinates of the point in the subgroup that is closest to the imaging position. The coordinate estimation method described in Appendix 21. (Note 23) In the coordinate estimation step, the point cloud is converted into mesh data, Based on the mesh data, a collision determination is performed between the projection line and the object. The coordinate estimation method described in Appendix 21. (Note 24) The aforementioned at least one specific pixel includes a plurality of specific pixels, In the aforementioned coordinate estimation step, A subgroup of points is extracted from the point group in which the distance from the projection line to each specific pixel is less than or equal to a predetermined value. Cluster all subgroups of points, Based on one of the clusters obtained by the clustering, the coordinates of the specific pixel are estimated for each specific pixel. The coordinate estimation method described in Appendix 21. (Note 25) On the computer, A positioning step of aligning a close-up image obtained by imaging a first imaging area of an object with at least one distant image obtained by imaging a second imaging area of the object that is larger than the first imaging area and includes the first imaging area, A position and orientation estimation step, based on the point cloud of the object and the at least one distant view image, estimates the imaging position and orientation of the imaging device that captured the at least one distant view image, including the imaging position and orientation in the point cloud coordinate system. A coordinate estimation step that estimates the coordinates of a specific pixel in the point cloud coordinate system of at least one specific pixel, which is any pixel of the foreground image, based on the point cloud, the alignment result obtained by the alignment step, and the imaging position and orientation. A program that executes the command. (Note 26) In the coordinate estimation step, the coordinates of the at least one specific pixel in the at least one distant image are calculated based on the alignment result, and the coordinates of the specific pixel are estimated based on the point cloud, the calculation result, and the imaging position and orientation. The program described in Appendix 25. (Note 27) The at least one distant view image includes a plurality of distant view images whose imaging positions and orientations differ from each other. In the aforementioned coordinate estimation step, For each of the aforementioned distant view images, the specific pixel coordinates are estimated. For each of the aforementioned distant view images, the angle formed by the normal of the object at the specific pixel coordinates and the line segment connecting the imaging position and the specific pixel coordinates is calculated. The specific pixel coordinates are estimated based on the distant view image in which the angle is smallest among the multiple distant view images. The program described in Appendix 25. (Note 28) The at least one distant view image includes a plurality of distant view images whose imaging positions and orientations differ from each other. In the coordinate estimation step, the coordinates of the specific pixel are estimated based on the distant image in which the foreground image occupies the largest proportion among the multiple distant images. The program described in Appendix 25. (Note 29) In the coordinate estimation step, a projection line is calculated from the imaging position toward at least one specific pixel based on the imaging position and orientation, a collision determination is performed between the projection line and the object, and the coordinates of the specific pixel are estimated based on the determination result. The program described in Appendix 25. (Note 30) In the coordinate estimation step, a subgroup of points whose distance from the projection line is less than or equal to a predetermined value is extracted from the point group, and the coordinates of the specific pixel are estimated based on the coordinates of the point in the subgroup that is closest to the imaging position. The program described in Appendix 29. (Note 31) In the coordinate estimation step, the point cloud is converted into mesh data, Based on the mesh data, a collision determination is performed between the projection line and the object. The program described in Appendix 29. (Note 32) The aforementioned at least one specific pixel includes a plurality of specific pixels, In the aforementioned coordinate estimation step, A subgroup of points is extracted from the point group in which the distance from the projection line to each specific pixel is less than or equal to a predetermined value. Cluster all subgroups of points, Based on one of the clusters obtained by the clustering, the coordinates of the specific pixel are estimated for each specific pixel. The program described in Appendix 29. [Industrial applicability]
[0095] This disclosure can be applied to techniques for estimating the positional relationship between point clouds and captured images that have significantly different resolutions. [Explanation of Symbols]
[0096] 1. Coordinate Estimation Device 2 Bridges 3 Close-up image 3a Close-up image 4. Distant view image 4a Distant view image 5. Deformation 10 Data Reception Department 11 Point cloud storage 12 Image storage unit 13 Alignment section 14 Position and orientation estimation section 15. Coordinate Estimation Unit 20 Superimposed Image Generation Unit 20a Superimposed image 21 Superimposed image output section 22 History DB 23 History DB Update Section 24 History DB Extraction Unit 25 History Image Output Unit C1 cluster C2 cluster H homography matrix M line segment L projection line L1 projection line L2 projection line L3 projection line Q Specific pixel Q1 Specific pixel Q2 Specific Pixel Q3 Specific pixels R1 Near-view imaging area R2 Distant View Imaging Area S normal T-segment θ angle
Claims
1. Alignment means for aligning a close-up image obtained by imaging a first imaging area of an object and at least one distant image obtained by imaging a second imaging area of the object that is larger than the first imaging area and includes the first imaging area, A position and orientation estimation means for estimating the imaging position and orientation, including the imaging position and orientation of the imaging device that captured the at least one distant image, based on the point cloud of the object and the at least one distant image, A coordinate estimation means for estimating the coordinates of a specific pixel in the point cloud coordinate system of at least one specific pixel, which is any pixel in the foreground image, based on the point cloud, the alignment result by the alignment means, and the imaging position and orientation. including, Coordinate estimation system.
2. The coordinate estimation means calculates the coordinates of the at least one specific pixel in the at least one distant image based on the alignment result, and estimates the coordinates of the specific pixel based on the point cloud, the calculation result, and the imaging position and orientation. The coordinate estimation system according to claim 1.
3. The at least one distant view image includes a plurality of distant view images whose imaging positions and orientations differ from each other. The coordinate estimation means is For each of the aforementioned distant view images, the specific pixel coordinates are estimated. For each of the aforementioned distant view images, the angle formed by the normal of the object at the specific pixel coordinates and the line segment connecting the imaging position and the specific pixel coordinates is calculated. The specific pixel coordinates are estimated based on the distant view image in which the angle is smallest among the multiple distant view images. The coordinate estimation system according to claim 1.
4. The at least one distant view image includes a plurality of distant view images whose imaging positions and orientations differ from each other. The coordinate estimation means estimates the specific pixel coordinates based on the distant image in which the foreground image occupies the largest proportion among the plurality of distant images. The coordinate estimation system according to claim 1.
5. The coordinate estimation means calculates a projection line emanating from the imaging position toward at least one specific pixel based on the imaging position orientation, performs collision determination between the projection line and the object, and estimates the coordinates of the specific pixel based on the determination result. The coordinate estimation system according to claim 1.
6. The coordinate estimation means extracts a subgroup of points from the point group whose distance from the projection line is less than or equal to a predetermined value, and estimates the specific pixel coordinates based on the coordinates of the point in the subgroup that is closest to the imaging position. The coordinate estimation system according to claim 5.
7. The coordinate estimation means converts the point cloud into mesh data, Based on the mesh data, a collision determination is performed between the projection line and the object. The coordinate estimation system according to claim 5.
8. The aforementioned at least one specific pixel includes a plurality of specific pixels, The coordinate estimation means is A subgroup of points is extracted from the point group in which the distance from the projection line to each specific pixel is less than or equal to a predetermined value. Cluster all subgroups of points, Based on one of the clusters obtained by the clustering, the coordinates of the specific pixel are estimated for each specific pixel. The coordinate estimation system according to claim 5.
9. Alignment means for aligning a close-up image obtained by imaging a first imaging area of an object and at least one distant image obtained by imaging a second imaging area of the object that is larger than the first imaging area and includes the first imaging area, A position and orientation estimation means for estimating the imaging position and orientation, including the imaging position and orientation of the imaging device that captured the at least one distant image, based on the point cloud of the object and the at least one distant image, A coordinate estimation means for estimating the coordinates of a specific pixel in the point cloud coordinate system of at least one specific pixel, which is any pixel in the foreground image, based on the point cloud, the alignment result by the alignment means, and the imaging position and orientation. including, Coordinate estimation device.
10. A computer performs a positioning step of aligning a close-up image obtained by imaging a first imaging area of an object with at least one distant image obtained by imaging a second imaging area of the object that is larger than the first imaging area and includes the first imaging area. The computer performs a position and orientation estimation step in which it estimates the imaging position and orientation of an imaging device that captured the at least one distant image, including the imaging position and orientation in the point cloud coordinate system, based on the point cloud of the object and the at least one distant image. The computer performs a coordinate estimation step in which it estimates the coordinates of a specific pixel in the point cloud coordinate system of at least one specific pixel, which is any pixel of the foreground image, based on the point cloud, the alignment result obtained by the alignment step, and the imaging position and orientation. including, Coordinate estimation method.