Geospatial information processing device and geospatial information processing method
A camera-based geospatial information processing system efficiently reconstructs and aligns 3D point clouds to estimate building heights, addressing the high cost of LiDAR by providing cost-effective digital map creation and updating.
Patent Information
- Application Number
- JP2025084234
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The LiDAR method for creating three-dimensional digital maps is costly, and the cost increases with the area of the digital map.
A geospatial information processing device and method that uses a camera to capture images, reconstructs a 3D point cloud, aligns it with real-world coordinates, detects buildings, divides building regions, identifies vertex candidates, and estimates building heights to create and update digital maps, potentially using satellite positioning for alignment and scale factor calculation.
Enables the creation and updating of accurate digital maps at lower costs, particularly benefiting larger areas.
Smart Images

Figure 0007738793000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a geospatial information processing device and a geospatial information processing method, and more particularly to a geospatial information processing device and a geospatial information processing method that are applicable to the creation and updating of digital maps. [Background technology]
[0002] In recent years, with the development of autonomous driving technology and advanced navigation systems, there has been an increasing demand for three-dimensional maps, a type of digital map. As a method for creating such three-dimensional maps, there is a method that uses LiDAR (Light Detection and Ranging), as disclosed in Non-Patent Document 1, for example, and in particular a method that uses a dedicated vehicle equipped with LiDAR. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] "Reiwa 4 Patent Application Technology Trends Survey Report - LiDAR -", Japan Patent Office, March 2023, pp. 1-7 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the LiDAR method is quite costly, and the larger the area of the digital map, the higher the cost.
[0005] Therefore, an object of the present invention is to provide a new geospatial information processing device and a new geospatial information processing method that will greatly contribute to realizing the creation and updating of accurate digital maps while reducing costs compared to conventional methods. [Means for solving the problem]
[0006] To achieve this object, the present invention includes first and second inventions relating to a geospatial information processing device, and third and fourth inventions relating to a geospatial information processing method.
[0007] A first aspect of the present invention, relating to a geospatial information processing device, includes an image data input accepting means, a 3D point cloud reconstructing means, a spatial matching means, a building detection means, a dividing means, a vertex candidate identifying means, a filter region setting means, a building vertex identifying means, and a building height estimating means. The image data input accepting means accepts input of image data consisting of multiple frames obtained by a camera moving on a road sequentially capturing images of the geospatial space around the camera at a predetermined period. The 3D point cloud reconstructing means estimates the position of the camera when each of the multiple frames comprising the image data was obtained based on the image data accepted by the image data input accepting means, and reconstructs a 3D model of the geospatial space from the 3D point cloud in association with the estimated position of the camera. The spatial matching means aligns the 3D point cloud reconstructed space reconstructed by the 3D point cloud reconstructing means with the real-world coordinate space, together with the estimated position of the camera. The building detection means detects buildings included in each of the multiple frames. For example, an object detection model is used to detect buildings by the building detection means. The division means divides a region of a building detected by the building detection means in a frame including the building. The division means divides the building region using, for example, a segmentation model. The vertex candidate identification means identifies, from the reconstructed 3D point cloud, a 3D point corresponding to the uppermost 2D feature point in the region divided by the division means as a vertex candidate of the building related to the region. The filter region setting means sets a predetermined 3D filter region centered on the vertex candidate. The building vertex identification means identifies, from the reconstructed 3D point cloud, a 3D point that is uppermost in the filter region as the vertex of the building related to the 3D point. The building height estimation means estimates the height of the building based on the vertical distance between the vertex of the building identified by the building vertex identification means and the estimated position of the camera when the frame including the building was obtained, preferably the estimated position of the camera closest to the building among the estimated positions of the camera when the frame including the building was obtained. The building height estimated in this way is used, for example, in creating and updating digital maps.
[0008] The first invention may further include a grouping means. When common vertex candidates are identified from a predetermined number of frames or more, the grouping means groups, from the reconstructed 3D point cloud, 3D points corresponding to 2D feature points located within a region of a building common to the vertex candidates, i.e., 3D points obtained from a region of a building common to the vertex candidates. In this case, the filter region setting means may set the filter region by using, as the center of the filter region, a 3D point that corresponds to an intermediate value in the horizontal direction among the 3D points grouped by the grouping means, instead of the vertex candidates.
[0009] Furthermore, in the first aspect of the present invention, a positioning data input accepting means may be further provided. This positioning data input accepting means accepts input of positioning data obtained by positioning using a positioning satellite receiver that moves on the road together with the camera. In this case, the spatial alignment means may align the 3D point cloud reconstruction space with the real world coordinate space together with the estimated position of the camera, using the positioning position represented by the positioning data as an index. The building height estimation means may estimate the building height by multiplying the mutual distance by a scale factor that is the mutual ratio between the movement distance of the estimated position of the camera after alignment by the spatial alignment means and the geodesic distance corresponding to the movement distance.
[0010] A second aspect of the present invention relates to a geospatial information processing device having a different configuration from the first aspect, and includes an image data input accepting means, a 3D point cloud reconstructing means, a spatial matching means, a building detection means, a segmenting means, a vertex candidate identifying means, a filter region setting means, and a building center position estimating means. The image data input accepting means accepts input of image data consisting of multiple frames obtained by a camera moving on a road sequentially capturing images of the geospatial space around the camera at a predetermined period. The 3D point cloud reconstructing means estimates the position of the camera when each of the multiple frames comprising the image data was obtained based on the image data accepted by the image data input accepting means, and reconstructs a 3D model of the geospatial space from the 3D point cloud in association with the estimated position of the camera. The spatial matching means aligns the 3D point cloud reconstructed space reconstructed by the 3D point cloud reconstructing means with the real-world coordinate space, together with the estimated position of the camera. The building detection means detects buildings included in each of the multiple frames. For example, an object detection model is used to detect buildings by the building detection means. The division means divides a region of a building detected by the building detection means in a frame including the building. The division means divides the region of the building using, for example, a segmentation model. The vertex candidate identification means identifies, from the reconstructed 3D point cloud, a 3D point corresponding to the uppermost 2D feature point in the region divided by the division means as a vertex candidate of the building related to the region. The filter region setting means sets a predetermined 3D filter region centered on the vertex candidate. Then, the building center position estimation means estimates, from the reconstructed 3D point cloud, a 3D point that corresponds to the median value in the horizontal direction within the filter region as the center position of the building related to the 3D point. The center position of the building estimated in this way is used, for example, to identify the position (i.e., latitude and longitude) of the building in real-world coordinate space, and is ultimately used to create and update digital maps.
[0011] In addition, the second invention may also further include a grouping means. As described above, when a common vertex candidate is identified from a predetermined number of frames or more, the grouping means groups, from the reconstructed 3D point cloud, 3D points corresponding to 2D feature points located within a region of a building common to the vertex candidate, i.e., 3D points obtained from a region of a building common to the vertex candidate. In this case, the filter region setting means may set the filter region by using, as the center of the filter region, a 3D point that corresponds to an intermediate value in the horizontal direction among the 3D points grouped by the grouping means, instead of the vertex candidate.
[0012] A third aspect of the present invention, relating to a geospatial information processing method, is a geospatial information processing method implemented by a computer, in which a processor of the computer executes an image data input receiving step, a 3D point cloud reconstructing step, a spatial alignment step, a building detection step, a segmentation step, a vertex candidate identifying step, a filter region setting step, a building vertex identifying step, and a building height estimation step. In the image data input receiving step, the processor receives input of image data consisting of multiple frames obtained by a camera moving on a road sequentially capturing images of the geospatial space around the camera at a predetermined period. In the 3D point cloud reconstructing step, the processor estimates the position of the camera when each of the multiple frames comprising the image data was obtained based on the image data received in the image data input receiving step, and reconstructs a 3D model of the geospatial space using the 3D point cloud in association with the estimated position of the camera. In the spatial alignment step, the processor aligns the 3D point cloud reconstructed space reconstructed in the 3D point cloud reconstructing step with the real world coordinate space, together with the estimated position of the camera. In the building detection step, the processor detects buildings included in each of a plurality of frames. The building detection step uses, for example, an object detection model. In the division step, the processor divides the area of the building detected in the building detection step in the frame including the building. The division step uses, for example, a segmentation model. In the vertex candidate identification step, the processor identifies, from the reconstructed 3D point cloud, a 3D point corresponding to the uppermost 2D feature point in the area divided by the division step, as a vertex candidate of the building related to the area. In the filter area setting step, the processor sets a predetermined 3D filter area centered on the vertex candidate. In the building vertex identification step, the processor identifies, from the reconstructed 3D point cloud, a 3D point located at the uppermost position in the filter area as the vertex of the building related to the 3D point.Then, in the building height estimation step, the processor estimates the height of the building based on the vertical distance between the vertex of the building identified in the building vertex identification step and the estimated position of the camera when the frame including the building was obtained, preferably the estimated position of the camera closest to the building among the estimated positions of the cameras when the frame including the building was obtained. The building height estimated in this way is applied to, for example, creating and updating a digital map.
[0013] A fourth aspect of the present invention relates to a geospatial information processing method using a computer, different from the third aspect. The computer processor executes the following steps: image data input reception, 3D point cloud reconstruction, spatial alignment, building detection, segmentation, vertex candidate identification, filter region setting, and building center position estimation. In the image data input reception step, the processor receives input of image data consisting of multiple frames obtained by a camera moving on a road sequentially capturing images of the geographical space around the camera at a predetermined interval. In the 3D point cloud reconstruction step, the processor estimates the position of the camera when each of the multiple frames comprising the image data was obtained based on the image data received in the image data input reception step, and reconstructs a 3D model of the geographical space using the 3D point cloud in association with the estimated position of the camera. In the spatial alignment step, the processor aligns the 3D point cloud reconstructed space reconstructed in the 3D point cloud reconstruction step with the real world coordinate space, together with the estimated position of the camera. In the building detection step, the processor detects buildings included in each of a plurality of frames. The building detection step uses, for example, an object detection model. In the division step, the processor divides the area of the building detected in the building detection step in the frame including the building. The division step uses, for example, a segmentation model. In the vertex candidate identification step, the processor identifies, from the reconstructed 3D point cloud, a 3D point corresponding to the uppermost 2D feature point in the area divided in the division step, as a vertex candidate of the building related to the area. In the filter area setting step, the processor sets a predetermined 3D filter area centered on the vertex candidate. Then, in the building center position estimation step, the processor estimates, from the reconstructed 3D point cloud, a 3D point that is the median in the horizontal direction within the filter area as the center position of the building related to the 3D point.The center position of a building estimated in this way can be used, for example, to determine the location (i.e., latitude and longitude) of the building in real-world coordinate space, and thus to create and update digital maps. [Effects of the Invention]
[0014] The present invention thus contributes greatly to the realization of creating and updating accurate digital maps at lower costs than conventional methods, and this effect is more pronounced the larger the area of the digital map. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a diagram showing the overall configuration of a geospatial information processing system according to an embodiment of the present invention. [Figure 2] 3 is a diagram showing an example of an image captured by a drive recorder in the embodiment; FIG. [Figure 3] FIG. 2 is a diagram showing a part of OSM in the same embodiment. [Figure 4] FIG. 2 is a block diagram showing the electrical configuration of the geospatial information processing device according to the embodiment. [Figure 5] FIG. 2 is a flowchart showing the overall flow of processing by the geospatial information processing device in the embodiment. [Figure 6] 10 is a diagram showing an example of visualization of a 3D point cloud reconstruction space and an estimated position of a camera in the same embodiment. FIG. [Figure 7] FIG. 10 is a diagram showing an example of a matching result by similarity transformation using Procrustes analysis in the same embodiment. [Figure 8] 10 is a diagram showing an example of a correction space used for correction based on a positioning position in the embodiment. FIG. [Figure 9] 10A and 10B are diagrams showing an example of a correction result based on the measured position in the embodiment; [Figure 10] FIG. 2 is a diagram showing the basic configuration of OSM roads in the embodiment. [Figure 11] 3A and 3B are diagrams for explaining the procedure for road matching in the embodiment; [Figure 12] 10A and 10B are diagrams for explaining how to increase the resolution of OSM road shapes in the embodiment. [Figure 13] FIG. 10 is another diagram for explaining how to increase the resolution of the OSM road shape in the same embodiment. [Figure 14] 10 is a diagram for explaining the procedure for additional correction based on OSM roads in the embodiment. FIG. [Figure 15] FIG. 10 is a diagram showing an example of a correction space used for correction based on an OSM road position in the embodiment. [Figure 16] FIG. 10 is a diagram showing an example of an additional correction result based on an OSM road in the embodiment. [Figure 17] 10A and 10B are diagrams showing an example of a captured image after lens distortion of the camera has been corrected in the embodiment. [Figure 18] FIG. 10 is a diagram showing an example of a building detection result using the object detection model in the embodiment. [Figure 19] FIG. 10 is a diagram showing an example of a segmentation result obtained by the segmentation model in the same embodiment. [Figure 20] FIG. 20 is an enlarged view of a part of FIG. 19. [Figure 21] 3 is a conceptual diagram showing a cylindrical filter region in the same embodiment. FIG. [Figure 22] FIG. 10 is a diagram for explaining the procedure for building matching in the embodiment. [Figure 23] FIG. 10 is another diagram for explaining the method of building matching in the embodiment. [Figure 24] FIG. 10 is yet another diagram for explaining the method of building matching in the embodiment. [Figure 25] FIG. 10 is yet another diagram for explaining the method of building matching in the embodiment. [Figure 26] FIG. 6 is a flowchart showing details of S5 in FIG. 5. [Figure 27] FIG. 6 is a flowchart showing details of S7 in FIG. 5. [Figure 28] FIG. 6 is a flowchart showing details of S9 in FIG. 5. [Figure 29] FIG. 6 is a flowchart showing details of a part of S11 in FIG. 5. [Figure 30] FIG. 6 is a flowchart showing the details of the remaining part of S11 in FIG. 5. [Figure 31] FIG. 6 is a flowchart showing details of S13 in FIG. 5. DETAILED DESCRIPTION OF THE INVENTION
[0016] An embodiment of the present invention will be described using a geospatial information processing system 10 shown in FIG. 1 as an example.
[0017] 1, a geospatial information processing system 10 according to this embodiment includes a data management server 30, a map database 50, and a geospatial information processing device 70, which are interconnected for mutual communication via a network 90. The network 90 is, for example, the Internet, but may also be a LAN or a WAN, or may include a LAN or a WAN.
[0018] The data management server 30 has a database (not shown), which stores data recorded by a large number of drive recorders. The data recorded by the drive recorders includes image data (video data) captured by a camera equipped in the drive recorder and positioning data captured by a positioning satellite receiver (GPS receiver) built into the drive recorder. The positioning data includes positioning position data indicating the position and positioning date and time data (timestamp) indicating the date and time of positioning. Note that the captured image data does not include data indicating the time of capture.
[0019] The map database 50 is a database of digital maps, such as an Open Street Map (hereinafter referred to as "OSM") database. OSM is well known, so a detailed description thereof will be omitted.
[0020] Then, when a geographic space represented by any photographed image data stored in the data management server 30, i.e., a photographed image of the geographic space, includes (an image of) a building such as a building, the geospatial information processing device 70 estimates the center position Pb and height Hb of the building. Furthermore, based on the estimation result of the center position Pb of the building, the geospatial information processing device 70 identifies which building on OSM (hereinafter referred to as an "OSM building") the building corresponds to, in other words, performs building matching. For example, when the photographed image includes a building as shown in FIG. 2, the geospatial information processing device 70 identifies the OSM building corresponding to the building in OSM as shown in FIG. 3, and more specifically, identifies the OSM_ID, which is identification information for the OSM building.
[0021] Here, various information about an OSM building is associated with the OSM building, including an OSM_ID, information about the land area (polygon) occupied by the OSM building, and information about the height Hb of the OSM building. However, because OSM is an open platform (user-generated content) that allows users to update, some OSM buildings may not have information about their height Hb registered, i.e., may be unregistered. The geospatial information processing device 70 estimates the building's height Hb to register (add) information about the height Hb of an OSM building that does not have information about its height Hb registered. In other words, the geospatial information processing device 70 estimates the building's center position Pb and height Hb to update the OSM. Furthermore, the geospatial information processing device 70 estimates the building's center position Pb and height Hb to update the OSM. Furthermore, the geospatial information processing device 70 estimates the building's center position Pb and height Hb to update the OSM. These estimates can also be applied to creating new digital maps other than OSM maps.
[0022] Such a geospatial information processing device 70 is configured, for example, by a personal computer (hereinafter referred to as "PC") as shown in Fig. 4. That is, the PC as the geospatial information processing device 70 has a control unit 72, an input / output interface (I / F) unit 74, an input device 76, a display device 78, an auxiliary storage unit 80, and a communication unit 82. Note that the PC as the geospatial information processing device 70 also has various other elements, but elements not directly related to the present invention are omitted from Fig. 4.
[0023] The control unit 72 is a control means responsible for overall control of the PC serving as the geospatial information processing device 70, and this control unit 72 has a CPU 72a as a control execution means. The control unit 72 also has a main memory unit 72b that is directly accessible by the CPU 72a. The main memory unit 72b includes, for example, a ROM and a RAM. Firmware including the BIOS is stored in the ROM. Meanwhile, the RAM provides a working area and a buffer area for the CPU 72a when it executes processes based on various programs included in the firmware, the operating system, and various pieces of software, such as the geospatial information processing software 80a described below.
[0024] The input / output interface unit 74 is an element that acts as a bridge between the control unit 72, particularly the CPU 72a, and each element such as the input device 76, and includes, for example, a chipset. Therefore, the input / output interface unit 74 is connected to the control unit 72 as well as each element such as the input device 76.
[0025] The input device 76 is an operation receiving means for receiving operations by an operator (not shown), and includes, for example, a keyboard and a mouse.
[0026] The display device 78 is a display means for displaying various information. This display means is, for example, a flat panel display such as a liquid crystal display or an organic EL display, but is not limited to these.
[0027] The auxiliary storage unit 80 is an auxiliary storage means and includes, for example, a hard disk drive. The auxiliary storage unit 80 stores an operating system and various software, including application software called geospatial information processing software 80a that causes the PC to function as the geospatial information processing device 70. The auxiliary storage unit 80 may include a rewritable nonvolatile memory such as a flash memory instead of or in addition to the hard disk drive.
[0028] The communication unit 82 is a communication connection means that is responsible for connection with the above-mentioned network 90. The connection between the communication unit 82 and the network 90 may be wired or wireless.
[0029] As described above, when an image captured by a drive recorder includes a building such as a building, the geospatial information processing device 70 estimates the center position Pb and height Hb of the building and then performs building matching to determine which OSM building the building corresponds to. The flow of a series of processes performed by the geospatial information processing device 70 for this purpose is shown in Figure 5. In Figure 5, symbols beginning with "S" identify each processing step, and in the following explanation, each processing step will be represented by this symbol.
[0030] As shown in Fig. 5, the geospatial information processing device 70 (more precisely, the CPU 72a) first acquires data recorded by any drive recorder from the data management server 30 in S1. As described above, the data recorded by the drive recorder includes image data captured by a camera provided in the drive recorder and positioning data captured by a positioning satellite receiver built into the drive recorder. The positioning data includes positioning position data indicating the position and positioning date and time data indicating the date and time of positioning. Note that the captured image data does not include data indicating the time of capture.
[0031] Here, the captured image data includes multiple frames. The frame rate of this captured image data is, for example, 7 fps. In other words, the frame update period of the captured image data is 1 / 7 second. On the other hand, the positioning data is acquired at a predetermined cycle, for example, every 3 seconds. That is, the frame update period of the captured image data (first cycle) is shorter than the update period of the positioning data (second cycle), i.e., different from the update period of the positioning data. Furthermore, these two update periods are asynchronous with each other.
[0032] After S1 is executed, in the subsequent S3, the geospatial information processing device 70 reconstructs a 3D model of the geospatial space represented by the captured image data, as shown in FIG. 6, using the captured image data included in the recording data acquired in S1. This reconstruction is performed, for example, using the well-known SfM (Structure from Motion) method, and more specifically, using the well-known COLMAP method. For this reason, the geospatial information processing software 40ab includes COLMAP. Furthermore, during this reconstruction, distortion due to the camera lens is corrected before the reconstruction is performed. Furthermore, during the reconstruction, the camera position when each frame is acquired is estimated, and the estimated camera position Pc[j] (j; frame number) is associated with the 3D point cloud reconstruction space, which is the space of the reconstructed 3D point cloud 100. Note that, in addition to estimating the camera position, the camera's pose (pan, tilt, roll) and parameters are also estimated, but the camera's pose and parameters are not particularly used in this embodiment.
[0033] FIG. 6 is an example of visualization of a 3D point cloud reconstruction space and a camera estimated position Pc[j] for the same geographic space as shown in FIGS. 2 and 3. As shown in FIG. 6, a 3D point cloud 100 is composed of a large number of 3D points 102. In addition, in FIG. 6, the camera estimated position Pc[j] is shown superimposed on the 3D point cloud 100, thereby expressing that the camera estimated position Pc[j] is associated with the 3D point cloud reconstruction space. Note that FIG. 6 is merely an example of visualization of the 3D point cloud reconstruction space and the camera estimated position Pc[j]; in reality, the coordinate values of each of the 3D points 102 and the coordinate values of the camera estimated position Pc[j] are obtained as processing results by SfM.
[0034] Referring again to FIG. 5, after execution of S3, the geospatial information processing device 70 aligns the 3D point cloud reconstruction space reconstructed in S3 with the real world coordinate space (WGS84) in the following S5. At the same time, the camera estimated position Pc[j] is also aligned with the real world coordinate space. This alignment is performed by a similarity transformation using Procrustes analysis, for example. The procedure will be described with reference to FIG. 7.
[0035] FIG. 7 shows an example of the matching results obtained by similarity transformation using Procrustes analysis. Specifically, the positioning positions Pg[k] (k: positioning order) represented by the positioning position data are plotted on the OSM with blue circles, and the camera estimated positions Pc[j] after matching are plotted on the OSM with red circles. The geographic space shown in FIG. 7 is different from the geographic spaces shown in FIGS. 2, 3, and 6. Furthermore, FIG. 7 plots the positioning positions Pg[k] for one minute, i.e., a total of 21 positioning positions Pg[k]. Additionally, FIG. 7 plots the camera estimated positions Pc[j] for one minute (strictly speaking, just under one minute) based on the captured image data obtained in parallel with the positioning positions Pg[k], i.e., a total of 420 camera estimated positions Pc[j]. 7, the positioning position Pg[k] at the bottom right is the first positioning position Pg[1], and the positioning position Pg[k] at the top left is the last positioning position Pg[K] (=Pg
[21] ). That is, the trajectory of the positioning positions Pg[k] shows that the vehicle (moving object) equipped with the drive recorder moved from the first positioning position Pg[1] to the last positioning position Pg[K]. The camera estimated position Pc[j] at the bottom right in FIG. 7 is the first camera estimated position Pc[1], and the camera estimated position Pc[j] at the top left in FIG. 7 is the last camera estimated position Pc[J] (=Pc
[0420] ).
[0036] This similarity transformation using Procrustes analysis is performed using each positioning position Pg[k] as an index, and therefore, a corresponding camera estimated position Pcg[k] corresponding to each positioning position Pg[k] is identified from all camera estimated positions Pc[j].
[0037] Specifically, the first positioning position Pg[1] (first positioning position) and the first camera estimated position Pc[1] (first estimated position) correspond to each other, and the last positioning position Pg[K] (second positioning position) and the last camera estimated position Pc[J] (second estimated position) correspond to each other. That is, the first camera estimated position Pc[1] is identified as the corresponding camera estimated position Pcg[1] corresponding to the first positioning position Pg[1], and the last camera estimated position Pc[J] is identified as the corresponding camera estimated position Pcg[K] (=Pcg
[21] ) corresponding to the last positioning position Pg[K].
[0038] Then, the corresponding camera estimated position Pcg[k] (third estimated position) corresponding to each of the positioning positions Pg[k] (third positioning positions) other than the first positioning position Pg[1] and the last positioning position Pg[k] is identified based on the idea that the degree of movement progress of the positioning position Pg[k] and the degree of movement progress of the camera estimated position Pc[j] are the same. More specifically, the corresponding camera estimated position Pcg[k] corresponding to the arbitrary positioning position Pg[k] is identified based on the idea that the spatial relationship of the arbitrary positioning position Pg[k] to the first positioning position Pg[1] and the last positioning position Pg[K] is the same as the spatial relationship of the corresponding camera estimated position Pcg[k] corresponding to the arbitrary positioning position Pg[k] to the first camera estimated position Pc[1] and the last camera estimated position Pc[J].
[0039] To this end, the path length Lcg[k] from the initial camera estimated position Pc[1] to the corresponding camera estimated position Pcg[k] corresponding to any positioning position Pg[k] is calculated based on the following equation 1. Here, Lg[k] is the path length from the initial positioning position Pg[1] to any positioning position Pg[k]. Lg_ALL is the total path length from the initial positioning position Pg[1] to the final positioning position Pg[K]. Lc_ALL is the total path length from the initial camera estimated position Pc[1] to the final camera estimated position Pc[J].
[0040] 《Formula 1》 Lcg[k]=(Lg[k] / Lg_ALL)·Lc_ALL
[0041] Then, the camera estimated position Pc[j] corresponding to the position Pt from the initial camera estimated position Pc[1] via the path length Lcg[k] based on Equation 1 is identified as the corresponding camera estimated position Pcg[k] corresponding to the arbitrary positioning position Pg[k]. Note that if there is no camera estimated position Pc[j] corresponding to the position Pt, the camera estimated position Pc[j] closest to the position Pt is identified as the corresponding camera estimated position Pcg[k] corresponding to the arbitrary positioning position Pg[k].
[0042] Then, a similarity transformation is performed using Procrustes analysis so that each corresponding camera estimated position Pcg[k] coincides with the positioning position Pg[k] corresponding to the corresponding camera estimated position Pcg[k]. As a result, the 3D point cloud reconstruction space, together with the camera estimated position Pc[j], is translated, rotated, and scaled to be aligned with the real world coordinate space.
[0043] When matching by similarity transformation using Procrustes analysis, preprocessing may be performed to correct variations in the positioning position Pg[k] and the camera estimated position Pc[j]. This preprocessing can be performed using, for example, a known Kalman filter. Performing such preprocessing improves the matching accuracy.
[0044] Here, each corresponding camera estimated position Pcg[k] after matching by similarity transformation using Procrustes analysis, that is, each corresponding camera estimated position Pcg[k] in the real world coordinate space, is defined by latitude and longitude, just like the positioning position Pg[k] corresponding to each corresponding camera estimated position Pcg[k]. Note that in the real world coordinate space, the vertical (perpendicular) axis is the Y axis, and the two horizontal axes are the X axis and the Y axis. In other words, coordinate values on the XZ plane of the real world coordinate space are defined by latitude and longitude.
[0045] On the other hand, coordinate values in the Y-axis direction of the real-world coordinate space are defined by, for example, altitude, but the Y-values that represent the height positions of the respective 3D points 102 that make up the aligned 3D point cloud 100, particularly the distance between two 3D points 102 and 102 in the Y-axis direction, cannot be calculated quantitatively (in meters) without some kind of standard. This is because the positioning data from a positioning satellite receiver does not include information about the height direction.
[0046] Therefore, the geospatial information processing device 70 derives a scale factor Sf as a criterion based on the mutual ratio between the distance between each corresponding camera estimated position Pcg[k] after alignment and the geodesic distance between each corresponding camera estimated position Pcg[k].
[0047] Specifically, the geospatial information processing device 70 calculates the shortest distance ΔLcg[k] between two consecutive corresponding camera estimated positions Pcg[k] and Pcg[k+1] after matching. The geospatial information processing device 70 also calculates the geodesic distance ΔLa[k] between the two corresponding camera estimated positions Pcg[k] and Pcg[k+1] in the real world coordinate space. The geospatial information processing device 70 then calculates the ratio R[k] (=ΔLcg[k] / ΔLa[k]) of the shortest distance ΔLcg[k] to the geodesic distance ΔLa[k]. The geospatial information processing device 70 then determines the average of the ratios R[k] over all sections as the scale factor Sf, i.e., derives the scale factor Sf based on the following equation 2:
[0048] 《Formula 2》 Sf=(ΣR[k]) / (K-1) where k=1~K-1
[0049] In the Y-axis direction of the real world coordinate space, the mutual distance between the two three-dimensional points 102 and 102 is multiplied by this scale factor Sf, and the mutual distance can be quantitatively determined.
[0050] As mentioned above, Fig. 7 shows an example of the matching result by similarity transformation using Procrustes analysis, in which the positioning position Pg[k] is marked on the OSM with a blue circle and the matching estimated camera position Pc[j] is marked on the OSM with a red circle. In Fig. 7, the corresponding camera estimated position Pcg[k] corresponding to each positioning position Pg[k] is marked with a green circle.
[0051] Here, ideally, the corresponding camera estimated position Pcg[k] (green circle in FIG. 7) and the positioning position Pg[k] (blue circle in FIG. 7) after matching should match each other. However, in the example shown in FIG. 7, the corresponding camera estimated position Pcg[k] and the positioning position Pg[k] after matching do not match locally, that is, there is a slight deviation between them. This deviation is due to the camera path, scene characteristics, etc. Furthermore, this deviation is considered to occur throughout the entire 3D point cloud reconstruction space after matching.
[0052] Therefore, referring again to FIG. 5, in S7 following S5, the geospatial information processing device 70 corrects the entire 3D point cloud reconstruction space with the positioned position Pg[k] as the reference.
[0053] Specifically, the geospatial information processing device 70 calculates an error vector (correction offset) Vg[k] for each of the corresponding camera estimated positions Pcg[k] after matching with respect to the positioning position Pg[k]. This error vector Vg[k] includes an error vector Vg_LAT[k] for latitude and an error vector Vg_LON[k] for longitude.
[0054] The geospatial information processing device 70 then generates a corrected space (first correction space) as shown in FIG. 8 by radial basis function (RBF) interpolation, using the corresponding camera estimated position Pcg[k] after matching as a control point and the error vector Vg[k] as an interpolated value. In FIG. 8, the red circle represents the corresponding camera estimated position Pcg[k] after matching, or in other words, the corresponding camera estimated position Pcg[k] before the correction by S7. In FIG. 8, the green circle represents the corresponding camera estimated position Pcg′[k] after the correction by S7, or in other words, the corresponding camera estimated position Pcg′[k] that is the target of the correction by S7. The horizontal axis of the correction space shown in FIG. 8 represents longitude, the vertical axis of the correction space represents longitude, and the color of the correction space represents the degree (strength) of correction.
[0055] 8 to the aligned 3D point cloud reconstructed space, thereby correcting the entire 3D point cloud reconstructed space. As a result, each pre-correction corresponding camera estimated position Pcg[k] is corrected to the target corresponding camera estimated position Pcg'[k], and accordingly, the entire 3D point cloud reconstructed space is corrected.
[0056] Fig. 9 shows an example of the correction result by S7, that is, an example of the correction result based on the positioning position Pg[k]. Note that Fig. 9 shows an extracted portion of the road on the OSM in Fig. 7 (hereinafter referred to as "OSM road"), the camera estimated position Pc[j] (red circle) and the positioning position Pg[k] (blue circle) written on the OSM, and the camera estimated position Pc'[j] after correction based on the positioning position Pg[k] (i.e., based on the correction space shown in Fig. 8) is shown with an orange circle.
[0057] As shown in Fig. 9, by performing correction based on the positioning position Pg[k], the corrected camera estimated position Pc'[j] follows the trajectory of the positioning position Pg[k]. Also, although it is not clear from Fig. 9, each corresponding camera estimated position Pcg'[k] coincides with the positioning position Pg[k]. At the same time, the entire 3D point cloud reconstruction space is corrected.
[0058] However, depending on the circumstances, an error may occur in the positioning position Pg[k]. For example, in FIG. 9, it is recognized that an error occurs in the positioning position Pg[k] within the area surrounded by the dashed rectangular frame 110 (the positioning position Pg[k] is slightly toward the center of the OSM road). In such a case, it is considered appropriate to further correct the entire 3D point cloud reconstruction space including the corrected camera estimated position Pc'[j] based on the positioning position Pg[k], specifically, to perform additional correction based on the OSM road.
[0059] Therefore, referring again to FIG. 5, in S9 following S7, the geospatial information processing device 70 performs additional correction of the entire 3D point cloud reconstructed space based on the OSM road.
[0060] To this end, the geospatial information processing device 70 acquires information about OSM roads required for additional correction, that is, OSM road information, from the map database 50. The acquisition of this OSM road information from the map database 50 is performed using Overpass_API.
[0061] The OSM road information for the OSM roads required for the additional correction refers to the OSM road information for the OSM roads that include one of the estimated camera positions Pc'[j] that are the subject of the additional correction, or the OSM road information for the OSM roads that are in the vicinity of the estimated camera position Pc'[j], more specifically, within a predetermined search radius Rs centered on the estimated camera position Pc'[j]. The search radius Rc is, for example, 10 m, but is not limited to this.
[0062] The geospatial information processing device 70 then identifies OSM roads that match each estimated camera position Pc'[j] from among the OSM roads represented by the OSM road information acquired from the map database 50, and more specifically, identifies the OSM_ID of the OSM road, performing so-called road matching. For example, if any estimated camera position Pc'[j] is included in any OSM road, the estimated camera position Pc'[j] is determined to match the OSM road that includes the estimated camera position Pc'[j]. On the other hand, for an estimated camera position Pc'[j] that is not included in any OSM road, road matching is performed based on the relationship between the shortest distance DS from the estimated camera position Pc'[j] and the movement direction DRc of the estimated camera position Pc'[j].
[0063] 10, an OSM road is defined by two nodes 120 and 120 and a linear edge 122 connecting these two nodes 120 and 120. If any of the camera estimated positions Pc'[j] is on any of the OSM roads (on the edge 122), it is determined that the OSM road matches the camera estimated position Pc'[j].
[0064] On the other hand, for a camera estimated position Pc'[j] that is not located on any OSM road, an OSM road that matches the camera estimated position Pc'[j] is identified based on a distance score Sds corresponding to the shortest distance DS from the camera estimated position Pc'[j] to each OSM road (represented by OSM road information obtained from the map database 50), and a direction score Sdr corresponding to the relationship between the movement direction DRc of the camera estimated position Pc'[j] and the orientation DRs of each OSM road.
[0065] The shortest distance DS here is the distance (orthogonal distance) from the estimated camera position Pc'[j] to the OSM road in a direction perpendicular to the trajectory of the estimated camera position Pc'[j], as shown in Figure 11. The distance score Sds is calculated based on the following equation 3. That is, the distance score Sds is a value obtained by normalizing the shortest distance DS by the above-mentioned search radius Rs.
[0066] 《Formula 3》 Sds=DS / Rs
[0067] The direction score Sdr is calculated based on the following equation 4. θdr in this equation 4 is the angle between the movement direction DRc of the camera estimated position Pc'[j] and the direction DRs of each OSM road, which is the error angle. In other words, the direction score Sdr is a value obtained by normalizing the error angle θdr by a predetermined angle of 90°.
[0068] 《Formula 4》 Sdr=θdr / 90°
[0069] It is possible to normalize the error angle θdr by 180° for the direction score Sdr, but to prevent erroneous road matching, it is more appropriate to normalize the error angle θdr by 90°, in other words, to treat 90° as the maximum error.The direction score Sdr also takes into account whether the OSM road is a one-way road or a two-way road.
[0070] For example, if an OSM road is a one-way road, the OSM road information is tagged with "oneway." For OSM roads represented by OSM road information tagged with "oneway," i.e., for one-way OSM roads, the direction score Sdr is calculated based on the above-mentioned formula 4.
[0071] On the other hand, for OSM roads represented by OSM road information that is not tagged as "oneway", i.e., for OSM roads with two-way traffic, the direction score Sdr is calculated based on Equation 4, and also based on the following Equation 5.
[0072] 《Formula 5》 Sdr=(180°-θdr) / 90°
[0073] Then, the smaller value of the direction score Sdr based on Equation 4 and the direction score Sdr based on Equation 5 is used for road matching.
[0074] After calculating the distance score Sds and direction score Sdr in this manner, the geospatial information processing device 70 identifies the OSM road with the smallest total score Sa, which is the sum of the distance score Sds and direction score Sdr, as the one that matches the camera estimated position Pc'[j].
[0075] To give a specific example, suppose the movement direction of a certain camera's estimated position Pc'[j] is north, and this certain camera's estimated position Pc'[j] is 9 m east of the actual (to-be-matched) OSM road A. Here, OSM road A is a one-way road heading north. Suppose there is another OSM road B, which is a one-way road heading southeast and is 1 m west of the certain camera's estimated position. In this case, the distance score Sds for OSM road A is 0.9 (= 9 m / 10 m), the direction score Sdr is 0.0 (= 0° / 90°), and the total score Sa is 0.9 (= 0.9 + 0.0). On the other hand, the distance score Sds for OSM road B is 0.1 (=1 m / 10 m), the direction score Sdr is 1.5 (=135° / 90°), and the total score Sa is 1.6 (=0.1 + 1.5). Therefore, a certain camera estimated position Pc'[j] is identified as matching with OSM road A, which has a smaller total score Sa.
[0076] After identifying the OSM roads that match each of the camera estimated positions Pc'[j] in this way, the geospatial information processing device 70 performs processing to increase the resolution (smooth) of the shape of the OSM roads.
[0077] That is, as explained with reference to Fig. 10, an OSM road is defined by two nodes 120 and 120 and a straight edge 122 connecting these two nodes 120 and 120. Therefore, as shown in Fig. 12, the resolution of the OSM road is low, especially in curved sections, and therefore the shape of the OSM road (road spline) tends to deviate from the shape of the actual road.
[0078] Therefore, the geospatial information processing device 70 performs processing to increase the resolution of the OSM road shapes. This processing is performed using the well-known cubic Hermite spline interpolation. As a result, the OSM road shapes are increased in resolution as shown in FIG. 13.
[0079] Then, the geospatial information processing device 70 identifies a position Ps[j] on the high-resolution OSM road that corresponds to each of the estimated camera positions Pc'[j], as shown in Fig. 14. Note that the position closest to each of the estimated camera positions Pc'[j] on the high-resolution OSM road is identified as the corresponding position Ps[j].
[0080] The geospatial information processing device 70 then calculates an error vector (correction offset) Vs[j] for each camera's estimated position Pc'[j] relative to its corresponding position Ps[j]. The error vector Vs[j] includes an error vector Vs_LAT[j] for latitude and an error vector Vs_LON[j] for longitude.
[0081] Furthermore, the geospatial information processing device 70 generates a correction space (second correction space) as shown in FIG. 15 by radial basis function interpolation, using each camera estimated position Pc'[j] as a control point and the error vector Vs[j] as an interpolated value. In FIG. 15, the red circle represents the camera estimated position Pc'[j] before the additional correction by S9, in other words, the camera estimated position Pc'[j] after the correction by S7. In addition, the green circle in FIG. 15 represents the camera estimated position Pc''[j] after the additional correction by S9, in other words, the camera estimated position Pc''[j] that is the target of the additional correction by S9. In addition, the horizontal axis of the correction space shown in FIG. 15 represents longitude, the vertical axis of the correction space represents longitude, and the color of the correction space represents the degree of additional correction.
[0082] The geospatial information processing device 70 applies the corrected space shown in FIG. 15 to the 3D point cloud reconstructed space after correction in S7, thereby additionally correcting the entire 3D point cloud reconstructed space. As a result, each camera estimated position Pc'[j] before the additional correction is corrected to the target camera estimated position Pc"[j], and accordingly, the entire 3D point cloud reconstructed space is corrected.
[0083] FIG. 16 shows an example of the result of additional correction by S9, that is, an example of the result of additional correction based on OSM roads. In FIG. 16, as in FIGS. 7 and 9, the positioning position Pg[k] is indicated by a blue circle, the corrected estimated camera position Pc'[j] is indicated by an orange circle, and the additionally corrected estimated camera position Pc"[j] is indicated by a pink circle.
[0084] As shown in Figure 16, by performing additional correction based on the OSM road, the camera estimated position Pc''[j] after this additional correction accurately follows the OSM road. In particular, the camera estimated position Pc''[j] within the area surrounded by the dashed rectangular frame 110 also accurately follows the OSM road. And, although it cannot be seen from Figure 16, the entire 3D point cloud reconstruction space is additionally corrected.
[0085] Referring again to FIG. 5, in S11 following S9, the geospatial information processing device 70 performs processing for estimating the center position Pb and height Hb of a building, in other words, for estimating the attributes of the building.
[0086] In this process, the captured image data is used after the lens distortion of the camera has been corrected by SfM. For example, Fig. 17 shows an example of captured image data, specifically a certain frame, after the lens distortion correction by SfM has been applied to the captured image data of the same geographical space as shown in Fig. 2.
[0087] The geospatial information processing device 70 performs building detection using an object detection model on the captured image data after SfM lens distortion has been applied, specifically for each frame. The object detection model used here is, for example, the well-known GroundingDINO. Therefore, GroundingDINO is set up in the geospatial information processing software 40ab.
[0088] As a result of this building detection using GroundingDINO, a bounding box 130 is attached to each detected building, as shown in Fig. 18. Note that the geographic space shown in Fig. 18 is different from the geographic space shown in Fig. 17. Furthermore, in the frame shown in Fig. 18, a two-dimensional coordinate system is set with the upper left corner as the origin, the horizontal axis as the X axis, and the vertical axis as the Y axis. This is also true for frames shown in other drawings, such as Fig. 17.
[0089] The geospatial information processing device 70 then performs filtering to select bounding boxes 130 that are suitable for estimating building attributes, in other words, to exclude bounding boxes 130 that are unsuitable for estimating building attributes. In this filtering, bounding boxes 130 that meet any of the following four exclusion conditions are excluded from subsequent processing.
[0090] The first exclusion condition is that the bounding box 130 is excessively large. Whether the bounding box 130 is excessively large is determined, for example, based on whether the area of the bounding box 130 is larger than a predetermined first threshold area. Note that the first threshold area is set appropriately based on experiments, experience, etc.
[0091] The second exclusion condition is that the bounding box 130 is excessively small. Whether the bounding box 130 is excessively small is determined, for example, based on whether the area of the bounding box 130 is larger than a predetermined second threshold area. Note that the second threshold area is smaller than the first threshold area, and like the first threshold area, this second threshold area is also set appropriately based on experimentation, experience, etc.
[0092] The third exclusion condition is that at least one of the top, bottom, left, and right edges of the bounding box 130 is outside the frame.
[0093] The fourth exclusion condition is that the upper end of the bounding box 130 is too close to the upper end of the frame. Whether the upper end of the bounding box 130 is too close to the upper end of the frame is determined based on whether the Y value of the upper end of the bounding box 130 is smaller than a predetermined threshold. This predetermined threshold is also set appropriately based on experimentation, experience, etc.
[0094] After the bounding box 130 to be subjected to subsequent processing is selected through such filtering, the geospatial information processing device 70 performs segmentation for each frame using a segmentation model to divide the building area. The segmentation model used here is, for example, the well-known SAM2. Therefore, SAM2 is set up in the geospatial information processing software 40ab. Furthermore, the bounding box 110 selected through the filtering is directly input as a prompt into SAM2. This divides the building area into buildings each having a bounding box 130 attached.
[0095] Fig. 19 is an example of a frame after segmentation by SAM 2. As shown in Fig. 19, for each building to which the bounding box 130 is attached, the building area is divided by a closed-loop division line 140. Furthermore, each division line 140 is assigned a different color.
[0096] 20 is an enlarged view of a portion of the frame shown in FIG. 19. As shown in FIG. 20, each building region includes two-dimensional feature points 142 detected from the frame when the three-dimensional point cloud is reconstructed using the SfM method described above. Each two-dimensional feature point 142 is colored the same as the division line 140 that divides the building region including the two-dimensional feature point 142. Note that, when the three-dimensional point cloud 100 is reconstructed using the SfM method described above, the three-dimensional point cloud 100 is reconstructed based on the coordinate values of the two-dimensional feature points 142 that are common between frames, the camera estimated position Pc[j] described above, and the like. Therefore, (basically), each two-dimensional feature point 142 is associated with the three-dimensional point 102 that corresponds to the two-dimensional feature point 142.
[0097] Furthermore, for each frame after segmentation by SAM2, the geospatial information processing device 70 extracts only those 2D feature points 142 within each building region that correspond to 3D points 102 in the 3D point cloud 100 reconstructed by the above-mentioned SfM (those associated with the 3D points 102 when the 3D point cloud 100 was reconstructed). This is also a type of filtering. That is, some of the 2D feature points 142 may not correspond to 3D points 102 (those not associated with the 3D points 102 when the 3D point cloud 100 was reconstructed), and such filtering is performed to exclude such 2D feature points 142 from the targets of subsequent processing.
[0098] After filtering to extract only the 2D feature points 142 corresponding to the 3D point cloud 100, the geospatial information processing device 70 determines for each building region in each frame whether the total number of 2D feature points 142 within that building region is less than a predetermined threshold number. Then, for building regions where the total number of 2D feature points 142 is less than the predetermined threshold number, the 2D feature points 142 within that building region are excluded from subsequent processing. This is also a type of filtering. Specifically, when the total number of 2D feature points 142 within a building region is extremely small, such 2D feature points 142 are likely to represent two-dimensional elements such as lines and surfaces rather than three-dimensional structures, and therefore do not contribute to accurate processing. The predetermined threshold number here is, for example, four, but is not limited to this.
[0099] Then, for each frame, the geospatial information processing device 70 identifies (detects) the 3D point 102 corresponding to the uppermost 2D feature point 144 (i.e., the one with the smallest Y value) among the 2D feature points 142 in each building region as the vertex candidate PTa of the building. As shown in Fig. 20, the uppermost 2D feature point 144 among the 2D feature points 142 in each building region is assigned a uniform color, for example, a yellow-green color.
[0100] When the same vertex candidate PTa is identified from multiple frames, more precisely, from a predetermined number of frames or more, the geospatial information processing device 70 determines that the vertex candidate PTa has been stably identified, that is, determines that the vertex candidate PTa has a high degree of certainty, and sets the vertex candidate PTa as a grouping reference point PTg.The geospatial information processing device 70 then considers that the 3D points 102 obtained from the same building area as the grouping reference point PTg represent the same building, and groups (integrates) them.The predetermined number referred to here is, for example, 3, but is not limited to this.
[0101] Furthermore, the geospatial information processing device 70 calculates the average value and standard deviation of the grouped 3D points 102 (point cloud) in the XZ plane of the 3D point cloud reconstruction space (real world coordinate space), and calculates the z-score (= (data value - average value) / standard deviation). The geospatial information processing device 70 then excludes 3D points 102 whose z-score in the XZ plane exceeds a predetermined value as outliers from subsequent processing. This method is a standard filtering method (z-score filtering method) using SciPy, a well-known Python (registered trademark) library. Note that the predetermined value here is, for example, ±2, that is, a value corresponding to a confidence range of approximately 97%, but is not limited to this.
[0102] After the outliers have been removed in this manner, the geospatial information processing device 70 further calculates the median value in the XZ plane for the 3D points 102, and sets this median value as the filter reference point PTf.
[0103] Additionally, the geospatial information processing device 70 calculates the standard deviation σ in the XZ plane with the filter reference point PTf as the reference for the 3D points 102 after the outliers have been removed.
[0104] Then, the geospatial information processing device 70 sets a cylindrical filter region 150 as shown in Fig. 21 in the 3D point cloud reconstruction space. In detail, the geospatial information processing device 70 sets the filter region 150 having a center O that is a point that is conjugate (has the same coordinate value) with the filter reference point PTf on the XZ plane and conjugate (has the same coordinate value) with the grouping reference point PTg in the Y-axis direction, a radius that is the standard deviation σ, and a predetermined dimension L in the Y-axis direction. Note that the predetermined dimension L referred to here is, for example, ±1.5 m with the center O as the reference, but is not limited to this.
[0105] Then, the geospatial information processing device 70 identifies the 3D point 102 at the top (with the largest Y value) within the filter region 150 as the building vertex Ptb.
[0106] Furthermore, the geospatial information processing device 70 identifies the camera estimated position Pc''[j] closest to the building for which the building vertex Ptb has been identified, or more precisely, the camera estimated position Pc''[j] closest to the building among the camera estimated positions Pc''[j] when the frame that contributed to the identification of the building vertex Ptb was obtained, as the shortest camera estimated position Pa.
[0107] Then, the geospatial information processing device 70 imagines the horizontal plane (XZ plane) including the nearest camera estimated position Pa as the ground plane, and estimates the height Hb of the building having the building vertex Ptb by multiplying the difference between the Y value of this nearest camera estimated position Pa and the Y value of the building vertex Ptb by the aforementioned scale factor Sf.
[0108] Note that for buildings with an excessively low estimated height Hb, specifically for buildings with an estimated height Hb less than a predetermined height threshold, the geospatial information processing device 70 determines that the building is not actually a building but some other object. The geospatial information processing device 70 then excludes such objects other than buildings from subsequent processing. This is also a type of filtering. The predetermined height threshold here is, for example, 3 m, but is not limited to this.
[0109] Then, the geospatial information processing device 70 estimates the center position Pb of a building whose height Hb is equal to or greater than a predetermined height threshold. Specifically, the geospatial information processing device 70 estimates the position corresponding to the median value on the XZ plane of the three-dimensional points within the filter region 150 as the center position Pb of the building.
[0110] The reason why the median (median absolute deviation) is used in identifying the building vertices Ptb and estimating the center position Pb as described above is that a characteristic of the 3D point cloud 100 reconstructed by SfM is that 3D points 102 tend to concentrate at the corners of a building, leaving a hollow area near the center of the building. In other words, if the average value were used instead of the median, there is a possibility that the estimated building vertices Ptb and center position Pb would deviate significantly from their actual positions, and the median is used to reduce this inconvenience as much as possible.
[0111] Referring again to FIG. 5, the geospatial information processing device 70 estimates the attributes of the building, namely the central position Pb and height Hb of the building, in S11, and then performs building matching in S13.
[0112] 22, if the estimated center position Pb of a building is included in any OSM building (polygon), that is, if such a first condition is satisfied, the geospatial information processing device 70 determines that the OSM building matches the building having the center position Pb. Then, the geospatial information processing device 70 identifies the OSM_ID of the OSM building that matches the building.
[0113] On the other hand, if the estimated building center position Pb is not included in any OSM building, i.e., if the first condition is not satisfied, the geospatial information processing device 70 performs a so-called ray cast by casting a ray 160 from the nearest camera estimated position Pa toward the estimated building center position Pb, as shown in FIG. 23 . Furthermore, the geospatial information processing device 70 extends the ray 160 as shown by the dashed line 160a in FIG. 23 . If an OSM building exists on the extension 160a of the ray 160—in other words, if an OSM building exists behind the estimated building center position Pb—if such a second condition is satisfied, so to speak, the geospatial information processing device 70 determines that the OSM building closest to the center position Pb matches the building having the center position Pb. The ray 160, including its extension 160a, is a virtual line used to calculate whether it hits any OSM building and does not need to be actually displayed.
[0114] On the other hand, if no OSM building exists on the extension line 160a of the ray 160, i.e., if the second condition is not satisfied, the geospatial information processing device 70 casts a ray 162 from the estimated building center position Pb as the base point toward the nearest camera estimated position Pa, as shown in FIG. 24 . That is, it casts a ray in the opposite direction to that shown in FIG. 23 . Then, if an OSM building exists on the ray 162, i.e., if an OSM building exists between the nearest camera estimated position Pa and the estimated building center position Pb, in other words, if such a third condition is satisfied, the geospatial information processing device 70 determines that the OSM building closest to the center position Pb matches the building having the center position Pb. Note that, due to the camera lens distortion correction using SfM described above, the estimated building center position Pb tends to be closer to the nearest camera estimated position Pa than its actual location. For this reason, the determination of whether the second condition is satisfied is performed before the determination of whether the third condition is satisfied. In other words, the second condition has a higher priority than the third condition.
[0115] If the third condition is not satisfied, i.e., if no OSM building exists between the estimated nearest camera position Pa and the estimated building center position Pb, the geospatial information processing device 70 determines whether any OSM building exists within a predetermined range based on the estimated building center position Pb, as shown in FIG. 25 . If any OSM building exists within this predetermined range, i.e., if the fourth condition is satisfied, the geospatial information processing device 70 determines that the OSM building closest to the center position Pb matches the building with the center position Pb. Note that the predetermined range here is, for example, a circular area with a predetermined radius centered on the estimated building center position Pb, and the predetermined radius here is, for example, 5 m. The shape and size of this predetermined range are not limited to this.
[0116] Furthermore, if the fourth condition is not satisfied, that is, if no OSM building exists within a predetermined range based on the estimated building center position Pb, the geospatial information processing device 70 determines that the building has not yet been registered in OSM, i.e., is unregistered. Such unregistered buildings are stored as candidates for registration in OSM. This storage destination may be, for example, the auxiliary storage unit 80, but is not limited to this.
[0117] With this, the geospatial information processing device 70 ends the series of processes shown in FIG.
[0118] FIG. 26 is a flow diagram showing details of S5 in FIG. 5. As shown in FIG. 26, the geospatial information processing device 70 first identifies, in S101, corresponding camera estimated positions Pcg[k] corresponding to each positioning position Pg[k] among the camera estimated positions Pc[j]. Then, in S103, the geospatial information processing device 70 performs similarity transformation using Procrustes analysis. This aligns the 3D point cloud reconstruction space with the real world coordinate space. Furthermore, in S105, the geospatial information processing device 70 derives a scale factor Sf. With this, the geospatial information processing device 70 ends S5 in FIG. 5.
[0119] FIG. 27 is a flow diagram showing details of S7 in FIG. 5. As shown in FIG. 27, the geospatial information processing device 70 first calculates, in S201, an error vector Vg[k] of the corresponding camera estimated position Pcg[k] after alignment with respect to the positioned position Pg[k] by similarity transformation using Procrustes analysis. Then, in subsequent S203, the geospatial information processing device 70 generates a corrected space as shown in FIG. 8 by radial basis function interpolation, using the corresponding camera estimated position Pcg[k] after alignment as a control point and the error vector Vg[k] as an interpolated value. Furthermore, in subsequent S205, the geospatial information processing device 70 applies the corrected space generated in S203 to the 3D point cloud reconstructed space after alignment, thereby correcting the entire 3D point cloud reconstructed space. With this, the geospatial information processing device 70 ends S7 in FIG. 5.
[0120] FIG. 28 is a flow diagram showing details of S9 in FIG. 5. As shown in FIG. 28, first, in S301, the geospatial information processing device 70 acquires OSM road information about OSM roads required for additional correction from the map database 50. Then, in S303, the geospatial information processing device 70 identifies OSM roads that match each estimated camera position Pc'[j], that is, identifies the OSM_ID of the OSM road. Furthermore, in S305, the geospatial information processing device 70 increases the resolution of the shape of the OSM road identified in S303 using known cubic Hermite spline interpolation. Then, in S307, the geospatial information processing device 70 identifies a position Ps on the OSM road that has been increased in resolution in S305, corresponding to each estimated camera position Pc'[j].
[0121] Furthermore, in S309 following S307, the geospatial information processing device 70 calculates the error vector Vs[j] of each camera's estimated position Pc'[j] relative to the corresponding position Ps[j] identified in S307. Then, in S311, the geospatial information processing device 70 generates a corrected space as shown in FIG. 15 by radial basis function interpolation, using each camera's estimated position Pc'[j] as a control point and the error vector Vs[j] as an interpolated value. Then, in S313, the geospatial information processing device 70 applies the corrected space generated in S311 to the 3D point cloud reconstructed space corrected in S7 (S205) described above, thereby additionally correcting the entire 3D point cloud reconstructed space. With this, the geospatial information processing device 70 ends S9 in FIG. 5.
[0122] In addition, FIGS. 29 and 30 are flow diagrams showing details of S11 in FIG. 5. First, as shown in FIG. 29, in S401, the geospatial information processing device 70 performs building detection using an object detection model for each frame of a captured image after lens distortion by SfM has been applied. The object detection model used here is, for example, the well-known GroundingDINO. Then, in S403, the geospatial information processing device 70 performs filtering to remove bounding boxes 130, which are the building detection results of S401, that are inappropriate for estimating building attributes. Furthermore, in S405, the geospatial information processing device 70 performs segmentation to divide building regions into segments for each frame of the captured image data using a segmentation model. The segmentation model used here is, for example, SAM2. The bounding box 110 selected by filtering in S403 is directly input to SAM2 as a prompt.
[0123] In S407 following S405, the geospatial information processing device 70 performs filtering for each frame after the segmentation in S405 to extract only those 2D feature points 142 within each building region that correspond to 3D points 102 in the 3D point cloud 100. Then, in S409, the geospatial information processing device 70 determines for each frame, for each building region, whether the total number of 2D feature points 142 within the building region is less than a predetermined threshold number, and excludes building regions whose total number of 2D feature points 142 is less than the predetermined threshold number from targets for subsequent processing, that is, performs filtering to do so. Furthermore, in S411, for each building region for each frame, the geospatial information processing device 70 identifies the 3D point 102 corresponding to the uppermost 2D feature point 144 (i.e., the 2D feature point with the smallest Y value) among the 2D feature points 142 within the building region as a vertex candidate PTa of the building.
[0124] Then, in S413, the geospatial information processing device 70 sets the vertex candidate PTa that is commonly identified in a predetermined number of frames or more as a grouping reference point PTg. Furthermore, in S415, the geospatial information processing device 70 groups the 3D points 102 obtained from the building area common to the grouping reference point PTg. Additionally, in S417, the geospatial information processing device 70 excludes, as outliers, 3D points 102 whose z-scores in the XZ plane of the 3D point cloud reconstruction space (real world coordinate space) exceed a predetermined value from subsequent processing, i.e., performs filtering to do so. Then, in S419, the geospatial information processing device 70 calculates the median in the XZ plane of the 3D points 102 after the outliers have been removed by the filtering in S417, and sets this median as a filter reference point PTf.
[0125] 30, in the next step S421, the geospatial information processing device 70 calculates the standard deviation σ in the XZ plane based on the filter reference point PTf for the 3D points 102 after the outliers have been removed by filtering in step S419. Then, in the next step S423, the geospatial information processing device 70 sets a cylindrical filter region 150. That is, the geospatial information processing device 70 sets the filter region 150 having a center O that is a point conjugate to the filter reference point PTf in the XZ plane and conjugate to the grouping reference point PTg in the Y-axis direction, a radius that is the standard deviation σ, and a predetermined dimension L in the Y-axis direction.
[0126] Then, in the next S425, the geospatial information processing device 70 identifies as the building vertex Ptb the 3D point 102 that is located at the top (maximum Y value) within the filter region 150. Furthermore, in the next S427, the geospatial information processing device 70 estimates the height Hb of the building that has the building vertex Ptb based on the difference between the Y value of the nearest camera estimated position Pa and the Y value of the building vertex Ptb, specifically by multiplying this difference by a scale factor Sf.
[0127] Additionally, in the following S429, the geospatial information processing device 70 performs filtering to exclude buildings whose height Hb estimated in S427 is less than a predetermined height threshold from targets for subsequent processing. Then, for buildings whose height Hb is equal to or greater than the predetermined height threshold, the geospatial information processing device 70 estimates the position corresponding to the median on the XZ plane of the three-dimensional points within the filter region 150 as the center position Pb of the building. With this, the geospatial information processing device 70 ends S11 in FIG. 5.
[0128] 31 is a flow diagram showing details of S13 in FIG. 5. As shown in FIG. 31, the geospatial information processing device 70 first determines in S501 whether the above-mentioned first condition is satisfied, that is, whether an OSM building including the estimated building center position Pb exists. If the first condition is satisfied, that is, if an OSM building including the estimated building center position Pb exists (S501: YES), the geospatial information processing device 70 proceeds to S503. On the other hand, if the first condition is not satisfied, that is, if an OSM building including the estimated building center position Pb does not exist (S501: NO), the geospatial information processing device 70 proceeds to S505, which will be described later.
[0129] In S503, the geospatial information processing device 70 identifies the corresponding OSM building, that is, the OSM building that satisfies the first condition, as the target OSM building, and more specifically, identifies its OSM_ID. With this, the geospatial information processing device 70 ends S13 in FIG. 5.
[0130] On the other hand, if the process proceeds from S501 to S505, the geospatial information processing device 70 determines in S505 whether the second condition is satisfied, i.e., whether an OSM building exists on the extension 160a of the ray 160. If the second condition is satisfied, i.e., if an OSM building exists on the extension 160a of the ray 160 (S505: YES), the geospatial information processing device 70 proceeds to S503. On the other hand, if the second condition is not satisfied, i.e., if an OSM building does not exist on the extension 160a of the ray 160 (S505: NO), the geospatial information processing device 70 proceeds to S507. Note that if the process proceeds from S505 to S503, the OSM building that satisfies the second condition and is closest to the estimated building center position Pb is identified as the target OSM building in S503.
[0131] In S507, the geospatial information processing device 70 determines whether the aforementioned third condition is satisfied, that is, whether an OSM building exists on the ray 160. If the third condition is satisfied, that is, if an OSM building exists on the ray 160 (S507: YES), the geospatial information processing device 70 proceeds to S503. On the other hand, if the third condition is not satisfied, that is, if an OSM building does not exist on the ray 160 (S507: NO), the geospatial information processing device 70 proceeds to S509. Note that when the process proceeds from S507 to S503, the OSM building that satisfies the third condition and is closest to the estimated building center position Pb is identified as the target OSM building in S503.
[0132] In S509, the geospatial information processing device 70 determines whether the aforementioned fourth condition is satisfied, i.e., whether an OSM building exists within a predetermined range based on the estimated building center position Pb. If the fourth condition is satisfied, i.e., if an OSM building exists within the predetermined range based on the estimated building center position Pb (S509: YES), the geospatial information processing device 70 proceeds to S503. On the other hand, if the fourth condition is not satisfied, i.e., if an OSM building does not exist within the predetermined range based on the estimated building center position Pb (S509: NO), the geospatial information processing device 70 proceeds to S519. Note that when the process proceeds from S509 to S503, the OSM building that satisfies the fourth condition and is closest to the estimated building center position Pb is identified as the target OSM building in S503.
[0133] In S511, the geospatial information processing device 70 stores the building having the estimated central position Pb as an unregistered building, and then ends S13.
[0134] As described above, according to this embodiment, the center position Pb and height Hb of a building included in the recorded data are estimated based on the data recorded by a drive recorder. That is, the center position Pb and height Hb of a building are estimated using a drive recorder, a device that is widely available on the market. The estimated results of the center position Pb and height Hb of a building can be applied to updating OSM or to creating new digital maps other than OSM. This significantly contributes to realizing the creation and updating of accurate digital maps at lower costs than conventional methods. Furthermore, this effect is more pronounced the larger the area of the digital map.
[0135] In this embodiment, the geospatial information processing device 70, specifically the CPU 72a that executes S1 in FIG. 5, is an example of an image data input accepting means according to the present invention and an example of a positioning data input accepting means according to the present invention. The CPU 72a that executes S3 in FIG. 5 is an example of a 3D point cloud reconstruction means according to the present invention, and the CPU 72a that executes S5 in FIG. 5 is an example of a spatial matching means according to the present invention. The CPU 72a that executes S401 in FIG. 29 is an example of a building detection means according to the present invention, and the CPU 72a that executes S405 in FIG. 29 is an example of a segmentation means according to the present invention. The CPU 72a that executes S411 in FIG. 29 is an example of a vertex candidate identification means according to the present invention, and the CPU 72a that executes S415 in FIG. 29 is an example of a grouping means according to the present invention. Additionally, the CPU 72a that executes S423 in FIG. 30 is an example of a filter region setting means according to the present invention, and the CPU 72a that executes S425 in FIG. 30 is an example of a building vertex identification means according to the present invention. Furthermore, the CPU 72a that executes S427 in FIG. 30 is an example of a building height estimation means according to the present invention, and the CPU 72a that executes S431 in FIG. 30 is an example of a building center position estimation means according to the present invention.
[0136] The present embodiment is merely a specific example of the present invention and does not limit the scope of the present invention. The present invention can be applied to various aspects other than the present embodiment.
[0137] For example, although the geospatial information processing device 70 is configured by a PC in the above embodiment, the present invention is not limited to this. The geospatial information processing device 70 may be configured by a device other than a PC, particularly a dedicated device.
[0138] Furthermore, although the data recorded by the drive recorder is assumed to be acquired from the data management server 30, this is not limitative. For example, the data recorded by the drive recorder may be acquired directly from the drive recorder.
[0139] Furthermore, the map database 50 may be a database of digital maps other than OSM.
[0140] Additionally, the center O of the filter region 150 is set to a point that is conjugate with the filter reference point PTf in the XZ plane and conjugate with the grouping reference point PTg in the Y-axis direction, but this is not limitative. For example, the grouping reference point PTg, that is, the stably identified vertex candidate PTa, may be set to the center O of the filter region 150. The filter region 150 may also have a shape other than a cylindrical shape.
[0141] The present invention is applicable not only to the creation and updating of digital maps, but also to fields other than digital maps, such as the metaverse and games.
[0142] Furthermore, the present invention is not limited to being provided in the form of a geospatial information processing device, but can also be provided in the form of a geospatial information processing method. [Explanation of symbols]
[0143] 10...Geospatial Information Processing System 30...Data management server 50 … Map database 70 ... Geospatial information processing device 72 ... Control section 72a...CPU 72b… Main memory section 80 … Auxiliary storage section 80a ... Geospatial information processing software 100...3D point cloud 102 … 3D point 12 ... Control section 12a … APP 12b... Main memory section Pa: Shortest estimated camera position Pc[j] … Estimated camera position Pc'[j] ... estimated camera position after correction Pc”[j] … Estimated camera position after additional correction Pcg[k] … Estimated corresponding camera position Pcg'[k] … Estimated corresponding camera position after correction Pg[k] … Positioning position Ps[j] … Corresponding position
Claims
1. an image data input receiving means for receiving input of image data consisting of a plurality of frames obtained by sequentially photographing the geographical space around the camera at a predetermined period by a camera moving on a road; a three-dimensional point cloud reconstruction means for estimating a position of the camera when each of the plurality of frames was obtained based on the image data, and for reconstructing a three-dimensional model of the geographical space using a three-dimensional point cloud in association with the estimated position of the camera; a space alignment means for aligning the three-dimensional point cloud reconstruction space reconstructed by the three-dimensional point cloud reconstruction means with a real world coordinate space together with the estimated position of the camera; a building detection means for detecting a building included in each of the plurality of frames; a division means for dividing an area of the building in the frame including the building detected by the building detection means; a vertex candidate specifying means for specifying, from the three-dimensional point cloud, a three-dimensional point corresponding to the uppermost two-dimensional feature point in the area divided by the dividing means, as a vertex candidate of the building related to the area; a filter region setting means for setting a predetermined three-dimensional filter region centered on the vertex candidate; a building vertex specifying means for specifying a three-dimensional point located at the top of the three-dimensional point cloud within the filter region as a vertex of the building related to the three-dimensional point; a building height estimation means for estimating the height of the building based on the vertical distance between the vertex of the building identified by the building vertex identification means and the estimated position of the camera when the frame including the building was obtained.
2. an image data input receiving means for receiving input of image data consisting of a plurality of frames obtained by sequentially photographing the geographical space around the camera at a predetermined period by a camera moving on a road; a three-dimensional point cloud reconstruction means for estimating a position of the camera when each of the plurality of frames was obtained based on the image data, and for reconstructing a three-dimensional model of the geographical space using a three-dimensional point cloud in association with the estimated position of the camera; a space alignment means for aligning the three-dimensional point cloud reconstruction space reconstructed by the three-dimensional point cloud reconstruction means with a real world coordinate space together with the estimated position of the camera; a building detection means for detecting a building included in each of the plurality of frames; a division means for dividing an area of the building in the frame including the building detected by the building detection means; a vertex candidate specifying means for specifying, from the three-dimensional point cloud, a three-dimensional point corresponding to the uppermost two-dimensional feature point in the area divided by the dividing means, as a vertex candidate of the building related to the area; a filter region setting means for setting a predetermined three-dimensional filter region centered on the vertex candidate; a building center position estimation means for estimating a three-dimensional point among the three-dimensional point cloud that corresponds to a median value in the horizontal direction within the filter area as the center position of the building related to the three-dimensional point.
3. further comprising a grouping means for, when a common vertex candidate is identified from a predetermined number or more of the frames, grouping, among the three-dimensional point cloud, three-dimensional points corresponding to the two-dimensional feature points that are in the area common to the vertex candidate, 3. The geospatial information processing device according to claim 1, wherein the filter area setting means sets the filter area by replacing the vertex candidate with a three-dimensional point that corresponds to an intermediate value in a horizontal direction among the three-dimensional points grouped by the grouping means, and setting the filter area as the center of the filter area.
4. a positioning data input receiving means for receiving input of positioning data obtained by positioning using a positioning satellite receiver that moves on the road together with the camera; the space alignment means aligns the three-dimensional point cloud reconstruction space with the real world coordinate space together with the estimated position of the camera using the positioning position represented by the positioning data as an index; 2. The geospatial information processing device according to claim 1, wherein the building height estimation means estimates the height of the building by multiplying the mutual distance by a scale factor that is a ratio between a movement distance of the estimated position of the camera after alignment by the spatial alignment means and a geodesic distance corresponding to the movement distance.
5. A geospatial information processing method by a computer, comprising: The processor of the computer an image data input receiving step of receiving input of image data consisting of a plurality of frames obtained by sequentially photographing a geographical space around the camera at a predetermined period by a camera moving on a road; a three-dimensional point cloud reconstruction step of estimating a position of the camera when each of the plurality of frames was obtained based on the image data, and reconstructing a three-dimensional model of the geographical space using a three-dimensional point cloud in association with the estimated position of the camera; a space alignment step of aligning the three-dimensional point cloud reconstruction space reconstructed by the three-dimensional point cloud reconstruction step with a real world coordinate space together with the estimated position of the camera; a building detection step of detecting a building included in each of the plurality of frames; a division step of dividing an area of the building in the frame including the building detected by the building detection step; a vertex candidate specifying step of specifying, from the three-dimensional point cloud, a three-dimensional point corresponding to the two-dimensional feature point at the top of the region divided by the dividing step, as a vertex candidate of the building related to the region; a filter region setting step of setting a predetermined three-dimensional filter region centered on the vertex candidate; a building vertex identification step of identifying a 3D point located at the top of the filter region among the 3D point cloud as a vertex of the building related to the 3D point; a building height estimation step of estimating the height of the building based on the vertical distance between the vertex of the building identified in the building vertex identification step and the estimated position of the camera when the frame including the building was obtained.
6. A geospatial information processing method by a computer, comprising: The processor of the computer an image data input receiving step of receiving input of image data consisting of a plurality of frames obtained by sequentially photographing a geographical space around the camera at a predetermined period by a camera moving on a road; a three-dimensional point cloud reconstruction step of estimating a position of the camera when each of the plurality of frames was obtained based on the image data, and reconstructing a three-dimensional model of the geographical space using a three-dimensional point cloud in association with the estimated position of the camera; a space alignment step of aligning the three-dimensional point cloud reconstruction space reconstructed by the three-dimensional point cloud reconstruction step with a real world coordinate space together with the estimated position of the camera; a building detection step of detecting a building included in each of the plurality of frames; a division step of dividing an area of the building in the frame including the building detected by the building detection step; a vertex candidate specifying step of specifying, from the three-dimensional point cloud, a three-dimensional point corresponding to the two-dimensional feature point at the top of the region divided by the dividing step, as a vertex candidate of the building related to the region; a filter region setting step of setting a predetermined three-dimensional filter region centered on the vertex candidate; a building center position estimation step of estimating a three-dimensional point among the three-dimensional point cloud that corresponds to a median value in the horizontal direction within the filter area as the center position of the building related to the three-dimensional point.
Citation Information
Patent Citations
Method for estimating height of urban building by utilizing streetscape image under serious shielding condition
CN118229755A
Navigation device
JP2019105581A
Program, method, system, road map, and road map creation method
JP2024013788A
Position estimation device, position estimation method, and position estimation computer program
JP2024168104A
For herbicide wheat cropping
JP1989000007A