Geospatial Information Processing Apparatus and Geospatial Information Processing Method
The geospatial information processing apparatus and method address the high cost of LiDAR-based 3D mapping by using a camera and satellite positioning to reconstruct and align 3D models with real-world coordinates, achieving cost-effective and accurate digital map updates.
Patent Information
- Application Number
- JP2025084233
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The high cost associated with creating and updating large-scale 3D maps using LiDAR technology is a significant challenge, as the cost increases with the area covered.
A geospatial information processing apparatus and method that utilizes a camera and positioning satellite receiver to capture images and data, reconstructs a 3D model, aligns it with real-world coordinates, and corrects spatial relationships using radial basis function interpolation and cubic Hermite spline interpolation to enhance accuracy and resolution, thereby reducing costs.
This approach enables the creation and update of accurate digital maps at a lower cost, particularly for large areas, by leveraging widely available devices like drive recorders and open platforms like OpenStreetMap, ensuring high-resolution and alignment with real-world coordinates.
Smart Images

Figure 0007706035000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a geospatial information processing apparatus and a geospatial information processing method, and particularly to a geospatial information processing apparatus and a geospatial information processing method applicable to the creation and update of digital maps.
Background Art
[0002] In recent years, with the development of autonomous driving technology and advanced navigation systems, the demand for 3D maps, which are a type of digital map, has been increasing. As a method for creating such 3D maps, for example, as disclosed in Non-Patent Document 1, there is a method using LiDAR (Light Detection and Ranging), and particularly, a method using a dedicated vehicle equipped with LiDAR.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the method using LiDAR, a corresponding cost is incurred, and the wider the area of the digital map, the higher the cost.
[0005] Therefore, an object of the present invention is to provide a novel geospatial information processing apparatus and a geospatial information processing method that greatly contribute to realizing the creation and update of accurate digital maps while suppressing costs more than in the past.
Means for Solving the Problems
[0006] To achieve this object, the present invention includes a first invention related to a geospatial information processing apparatus and a second invention related to a geospatial information processing method.
[0007] Among these, the first invention related to the geospatial information processing apparatus includes an image data input receiving means, a positioning data input receiving means, a three-dimensional reconstruction means, a spatial alignment means, and a first spatial correction means. The image data input receiving means receives the input of image data obtained by sequentially photographing the geospatial around the camera by a camera moving on a road at a first cycle. Then, the positioning data input receiving means receives the input of positioning data obtained by sequentially positioning at a second cycle different from the first cycle by a positioning satellite receiver moving on the road together with the camera. The three-dimensional reconstruction means estimates the position of the camera for each first cycle based on the image data, and reconstructs (restores) a three-dimensional model of the geospatial in a state associated with the estimated position of the camera. The spatial alignment means aligns the three-dimensional reconstructed space reconstructed by the three-dimensional reconstruction means with the estimated position of the camera in the real-world coordinate space, using the positioning position for each second cycle represented by the positioning data as an index. And the first spatial correction means corrects the three-dimensional reconstructed space after alignment by the spatial alignment means together with the estimated position of the camera so that the positioning position for each second cycle and the corresponding camera estimated position corresponding to the positioning position among the estimated positions of the camera for each first cycle after alignment by the spatial alignment means match each other. The alignment of the three-dimensional reconstructed space to the real-world coordinate space and the correction based on the positioning position of the three-dimensional reconstructed space after alignment in this manner are applied to the creation and update of digital maps.
[0008] Note that when the spatial alignment means aligns the three-dimensional reconstructed space with the real-world coordinate space, it regards a first measurement position, which is a specific one among the measurement positions for each second period, and a first estimated position, which is a specific one among the estimated positions of the camera for each first period, as corresponding to each other. At the same time, the spatial alignment means regards a second measurement position, which is a specific measurement position different from the first measurement position, and a second estimated position, which is a specific estimated position of the camera different from the first estimated position, as corresponding to each other. Further, the spatial alignment means regards a third measurement position, which is at least one measurement position other than the first measurement position and the second measurement position, and a third estimated position, which is at least one estimated position of the camera other than the first estimated position and the second estimated position, as corresponding to each other. Then, the spatial alignment means aligns the three-dimensional reconstructed space with the real-world coordinate space. At this time, the third estimated position is determined so that the spatial relationship of the third measurement position with respect to each of the first measurement position and the second measurement position is common (the same) to the spatial relationship of the third estimated position with respect to each of the first estimated position and the second estimated position.
[0009] Furthermore, the first space correction means includes, for example, a first correction space generation means and a first correction execution means. Among these, the first correction space generation means generates a first correction space based on the corresponding camera estimated position and the error of the corresponding camera estimated position with respect to the measurement position corresponding to the corresponding camera estimated position. Then, the first correction execution means corrects the three-dimensional reconstructed space after alignment by the spatial alignment means with the first correction space.
[0010] Here, the first correction space generation means may generate the first correction space by radial basis function interpolation, using the corresponding camera estimated position as a control point and the aforementioned error as an interpolation value.
[0011] In addition, in the present first invention, a corresponding position specifying means and a second space correction means may be further provided. Among these, the corresponding position specifying means specifies a corresponding position on a map road included in predetermined map information corresponding to the estimated position of the camera after correction by the first space correction means. Then, the second space correction means, so that the estimated position of the camera after correction by the first space correction means and the corresponding position on the map road corresponding to the estimated position of the camera coincide with each other, so to speak, based on the map road, the three-dimensional reconstruction space after correction by the first space correction means is additionally corrected.
[0012] Here, the second space correction means includes, for example, a second correction space generation means and a second correction execution means. Among these, the second correction space generation means generates a second correction space based on the estimated position of the camera after correction by the first space correction means and the error of the estimated position of the camera with respect to the corresponding position on the map road corresponding to the estimated position of the camera. Then, the second correction execution means additionally corrects the three-dimensional reconstruction space after correction by the first space correction means with the second correction space.
[0013] Further, the second correction space generation means may generate the second correction space by radial basis function interpolation, with the estimated position of the camera after correction by the first space correction means as a control point and the aforementioned error as an interpolation value.
[0014] In addition, a high-resolution means may be further provided. This high-resolution means increases the resolution of the shape of the map road. Then, the corresponding position specifying means may specify the corresponding position on the high-resolution road, which is the map road with increased resolution by the high-resolution means.
[0015] Here, the high-resolution means may increase the resolution of the shape of the map road by, for example, cubic Hermite spline interpolation.
[0016] The second invention related to the method for processing geospatial information in the present invention is a method for processing geospatial information by a computer. The processor of the computer executes an image data input reception step, a positioning data input reception step, a three-dimensional reconstruction step, a spatial alignment step, and a spatial correction step. In the image data input reception step, the processor receives the input of image data obtained by sequentially photographing the geospatial area around the camera by a camera moving on a road in a first cycle. Then, in the positioning data input reception step, the processor receives the input of positioning data obtained by sequentially positioning by a positioning satellite receiver moving on the road together with the camera in a second cycle different from the first cycle. In the subsequent three-dimensional reconstruction step, the processor estimates the position of the camera for each first cycle based on the image data, and reconstructs a three-dimensional model of the geospatial area in a state associated with the estimated position of the camera. Further, in the spatial alignment step, the processor aligns the three-dimensional reconstructed space reconstructed in the three-dimensional reconstruction step with the real-world coordinate space together with the estimated position of the camera, using the positioning position for each second cycle represented by the positioning data as an index. And in the spatial correction step, the processor makes the positioning position for each second cycle and the corresponding camera estimated position corresponding to the positioning position among the estimated positions of the camera for each first cycle after alignment in the spatial alignment step coincide with each other. That is, with the positioning position as a reference, the three-dimensional reconstructed space after alignment in the spatial alignment step is corrected together with the estimated position of the camera. The alignment of the three-dimensional reconstructed space to the real-world coordinate space and the correction based on the positioning position of the three-dimensional reconstructed space after alignment in such a manner are applied to the creation and update of digital maps.
Advantages of the Invention
[0017] Such a present invention greatly contributes to realizing the creation and update of accurate digital maps while suppressing costs compared with the prior art. This effect is more remarkable as the area of the digital map is larger.
Brief Description of the Drawings
[0018]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Embodiments for Carrying Out the Invention
[0019] An embodiment of the present invention will be described by taking the geospatial information processing system 10 shown in FIG. 1 as an example.
[0020] As shown in FIG. 1, the geospatial information processing system 10 according to this embodiment includes a data management server 30, a map database 50, and a geospatial information processing device 70, which are communicably connected to each other via a network 90. Note that the network 90 is, for example, the Internet, but it may also be a LAN or a WAN, or may include a LAN and a WAN.
[0021] The data management server 30 has a database (not shown), and a large number of recording data from a plurality of drive recorders are stored in this database. The recording data from the drive recorder includes captured image data (video data) by a camera included in the drive recorder and positioning data by a positioning satellite receiver, a so-called GPS receiver, built in the drive recorder. Further, the positioning data includes positioning position data representing the positioning position and positioning date and time data (timestamp) representing the positioning date and time. Note that the captured image data does not include data representing the shooting time.
[0022] The map database 50 is a database of digital maps, for example, a database of OpenStreetMap (hereinafter referred to as "OSM"). Since OSM is well known, a detailed description thereof will be omitted.
[0023] Then, when a building (image) such as a building is included in the geospatial area represented by any captured image data stored in the data management server 30, that is, in the captured image of the geospatial area, the geospatial information processing device 70 estimates the center position Pb and height Hb of the building. Further, based on the estimated center position Pb of the building, the geospatial information processing device 70 identifies which building (hereinafter referred to as "OSM building") on the OSM the building corresponds to, and performs so-called building matching. For example, when the captured image includes a building as shown in FIG. 2, the geospatial information processing device 70 identifies the OSM building corresponding to the building in the OSM as shown in FIG. 3, and specifically identifies the OSM_ID, which is the identification information of the OSM building.
[0024] Here, for OSM buildings in OSM, various pieces of information related to the OSM building are associated, including the OSM_ID, information about the range (polygon) of the land occupied by the OSM building, information about the height Hb of the OSM building, and the like. On the other hand, since OSM is an open platform (user-generated content) that can be updated by users, for some OSM buildings, information about their height Hb may not be registered, that is, it may be unregistered. To register (add) the information about the height Hb of an OSM building for which the information about the height Hb is unregistered, the estimation result of the height Hb of the building by the geospatial information processing device 70 is applied. That is, the estimation results of the center position Pb and the height Hb of the building by the geospatial information processing device 70 are applied to the update of OSM. Also, the estimation results of the center position Pb and the height Hb of the building by the geospatial information processing device 70 can be applied not only to the update of OSM but also to the creation of a new digital map other than OSM.
[0025] Such a geospatial information processing device 70 is constituted by, for example, a personal computer (hereinafter referred to as "PC") as shown in FIG. 4. That is, the PC as the geospatial information processing device 70 has a control unit 72, an input / output interface (I / F) unit 74, an input device 76, a display device 78, an auxiliary storage unit 80, and a communication unit 82. Note that the PC as the geospatial information processing device 70 has various other elements, but in FIG. 4, the illustration of elements not directly related to the present invention is omitted.
[0026] The control unit 72 is a control means responsible for controlling the entire PC as the geospatial information processing device 70. This control unit 72 has a CPU 72a as control execution means. Additionally, the control unit 72 has a main memory unit 72b that can be directly accessed by the CPU 72a. The main memory unit 72b includes, for example, ROM and RAM. Firmware including BIOS is stored in the ROM among them. On the other hand, the RAM constitutes a work area and a buffer area when the CPU 72a executes processes based on various programs included in various software such as firmware, an operating system, and the geospatial information processing software 80a described later.
[0027] The input / output interface unit 74 is an element that bridges the control unit 72, particularly the CPU 72a, and each element such as the input device 76, and includes, for example, a chipset. For this reason, the control unit 72 is connected to the input / output interface unit 74, and each element such as the input device 76 is connected to the input / output interface unit 74.
[0028] The input device 76 is an operation reception means for receiving operations by an operator (not shown). This input device 76 includes, for example, a keyboard and a mouse.
[0029] The display device 78 is a display means for displaying various information. This display means is, for example, a flat panel display such as a liquid crystal display or an organic EL display, but is not limited to this.
[0030] The auxiliary storage unit 80 is auxiliary storage means and has, for example, a hard disk drive. An operating system is stored in this auxiliary storage unit 80, and various software is stored, and in particular, application software called geospatial information processing software 80a for causing the PC to function as the geospatial information processing device 70 is stored. Note that the auxiliary storage unit 80 may have a rewritable non-volatile memory such as a flash memory instead of or in addition to the hard disk drive.
[0031] The communication unit 82 is a communication connection means responsible for connecting to the aforementioned network 90. The connection between this communication unit 82 and the network 90 may be wired or wireless.
[0032] As described above, when a building such as a building is included in the captured image by the drive recorder, the geospatial information processing apparatus 70 estimates the center position Pb and the height Hb of the building, and further performs building matching to determine which OSM building on the OSM the building corresponds to. A series of processing flows by the geospatial information processing apparatus 70 for this purpose are shown in FIG. 5. In FIG. 5, the symbols starting with "S" are symbols for identifying each processing step. In the following description, each processing step is represented by the symbol.
[0033] As shown in FIG. 5, the geospatial information processing apparatus 70 (strictly speaking, the CPU 72a) first acquires, in S1, recording data from an arbitrary drive recorder from the data management server 30. As described above, the recording data by the drive recorder includes captured image data by a camera provided in the drive recorder and positioning data by a positioning satellite receiver built in the drive recorder. The positioning data includes positioning position data representing the positioning position and positioning time data representing the positioning time. Note that the captured image data does not include data (time stamp) representing the shooting time.
[0034] Here, the captured image data includes a plurality of frames. The frame rate of this captured image data is, for example, 7 fps. In other words, the frame update period of the captured image data is 1 / 7 second. On the other hand, the positioning data is acquired at a predetermined period, for example, at a period of 3 seconds. That is, the frame update period (the first period) of the captured image data is shorter than the update period (the second period) of the positioning data, that is, different from the update period of the positioning data. Also, these two update periods are asynchronous with each other.
[0035] After the execution of S1, in subsequent S3, based on the captured image data included in the recorded data acquired in S1, the geospatial information processing apparatus 70 reconstructs a three-dimensional model of the geospatial area represented by this captured image data with a three-dimensional point cloud 100 as shown in FIG. 6. This reconstruction is performed, for example, by known SfM (Structure from Motion), and more specifically, by known COLMAP. Therefore, the aforementioned geospatial information processing software 40ab includes COLMAP. Further, during this reconstruction, distortion caused by the camera lens is corrected, and on top of that, the reconstruction is performed. Furthermore, in the reconstruction, the position of the camera when each frame is obtained is estimated, and the estimated camera position Pc[j] (j; frame number), which is the estimation result, is associated with the three-dimensional point cloud reconstruction space that is the space of the reconstructed three-dimensional point cloud 100. Incidentally, together with the estimation of the camera position, the camera's attitude (pan, tilt, roll) and parameters are also estimated, but in this embodiment, the camera's attitude and parameters are not particularly utilized.
[0036] Incidentally, FIG. 6 is an example in which the three-dimensional point cloud reconstruction space and the estimated camera position Pc[j] for the same geospatial area shown in FIGS. 2 and 3 are visualized. As shown in FIG. 6, the three-dimensional point cloud 100 is composed of a large number of three-dimensional points 102. Also, in FIG. 6, the estimated camera position Pc[j] is shown superimposed on the three-dimensional point cloud 100, expressing that the estimated camera position Pc[j] is associated with the three-dimensional point cloud reconstruction space. Note that FIG. 6 is merely an example in which the three-dimensional point cloud reconstruction space and the estimated camera position Pc[j] are visualized, and actually, the coordinate values of each three-dimensional point 102 and the coordinate values of the estimated camera position Pc[j] are obtained as the processing result by SfM.
[0037] Referring back to FIG. 5, after the execution of S3, in subsequent S5, the geospatial information processing apparatus 70 aligns the three-dimensional point cloud reconstruction space reconstructed in S3 with the real-world coordinate space (WGS84). At this time, the camera estimated position Pc[j] is also aligned with the real-world coordinate space. This alignment is performed, for example, by similarity transformation using Procrustes analysis. The details thereof will be described with reference to FIG. 7.
[0038] Note that FIG. 7 is a diagram showing an example of the alignment result by similarity transformation using Procrustes analysis. Specifically, the positioning position Pg[k] (k; positioning order) represented by the positioning position data is marked (plotted) on the OSM with a blue circle, and the camera estimated position Pc[j] after alignment is marked on the OSM with a red circle. Also, the geospatial shown in FIG. 7 is different from the geospatial shown in FIGS. 2, 3, and 6. Further, in FIG. 7, the positioning positions Pg[k] for one minute are shown, that is, a total of 21 positioning positions Pg[k] are shown. At the same time, in FIG. 7, the camera estimated positions Pc[j] for one minute (strictly speaking, slightly less than one minute) based on the captured image data obtained in parallel with the positioning positions Pg[k] are shown, specifically, a total of 420 camera estimated positions Pc[j] are shown. Also, in FIG. 7, the positioning position Pg[k] at the lower right end is the first positioning position Pg[1], and the positioning position Pg[k] at the upper left end is the last positioning position Pg[K] (=Pg
[21] ). That is, the movement of the vehicle (moving body) equipped with the drive recorder from the first positioning position Pg[1] to the last positioning position Pg[K] is represented by the trajectory of the positioning positions Pg[k]. And the camera estimated position Pc[j] at the lower right end in FIG. 7 is the first camera estimated position Pc[1], and the camera estimated position Pc[j] at the upper left end in FIG. 7 is the last camera estimated position Pc[J] (=Pc
[0420] ).
[0039] This similarity transformation using Procrustes analysis is performed using each positioning position Pg[k] as an index. For this purpose, the corresponding camera estimated position Pcg[k] corresponding to each positioning position Pg[k] is specified from among all the camera estimated positions Pc[j].
[0040] Specifically, it is considered that the first positioning position Pg[1] (the first positioning position) and the first camera estimated position Pc[1] (the first estimated position) correspond to each other, and the last positioning position Pg[K] (the second positioning position) and the last camera estimated position Pc[J] (the second estimated position) correspond to each other. That is, the first camera estimated position Pc[1] is specified as the corresponding camera estimated position Pcg[1] corresponding to the first positioning position Pg[1], and the last camera estimated position Pc[J] is specified as the corresponding camera estimated position Pcg[K] (= Pcg
[21] ) corresponding to the last positioning position Pg[K].
[0041] Then, for the corresponding camera estimated position Pcg[k] (the third estimated position) corresponding to each positioning position Pg[k] (the third positioning position) other than the first positioning position Pg[1] and the last positioning position Pg[k], it is specified based on the idea that the degree of movement progress of the positioning position Pg[k] and the degree of movement progress of the camera estimated position Pc[j] are the same as each other. Specifically, based on the idea that the spatial relationship of an arbitrary positioning position Pg[k] with respect to each of the first positioning position Pg[1] and the last positioning position Pg[K] and the spatial relationship of the corresponding camera estimated position Pcg[k] corresponding to the arbitrary positioning position Pg[k] with respect to each of the first camera estimated position Pc[1] and the last camera estimated position Pc[J] are common to each other, the corresponding camera estimated position Pcg[k] corresponding to the arbitrary positioning position Pg[k] is specified.
[0042] Therefore, based on the following formula 1, the path length Lcg[k] from the first camera estimated position Pc[1] to the corresponding camera estimated position Pcg[k] corresponding to an arbitrary positioning position Pg[k] is obtained. Here, Lg[k] is the path length from the first positioning position Pg[1] to the arbitrary positioning position Pg[k]. And Lg_ALL is the total path length from the first positioning position Pg[1] to the last positioning position Pg[K]. Also, Lc_ALL is the total path length from the first camera estimated position Pc[1] to the last camera estimated position Pc[J].
[0043] Formula 1 Lcg[k]=(Lg[k] / Lg_ALL)·Lc_ALL
[0044] Then, the camera estimated position Pc[j] corresponding to the position Pt after passing through the path length Lcg[k] based on Formula 1 from the first camera estimated position Pc[1] is specified as the corresponding camera estimated position Pcg[k] corresponding to any positioning position Pg[k]. If there is no camera estimated position Pc[j] corresponding to the position Pt, the camera estimated position Pc[j] closest to the position Pt is specified as the corresponding camera estimated position Pcg[k] corresponding to any positioning position Pg[k].
[0045] Then, a similarity transformation using Procrustes analysis is performed so that each corresponding camera estimated position Pcg[k] matches the positioning position Pg[k] corresponding to the corresponding camera estimated position Pcg[k]. As a result, the three-dimensional point cloud reconstruction space is translated, rotated, and scaled, together with the camera estimated position Pc[j], to be aligned with the real-world coordinate space.
[0046] Note that preprocessing may be performed to correct the variations of each of the positioning position Pg[k] and the camera estimated position Pc[j] during the alignment by the similarity transformation using Procrustes analysis. This preprocessing can be performed, for example, by a known Kalman filter. By performing such preprocessing, the alignment accuracy can be improved.
[0047] Here, each corresponding camera estimated position Pcg[k] after the alignment by the similarity transformation using Procrustes analysis, that is, each corresponding camera estimated position Pcg[k] in the real-world coordinate space, is defined by latitude and longitude, similar to the positioning position Pg[k] corresponding to each corresponding camera estimated position Pcg[k]. In the real-world coordinate space, the vertical (plumb) axis is the Y-axis, and the two horizontal axes are the X-axis and the Y-axis. That is, the coordinate values in the XZ plane of the real-world coordinate space are defined by latitude and longitude.
[0048] On the one hand, the coordinate value in the Y-axis direction of the real-world coordinate space is defined by, for example, elevation. However, for the Y value representing the height position of each 3D point 102 that makes up the aligned 3D point cloud 100, especially for the mutual distance between two 3D points 102 and 102 in the Y-axis direction, without some criteria, this cannot be quantitatively determined (how many meters it is). This is because the positioning data by the positioning satellite receiver does not contain information regarding the height direction.
[0049] Therefore, the geospatial information processing device 70 derives a scale factor Sf as a criterion based on the mutual ratio between the distance between each pair of aligned corresponding camera estimated positions Pcg[k] and the geodesic distance between each pair of corresponding camera estimated positions Pcg[k].
[0050] Specifically, the geospatial information processing device 70 obtains the shortest distance ΔLcg[k] between two consecutive aligned corresponding camera estimated positions Pcg[k] and Pcg[k + 1]. At the same time, the geospatial information processing device 70 obtains the geodesic distance ΔLa[k] between these two corresponding camera estimated positions Pcg[k] and Pcg[k + 1] in the real-world coordinate space. Then, the geospatial information processing device 70 obtains the ratio R[k] (=ΔLcg[k] / ΔLa[k]) of the shortest distance ΔLcg[k] to the geodesic distance ΔLa[k]. And the geospatial information processing device 70 takes the average of the ratios R[k] in all intervals as the scale factor Sf, that is, based on the following formula 2, the scale factor Sf is derived.
[0051] 《Formula 2》 Sf=(ΣR[k]) / (K - 1) where k = 1~K - 1
[0052] In the Y-axis direction of the real-world coordinate space, especially for the mutual distance between two 3D points 102 and 102, by multiplying this scale factor Sf, the mutual distance can be quantitatively determined.
[0053] As described above, FIG. 7 is a diagram showing an example of the alignment result by similarity transformation using Procrustes analysis. Specifically, the positioning position Pg[k] is marked on the OSM with a blue circle, and the camera estimated position Pc[j] after alignment is marked on the OSM with a red circle. In FIG. 7, the corresponding camera estimated position Pcg[k] corresponding to each positioning position Pg[k] is marked with a green circle.
[0054] Here, it is ideal that the corresponding camera estimated position Pcg[k] (green circle in FIG. 7) after alignment and the positioning position Pg[k] (blue circle in FIG. 7) coincide with each other. However, in the example shown in FIG. 7, the corresponding camera estimated position Pcg[k] after alignment and the positioning position Pg[k] do not locally coincide, that is, there is some deviation between the two. This deviation is caused by factors such as the camera path and scene characteristics. Also, this deviation is considered to occur in the entire 3D point cloud reconstruction space after alignment.
[0055] Therefore, referring back to FIG. 5, in S7 following S5, the geospatial information processing apparatus 70 corrects the entire 3D point cloud reconstruction space based on the positioning position Pg[k].
[0056] Specifically, the geospatial information processing apparatus 70 calculates an error vector (correction offset) Vg[k] of the corresponding camera estimated position Pcg[k] after alignment with respect to the positioning position Pg[k]. This error vector Vg[k] includes an error vector Vg_LAT[k] for latitude and an error vector Vg_LON[k] for longitude.
[0057] Then, the geospatial information processing device 70 generates a correction space (first correction space) as shown in FIG. 8 by radial basis function (RBF) interpolation, using the aligned corresponding camera estimated position Pcg[k] as control points and the error vector Vg[k] as interpolation values. In FIG. 8, the red circles represent the aligned corresponding camera estimated positions Pcg[k], which, in other words, represent the corresponding camera estimated positions Pcg[k] before correction by S7. The green circles in FIG. 8 represent the corresponding camera estimated positions Pcg’[k] after correction by S7, which, in other words, represent the target corresponding camera estimated positions Pcg’[k] for the correction by S7. Also, the horizontal axis of the correction space shown in FIG. 8 represents longitude, the vertical axis of the correction space represents latitude, and the color of the correction space represents the degree (strength) of correction.
[0058] Then, the geospatial information processing device 70 corrects the entire 3D point cloud reconstruction space by applying the correction space shown in FIG. 8 to the 3D point cloud reconstruction space after alignment. As a result, each of the corresponding camera estimated positions Pcg[k] before correction is corrected to the target corresponding camera estimated position Pcg’[k], and accordingly, the entire 3D point cloud reconstruction space is corrected.
[0059] FIG. 9 shows an example of the correction result by S7, that is, an example of the correction result based on the positioning position Pg[k]. FIG. 9 extracts and shows a part of the road on OSM in FIG. 7 (hereinafter referred to as “OSM road”), the camera estimated positions Pc[j] (red circles) and the positioning positions Pg[k] (blue circles) marked on the OSM, and also marks the corrected camera estimated positions Pc’[j] (orange circles) based on the positioning position Pg[k] (that is, by the correction space shown in FIG. 8).
[0060] As shown in FIG. 9, by performing the correction based on the positioning position Pg[k], the corrected camera estimated position Pc’[j] comes to follow the trajectory of the positioning position Pg[k]. Also, although not visible from FIG. 9, each of the corresponding camera estimated positions Pcg’[k] comes to coincide with the positioning position Pg[k]. Along with this, the entire 3D point cloud reconstruction space is corrected.
[0061] By the way, regarding the positioning position Pg[k], errors may occur depending on the situation. For example, in FIG. 9, it is recognized that an error has occurred in the positioning position Pg[k] within the region surrounded by the dashed rectangular frame 110 (the positioning position Pg[k] is slightly closer to the center of the OSM road). In such a case, it is considered appropriate to perform further correction on the entire three-dimensional point cloud reconstruction space including the corrected camera estimation position Pc’[j] based on the positioning position Pg[k], specifically, to perform additional correction based on the OSM road.
[0062] Therefore, referring back to FIG. 5, in S9 following S7, the geospatial information processing device 70 performs additional correction on the entire three-dimensional point cloud reconstruction space based on the OSM road.
[0063] For this purpose, the geospatial information processing device 70 acquires information regarding the OSM road necessary for the additional correction, so-called OSM road information, from the map database 50. The acquisition of the OSM road information from the map database 50 is performed using the Overpass_API.
[0064] The OSM road information regarding the OSM road necessary for the additional correction mentioned here refers to the OSM road information regarding the OSM road including any one of the camera estimation positions Pc’[j] that is the target of the additional correction, or the OSM road information regarding the OSM road in the vicinity of the camera estimation position Pc’[j], specifically, within a predetermined search radius Rs centered on the camera estimation position Pc’[j]. Note that the search radius Rc is, for example, 10 m, but is not limited to this.
[0065] Then, the geospatial information processing device 70 identifies an OSM road that matches each camera estimated position Pc’[j] from the OSM roads represented by the OSM road information acquired from the map database 50. Specifically, it identifies the OSM_ID of the OSM road, which is so-called road matching. For example, when any camera estimated position Pc’[j] is included in any OSM road, it is determined that the camera estimated position Pc’[j] matches the OSM road that includes the camera estimated position Pc’[j]. On the other hand, for a camera estimated position Pc’[j] that is not included in any OSM road, road matching is performed based on the relationship between the shortest distance DS from the camera estimated position Pc’[j] and the moving direction DRc of the camera estimated position Pc’[j].
[0066] Specifically, as shown in FIG. 10, the OSM road is defined by two nodes 120 and 120 and a linear edge 122 connecting these two nodes 120 and 120. When any camera estimated position Pc’[j] exists on any OSM road (on the edge 122), it is determined that the OSM road matches the camera estimated position Pc’[j].
[0067] On the other hand, for a camera estimated position Pc’[j] that does not exist on any OSM road, an OSM road that matches the camera estimated position Pc’[j] is identified based on a distance score Sds corresponding to the shortest distance DS from the camera estimated position Pc’[j] to each OSM road (represented by the OSM road information acquired from the map database 50) and a direction score Sdr corresponding to the relationship between the moving direction DRc of the camera estimated position Pc’[j] and the direction DRs of each OSM road.
[0068] The shortest distance DS mentioned here is, as shown in Fig. 11, the distance (orthogonal distance) from the camera estimated position Pc’[j] to the OSM road in the direction perpendicular to the trajectory of the camera estimated position Pc’[j]. And the distance score Sds is calculated based on the following Equation 3. That is, the distance score Sds is a value obtained by normalizing the shortest distance DS by the aforementioned search radius Rs.
[0069] 《Equation 3》 Sds = DS / Rs
[0070] Also, the direction score Sdr is calculated based on the following Equation 4. Θdr in this Equation 4 is the angle formed between the moving direction DRc of the camera estimated position Pc’[j] and the direction DRs of each OSM road, which is so-called the error angle. That is, the direction score Sdr is a value obtained by normalizing the error angle Θdr by a predetermined angle of 90°.
[0071] 《Equation 4》 Sdr = Θdr / 90°
[0072] Note that for the direction score Sdr, it is also conceivable to use a value obtained by normalizing the error angle Θdr by an angle of 180°. However, in order to suppress misjudgment in road matching, it is more appropriate to normalize the error angle Θdr by an angle of 90°, in other words, to handle the angle of 90° as the maximum error. Also, for the direction score Sdr, whether the OSM road is a one-way road or a two-way road is taken into account.
[0073] That is, for example, when the OSM road is a one-way road, the OSM road information is tagged with “oneway”. For the OSM road represented by the OSM road information tagged with “oneway”, that is, for the one-way OSM road, the direction score Sdr is calculated based on the aforementioned Equation 4.
[0074] On the other hand, for OSM roads represented by OSM road information without the tag "oneway", that is, for two-way OSM roads, the direction score Sdr is calculated based on Equation 4 and also calculated based on the following Equation 5.
[0075] 《Equation 5》 Sdr=(180°-θdr) / 90°
[0076] Then, the smaller value of the direction score Sdr based on Equation 4 and the direction score Sdr based on Equation 5 is used for road matching.
[0077] After calculating the distance score Sds and the direction score Sdr in this way, the geospatial information processing device 70 identifies the OSM road with the smallest total score Sa, which is the sum of the distance score Sds and the direction score Sdr, as the one that matches the camera estimated position Pc’[j].
[0078] To explain with a specific example, assume that the moving direction of a certain camera estimated position Pc’[j] is north, and this certain camera estimated position Pc’[j] is 9 m away from the east side of an actual (to-be-matched) OSM road named A. Here, assume that the OSM road A is a one-way road in the north direction. And there is another OSM road named B different from the OSM road A, and this OSM road B is a one-way road in the southeast direction and is 1 m away from the west side of the certain camera estimated position mentioned here. In this case, the distance score Sds for the OSM road A is 0.9 (=9 m / 10 m), the direction score Sdr is 0.0 (=0° / 90°), and the total score Sa is 0.9 (=0.9 + 0.0). On the other hand, the distance score Sds for the OSM road B is 0.1 (=1 m / 10 m), the direction score Sdr is 1.5 (=135° / 90°), and the total score Sa is 1.6 (=0.1 + 1.5). Therefore, the certain camera estimated position Pc’[j] mentioned here is identified as the one that matches the OSM road A with the smaller total score Sa.
[0079] After thus identifying the OSM road that matches each camera estimated position Pc’[j], the geospatial information processing device 70 performs a process for increasing the resolution (smoothing) of the shape of the OSM road.
[0080] That is, as described with reference to FIG. 10, the OSM road is defined by two nodes 120 and 120 and a linear edge 122 connecting these two nodes 120 and 120. Therefore, particularly in a curved section, as shown in FIG. 12, the resolution of the OSM road is low, and thus the shape (road spline) of the OSM road tends to deviate from the shape of the actual road.
[0081] Therefore, the geospatial information processing device 70 performs a process for increasing the resolution of the shape of the OSM road. This process is performed by known cubic Hermite spline interpolation. As a result, the shape of the OSM road is increased in resolution as shown in FIG. 13.
[0082] Then, as shown in FIG. 14, the geospatial information processing device 70 identifies a position Ps[j] corresponding to each camera estimated position Pc’[j] on the OSM road with increased resolution. Note that the position closest to each camera estimated position Pc’[j] on the OSM road with increased resolution is identified as the corresponding position Ps[j].
[0083] And the geospatial information processing device 70 calculates an error vector (correction offset) Vs[j] of each camera estimated position Pc’[j] with respect to the corresponding position Ps[j]. This error vector Vs[j] includes an error vector Vs_LAT[j] with respect to latitude and an error vector Vs_LON[j] with respect to longitude.
[0084] Furthermore, the geospatial information processing device 70 generates a correction space (second correction space) as shown in FIG. 15 by radial basis function interpolation, using each camera estimated position Pc’[j] as a control point and the error vector Vs[j] as an interpolation value. In FIG. 15, the red circles represent the camera estimated positions Pc’[j] before the additional correction by S9, or in other words, the camera estimated positions Pc’[j] after the correction by S7. And the green circles in FIG. 15 represent the camera estimated positions Pc”[j] after the additional correction by S9, or in other words, the target camera estimated positions Pc”[j] of the additional correction by S9. Also, the horizontal axis of the correction space shown in FIG. 15 represents longitude, the vertical axis of the correction space represents latitude, and the color of the correction space represents the degree of additional correction.
[0085] The geospatial information processing device 70 applies the correction space shown in FIG. 15 to the three-dimensional point cloud reconstruction space after the correction by S7, thereby performing additional correction on the entire three-dimensional point cloud reconstruction space. As a result, each camera estimated position Pc’[j] before the additional correction is corrected to the target camera estimated position Pc”[j], and accordingly, the entire three-dimensional point cloud reconstruction space is corrected.
[0086] FIG. 16 shows an example of the additional correction result by S9, that is, an example of the additional correction result based on the OSM road. In this FIG. 16, similar to FIGS. 7 and 9, the positioning position Pg[k] is marked with a blue circle, the corrected camera estimated position Pc’[j] is marked with an orange circle, and further, the camera estimated position Pc”[j] after the additional correction is marked with a pink circle.
[0087] As shown in this FIG. 16, by performing the additional correction based on the OSM road, the camera estimated position Pc”[j] after the additional correction comes to accurately follow the OSM road. In particular, the camera estimated position Pc”[j] within the region surrounded by the dashed rectangular frame 110 also comes to accurately follow the OSM road. And although not visible from FIG. 16, the entire three-dimensional point cloud reconstruction space is additionally corrected.
[0088] Referring again to FIG. 5, in S11 following S9, the geospatial information processing apparatus 70 performs processing for estimating the center position Pb and height Hb of the building, that is, for estimating the attributes of the building so to speak.
[0089] In the processing of S11, the captured image data after the lens distortion of the camera has been corrected by the aforementioned SfM is used. For example, FIG. 17 shows an example of the captured image data after lens distortion correction by SfM for the captured image data of the same geospatial area shown in FIG. 2, specifically, of a certain frame.
[0090] The geospatial information processing apparatus 70 performs building detection using an object detection model on the captured image data after lens distortion correction by SfM, specifically, for each frame. As the object detection model mentioned here, for example, the known GroundingDINO is used. For this reason, GroundingDINO is set up in the aforementioned geospatial information processing software 40ab.
[0091] By performing building detection using this GroundingDINO, as shown in FIG. 18, a bounding box 130 is attached to each detected building. Note that the geospatial area shown in FIG. 18 is different from the geospatial area shown in FIG. 17. Also, in the frame shown in FIG. 18, a two-dimensional coordinate system is set with the upper left corner as the origin, the horizontal axis as the X-axis, and the vertical axis as the Y-axis. This is the same for the frames shown in other drawings such as FIG. 17.
[0092] Then, the geospatial information processing apparatus 70 performs filtering to select a bounding box 130 that is appropriate for estimating the building attributes from among the bounding boxes 130, that is, to exclude those that are inappropriate for estimating the building attributes. In this filtering, the bounding box 130 that meets any of the following four exclusion conditions is excluded from the target of subsequent processing.
[0093] The first exclusion condition is that the bounding box 130 is excessively large. Whether the bounding box 130 is excessively large is determined based on, for example, whether the area of the bounding box 130 is larger than a predetermined first threshold area. The first threshold area is appropriately set based on experiments, experience, etc.
[0094] The second exclusion condition is that the bounding box 130 is excessively small. Whether the bounding box 130 is excessively small is determined based on, for example, whether the area of the bounding box 130 is larger than a predetermined second threshold area. The second threshold area is smaller than the first threshold area, and this second threshold area is also appropriately set based on experiments, experience, etc., similar to the first threshold area.
[0095] The third exclusion condition is that at least one of the upper, lower, left, and right ends of the bounding box 130 is outside the frame.
[0096] And the fourth exclusion condition is that the upper end of the bounding box 130 is too close to the upper end of the frame. Whether the upper end of the bounding box 130 is too close to the upper end of the frame is determined based on whether the Y value of the upper end of the bounding box 130 is smaller than a predetermined threshold. The predetermined threshold mentioned here is also appropriately set based on experiments, experience, etc.
[0097] After such filtering selects the bounding box 130 to be used in subsequent processing, the geospatial information processing apparatus 70 performs segmentation for dividing the building area by a segmentation model for each frame. As the segmentation model mentioned here, for example, a known SAM2 is used. For this reason, SAM2 is set up in the aforementioned geospatial information processing software 40ab. Further, the bounding box 110 selected by the aforementioned filtering is input as it is as a prompt to SAM2. As a result, the building area is divided for each building with each bounding box 130 attached.
[0098] FIG. 19 is an example of a frame after segmentation by SAM2. As shown in this FIG. 19, for each building with the aforementioned bounding box 130 attached, the building area is divided by a closed-loop dividing line 140. Further, different colors are attached to each dividing line 140.
[0099] Further, FIG. 20 is a diagram showing an enlarged part of the frame shown in FIG. 19. As shown in this FIG. 20, each building area includes the two-dimensional feature points 142 detected from the frame during the reconstruction of the three-dimensional point cloud by the aforementioned SfM. The same color as the dividing line 140 that divides the building area including the two-dimensional feature point 142 is attached to each two-dimensional feature point 142. In the reconstruction of the three-dimensional point cloud 100 by the aforementioned SfM, the three-dimensional point cloud 100 is reconstructed based on the coordinate values of the two-dimensional feature points 142 common among the frames and the aforementioned camera estimated position Pc[j], etc. Therefore (basically), each two-dimensional feature point 142 is associated with the three-dimensional point 102 corresponding to the two-dimensional feature point 142.
[0100] Furthermore, for each frame after segmentation by SAM2, the geospatial information processing device 70 extracts only those of the two-dimensional feature points 142 within each building area that correspond to the three-dimensional points 102 of the three-dimensional point cloud 100 reconstructed by the aforementioned SfM (i.e., those associated with the three-dimensional points 102 during the reconstruction of the three-dimensional point cloud 100). This is also a kind of filtering. That is, among the two-dimensional feature points 142, there may be those that do not correspond to the three-dimensional points 102 (i.e., those not associated with the three-dimensional points 102 during the reconstruction of the three-dimensional point cloud 100), and such filtering is performed to exclude such two-dimensional feature points 142 from the targets of subsequent processing.
[0101] After performing the filtering to extract only the two-dimensional feature points 142 corresponding to this three-dimensional point cloud 100, the geospatial information processing device 70 determines, for each frame and for each individual building area, whether the total number of two-dimensional feature points 142 within the building area is less than a predetermined threshold number. And for the building areas where the total number of two-dimensional feature points 142 is less than the predetermined threshold number, in other words, for the two-dimensional feature points 142 within the building area, they are excluded from the targets of subsequent processing. This is also a kind of filtering. That is, when the total number of two-dimensional feature points 142 within a building area is extremely small, such two-dimensional feature points 142 are likely to show two-dimensional elements such as lines and planes rather than a three-dimensional structure and do not contribute to accurate processing, so such filtering is performed. Note that the predetermined threshold number mentioned here is, for example, 4, but it is not limited to this.
[0102] Then, for each frame and for each individual building area, the geospatial information processing device 70 identifies (detects) the three-dimensional point 102 corresponding to the uppermost (i.e., the one with the minimum Y value) two-dimensional feature point 144 among the two-dimensional feature points 142 within the building area as the vertex candidate PTa of the building. As shown in FIG. 20, the uppermost two-dimensional feature point 144 among the two-dimensional feature points 142 within each building area is colored uniformly, for example, colored yellow-green.
[0103] Then, when the same vertex candidate PTa is identified from a plurality of frames, strictly speaking, from a predetermined number or more of frames, the geospatial information processing device 70 determines that the vertex candidate PTa is stably identified, so to speak, determines that the vertex candidate PTa has a high degree of certainty, and sets the vertex candidate PTa as the grouping reference point PTg. Then, the geospatial information processing device 70 regards the three-dimensional points 102 obtained from the building area common to the grouping reference point PTg as representing the same building, and groups (integrates) them. Note that the predetermined number mentioned here is, for example, 3, but is not limited to this.
[0104] Furthermore, for the grouped three-dimensional points 102 (point cloud), the geospatial information processing device 70 obtains the average value and standard deviation in the XZ plane of the three-dimensional point cloud reconstruction space (real-world coordinate space), and obtains the z-score (=(data value - average value) / standard deviation). Then, for the three-dimensional points 102 whose z-score in the XZ plane exceeds a predetermined value, the geospatial information processing device 70 excludes them from the target of subsequent processing as outliers. This method is a standard filtering method (z-score filtering method) using SciPy, which is one of the known Python (registered trademark) libraries. Note that the predetermined value mentioned here is, for example, ±2, that is, a value corresponding to a confidence range of about 97%, but is not limited to this.
[0105] For the three-dimensional points 102 after excluding outliers in this way, further, the geospatial information processing device 70 obtains the median in the XZ plane and sets this median as the filter reference point PTf.
[0106] In addition, the geospatial information processing device 70 calculates the standard deviation σ in the XZ plane with respect to the filter reference point PTf for the three-dimensional points 102 after excluding outliers.
[0107] On top of that, the geospatial information processing device 70 sets a cylindrical filter region 150 as shown in FIG. 21 in the three-dimensional point cloud reconstruction space. Specifically, the geospatial information processing device 70 sets a point that is conjugate (has the same coordinate value) to the filter reference point PTf in the XZ plane and conjugate (has the same coordinate value) to the grouping reference point PTg in the Y-axis direction, with the center O, a standard deviation σ as the radius, and a filter region 150 having a predetermined dimension L in the Y-axis direction. Note that the predetermined dimension L mentioned here is, for example, ±1.5 m with respect to the center O, but is not limited to this.
[0108] Then, the geospatial information processing device 70 identifies the three-dimensional point 102 at the uppermost position (with the maximum Y value) within the filter region 150 as the building vertex Ptb.
[0109] Furthermore, the geospatial information processing device 70 specifically identifies the camera estimated position Pc”[j] closest to the building for which the building vertex Ptb has been identified as the shortest camera estimated position Pa, which is strictly the camera estimated position Pc”[j] closest to the building among those obtained when the frame that contributed to the identification of the building vertex Ptb was obtained.
[0110] Then, the geospatial information processing device 70 virtually sets the horizontal plane (XZ plane) including the shortest camera estimated position Pa as the ground plane, and multiplies the mutual difference between the Y value of the shortest camera estimated position Pa and the Y value of the building vertex Ptb by the aforementioned scale factor Sf to estimate the height Hb of the building having the building vertex Ptb.
[0111] For buildings with an estimated height Hb that is excessively low, specifically, for buildings where the estimated height Hb is less than a predetermined height threshold, the geospatial information processing device 70 determines that it is actually not a building but some object other than a building. Then, the geospatial information processing device 70 excludes such objects other than buildings from the targets of subsequent processing. This is also a kind of filtering. The predetermined height threshold mentioned here is, for example, 3 m, but is not limited to this.
[0112] Then, for a building whose height Hb is equal to or greater than a predetermined height threshold, the geospatial information processing apparatus 70 estimates its central position Pb. Specifically, the geospatial information processing apparatus 70 estimates the position corresponding to the median value in the XZ plane of the three-dimensional points within the aforementioned filter region 150 as the central position Pb of the building.
[0113] In addition, in the identification of the building vertex Ptb and the estimation of the central position Pb, the median value (median absolute deviation) is used as described above because, as a feature of the three-dimensional point cloud 100 reconstructed by SfM, there is a characteristic that 102 three-dimensional points tend to concentrate at the corners of the building and the vicinity of the center of the building becomes a cavity. That is, if, for example, the average value is used instead of the median value, the estimated building vertex Ptb and the central position Pb may deviate significantly from their actual positions. Therefore, the median value is used to reduce such inconveniences as much as possible.
[0114] Referring back to FIG. 5, after the geospatial information processing apparatus 70 estimates the attributes of the building, namely the central position Pb and the height Hb of the building, in S11, in the subsequent S13, the geospatial information processing apparatus 70 performs building matching.
[0115] Specifically, for example, as shown in FIG. 22, when the estimated central position Pb of the building is included in any OSM building (polygon), that is, when such a first condition is satisfied, the geospatial information processing apparatus 70 determines that the OSM building matches the building having the central position Pb. Then, the geospatial information processing apparatus 70 identifies the OSM_ID of the OSM building that matches the building.
[0116] On the other hand, when the presumed center position Pb of the building is not included in any of the OSM buildings, that is, when the first condition is not satisfied, as shown in FIG. 23, the geospatial information processing device 70 emits a ray 160 from the shortest camera estimated position Pa toward the presumed center position Pb of the building and performs so-called ray casting. Further, the geospatial information processing device 70 extends the ray 160 as shown by the dashed line 160a in FIG. 23. Then, when an OSM building exists on the extension line 160a of the ray 160, in other words, when an OSM building exists behind the presumed center position Pb of the building, that is, when such a second condition is satisfied, the geospatial information processing device 70 determines that the OSM building closest to the center position Pb among them matches the building having the center position Pb. Note that the ray 160 is a virtual straight line used to calculate whether it hits any of the OSM buildings, including its extended portion 160a, and does not need to be actually displayed.
[0117] On the contrary, when no OSM building exists on the extension line 160a of the ray 160, that is, when the second condition is not satisfied, as shown in FIG. 24, the geospatial information processing device 70 now emits a ray 162 from the presumed center position Pb of the building toward the shortest camera estimated position Pa, that is, performs ray casting in the opposite direction to that shown in FIG. 23. Then, when an OSM building exists on the ray 162, that is, when an OSM building exists between the shortest camera estimated position Pa and the presumed center position Pb of the building, that is, when such a third condition is satisfied, the geospatial information processing device 70 determines that the OSM building closest to the center position Pb among them matches the building having the center position Pb. Note that due to the lens distortion correction of the camera by the aforementioned SfM, the presumed center position Pb of the building tends to be closer to the shortest camera estimated position Pa than the actual position. Therefore, the determination of whether the second condition is satisfied is performed before the determination of whether the third condition is satisfied. In other words, the second condition has a higher priority than the third condition.
[0118] And when the third condition is not satisfied, that is, when there is no OSM building between the shortest camera estimated position Pa and the estimated center position Pb of the building, the geospatial information processing apparatus 70 determines whether there is any OSM building within a predetermined range centered on the estimated center position Pb of the building, as shown in FIG. 25. And when there is any OSM building within the predetermined range mentioned here, that is, when the fourth condition is satisfied, the geospatial information processing apparatus 70 determines that the OSM building closest to the center position Pb among them matches the building having the center position Pb. Here, the predetermined range mentioned here is, for example, a circular area with a predetermined radius centered on the estimated center position Pb of the building, and the predetermined radius mentioned here is, for example, 5 m, but is not limited thereto.
[0119] Furthermore, when the fourth condition is not satisfied, that is, when there is no OSM building within the predetermined range centered on the estimated center position Pb of the building, the geospatial information processing apparatus 70 determines that the building is not yet registered in OSM, that is, it is unregistered. Such an unregistered building is stored as a registration candidate in OSM. The storage destination is, for example, the auxiliary storage unit 80, but is not limited thereto.
[0120] With this, the geospatial information processing apparatus 70 ends the series of processes shown in FIG. 5.
[0121] Note that FIG. 26 is a flowchart showing details of S5 in FIG. 5. As shown in FIG. 26, the geospatial information processing apparatus 70 first specifies, in S101, the corresponding camera estimated position Pcg[k] corresponding to each positioning position Pg[k] among the camera estimated positions Pc[j]. Then, the geospatial information processing apparatus 70 performs a similarity transformation using Procrustes analysis in subsequent S103. Thereby, the three-dimensional point cloud reconstruction space is aligned with the real-world coordinate space. Furthermore, the geospatial information processing apparatus 70 derives a scale factor Sf in subsequent S105. With this, the geospatial information processing apparatus 70 ends S5 in FIG. 5.
[0122] Also, FIG. 27 is a flowchart showing the details of S7 in FIG. 5. As shown in this FIG. 27, the geospatial information processing apparatus 70 first calculates, in S201, an error vector Vg[k] with respect to the measured position Pg[k] of the corresponding camera estimated position Pcg[k] after alignment by similarity transformation using Procrustes analysis. Then, in subsequent S203, the geospatial information processing apparatus 70 generates a correction space as shown in FIG. 8 by radial basis function interpolation, with the corresponding camera estimated position Pcg[k] after alignment as the control point and the error vector Vg[k] as the interpolation value. Further, in subsequent S205, the geospatial information processing apparatus 70 corrects the entire 3D point cloud reconstruction space by applying the correction space generated in S203 to the 3D point cloud reconstruction space after alignment. With this, the geospatial information processing apparatus 70 ends S7 in FIG. 5.
[0123] And FIG. 28 is a flowchart showing the details of S9 in FIG. 5. As shown in this FIG. 28, the geospatial information processing apparatus 70 first acquires, in S301, OSM road information regarding OSM roads necessary for additional correction from the map database 50. Then, in subsequent S303, the geospatial information processing apparatus 70 identifies the OSM roads that match each camera estimated position Pc’[j], that is, identifies the OSM_ID of the OSM roads. Further, in subsequent S305, the geospatial information processing apparatus 70 increases the resolution of the shape of the OSM roads identified in S303 by known cubic Hermite spline interpolation. Then, in subsequent S307, the geospatial information processing apparatus 70 identifies the position Ps corresponding to each camera estimated position Pc’[j] in the OSM roads whose resolution has been increased in S305.
[0124] Furthermore, in S309 following S307, the geospatial information processing apparatus 70 calculates an error vector Vs[j] of each camera estimated position Pc’[j] with respect to the corresponding position Ps[j] specified in S307. Then, in subsequent S311, the geospatial information processing apparatus 70 generates a correction space as shown in FIG. 15 by radial basis function interpolation, using each camera estimated position Pc’[j] as a control point and the error vector Vs[j] as an interpolation value. And then, in subsequent S313, the geospatial information processing apparatus 70 applies the correction space generated in S311 to the three-dimensional point cloud reconstruction space after correction by the aforementioned S7 (S205), thereby additionally correcting the entire three-dimensional point cloud reconstruction space. With this, the geospatial information processing apparatus 70 ends S9 in FIG. 5.
[0125] In addition, FIGS. 29 and 30 are flowcharts showing details of S11 in FIG. 5. First, as shown in FIG. 29, the geospatial information processing apparatus 70 performs building detection using an object detection model for each frame of the captured image after lens distortion by SfM. As the object detection model mentioned here, for example, a known GroundingDINO is used. Then, in subsequent S403, the geospatial information processing apparatus 70 performs filtering to exclude those that are inappropriate for estimating building attributes among the bounding boxes 130 that are the building detection results in S401. Furthermore, in subsequent S405, the geospatial information processing apparatus 70 performs segmentation for dividing the building area using a segmentation model for each frame of the captured image data. As the segmentation model mentioned here, for example, SAM2 is used. Also, the bounding box 110 selected by the filtering in S403 is input as it is to SAM2 as a prompt.
[0126] In S407 following S405, the geospatial information processing apparatus 70 performs filtering to extract only those among the two-dimensional feature points 142 within each individual building area that correspond to the three-dimensional points 102 of the three-dimensional point cloud 100, for each frame after segmentation in S405. Then, in subsequent S409, the geospatial information processing apparatus 70 determines, for each frame and for each individual building area, whether the total number of two-dimensional feature points 142 within the building area is less than a predetermined threshold number, and excludes from the target of subsequent processing the building areas where the total number of two-dimensional feature points 142 is less than the predetermined threshold number, that is, performs filtering for that purpose. Further, in subsequent S411, the geospatial information processing apparatus 70, for each frame and for each individual building area, identifies, as the vertex candidate PTa of the building, the three-dimensional point 102 corresponding to the uppermost (that is, the one with the minimum Y value) two-dimensional feature point 144 among the two-dimensional feature points 142 within the building area.
[0127] Then, in subsequent S413, the geospatial information processing apparatus 70 sets the vertex candidate PTa commonly identified by a predetermined number or more of frames as the grouping reference point PTg. Further, in subsequent S415, the geospatial information processing apparatus 70 groups the three-dimensional points 102 obtained from the building areas common to the grouping reference point PTg. In addition, in subsequent S417, for the three-dimensional points 102 grouped in S415, the geospatial information processing apparatus 70 excludes, as outliers, the three-dimensional points 102 whose z-score in the XZ plane of the three-dimensional point cloud reconstruction space (real-world coordinate space) exceeds a predetermined value from the target of subsequent processing. Then, in subsequent S419, for the three-dimensional points 102 after outliers are removed by the filtering in S417, the geospatial information processing apparatus 70 obtains the median in the XZ plane and sets this median as the filtering reference point PTf.
[0128] Referring to FIG. 30, in subsequent S421, for the three-dimensional points 102 after outliers are removed by the filtering in S419, the geospatial information processing apparatus 70 calculates the standard deviation σ in the XZ plane with respect to the filter reference point PTf. Then, in subsequent S423, the geospatial information processing apparatus 70 sets a cylindrical filter region 150. That is, the geospatial information processing apparatus 70 sets a filter region 150 having a center O at a point that is conjugate to the filter reference point PTf in the XZ plane and conjugate to the grouping reference point PTg in the Y-axis direction, with the standard deviation σ as the radius and a predetermined dimension L in the Y-axis direction.
[0129] Then, in subsequent S425, the geospatial information processing apparatus 70 identifies the three-dimensional point 102 at the uppermost position (with the maximum Y value) within the filter region 150 as the building vertex Ptb. Further, in subsequent S427, based on the difference between the Y value of the shortest camera estimation position Pa and the Y value of the building vertex Ptb, specifically by multiplying this difference by the scale factor Sf, the geospatial information processing apparatus 70 estimates the height Hb of the building having the building vertex Ptb.
[0130] In addition, in subsequent S429, for buildings where the height Hb estimated in S427 is less than a predetermined height threshold, the geospatial information processing apparatus 70 performs filtering to exclude them from the targets of subsequent processing. Then, for buildings where the height Hb is greater than or equal to the predetermined height threshold, the geospatial information processing apparatus 70 estimates the position corresponding to the median value in the XZ plane of the three-dimensional points within the aforementioned filter region 150 as the center position Pb of the building. With this, the geospatial information processing apparatus 70 ends S11 in FIG. 5.
[0131] Also, FIG. 31 is a flowchart showing the details of S13 in FIG. 5. As shown in this FIG. 31, the geospatial information processing apparatus 70 first determines in S501 whether the above-described first condition is satisfied, that is, determines whether there is an OSM building including the center position Pb of the estimated building. Here, when the first condition is satisfied, that is, when there is an OSM building including the center position Pb of the estimated building (S501: YES), the geospatial information processing apparatus 70 proceeds with the process to S503. On the other hand, when the first condition is not satisfied, that is, when there is no OSM building including the center position Pb of the estimated building (S501: NO), the geospatial information processing apparatus 70 proceeds with the process to S505 described later.
[0132] In S503, the geospatial information processing apparatus 70 specifies the corresponding OSM building, that is, the OSM building that satisfies the first condition, as the target OSM building, and specifically specifies its OSM_ID. With this, the geospatial information processing apparatus 70 ends S13 in FIG. 5.
[0133] On the other hand, when the process proceeds from S501 to S505, the geospatial information processing apparatus 70 determines in S505 whether the above-described second condition is satisfied, that is, determines whether there is an OSM building on the extension line 160a of the ray 160. Here, when the second condition is satisfied, that is, when there is an OSM building on the extension line 160a of the ray 160 (S505: YES), the geospatial information processing apparatus 70 proceeds with the process to S503. On the other hand, when the second condition is not satisfied, that is, when there is no OSM building on the extension line 160a of the ray 160 (S505: NO), the geospatial information processing apparatus 70 proceeds with the process to S507. Incidentally, when the process proceeds from S505 to S503, in the S503, the OSM building that satisfies the second condition and is closest to the center position Pb of the estimated building is specified as the target OSM building.
[0134] In S507, the geospatial information processing device 70 determines whether the above-described third condition is satisfied, that is, determines whether there is an OSM building on the ray 160. Here, when the third condition is satisfied, that is, when there is an OSM building on the ray 160 (S507: YES), the geospatial information processing device 70 proceeds with the process to S503. On the other hand, when the third condition is not satisfied, that is, when there is no OSM building on the ray 160 (S507: NO), the geospatial information processing device 70 proceeds with the process to S509. Note that when the process proceeds from S507 to S503, in the S503, the OSM building that satisfies the third condition and is closest to the estimated center position Pb of the building is specified as the target OSM building.
[0135] In S509, the geospatial information processing device 70 determines whether the above-described fourth condition is satisfied, that is, determines whether there is an OSM building within a predetermined range based on the estimated center position Pb of the building. Here, when the fourth condition is satisfied, that is, when there is an OSM building within a predetermined range based on the estimated center position Pb of the building (S509: YES), the geospatial information processing device 70 proceeds with the process to S503. On the other hand, when the fourth condition is not satisfied, that is, when there is no OSM building within a predetermined range based on the estimated center position Pb of the building (S509: NO), the geospatial information processing device 70 proceeds with the process to S519. Note that when the process proceeds from S509 to S503, in the S503, the OSM building that satisfies the fourth condition and is closest to the estimated center position Pb of the building is specified as the target OSM building.
[0136] In S511, the geospatial information processing device 70 stores the building having the estimated center position Pb as an unregistered building. With this, the geospatial information processing device 70 ends S13.
[0137] As described above, according to this embodiment, based on the recording data by the drive recorder, a three-dimensional model of the geospatial is reconstructed by the three-dimensional point cloud. Then, the reconstructed three-dimensional point cloud reconstruction space is aligned with the real-world coordinate space, using the positioning position Pg[k] by the positioning satellite receiver built in the drive recorder as an index, together with the camera estimation position Pc[j] representing the position of the drive recorder. Further, the corresponding camera estimation position Pcg[k] corresponding to the positioning position Pg[k] among the aligned positioning position Pg[k] and the aligned camera estimation position Pc[j] is made to coincide with each other, that is, based on the positioning position Pg[k], the entire three-dimensional point cloud reconstruction space is corrected. Thereby, a three-dimensional point cloud reconstruction space faithful to the world coordinate space is obtained. That is, according to this embodiment, by using a device widely popular in the market, such as a drive recorder, a three-dimensional point cloud reconstruction space faithful to the real-world coordinate space is obtained. Such a three-dimensional point cloud reconstruction space can be applied to the update of OSM or the creation of a new digital map other than OSM. This greatly contributes to realizing the creation and update of an accurate digital map while suppressing costs compared to the conventional method. And this effect is more remarkable as the area of the digital map is larger.
[0138] Also, according to this embodiment, based on the idea that the degree of movement progress of the positioning position Pg[k] and the degree of movement progress of the camera estimated position Pc[j] are the same as each other, the corresponding camera estimated position Pcg[k] corresponding to the positioning position Pg[k] is specified. Specifically, based on the idea that the spatial relationship of an arbitrary positioning position Pg[k] with respect to each of the first positioning position Pg[1] and the last positioning position Pg[K] and the spatial relationship of the corresponding camera estimated position Pcg[k] corresponding to the arbitrary positioning position Pg[k] with respect to each of the first camera estimated position Pc[1] and the last camera estimated position Pc[J] are common to each other, the corresponding camera estimated position Pcg[k] is specified. Then, when the three-dimensional point cloud reconstruction space is aligned with the real-world coordinate space, the alignment of the three-dimensional point cloud reconstruction space to the real-world coordinate space is performed using each positioning position Pg[k] as an index so that each positioning position Pg[k] and the corresponding camera estimated position Pcg[k] corresponding to the positioning position Pg[k] coincide with each other. Thereby, the alignment accuracy is improved. This also greatly contributes to realizing the creation and update of an accurate digital map while suppressing costs more than before.
[0139] In addition, according to this embodiment, further additional correction based on OSM roads is performed on the three-dimensional point cloud reconstruction space corrected based on the positioning position Pg[k]. As a result, a three-dimensional point cloud reconstruction space more faithful to the real-world coordinate space is obtained. This greatly contributes to realizing the creation and update of a more accurate digital map.
[0140] Note that in this embodiment, the geospatial information processing apparatus 70 that executes S1 in FIG. 5 is strictly speaking, the CPU 72a is an example of the image data input receiving means according to the present invention, and is also an example of the positioning data input receiving means according to the present invention. Further, the CPU 72a that executes S3 in FIG. 5 is an example of the three-dimensional reconstruction means according to the present invention, and the CPU 72a that executes S5 in FIG. 5 is an example of the spatial alignment means according to the present invention. Furthermore, the CPU 72a that executes S7 in FIG. 5 is an example of the first spatial correction means according to the present invention. In particular, the CPU 72a that executes S203 in FIG. 27 is an example of the first corrected space generation means according to the present invention, and the CPU 72a that executes S205 in FIG. 27 is an example of the first correction execution means according to the present invention. In addition, the CPU 72a that executes S9 in FIG. 5 is an example of the second spatial correction means according to the present invention, and in that case, the CPU 72a that executes S307 in FIG. 28 is an example of the corresponding position specifying means according to the present invention. And the CPU 72a that executes S311 in FIG. 28 is an example of the second corrected space generation means according to the present invention, and the CPU 72a that executes S311 in FIG. 28 is an example of the second correction execution means according to the present invention. Additionally, the CPU 72a that executes S305 in FIG. 28 is an example of the high-resolution means according to the present invention.
[0141] This embodiment is merely a specific example of the present invention and does not limit the scope of the present invention. The present invention can be applied in various aspects other than this embodiment.
[0142] For example, although the geospatial information processing apparatus 70 is configured by a PC, it is not limited thereto. The geospatial information processing apparatus 70 may be configured by a device other than a PC, especially by a dedicated device.
[0143] Also, although the recorded data by the drive recorder is to be obtained from the data management server 30, it is not limited thereto. For example, the recorded data by the drive recorder may be directly obtained from the drive recorder.
[0144] Furthermore, as the map database 50, a database of digital maps other than OSM may be adopted.
[0145] In addition, in S5 in FIG. 5, the alignment of the three-dimensional point cloud reconstruction space to the real-world coordinate space was performed by similarity transformation using Procrustes analysis, but it is not limited to this. For example, the translation and rotation of the three-dimensional point cloud reconstruction space are performed by known rigid body transformation, and the scaling of the three-dimensional point cloud reconstruction space is performed by known scaling transformation, so that the alignment of the three-dimensional point cloud reconstruction space to the real-world coordinate space may be performed. Alternatively, the alignment of the three-dimensional point cloud reconstruction space to the real-world coordinate space may be performed by known affine transformation.
[0146] The present invention is applicable not only to the creation and update of digital maps but also to fields other than such digital maps, such as the metaverse and games.
[0147] In addition, the present invention can be provided not only in the form of an apparatus called a geospatial information processing apparatus but also in the form of a method called a geospatial information processing method.
Explanation of Reference Numerals
[0148] 10 … Geospatial information processing system 30 … Data management server 50 … Map database 70 … Geospatial information processing apparatus 72 … Control unit 72a … CPU 72b … Main memory unit 80 … Auxiliary storage unit 80a … Geospatial information processing software 100 … Three-dimensional point cloud 102 … Three-dimensional point 12 … Control unit 12a … APP 12b … Main memory unit Pc[j] … Camera estimated position Pc’[j] … Corrected camera estimated position Pc”[j] … Estimated camera position after additional correction Pcg[k] … Corresponding estimated camera position Pcg’[k] … Corrected corresponding estimated camera position Pg[k] … Positioning position Ps[j] … Corresponding position
Claims
1. Image data input receiving means for receiving input of image data obtained by sequentially photographing the geospatial area around the camera at a first period by a camera moving on a road; Positioning data input receiving means for receiving input of positioning data obtained by sequentially positioning at a second period different from the first period by a positioning satellite receiver moving on the road together with the camera; 3D reconstruction means for estimating the position of the camera for each first period based on the image data and reconstructing a 3D model of the geospatial area in a state associated with the estimated position of the camera; Spatial alignment means for aligning the 3D reconstructed space reconstructed by the 3D reconstruction means with the real-world coordinate space using the positioning positions for each second period represented by the positioning data as indices together with the estimated position of the camera; A geospatial information processing apparatus comprising: first space correction means for correcting the 3D reconstructed space after alignment by the spatial alignment means together with the estimated position of the camera so that the positioning position for each second period and the corresponding camera estimated position corresponding to the positioning position among the estimated positions of the camera for each first period after alignment by the spatial alignment means coincide with each other.
2. The spatial alignment means includes a first positioning position which is a specific one of the positioning positions for each second period and a first estimated position which is a specific one of the estimated positions of the camera for each first period corresponding to each other, and a second positioning position which is a specific positioning position different from the first positioning position and a second estimated position which is a specific estimated position of the camera different from the first estimated position corresponding to each other. Further, a third positioning position which is at least one of the positioning positions other than the first positioning position and the second positioning position and a third estimated position which is at least one of the estimated positions of the camera other than the first estimated position and the second estimated position are regarded as corresponding to each other, and the 3D reconstructed space is aligned with the real-world coordinate space. The geospatial information processing apparatus according to claim 1, wherein the third estimated position is determined such that the spatial relationship of the third positioning position with respect to each of the first positioning position and the second positioning position and the spatial relationship of the third estimated position with respect to each of the first estimated position and the second estimated position are common to each other.
3. The first space correction means First correction space generation means for generating a first correction space based on the corresponding camera estimated position and the error of the corresponding camera estimated position with respect to the positioning position corresponding to the corresponding camera estimated position; First correction execution means for correcting the three-dimensional reconstruction space after alignment by the space alignment means with the first correction space; The geospatial information processing apparatus according to claim 1.
4. The geospatial information processing apparatus according to claim 3, wherein the first correction space generation means generates the first correction space by radial basis function interpolation, using the corresponding camera estimated position as a control point and the error as an interpolation value.
5. Corresponding position specifying means for specifying a corresponding position on a map road included in predetermined map information corresponding to the estimated position of the camera after correction by the first space correction means; The geospatial information processing apparatus according to claim 1, further comprising second space correction means for additionally correcting the three-dimensional reconstruction space after correction by the first space correction means so that the estimated position of the camera after correction by the first space correction means and the corresponding position on the map road corresponding to the estimated position of the camera coincide with each other.
6. The second space correction means includes: Second correction space generation means for generating a second correction space based on the estimated position of the camera after correction by the first space correction means and the error of the estimated position of the camera with respect to the corresponding position on the map road corresponding to the estimated position of the camera; Second correction execution means for additionally correcting the three-dimensional reconstruction space after correction by the first space correction means with the second correction space; The geospatial information processing apparatus according to claim 5.
7. The geospatial information processing apparatus according to claim 6, wherein the second correction space generation means generates the second correction space by radial basis function interpolation, using the estimated position of the camera after correction by the first space correction means as a control point and the error as an interpolation value.
8. Further comprising high-resolution means for increasing the resolution of the shape of the map road; The geospatial information processing apparatus according to claim 5, wherein the corresponding position specifying means specifies the corresponding position on a high-resolution road, which is the map road after being increased in resolution by the high-resolution means.
9. The geospatial information processing apparatus according to claim 8, wherein the high-resolution means increases the resolution of the shape of the map road by cubic Hermite spline interpolation.
10. A method for processing geospatial information by a computer, comprising: The processor of the computer: An image data input receiving step of receiving input of image data obtained by sequentially photographing the geospatial area around the camera by a camera moving on a road in a first cycle; A positioning data input receiving step of receiving input of positioning data obtained by sequentially performing positioning by a positioning satellite receiver moving on the road together with the camera in a second cycle different from the first cycle; A three-dimensional reconstruction step of estimating the position of the camera for each first cycle based on the image data and reconstructing a three-dimensional model of the geospatial area in a state associated with the estimated position of the camera; A space alignment step of aligning the three-dimensional reconstructed space reconstructed in the three-dimensional reconstruction step with the estimated position of the camera in the real-world coordinate space using the positioning position for each second cycle represented by the positioning data as an index; A space correction step of correcting the three-dimensional reconstructed space after alignment in the space alignment step together with the estimated position of the camera so that the positioning position for each second cycle and the corresponding camera estimated position corresponding to the positioning position among the estimated positions of the camera for each first cycle after alignment in the space alignment step match each other. A geospatial information processing method.
Citation Information
Patent Citations
Information processing device and information processing program
JP2018017668A
Video-based localization and mapping method and system
JP2020516853A
Information processor and program
JP2024073712A
Data correction method and data correction device
WO2022024259A1
For herbicide wheat cropping
JP1989000007A
Cited By
Building image clustering device and building footprint generation device
JP7864236B1