Physical Region 3D Spatial Reconstruction Method and Electronic Equipment Based on Panoramic Video

By performing time-based segmentation and feature point detection on panoramic videos, and combining incremental structure restoration algorithms with progressive parameter tuning and multi-dimensional quality verification, the problems of image distortion, computational resource consumption, and coordinate system alignment in panoramic video 3D reconstruction are solved, achieving efficient and accurate 3D spatial reconstruction and data management.

CN122492972APending Publication Date: 2026-07-31KE COM (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
KE COM (BEIJING) TECHNOLOGY CO LTD
Filing Date
2026-04-24
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing 3D reconstruction methods based on panoramic video have shortcomings in handling image distortion characteristics, computational resource consumption, reconstruction accuracy and efficiency, coordinate system alignment, and data management, making it difficult to meet the needs of high-precision location services and large-scale scene reconstruction.

Method used

A physical region 3D spatial reconstruction method based on panoramic video is adopted. Through an incremental structure recovery algorithm that combines time division, feature point detection and masking filtering, progressive parameter tuning and multi-dimensional quality verification with GPS trajectory information and coordinate system transformation, the integration of local 3D spatial reconstruction results and efficient data management are achieved.

Benefits of technology

It effectively reduces computational complexity, improves reconstruction accuracy and efficiency, ensures precise alignment of reconstruction results with real geographic coordinates, supports dynamic data querying and cross-video spatial correlation analysis, and enhances the reusability value of reconstruction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492972A_ABST
    Figure CN122492972A_ABST
Patent Text Reader

Abstract

This disclosure provides a method and electronic device for physical region 3D spatial reconstruction based on panoramic video. The method includes: dividing the panoramic video of the physical region into multiple sets of reconstructed video frame sequences according to a preset time division step, with each set of reconstructed video frame sequences corresponding to a reconstruction sub-project; extracting static feature points from the reconstructed input images under each reconstruction sub-project, performing feature matching and geometric verification on the static feature points to obtain image matching relationships under the reconstruction sub-project; processing the image matching relationships using an incremental structure restoration algorithm based on progressive parameter tuning and multi-dimensional quality verification to obtain the local 3D spatial reconstruction results corresponding to the reconstruction sub-project; and integrating the reconstruction results of each reconstruction sub-project into a unified global geographic coordinate system to obtain the target 3D spatial reconstruction result of the physical region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of three-dimensional spatial reconstruction technology, and more particularly to a method and electronic device for three-dimensional spatial reconstruction of physical regions based on panoramic video. Background Technology

[0002] With the accelerated digitalization of smart city infrastructure, 3D spatial reconstruction technology based on panoramic video has become a key means to support high-precision location services. Panoramic video can capture 360° surround view scene information in one go, significantly improving data acquisition efficiency compared to traditional narrow-angle cameras, and has unique advantages in large-scale urban scene mapping.

[0003] However, existing 3D reconstruction methods based on panoramic video have significant limitations in handling image distortion characteristics, and direct processing of high-resolution panoramic images (such as 8K and above) consumes huge computational resources. Existing technologies lack geometric projection transformation mechanisms for panoramic characteristics, making it difficult to reduce the computational complexity of subsequent feature matching while maintaining the integrity of visual information.

[0004] For large-scale scene reconstruction of long-term panoramic videos, existing technologies generally adopt a single reconstruction project processing mode. As the video duration increases, the number of camera pose parameters and 3D spatial points that need to be optimized simultaneously during the reconstruction process expands dramatically, leading to an exponential increase in memory consumption and easily triggering computational resource bottlenecks. In addition, the cumulative effect of errors in long video sequences makes the reconstruction trajectory drift problem more and more serious. Existing technologies lack effective segmented reconstruction and integration mechanisms, making it difficult to control computational complexity while ensuring reconstruction accuracy.

[0005] Regarding the alignment of reconstruction results with real geospatial data, existing methods are mostly limited to geometric reconstruction in local coordinate systems, lacking precise registration methods from the reconstructed coordinate system to the geocentric-fixed coordinate system (ECEF) and even the local engineering coordinate system (ENU). This leads to systematic deviations between the reconstruction results and synchronously acquired GPS tracks and Geographic Information System (GIS) data, making it difficult to meet surveying-grade accuracy requirements. Furthermore, existing technologies lack efficient spatial indexing mechanisms for reconstructed data. When it is necessary to support dynamic registration of new videos, query existing reconstructed areas, or perform cross-video spatial correlation analysis, data retrieval efficiency is low, severely limiting the reusability of reconstruction results in downstream tasks such as location services and spatial analysis. Summary of the Invention

[0006] This disclosure provides a method and electronic device for three-dimensional spatial reconstruction of physical regions based on panoramic video, in order to solve at least one of the aforementioned technical problems.

[0007] According to one aspect of this disclosure, a method for reconstructing a three-dimensional spatial physical region based on panoramic video is provided, comprising: The panoramic video of the physical region is divided according to the preset time division step, resulting in multiple sets of reconstructed video frame sequences, each set of reconstructed video frame sequences corresponding to a reconstructed sub-project. For each reconstruction sub-project, static feature points are extracted from the reconstruction input image under the reconstruction sub-project based on feature point detection and mask filtering algorithms. Feature matching and geometric verification are performed on the static feature points to obtain the image matching relationship under the reconstruction sub-project. The reconstruction input image under the reconstruction sub-project is obtained by preprocessing the reconstruction video frame sequence corresponding to the reconstruction sub-project. An incremental structure restoration algorithm based on progressive parameter tuning and multi-dimensional quality verification is used to process image matching relationships and obtain local three-dimensional spatial reconstruction results corresponding to the reconstruction sub-items. The local 3D spatial reconstruction results of each reconstruction sub-project are integrated into a unified global geographic coordinate system through coordinate system transformation and geographic registration to obtain the target 3D spatial reconstruction results of the physical area.

[0008] In some embodiments, the method further includes: for each reconstruction sub-project, projecting each video frame in the reconstruction video frame sequence corresponding to the reconstruction sub-project into a cube projection image with multiple perspectives, and selecting a target perspective image from the cube projection images with multiple perspectives as the reconstruction input image under the reconstruction sub-project.

[0009] In some embodiments, an incremental structure restoration algorithm based on progressive parameter tuning and multi-dimensional quality verification is used to process image matching relationships to obtain local three-dimensional spatial reconstruction results corresponding to the reconstruction sub-items, including: Obtain multiple preset parameter combinations, each including: camera triangulation angle threshold and camera forward motion threshold; wherein, the camera triangulation angle threshold and camera forward motion threshold are negatively correlated; Following the order of decreasing camera triangulation angle thresholds, the following incremental reconstruction steps are performed on the current parameter combination sequentially until multi-dimensional quality verification is passed or all parameter combinations have been traversed: From the image matching relationships, the reconstructed input image pair with the highest number of target feature point matching pairs and the smallest reprojection error is selected as the seed image pair to perform the initial 3D reconstruction. Based on the incremental structure restoration strategy, new reconstructed input images and corresponding 3D spatial points are added step by step, and bundle adjustment optimization is performed after each addition. Bundle adjustment optimization includes: simultaneously optimizing the poses of all reconstructed cameras and the positions of the generated 3D spatial points to minimize reprojection error. In this process, after the initial 3D reconstruction and each time a new image is added, multi-dimensional quality verification is performed. Multi-dimensional quality verification includes: model continuity verification, camera trajectory smoothness verification, and frame coverage verification. When the verification fails and there are untried parameter combinations, the process switches to the next set of parameter combinations and re-executes the incremental reconstruction step. When all verifications pass, the current reconstruction result is determined to be the local 3D spatial reconstruction result corresponding to the reconstruction sub-project.

[0010] In some embodiments, before dividing the panoramic video of the physical region according to a preset time division step, the method further includes: The panoramic video is sampled according to a preset frame sampling step size to obtain a key video frame sequence; The panoramic video of the physical region is divided according to the preset time division step to obtain multiple sets of reconstructed video frame sequences, including: grouping the key video frame sequence according to the preset time window length to obtain multiple sets of reconstructed video frame sequences, each set of reconstructed video frame sequences containing some video frames from the key video frame sequence.

[0011] In some embodiments, static feature points in the reconstructed input image under the reconstruction sub-project are extracted based on feature point detection and mask filtering algorithms, including: For the reconstructed input image, the scale-invariant feature transform algorithm is used to detect key points in the reconstructed input image; Identify dynamic object regions and image regions that do not meet quality requirements in the reconstructed input image, and use them as masking regions in the reconstructed input image. Key points are filtered based on the masked region, excluding key points located within the masked region, and the remaining key points are used as static feature points in the reconstructed input image.

[0012] In some embodiments, feature matching and geometric verification are performed on static feature points to obtain image matching relationships under the reconstructed sub-project, including: Obtain the feature descriptor vectors of static feature points, and determine the feature point matching pairs between each pair of reconstructed input images under the reconstruction sub-project by calculating the Euclidean distance between the feature descriptor vectors; Using the random sampling consensus algorithm and the essential matrix constraint method, geometric verification filtering is performed on the feature point matching pairs between each pair of reconstructed input images under the reconstruction sub-project to obtain the target feature point matching pairs between each pair of reconstructed input images under the reconstruction sub-project. Remove reconstructed input image pairs whose number of target feature point matching pairs under the reconstruction sub-project is less than a preset threshold, and construct image matching relationships under the reconstruction sub-project based on the target feature point matching pairs between the remaining reconstructed input image pairs.

[0013] In some embodiments, the method further includes: Acquire GPS trajectory information that is synchronously collected with the panoramic video. The GPS trajectory information includes multiple trajectory points and the timestamp and geographic coordinates of each trajectory point. The key video frame sequence is timestamped and aligned with the GPS trajectory information to generate a video frame geographic mapping table. The video frame geographic mapping table is used to record the mapping relationship between the frame number of each key video frame and the corresponding geographic coordinates. The timestamp matching and alignment process includes: calculating the time deviation between the video frame timestamp and the trajectory point timestamp; when the time deviation is less than a preset time threshold, the alignment is confirmed to be successful.

[0014] In some embodiments, the local 3D spatial reconstruction results of each reconstruction sub-project are integrated into a unified global geographic coordinate system through coordinate system transformation and georegistration to obtain the target 3D spatial reconstruction result of the physical region, including: For each reconstruction sub-project, based on the iterative nearest point algorithm, the local 3D spatial reconstruction results of the reconstruction sub-project are registered from the local coordinate system to the geocentric coordinate system using the video frame geographic mapping table; The reconstruction results in the geocentric-fixed coordinate system are transformed to the northeast-sky local rectangular coordinate system to obtain the three-dimensional spatial reconstruction results of the reconstruction sub-project in the global geographic coordinate system. The origin of the northeast-sky local rectangular coordinate system is determined based on the geographic center of the reconstruction sub-project, the East axis points eastward, the North axis points northward, and the Up axis points to the sky. The 3D spatial reconstruction results of each reconstruction sub-project in the global geographic coordinate system are merged to obtain the target 3D spatial reconstruction result of the physical area.

[0015] In some embodiments, after obtaining the target three-dimensional spatial reconstruction result of the physical region, the method further includes: A two-dimensional spatial mesh with a preset mesh size is constructed, and each pose in the target three-dimensional spatial reconstruction result is used as a candidate pose and mapped to the corresponding mesh cell. Perform a neighborhood occupancy check on each candidate pose sequentially to determine the pose occupancy status of its own grid cell and adjacent grid cells; Based on spatial distance calculation, each candidate pose is screened, and poses with spatial distance greater than the preset mesh size are retained. This ensures that the spatial distance between any two ultimately retained poses is greater than the preset mesh size, thus obtaining the sparsed target 3D spatial reconstruction result.

[0016] In some embodiments, the method further includes: Following a four-level tree-like hierarchical structure of reconstruction task, panoramic video source, reconstruction sub-project, and camera pose, the pose data in the sparsed target 3D space reconstruction result is reorganized. Here, the reconstruction task is the root node, corresponding to the complete 3D space reconstruction process of the physical region; the panoramic video source is the first-level child node under the root node, corresponding to the original video file of the acquired panoramic video; the reconstruction sub-project is the second-level child node under the first-level child node, corresponding to each reconstruction sub-project; and the camera pose is the leaf node under the second-level child node, corresponding to the position and pose information of each camera retained in the sparsed target 3D space reconstruction result. Calculate and record statistical information for nodes at each level in a four-level tree hierarchy. The statistical information includes: the total number of original poses under the node, the number of poses retained after the pose filtering process, and the proportion of poses removed after the pose filtering process to the total number of original poses. Record the grid position information of each leaf node corresponding to the pose of each camera in the two-dimensional spatial grid to build a spatial index, and maintain the associated metadata of each leaf node with the original video file, the project storage directory of the reconstructed sub-project, and the coordinate transformation parameters between the geocentric coordinate system and the northeast-sky local rectangular coordinate system.

[0017] In some embodiments, the method further includes: For the new panoramic video to be processed, new GPS trajectory information acquired synchronously with the new panoramic video is obtained, and based on the geographic coordinates in the new GPS trajectory information, candidate reconstruction sub-projects are selected from the reconstruction sub-projects in the four-level tree hierarchy. The four-level tree hierarchy includes: reconstruction task, panoramic video source, reconstruction sub-project, and camera pose. Based on the visual feature matching algorithm, the new panoramic video is matched with the reconstruction input image under the candidate reconstruction sub-project to determine the localization result; The localization results are geometrically verified using the random sampling consensus algorithm and the essential matrix constraint method. After removing outliers, the camera pose estimate corresponding to the new panoramic video is obtained. The camera pose estimation is transformed from the local coordinate system to the northeast-sky local rectangular coordinate system, and the grid position information of the new panoramic video in the two-dimensional spatial grid is calculated. The camera pose corresponding to the new panoramic video is then used as a leaf node to associate the corresponding panoramic video source in the four-level tree hierarchy.

[0018] According to one aspect of this disclosure, an electronic device is provided, comprising: The memory stores execution instructions; and the processor executes the execution instructions stored in the memory, causing the processor to perform the above-described method for physical region 3D spatial reconstruction based on panoramic video.

[0019] According to one aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described method for three-dimensional spatial reconstruction of physical regions based on panoramic video. Attached Figure Description

[0020] The accompanying drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, serve to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification.

[0021] Figure 1 This is a flowchart illustrating a method for reconstructing a three-dimensional physical region based on panoramic video according to an embodiment of the present disclosure. Figure 2 It is the camera movement path when capturing panoramic video according to the embodiments of this disclosure; Figure 3 This is a schematic diagram of a cube-shaped projection image according to an embodiment of the present disclosure; Figure 4 This is a schematic diagram of reconstructing an input image according to an embodiment of the present disclosure; Figure 5 This is a schematic diagram of the masking region of a reconstructed input image according to an embodiment of the present disclosure; Figure 6 This is a schematic diagram illustrating the coordinate registration principle of the iterative nearest point algorithm according to the embodiments of this disclosure; Figure 7 This is a schematic diagram illustrating the principle of pose selection in one embodiment of this disclosure; Figure 8 This is a schematic diagram of the three-dimensional spatial reconstruction result of one embodiment of the present disclosure; Figure 9 This is a schematic block diagram of an electronic device according to one embodiment of the present disclosure. Detailed Implementation

[0022] The present disclosure will now be described in further detail with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are for illustrative purposes only and are not intended to limit the scope of the disclosure. Furthermore, it should be noted that, for ease of description, only the parts relevant to the present disclosure are shown in the accompanying drawings.

[0023] It should be noted that, where there is no conflict, the embodiments and features described in this disclosure can be combined with each other. The technical solutions of this disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0024] This solution aims to address a series of specific technical challenges encountered in large-scale 3D spatial reconstruction at the community level based on panoramic video. It is primarily applied to high-precision digital acquisition and long-term operation and maintenance management in medium-to-large-scale scenarios such as urban residential communities and commercial districts. This solution can also be applied to other scenarios. For example, the physical area in this solution can be a residential community, school, hospital, or other similar area. Optionally, the physical area can also be a pre-defined area.

[0025] Some of the main technical problems in the related technologies are: The technical contradiction between continuous reconstruction of large-scale scenes and limited computing resources. In community-level road inspections or surveying operations, the duration of a single panoramic video acquisition often reaches tens of minutes or even hours, covering an area of ​​several square kilometers. When existing technologies process such long-sequence video data using a single reconstruction project, memory consumption expands rapidly as the number of camera poses and 3D points accumulates, easily triggering computing resource bottlenecks and causing reconstruction interruptions. At the same time, the accumulation of errors during long-sequence reconstruction processes can lead to trajectory drift and model deformation, making it difficult to guarantee global geometric consistency at the community level. Therefore, there is an urgent need for a technical solution that can both reasonably divide large-scale video into blocks to control the scale of a single computation and ensure seamless integration of the reconstruction results of each block, maintaining topological continuity.

[0026] The contradiction between geometric distortion and feature extraction accuracy in panoramic images. During handheld panoramic camera acquisition, the device inevitably passes through complex areas such as building corners and narrow passages. In these areas, the equidistant cylindrical projection image exhibits severe distortion at the top and bottom, leading to numerous mismatches in traditional feature extraction algorithms. Simultaneously, direct processing of 8K high-resolution panoramic images consumes enormous computational resources, while simple downsampling results in the loss of fine texture information, affecting subsequent reconstruction accuracy. Therefore, a geometric projection transformation mechanism tailored to the characteristics of panoramic images is needed to eliminate extreme distortion, reduce computational complexity, and preserve key visual features to support high-precision reconstruction.

[0027] The contradiction between the dynamic changes in complex urban scenes and the robustness of reconstruction algorithms. Real-world residential environments are subject to dynamic interference factors such as pedestrian movement, vehicle parking, lighting changes, and reflections from glass curtain walls. Furthermore, the texture richness varies significantly across different areas (e.g., green areas, paved roads, building facades). Existing fixed-parameter reconstruction strategies struggle to adapt to this scene diversity, often resulting in reconstruction interruptions due to strict parameters or the introduction of numerous outliers due to lenient parameters. In addition, the lack of effective quality monitoring methods during long video reconstruction makes it difficult to automatically identify local reconstruction failures caused by feature mismatches or geometric degradation, leading to a low overall reconstruction success rate. Therefore, a progressive reconstruction mechanism that can adaptively adjust reconstruction parameters based on scene complexity and possess multi-dimensional quality verification capabilities is urgently needed.

[0028] The technical contradiction between the fusion of locally reconstructed coordinate systems and real geographic coordinates. In applications such as smart property management and emergency navigation, it is necessary to accurately align the reconstructed results with BeiDou / GPS tracks and urban GIS base maps to achieve coordinate unification between "video and map". Existing reconstruction methods are mostly limited to their own local coordinate systems and lack precise registration methods with the Geocentric Geofixed Coordinate System (ECEF) and the Northeastern Sky Engineering Coordinate System (ENU), resulting in systematic deviations between the reconstructed model and real geographic data, making it difficult to meet positioning accuracy requirements. At the same time, due to the lack of traceable transformation parameters from reconstructed coordinates to geographic coordinates, it is difficult to quickly register newly added videos into the existing reconstruction system. Therefore, a scale alignment and coordinate transformation method based on synchronized GPS tracks is needed to establish a precise mapping relationship between local reconstruction and global geographic coordinates.

[0029] There is a technical contradiction between the static management and dynamic application of reconstructed data. Community-level reconstruction results contain massive amounts of pose data and spatial point clouds. Existing technologies mostly employ discrete file storage, lacking an efficient hierarchical indexing structure. When it is necessary to support the dynamic registration of newly added inspection videos, query historical images of specific buildings, or perform cross-time period change detection, the current lack of a spatial location-based data organization mechanism leads to low data retrieval efficiency and severely restricts the reusability of reconstruction results. Therefore, there is an urgent need for a data management method that supports a four-level tree-structured hierarchical organization and a two-dimensional spatial grid index to achieve long-term archiving and efficient querying of community-level reconstruction data.

[0030] For ease of description and to make the technical solutions of the specific embodiments of this disclosure easier to understand, the technical terms involved in the specific embodiments of this disclosure are explained as follows: INSV: Panoramic video format; GPX: GPS Exchange Format; H.265: High Efficiency Video Coding; SDK: Software Development Kit; OpenCV: Open Source Computer Vision Library; COLMAP: Combination of Structure-from-Motion and Multi-View Stereo; SIFT: Scale-Invariant Feature Transform; ORB: Oriented Fast and Rotated BRIEF, characterized by its orientation, fast rotation, invariance, robust binary, and independent fundamental features. RANSAC: Random Sample Consensus; SfM: Structure from Motion; MVS: Multi-View Stereo; SLAM: Simultaneous Localization and Mapping; BA: Bundle Adjustment; ICP: Iterative Closest Point; GPS: Global Positioning System; GIS: Geographic Information System; ECEF: Earth-Centered, Earth-Fixed coordinate system; WGS84: World Geodetic System 1984; ENU: East-North-Up, the coordinate system for the northeast celestial sphere. FPS: Frames Per Second; 8K: 8 Kilo (pixels), referring to an ultra-high-definition standard with a resolution of 7680×4320 or similar.

[0031] Figure 1 This is a flowchart illustrating a method for reconstructing a physical region's three-dimensional space based on panoramic video, as proposed in this disclosure. Method M100 includes the following steps S101-S104: S101. Divide the panoramic video of the physical area according to the preset time division step to obtain multiple sets of reconstructed video frame sequences, each set of reconstructed video frame sequences corresponding to a reconstructed sub-project. The panoramic video in this disclosure can be video captured by a camera along multiple paths, and the camera's movement path can be found in [reference needed]. Figure 2 As shown, Figure 2 In the diagram, different colors represent different paths.

[0032] Specifically, the aforementioned panoramic video is a transcoded panoramic video.

[0033] Optionally, the SDK compilation image can be packaged first. Containerization ensures environmental consistency and reproducibility of the panoramic video decoding process, thereby eliminating inconsistencies in SDK interface calls caused by differences in development environments and ensuring stable basic data quality in subsequent reconstruction processes. Then, overall debugging of the SDK processing is performed to determine the optimal combination of optical flow stitching parameters, color enhancement parameters, and H.265 encoding parameters. This ensures a balance between I / O efficiency and storage costs in subsequent frame sampling processes while eliminating multi-lens parallax, optimizing dynamic range and color reproduction.

[0034] When transcoding panoramic video, INSV format can be converted to standard 8K resolution MP4 format. The conversion process includes advanced processing such as optical flow stitching, color enhancement, and H.265 encoding optimization. Transcoding panoramic video generates a standardized video stream that meets the input requirements of reconstruction algorithms, enabling subsequent cubespan transformations to obtain seamless, low-noise, and high dynamic range input images, thereby ensuring the accuracy of SIFT keypoint detection.

[0035] In some optional embodiments of this disclosure, before dividing the panoramic video of the physical region according to a preset time division step, the method further includes: sampling the panoramic video according to a preset frame sampling step to obtain a key video frame sequence; in S101, dividing the panoramic video of the physical region according to the preset time division step to obtain multiple sets of reconstructed video frame sequences includes: grouping the key video frame sequences according to a preset time window length to obtain multiple sets of reconstructed video frame sequences, each set of reconstructed video frame sequences containing a portion of the video frames in the key video frame sequence.

[0036] The panoramic video is sampled according to the preset frame sampling step size to obtain the key video frame sequence. The reconstruction quality and computational efficiency can be balanced by interval sampling, avoiding redundant calculations caused by overly dense sampling or loss of key motion information by overly sparse sampling.

[0037] Optionally, the panoramic video file can be read using OpenCV to obtain the frame rate of the panoramic video and verify whether it is divisible by a preset number of samples per second. If so, the quotient obtained by dividing the frame rate by the number of samples per second is used as the preset frame sampling step size. If not, the preset number of samples per second is adjusted so that the frame rate is divisible by the number of samples per second.

[0038] By sampling the panoramic video based on its total number of frames and a preset frame sampling step size, a key video frame sequence can be obtained. These key video frames are then saved as JPEG images, and a final frame check is performed to ensure no data is lost at the end of the video, providing a complete frame sequence for subsequent reconstruction.

[0039] Optionally, after confirming that the video path of the panoramic video is valid, the video capture object can be initialized, and video frames can be extracted cyclically according to the preset frame sampling step size. The extracted frame images are saved in JPEG format, and the frame number corresponding to each frame is recorded synchronously. After all frames are extracted, the system resources occupied by the video capture object are released to prevent memory leaks.

[0040] Optionally, GPS trajectory information (GPX format) corresponding to the panoramic video can also be extracted simultaneously to obtain the geographic coordinates and timestamp data corresponding to each trajectory point, and to establish a mapping basis between video frames and geographic coordinates.

[0041] In some optional embodiments of this disclosure, the method further includes: matching and aligning the key video frame sequence with GPS trajectory information using timestamps to generate a video frame geographic mapping table. Specifically, the time deviation between the video frame timestamp and the trajectory point timestamp is calculated, and alignment is confirmed to be successful when the time deviation is less than a preset time threshold. Through strict time deviation control (e.g., less than 1.05 seconds), high-precision matching between video frames and geographic coordinates is ensured, providing a reliable scale alignment benchmark for subsequent ICP registration and avoiding coordinate registration failures caused by time misalignment.

[0042] Optionally, the aforementioned time window length can be 60 seconds. If the duration corresponding to the remaining frames after division is greater than a preset threshold (e.g., 30 seconds), then an additional reconstruction sub-item is added to cover the tail data. If the number of reconstructed sub-items calculated is less than 1, then it is corrected to 1 to ensure that at least one reconstruction sub-item exists. The reconstruction sub-project in this solution can refer to the COLMAP project, specifically the independent reconstruction work unit created for each time period when using COLMAP (an open-source 3D reconstruction software) for reconstruction.

[0043] Optionally, the above method also includes initializing the workspace for each reconstruction sub-project: creating an independent working directory for each reconstruction sub-project and mapping the directory path to the corresponding project index, thus constructing a complete reconstruction sub-project directory structure. Each reconstruction sub-project corresponds one-to-one with a specific time period (usually set to a 60-second time window) in the reconstructed video frame sequence. By employing a time-dimension segmentation strategy, the long video reconstruction task is decomposed into multiple parallelizable subtasks, effectively solving the problems of excessive memory consumption and exponential growth in computational complexity with the length of the time series during long video reconstruction.

[0044] This solution can accurately calculate the frame range that each reconstruction sub-project should process based on the video frame rate and the reconstruction sub-project index. It employs a preset fixed time window length to ensure that each reconstruction sub-project contains sufficient frame data to support stable reconstruction, while avoiding computational bottlenecks caused by excessive data volume in a single project. Special processing is applied to the last reconstruction sub-project to ensure that no data is lost at the end of the video. It implements a temporal segmentation processing strategy for the video, optimizing the allocation of computational resources while ensuring reconstruction stability, and ensuring the integrity of the video data through an end-of-line data processing mechanism.

[0045] S102. For each reconstruction sub-project, static feature points are extracted from the reconstruction input image under the reconstruction sub-project based on feature point detection and mask filtering algorithms. Feature matching and geometric verification are performed on the static feature points to obtain the image matching relationship under the reconstruction sub-project. The reconstruction input image under the reconstruction sub-project is obtained by preprocessing the reconstruction video frame sequence corresponding to the reconstruction sub-project. In some optional embodiments of this disclosure, the reconstruction input image under the reconstruction sub-project is obtained by preprocessing the reconstruction video frame sequence corresponding to the reconstruction sub-project. The aforementioned method further includes: for each reconstruction sub-project, projecting each video frame in the reconstruction video frame sequence corresponding to the reconstruction sub-project into a cube projection image with multiple perspectives, and selecting the target perspective image from the cube projection images with multiple perspectives as the reconstruction input image under the reconstruction sub-project.

[0046] In some optional embodiments of this disclosure, projecting each video frame in the reconstructed video frame sequence corresponding to the reconstructed sub-project into a cube projection image with multiple perspectives includes: performing a coordinate transformation of cube projection on each video frame in the reconstructed video frame sequence corresponding to the reconstructed sub-project to generate 6 cube projection images, the 6 cube projection images corresponding to the front, back, left, right, top and bottom perspectives respectively. Figure 3 This is a schematic diagram of a cube's surface projection image.

[0047] Panoramic image distortion can be eliminated by using cube projection transformation, providing more accurate geometric information for COLMAP's feature detection algorithm.

[0048] When selecting the target viewpoint image from multiple cube projection images, the optimal faces from the six cube faces can be chosen for reconstruction based on the `cube_num` parameter. Typically, the four horizontal faces (front, back, left, and right) are selected because these faces contain more ground feature points, eliminating weak texture areas such as the sky, which is beneficial for camera pose estimation and reduces the computational cost of subsequent feature matching.

[0049] In some optional embodiments of this disclosure, interval sampling can be performed again according to the frame_step parameter as needed to balance reconstruction quality and computational efficiency. Interval sampling balances reconstruction quality and computational efficiency, avoiding redundant calculations caused by overly dense sampling, while preventing the loss of crucial motion information due to overly sparse sampling.

[0050] After selecting the target viewpoint image, the extracted cube face image can be preprocessed with measures such as size standardization and brightness equalization to generate the reconstructed input image. This ensures the stability of the feature detection and matching algorithm under different lighting conditions and resolutions, and improves the reliability of static feature point extraction.

[0051] In some optional embodiments of this disclosure, step S102, which involves extracting static feature points from the reconstructed input image under the reconstructed sub-project based on the feature point detection and mask filtering algorithm, includes the following steps S1021-S1023: S1021. For the reconstructed input image, the scale-invariant feature transform algorithm is used to detect key points in the reconstructed input image; S1022. Identify dynamic object regions and image regions that do not meet quality requirements in the reconstructed input image, and use them as masking regions in the reconstructed input image. For details, please refer to Figure 4 and Figure 5 As shown, Figure 4 For a reconstructed input image, Figure 5 The black area in the image is the masked area.

[0052] S1023. Filter key points based on the masking region, exclude key points located within the masking region, and use the remaining key points as static feature points in the reconstructed input image.

[0053] Specifically, it can automatically identify and reconstruct dynamic objects (pedestrians, vehicles, etc.) in the input image and generate corresponding masking regions. These dynamic elements can interfere with COLMAP's feature matching and affect reconstruction accuracy.

[0054] The aforementioned quality requirements may include sharpness requirements, exposure requirements, and regional color change frequency requirements. For example, blurry areas, overexposed areas, and solid color areas in the reconstructed input image that do not meet the quality requirements are considered image regions that do not meet the quality requirements. Removing low-quality regions that are detrimental to feature extraction can ensure the reliability of subsequent feature detection.

[0055] It is important to ensure that the masks in adjacent frames have temporal continuity to avoid reconstruction instability caused by mask jumps. This will help avoid reconstruction instability caused by mask jumps and improve the smoothness and consistency of the reconstruction results.

[0056] Optionally, no more than 2048 SIFT keypoints can be detected in each reconstructed input image. The feature descriptor is represented by a 128-dimensional vector to encode the gradient distribution information of the image patch around the keypoint. By utilizing the scale invariance and rotation invariance of the SIFT algorithm, the same three-dimensional points can be stably identified under different viewpoints and lighting conditions, providing highly discriminative descriptive information for subsequent feature matching.

[0057] Key points are filtered based on the masked region, excluding key points located within the masked region, and the remaining key points are used as static feature points in the reconstructed input image. Feature points of dynamic objects, blurred regions, and distorted regions can be automatically excluded through the mask file, ensuring that only high-quality static features participate in the reconstruction process, which significantly improves the robustness of the reconstruction algorithm.

[0058] Optionally, static feature points and their feature descriptors, image metadata, and other information can be structured and stored in a database file (such as database.db in SQLite format); this provides an efficient data access interface for subsequent feature matching and geometric verification, and supports rapid retrieval and management of large-scale feature data.

[0059] Optionally, the analysis of image features can be implemented based on the colmap_feature_extractor function.

[0060] In some optional embodiments of this disclosure, step S102, which involves performing feature matching and geometric verification on static feature points to obtain image matching relationships under the reconstructed sub-project, includes the following steps S1024-S1026: S1024. Obtain the feature descriptor vectors of static feature points, and determine the feature point matching pairs between each pair of reconstructed input images under the reconstruction sub-project by calculating the Euclidean distance between the feature descriptor vectors. S1025. Using the random sampling consensus algorithm and the essential matrix constraint method, geometric verification filtering is performed on the feature point matching pairs between each pair of reconstructed input images under the reconstruction sub-project to obtain the target feature point matching pairs between each pair of reconstructed input images under the reconstruction sub-project. S1026. Remove reconstructed input image pairs whose number of target feature point matching pairs under the reconstruction sub-project is less than a preset threshold, and construct image matching relationships under the reconstruction sub-project based on the target feature point matching pairs between the remaining reconstructed input image pairs.

[0061] Optionally, feature association across images can be achieved using the `colmap_exhaustive_match` function to establish visual correspondences: Optionally, in practice, feature matching can be performed on each reconstructed input image pair, traversing and calculating the distance between feature descriptor vectors to find potential corresponding points. Although computationally intensive, this ensures that no valid matching relationships are missed, establishing a complete foundation for visual correspondences.

[0062] This scheme filters out false matches by using epipolar geometric constraints, retaining only feature point pairs that satisfy geometric consistency, thus significantly improving the reliability of the matching.

[0063] In this scheme, reconstructed input image pairs with fewer than a preset threshold (e.g., 100) of matching target feature points are discarded. Based on the matching target feature points among the remaining reconstructed input image pairs, image matching relationships under the reconstructed sub-project are constructed. This ensures that there are sufficient feature associations between image pairs to support reliable relative pose estimation, excludes weakly connected image pairs, and avoids reconstruction failures caused by weak connections.

[0064] Furthermore, all valid image matching relationships can be recorded in the database to form a connected matching graph structure, providing topological support for subsequent incremental structure restoration and ensuring the coherence of the reconstruction process.

[0065] S103. Using an incremental structure restoration algorithm based on progressive parameter tuning and multi-dimensional quality verification, the image matching relationship is processed to obtain the local three-dimensional space reconstruction results corresponding to the reconstruction sub-items. In some optional embodiments of this disclosure, in the aforementioned step S103, the image matching relationship is processed by an incremental structure restoration algorithm based on progressive parameter tuning and multi-dimensional quality verification to obtain the local three-dimensional spatial reconstruction result corresponding to the reconstruction sub-item, including the following steps S1031-S1034: S1031. Obtain multiple preset parameter combinations, each parameter combination including: camera triangulation angle threshold and camera forward motion threshold; wherein, the camera triangulation angle threshold and the camera forward motion threshold are negatively correlated, the larger the camera triangulation angle threshold, the smaller the camera forward motion threshold; S1032. Following the order of decreasing camera triangulation angle thresholds, select the current parameter combination sequentially and execute the following incremental reconstruction steps until multi-dimensional quality verification is passed or all parameter combinations have been traversed: In practice, seven parameter combinations can be used. Following the order of decreasing camera triangulation angle thresholds, the current parameter combination is selected sequentially to execute the incremental reconstruction steps until multi-dimensional quality verification is passed or all parameter combinations have been traversed. Starting with stringent parameters and gradually downgrading them ensures high quality during successful reconstruction while lowering the threshold to improve the overall success rate, achieving a dynamic balance between reconstruction quality and success rate.

[0066] Optionally, the most stringent parameter combination is: triangulation angle 60 degrees, forward motion threshold 0.5.

[0067] S1033. From the image matching relationship, select the reconstructed input image pair with the highest number of target feature point matching pairs and the smallest reprojection error as the seed image pair, and perform the initial three-dimensional reconstruction. Choosing the image pair with the strongest geometric constraints and the smallest reprojection error as the starting point for reconstruction ensures the reliability of the initial reconstruction and lays the foundation for subsequent incremental expansion. Furthermore, starting with the two image pairs with the highest number of matching interior points and the smallest reprojection error, new images and 3D points are added incrementally. This strategy can handle large-scale scenes and performs global optimization at each step to ensure reconstruction quality.

[0068] S1034. Based on the incremental structure restoration strategy, new reconstructed input images and corresponding 3D spatial points are added step by step, and bundle adjustment optimization is performed after each addition. Bundle adjustment optimization includes: simultaneously optimizing the poses of all reconstructed cameras and the positions of the generated 3D spatial points to minimize reprojection error. Optionally, a three-dimensional spatial point refers to a three-dimensional spatial coordinate point representing the geometric structure of a scene, which is recovered from the feature matching pairs between the reconstructed image and the newly added image through triangulation calculation. The three-dimensional spatial points and the camera pose together constitute the sparse point cloud data of the local three-dimensional spatial reconstruction result.

[0069] In practice, nonlinear least squares optimization (Bundle Adjustment) can be performed. By gradually adding images and global optimization, large-scale scenes can be handled, error accumulation can be effectively controlled, and the global consistency of the reconstruction model can be ensured.

[0070] In this process, after the initial 3D reconstruction and each time a new image is added, multi-dimensional quality verification is performed. Multi-dimensional quality verification includes: model continuity verification, camera trajectory smoothness verification, and frame coverage verification. When the verification fails and there are untried parameter combinations, the process switches to the next set of parameter combinations and re-executes the incremental reconstruction step. When all verifications pass, the current reconstruction result is determined to be the local 3D spatial reconstruction result corresponding to the reconstruction sub-project.

[0071] Specifically, model continuity verification ensures the generation of a unique, continuous sparse model rather than multiple fragmented segments; camera trajectory smoothness verification detects abnormal jumps by analyzing the motion velocity between adjacent frames; and frame coverage verification ensures that the number of successfully reconstructed frames reaches the expected proportion (e.g., 90%). Reconstruction quality is comprehensively monitored from three dimensions: model integrity, trajectory continuity, and coverage integrity, allowing for the timely detection and correction of reconstruction defects caused by feature mismatches or geometric degradation.

[0072] Before each reconstruction, previous failed results are automatically cleaned up, detailed attempt logs are recorded, and degradation is achieved by lowering parameter requirements. When all verifications pass, the current reconstruction result is determined to be a local 3D spatial reconstruction result. It can achieve intelligent fault tolerance and automatic retry. Through parameter degradation and complete state tracking, the originally unstable reconstruction process is transformed into a highly reliable automated process, which significantly improves the reconstruction success rate in complex urban scenarios.

[0073] Optionally, the core SfM (Structure from Motion) algorithm can be executed via the retry_maper function.

[0074] Furthermore, it can monitor key indicators during the reconstruction process, including the number of reconstructed images and the average reprojection error. When an anomaly is detected, it can automatically adjust parameters and retry to improve the reconstruction success rate.

[0075] Furthermore, it can generate reconstruction results in binary (.bin) and text (.txt) formats for efficient storage and manual inspection, respectively; it outputs reconstruction statistics, including the number of successfully reconstructed images, the number of 3D points, and reprojection errors, to evaluate reconstruction quality. It can provide reconstruction results in multiple formats to support different application scenarios and achieve objective evaluation and monitoring of reconstruction quality through quantitative indicators, providing data support for reconstruction algorithm optimization and data selection.

[0076] S104. The local three-dimensional spatial reconstruction results of each reconstruction sub-project are integrated into a unified global geographic coordinate system through coordinate system transformation and geographic registration to obtain the target three-dimensional spatial reconstruction results of the physical area.

[0077] In some optional embodiments of this disclosure, the method further includes the following steps S31-S32: S31. Obtain GPS trajectory information that is synchronously collected with the panoramic video. The GPS trajectory information includes multiple trajectory points and the timestamp and geographic coordinates corresponding to each trajectory point. S32. Match and align the key video frame sequence with the GPS trajectory information using timestamps to generate a video frame geographic mapping table. The video frame geographic mapping table is used to record the mapping relationship between the frame number of each key video frame and its corresponding geographic coordinates. The timestamp matching and alignment process includes: calculating the time deviation between the video frame timestamp and the trajectory point timestamp; when the time deviation is less than a preset time threshold, the alignment is confirmed to be successful.

[0078] Optionally, GPS track information (GPX format) can be initialized and loaded, and the system can check whether the track data points have been loaded correctly. If not, an exception handling mechanism is triggered. An output file (such as icp.txt) for storing alignment data is created and opened, and a data writing channel is established. This ensures the integrity and availability of GPS track data, providing a reliable data foundation for subsequent timestamp matching.

[0079] Specifically, the key video frame sequence can be traversed, and the sequence number of each frame can be converted into a corresponding timestamp (second). Based on the timestamp, the matching geographical location data points in the GPS trajectory information can be queried. This establishes a temporal correspondence between video frames and geographical coordinates, achieving a preliminary association between visual data and spatial location.

[0080] Optionally, the aforementioned time threshold can be 1.05 seconds. If the deviation exceeds the preset time threshold, the frame is skipped and a warning message is output. Through strict time deviation control, abnormal matching points with excessive time misalignment are eliminated, ensuring high-precision alignment between video frames and geographic coordinates and avoiding coordinate registration deviations caused by time asynchrony.

[0081] Furthermore, the verified frame data is written to an output file according to a predetermined format. The frame data includes the frame file name, latitude, longitude, and altitude. After the data writing is complete, the file is closed, generating a video frame geographic mapping table. A frame mapping file containing complete geographic attributes (latitude, longitude, and elevation) is established to provide standardized input data for subsequent iterative point-to-close (ICP) registration and coordinate system transformation, supporting high-precision alignment between the reconstruction results and the real geographic space.

[0082] In some optional embodiments of this disclosure, in the aforementioned step S104, the local three-dimensional spatial reconstruction results of each reconstruction sub-project are integrated into a unified global geographic coordinate system through coordinate system transformation and geographic registration to obtain the target three-dimensional spatial reconstruction results of the physical region, including the following steps S1041-S1043: S1041. For each reconstruction sub-project, based on the iterative nearest point algorithm, the local three-dimensional spatial reconstruction results of the reconstruction sub-project are registered from the local coordinate system to the geocentric coordinate system using the video frame geographic mapping table. A schematic diagram illustrating the coordinate registration principle of the Iterative Closest Point (ICP) algorithm can be found in [link to diagram]. Figure 6As shown in the figure. Yellow points (Source points): represent the set of source points to be registered, corresponding to the local 3D spatial reconstruction results of the reconstruction sub-project in this scheme (including camera pose or 3D spatial point cloud); Green points (Destination points): represent the set of target points, corresponding to the geographic coordinate points in the GPS trajectory information (trajectory points after timestamp matching and alignment) in this scheme; Coordinate grid: shows a two-dimensional plane coordinate system (xy axis), which is expanded to a three-dimensional spatial coordinate system (xyz) in actual implementation.

[0083] S1042. Transform the reconstruction results in the geocentric-fixed coordinate system to the northeast-sky local rectangular coordinate system to obtain the three-dimensional spatial reconstruction results of the reconstruction sub-project in the global geographic coordinate system; where the origin of the northeast-sky local rectangular coordinate system is determined based on the geographic center of the reconstruction sub-project, the East axis points eastward, the North axis points northward, and the Up axis points to the sky; S1043. Merge the three-dimensional spatial reconstruction results of each reconstruction sub-project in the global geographic coordinate system to obtain the target three-dimensional spatial reconstruction result of the physical area.

[0084] Optionally, the reconstructed camera pose data can first be converted from its original format to a standard format (e.g., from images.txt to pose.json), containing the camera's quaternion representation of the rotation matrix and translation vector. Standardizing the pose data format facilitates subsequent processing, visualization, and coordinate transformation calculations.

[0085] Optionally, the `geo_model_align` function can be used to perform critical coordinate system alignment operations, aligning the reconstruction results with the GPS trajectory using the ICP (Iterative Closest Point) algorithm. Specifically, for each reconstruction sub-project, based on the iterative closest point algorithm and utilizing a video frame geographic mapping table, the registration process from the local coordinate system to the geocentric geofixed coordinate system for the local 3D spatial reconstruction results of the sub-project includes: performing scale estimation to calculate the scale ratio between the reconstructed coordinate system and the real world; performing rotation alignment to rotate the reconstructed coordinate system to the geographic coordinate system orientation through quaternion transformation; performing translation alignment to translate the reconstruction origin to the correct geographic location; and minimizing the alignment error through iterative optimization. This achieves high-precision spatial alignment between the reconstruction results and the GPS trajectory, solving the problem of scale uniformity between local reconstruction and the global coordinate system, and laying the foundation for subsequent geospatial applications.

[0086] This stage involves standardizing the global coordinate system, laying the foundation for subsequent spatial analysis.

[0087] Optionally, the geographical distribution of each reconstruction sub-project can be analyzed to intelligently select the optimal reference point as the origin of the local coordinate system. In practice, the geometric center of the dataset or the starting point of the GPS trajectory is usually selected as the origin of the local rectangular coordinate system. This ensures the numerical stability of the coordinate transformation and avoids the loss of floating-point precision caused by calculations far from the origin.

[0088] During multi-level coordinate transformations, a complete transformation link from local reconstructed coordinates to global standard coordinates can be established, registering the local 3D spatial reconstruction results of the reconstruction sub-project from the local coordinate system to the geocentric-fixed coordinate system (ECEF). Specifically, the local coordinates are first transformed to the ECEF using `_apply_transform_to_pose`, then to WGS84 geographic coordinates (latitude and longitude), and finally to the northeast-sky local rectangular coordinate system using `geodetic2enu`. The origin of the northeast-sky local rectangular coordinate system is determined based on the geographic center of the reconstruction sub-project, with the East axis pointing eastward, the North axis pointing northward, and the Up axis pointing upward. This achieves a precise mapping between the local reconstructed coordinates and the global geographic coordinate system, establishing a local planar coordinate system within the physical area, allowing Euclidean distance to directly correspond to the actual distance, meeting the engineering measurement requirements for meter-level or even sub-meter-level accuracy.

[0089] Regarding coordinate system accuracy optimization, the `lonlat_to_epsg3_zone` function can be used to automatically select the most suitable projected coordinate system for the current location area (such as CGCS2000 3-degree zone) to minimize projection distortion. This ensures meter-level accuracy in spatial calculations, meeting the needs of precision measurement and navigation applications.

[0090] Furthermore, multi-dimensional coordinate information can be added to each pose data point, including original local coordinates, geocentric coordinates, northeast-sky coordinates, geographic coordinates, as well as associated video information and coordinate transformation matrices. The 3D spatial reconstruction results of each reconstruction sub-project in the global geographic coordinate system are merged to obtain the target 3D spatial reconstruction result of the physical area. A complete data system is constructed to maintain the traceability and convertibility of reconstruction data across different coordinate systems, forming a unified 3D spatial model covering the entire physical area, providing a standardized data foundation for subsequent spatial queries, data management, and incremental updates.

[0091] In some optional embodiments of this disclosure, after obtaining the target three-dimensional spatial reconstruction result of the physical region, the method further includes the following steps S105-S107: S105. Construct a two-dimensional spatial mesh with a preset mesh size as the unit, and map each pose in the target three-dimensional space reconstruction result as a candidate pose to the corresponding mesh unit. S106. Perform neighborhood occupancy checks on each candidate pose sequentially to determine the pose occupancy status of its own grid cell and adjacent grid cells. Optionally, pose occupancy status refers to whether a selected and reserved camera pose already exists within the grid cell, and the specific identification information of the pose (including pose index or spatial coordinates). By checking the pose occupancy status of the grid cell containing the candidate pose and its adjacent grid cells, reserved poses with similar spatial positions can be quickly identified, providing neighborhood basic data for subsequent filtering based on spatial distance calculation, and realizing rapid query of grid occupancy status and preliminary conflict detection of spatial position.

[0092] S107. Based on spatial distance calculation, each candidate pose is screened, and poses with spatial distance greater than the preset mesh size are retained to ensure that the spatial distance between any two finally retained poses is greater than the preset mesh size, thereby obtaining the sparsed target 3D spatial reconstruction result.

[0093] Specifically, a two-dimensional spatial grid can be constructed with a preset grid size (grid_size) as the unit. Each pose in the target 3D spatial reconstruction result is used as a candidate pose and mapped to the corresponding grid unit. Each grid unit in the two-dimensional spatial grid corresponds to a square spatial region in the real world. The continuous spatial problem is discretized through the gridding strategy. This can significantly reduce computational complexity and provide a basic data structure for rapid spatial positioning.

[0094] Furthermore, a neighborhood occupancy check is performed sequentially on each candidate pose to determine the pose occupancy status of its own mesh cell and adjacent mesh cells. In practice, not only the main mesh cell where the candidate pose is located is checked, but also the pose distribution in the eight surrounding adjacent mesh cells. This achieves a more precise distance control mechanism than simple mesh filtering, handles edge cases at mesh boundaries, and ensures geometric consistency.

[0095] Furthermore, candidate poses are filtered based on spatial distance calculations, retaining those with a spatial distance greater than a preset mesh size. The first layer of filtering (rapid mesh occupancy elimination) quickly eliminates obviously redundant poses based on the occupancy status of mesh cells, achieving preliminary screening. The second layer of filtering (precise distance calculation) ensures that the spatial distance between any two ultimately retained poses is greater than the preset mesh size based on precise Euclidean distance calculations. This dual mechanism guarantees both algorithm efficiency and geometric consistency of the filtering results. Simultaneously, a first-come, first-served retention strategy naturally leads to sparsification in data-dense regions and maintains integrity in sparse regions, achieving global density balance and obtaining the sparsified target 3D spatial reconstruction result.

[0096] A schematic diagram illustrating the process of filtering candidate poses and retaining those with a spatial distance greater than the preset mesh size can be found here. Figure 7 As shown, the grid lines in the figure represent a two-dimensional spatial grid, with each grid cell corresponding to a square region in the real world. Dots marked "Don't" indicate candidate poses that have been rejected because their spatial distance from neighboring retained poses is less than the preset grid size (marked "less than 6M" in the figure), thus failing to meet the minimum spacing requirement. Dots marked "Keep" indicate poses that have been retained after filtering, provided that the Euclidean distance between this pose and its own grid cell, as well as the retained poses in adjacent grid cells, is greater than the preset grid size. Arrows indicate the search direction for neighborhood occupancy checks, showing that the check area covers the current grid cell and its eight surrounding neighboring grid cells (eight-neighborhood).

[0097] In some optional embodiments of this disclosure, the method further includes the following steps S108-S110: S108. According to the four-level tree hierarchy of reconstruction task, panoramic video source, reconstruction sub-project, and camera pose, the pose data in the sparsed target 3D space reconstruction result is reorganized. Among them, the reconstruction task is the root node, corresponding to the complete 3D space reconstruction process of the physical region; the panoramic video source is the first-level child node under the root node, corresponding to the original video file of the acquired panoramic video; the reconstruction sub-project is the second-level child node under the first-level child node, corresponding to each reconstruction sub-project; the camera pose is the leaf node under the second-level child node, corresponding to the position and pose information of each camera retained in the sparsed target 3D space reconstruction result. S109. Calculate and record statistical information for each level of nodes in the four-level tree hierarchy. The statistical information includes: the total number of original poses under the node, the number of poses retained after the pose filtering process, and the proportion of poses removed after the pose filtering process to the total number of original poses. S110. Record the grid position information of each leaf node corresponding to the pose of each camera in the two-dimensional spatial grid to build a spatial index, and maintain the associated metadata of each leaf node with the original video file, the project storage directory of the reconstructed sub-project, and the coordinate transformation parameters between the geocentric coordinate system and the northeast-sky local rectangular coordinate system.

[0098] Based on a four-level tree-like hierarchical structure of reconstruction task, panoramic video source, reconstruction sub-project, and camera pose, the pose data in the sparsed target 3D space reconstruction results is reorganized. The filtered discrete data can be reorganized into structured data packets. Each layer contains complete statistical information and index relationships, supporting efficient hierarchical queries.

[0099] This system calculates and records statistical information for nodes at each level of a four-level tree hierarchy. The statistical information includes: the total number of original poses under that node, the number of poses retained after the pose filtering process, and the proportion of poses removed after the pose filtering process to the total number of original poses. Real-time calculation of data statistics for each level provides a quantitative basis for quality assessment and performance optimization.

[0100] For each camera pose, the leaf nodes record their grid position information in a 2D spatial grid to construct a spatial index. It also maintains the associated metadata between each leaf node and the original video file, the project storage directory of the reconstructed sub-project, and the coordinate transformation parameters between the geocentric-fixed coordinate system and the northeast-sky local rectangular coordinate system. This allows for the construction of a lightweight spatial index, supporting subsequent spatial queries and neighborhood analysis. Simultaneously, it ensures that each pose carries complete associated information, including the source video, project file, and transformation parameters, maintaining data traceability and integrity.

[0101] In some optional embodiments of this disclosure, the method further includes the steps S111-S114: S111. For the new panoramic video to be processed, acquire new GPS trajectory information that is synchronously collected with the new panoramic video, and based on the geographic coordinates in the new GPS trajectory information, select candidate reconstruction sub-projects in the reconstruction sub-projects in the four-level tree hierarchy. S112. Based on the visual feature matching algorithm, the new panoramic video is matched with the reconstruction input image under the candidate reconstruction sub-project to determine the localization result. S113. Use the random sampling consensus algorithm and the essential matrix constraint method to perform geometric verification on the localization results, and obtain the camera pose estimate corresponding to the new panoramic video after removing outliers. Optionally, outliers refer to: in the localization results, there is a significant deviation between the initial pose estimate obtained based on visual feature matching and the theoretical pose calculated based on geometric constraints (essential matrix and epipolar geometry), which are abnormal matching points or abnormal pose estimates. Outliers are usually caused by feature mismatch, dynamic object occlusion, repetitive textures, or lighting changes. The random sampling consensus algorithm identifies and removes outliers through iterative sampling and model verification, ensuring that the final retained camera pose estimate meets the epipolar geometry constraints and has geometric consistency and reliability.

[0102] S114. Transform the camera pose estimation from the local coordinate system to the northeast-sky local rectangular coordinate system, and calculate the grid position information of the new panoramic video in the two-dimensional spatial grid, so as to associate the camera pose corresponding to the new panoramic video as a leaf node with the corresponding panoramic video source in the four-level tree hierarchy.

[0103] Optionally, the three-dimensional spatial reconstruction results in this disclosure can be found in [reference needed]. Figure 8 As shown, Figure 8 In the diagram, the red portion represents the camera pose trajectory, which can be the camera motion path in the local 3D space reconstruction result. Each node on the trajectory corresponds to the camera position (translation) and pose (rotation) at a reconstruction moment. The gray point cloud represents the 3D point cloud, which is the geometric structure that constitutes the physical area scene (such as building facades, roads, vegetation, etc.).

[0104] This scheme establishes a unified spatial reference framework to lay the foundation for subsequent spatial queries. Specifically, it analyzes the geographical distribution of the entire reconstructed dataset and intelligently selects the optimal origin of the Northeastern Sky Local Rectangular Coordinate System (ENU). The selection strategy comprehensively considers multiple factors such as data density, geographic center, and projection accuracy. A complete transformation chain from the original GPS coordinates to the unified spatial coordinates is established, including the transformation from WGS84 to the GCJ-02 Chinese coordinate system, the transformation from geographic coordinates to the geocentric-fixed coordinate system, the precise transformation from the geocentric-fixed coordinate system to the local Northeastern Sky Local Rectangular Coordinate System, and local coordinate system registration based on the reconstructed data. This ensures the numerical stability of subsequent spatial calculations. The entire transformation process employs high-precision geodetic algorithms to ensure meter-level or even sub-meter-level spatial positioning accuracy, meeting the needs of precision measurement and navigation applications.

[0105] For a new panoramic video to be processed, GPS trajectory information collected synchronously with the new panoramic video is obtained. Based on the geographic coordinates in the GPS trajectory information, candidate reconstruction sub-projects are selected from the reconstruction sub-projects in a four-level tree hierarchy. This can achieve coarse positioning based on GPS coordinates and quickly filter candidate reconstruction projects by utilizing geographic distance, thus significantly narrowing the search range.

[0106] Based on visual feature matching algorithms, new panoramic videos are matched with the reconstructed input images corresponding to candidate reconstructed sub-projects to determine precise localization results. Specifically, sub-meter level localization is achieved through Scale Invariant Feature Transform (SIFT) or ORB feature matching. The precise localization results are geometrically verified using the Random Sample Consensus algorithm and the Essential Matrix Constraint method. After removing outliers, the camera pose estimate corresponding to the new panoramic video is obtained. Multimodal feature matching achieves precise matching between the query video and the reconstructed database, and geometric constraints are used to verify the reliability of the matching, ensuring the accuracy of the localization results.

[0107] The camera pose estimation is transformed from the local coordinate system to the northeast-sky local rectangular coordinate system. The local coordinates of the query results are transformed to the global coordinate system to ensure that all query results are comparable and operable within a unified spatial reference frame.

[0108] The algorithm calculates the grid position information of new panoramic videos in a 2D spatial grid, and associates the camera pose corresponding to the new panoramic video as a leaf node with the corresponding panoramic video source in a four-level tree hierarchy. By attaching the pose data of the new video to the corresponding hierarchical node and updating the spatial grid index, it achieves rapid registration and fusion of new videos in the existing reconstruction system, improves the four-level tree hierarchical data organization, and ensures that all data are comparable and operable under a unified global coordinate framework (northeast-sky local rectangular coordinate system). It supports rapid neighborhood query, spatial clustering analysis, density statistics and visualization, as well as cross-time period change detection, meeting the needs of smart property management, emergency navigation and other application scenarios for dynamic fusion and long-term archiving management of incremental data.

[0109] This disclosed scheme divides long, cell-level videos into multiple reconstruction sub-projects by segmenting the video over time steps. This decomposes large-scale reconstruction tasks into sub-task units that can be processed independently or in parallel, effectively solving the problems of excessive memory consumption and exponentially increasing computational complexity during long video reconstruction. Simultaneously, through reasonable time window control, it ensures that each sub-project contains sufficient frame data to support stable reconstruction, achieving a balance between the feasibility and computational efficiency of large-scale scene reconstruction. By combining feature point detection and mask filtering with feature matching, mask filtering automatically excludes feature points of dynamic objects (pedestrians, vehicles) and low-quality regions (blurred, overexposed), ensuring that only high-quality static features participate in reconstruction, significantly reducing feature mismatches and improving reconstruction accuracy. Furthermore, geometric verification (RANSAC and essential matrix constraints) filters out false matches, establishing reliable image matching relationships and providing high-quality input data for subsequent structure restoration. Incremental structure restoration through progressive parameter tuning and multi-dimensional quality verification employs a progressive reconstruction using parameter combinations ranging from strict to lenient. Combined with multi-dimensional quality verification of model continuity, trajectory smoothness, and frame coverage, a dynamic balance between reconstruction quality and success rate is achieved. This approach adapts to varying scene complexities (lighting variations, sparse textures, dynamic interference), automatically identifies and corrects reconstruction defects, transforming the previously unstable reconstruction process into a highly reliable automated workflow, significantly improving the reconstruction success rate in complex urban scenarios. Through coordinate system transformation and georegistration integration, each reconstruction sub-project is registered from a local coordinate system to the Geocentric-Earth-Fixed (ECEF) coordinate system and then transformed to the Northeast-Sky Local Cartesian (ENU) coordinate system. This achieves high-precision alignment (meter-level or even sub-meter-level) between the reconstruction results and real geographic coordinates, resolving the scale uniformity issue between local reconstruction and the global coordinate system. This gives the reconstruction model realistic geospatial attributes, allowing direct fusion with GPS trajectories and GIS data, supporting surveying-grade accuracy requirements and engineering application needs. Through the synergy of the above four steps, efficient, high-precision, and robust 3D reconstruction of large-scale panoramic videos at the community level is achieved. This not only overcomes the computational resource bottleneck of long video reconstruction but also establishes an accurate mapping with the real world through georeferencing, ultimately forming a unified 3D spatial model covering the entire physical area, which meets the technical requirements for building a digital twin foundation in smart city construction.

[0110] The execution subject of the physical region three-dimensional spatial reconstruction method based on panoramic video in the specific embodiments of this disclosure can be an electronic device such as a mobile phone or computer.

[0111] Therefore, based on any of the above embodiments, this disclosure also provides an electronic device that can execute the three-dimensional spatial reconstruction method of physical regions based on panoramic video according to any of the embodiments described above.

[0112] Figure 9This is a schematic block diagram of an electronic device 1000 according to one embodiment of the present disclosure.

[0113] The hardware architecture of the electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnect buses and bridges, depending on the specific application of the hardware and overall design constraints. Bus 1100 connects various circuits, including one or more processors 1200, memory 1300, and / or hardware modules. Bus 1100 can also connect various other circuits 1400, such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.

[0114] Bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, only one connection line is used in this diagram, but this does not imply that there is only one bus or only one type of bus.

[0115] This disclosure also provides a readable storage medium storing a computer program that, when executed by a processor, is used to implement the methods described above. A "readable storage medium" can be any means capable of containing, storing, communicating, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples of a readable storage medium include: an electrical connection with one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable read-only memory (CDROM), etc.

[0116] This disclosure also provides a computer program product, the methods of which can be implemented wholly or partially through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented wholly or partially as a computer program product. A computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, all or part of the processes or functions of this disclosure are performed.

[0117] Computer programs or instructions can be stored in a readable storage medium or transferred from one readable storage medium to another. For example, a computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. A readable storage medium can be any accessible medium or a data storage device such as a server or data center that integrates one or more accessible media. The accessible medium can be magnetic media, such as floppy disks, hard disks, and magnetic tapes; optical media, such as digital video discs; or semiconductor media, such as solid-state drives. The computer-readable storage medium can be volatile or non-volatile, or it can include both types of storage media.

[0118] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0119] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0120] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0121] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0122] In the description of this specification, the references to terms such as "one embodiment / mode," "some embodiments / modes," "example," "specific example," or "some examples," etc., refer to specific features, structures, or characteristics described in connection with that embodiment / mode or example, which are included in at least one embodiment / mode or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment / mode or example. Moreover, the specific features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments / modes or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments / modes or examples described in this specification, as well as the features of different embodiments / modes or examples.

[0123] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0124] Those skilled in the art should understand that the above embodiments are merely for illustrating the present disclosure and are not intended to limit the scope of the disclosure. Those skilled in the art can make other changes or modifications based on the above disclosure, and these changes or modifications still fall within the scope of the present disclosure.

Claims

1. A method for three-dimensional spatial reconstruction of physical regions based on panoramic video, characterized in that, include: The panoramic video of the physical region is divided according to the preset time division step, resulting in multiple sets of reconstructed video frame sequences, each set of reconstructed video frame sequences corresponding to a reconstructed sub-project. For each reconstruction sub-project, static feature points are extracted from the reconstruction input image under the reconstruction sub-project based on feature point detection and mask filtering algorithms. Feature matching and geometric verification are performed on the static feature points to obtain the image matching relationship under the reconstruction sub-project. The reconstruction input image under the reconstruction sub-project is obtained by preprocessing the reconstruction video frame sequence corresponding to the reconstruction sub-project. The image matching relationship is processed by an incremental structure restoration algorithm based on progressive parameter tuning and multi-dimensional quality verification to obtain the local three-dimensional space reconstruction result corresponding to the reconstruction sub-item. The local 3D spatial reconstruction results of each reconstruction sub-project are integrated into a unified global geographic coordinate system through coordinate system transformation and geographic registration to obtain the target 3D spatial reconstruction result of the physical region.

2. The method according to claim 1, characterized in that, The incremental structure restoration algorithm, based on progressive parameter tuning and multi-dimensional quality verification, processes the image matching relationship to obtain the local three-dimensional spatial reconstruction result corresponding to the reconstruction sub-item, including: Obtain multiple preset parameter combinations, each parameter combination including: camera triangulation angle threshold and camera forward motion threshold; wherein, the camera triangulation angle threshold and the camera forward motion threshold are negatively correlated; Following the order of decreasing camera triangulation angle thresholds, the following incremental reconstruction steps are performed on the current parameter combination sequentially until multi-dimensional quality verification is passed or all parameter combinations have been traversed: From the image matching relationships, the reconstructed input image pair with the highest number of target feature point matching pairs and the smallest reprojection error is selected as the seed image pair, and the initial three-dimensional reconstruction is performed. Based on the incremental structure restoration strategy, new reconstructed input images and corresponding 3D spatial points are added step by step, and bundle adjustment optimization is performed after each addition. The bundle adjustment optimization includes: simultaneously optimizing the poses of all reconstructed cameras and the positions of the generated 3D spatial points to minimize the reprojection error. Specifically, after the initial 3D reconstruction and each addition of a new image, the multi-dimensional quality verification is performed. The multi-dimensional quality verification includes: model continuity verification, camera trajectory smoothness verification, and frame coverage verification. When the verification fails and there are untried parameter combinations, the incremental reconstruction step is re-executed to the next set of parameter combinations. When all verifications pass, the current reconstruction result is determined to be the local 3D spatial reconstruction result corresponding to the reconstruction sub-project.

3. The method according to claim 1, characterized in that, Before dividing the panoramic video of the physical region according to a preset time division step, the method further includes: The panoramic video is sampled according to a preset frame sampling step size to obtain a key video frame sequence; The step of dividing the panoramic video of the physical region according to a preset time division step to obtain multiple sets of reconstructed video frame sequences includes: grouping the key video frame sequences according to a preset time window length to obtain the multiple sets of reconstructed video frame sequences, each set of reconstructed video frame sequences containing a portion of the video frames in the key video frame sequences.

4. The method according to claim 1, characterized in that, The algorithm based on feature point detection and mask filtering extracts static feature points from the reconstructed input image under the reconstruction sub-project, including: For the reconstructed input image, the scale-invariant feature transform algorithm is used to detect key points in the reconstructed input image; Identify dynamic object regions and image regions that do not meet quality requirements in the reconstructed input image, and use them as masking regions in the reconstructed input image; The key points are filtered based on the masked area, and the key points located within the masked area are excluded. The remaining key points are used as static feature points in the reconstructed input image.

5. The method according to claim 1, characterized in that, The step of performing feature matching and geometric verification on the static feature points to obtain the image matching relationship under the reconstruction sub-project includes: Obtain the feature descriptor vectors of the static feature points, and determine the feature point matching pairs between each pair of reconstructed input images under the reconstructed sub-project by calculating the Euclidean distance between the feature descriptor vectors; Using the random sampling consensus algorithm and the essential matrix constraint method, geometric verification filtering is performed on the feature point matching pairs between each pair of reconstructed input images under the reconstruction sub-project to obtain the target feature point matching pairs between each pair of reconstructed input images under the reconstruction sub-project. Remove reconstructed input image pairs whose number of target feature point matching pairs under the reconstruction sub-project is less than a preset threshold, and construct image matching relationships under the reconstruction sub-project based on the target feature point matching pairs between the remaining reconstructed input image pairs.

6. The method according to claim 3, characterized in that, The method further includes: Obtain GPS trajectory information synchronously acquired with the panoramic video, the GPS trajectory information including multiple trajectory points and the timestamp and geographic coordinates corresponding to each trajectory point; The key video frame sequence is timestamped and aligned with the GPS trajectory information to generate a video frame geographic mapping table. The video frame geographic mapping table is used to record the mapping relationship between the frame number of each key video frame and the corresponding geographic coordinates. The timestamp matching and alignment includes: calculating the time deviation between the video frame timestamp and the trajectory point timestamp; when the time deviation is less than a preset time threshold, the alignment is confirmed to be successful.

7. The method according to claim 6, characterized in that, The process of integrating the local 3D spatial reconstruction results of each reconstruction sub-project into a unified global geographic coordinate system through coordinate system transformation and geographic registration to obtain the target 3D spatial reconstruction result of the physical region includes: For each reconstruction sub-project, based on the iterative nearest point algorithm, the local three-dimensional spatial reconstruction results of the reconstruction sub-project are registered from the local coordinate system to the geocentric coordinate system using the video frame geographic mapping table; The reconstruction results under the Earth-centered Earth-fixed coordinate system are transformed to the Northeast-Sky Local Cartesian coordinate system to obtain the three-dimensional spatial reconstruction results of the reconstruction sub-project under the global geographic coordinate system; wherein, the origin of the Northeast-Sky Local Cartesian coordinate system is determined based on the geographic center of the reconstruction sub-project, the East axis points eastward, the North axis points northward, and the Up axis points upward; The 3D spatial reconstruction results of each reconstruction sub-project under the global geographic coordinate system are merged to obtain the target 3D spatial reconstruction result of the physical region.

8. The method according to any one of claims 6-7, characterized in that, The method further includes: For a new panoramic video to be processed, new GPS trajectory information acquired synchronously with the new panoramic video is obtained, and based on the geographic coordinates in the new GPS trajectory information, candidate reconstruction sub-projects are selected from the reconstruction sub-projects in a four-level tree hierarchy. The four-level tree hierarchy includes: reconstruction task, panoramic video source, reconstruction sub-project, and camera pose. Based on the visual feature matching algorithm, the new panoramic video is matched with the reconstruction input image under the candidate reconstruction sub-project to determine the positioning result; The localization results are geometrically verified using the random sampling consensus algorithm and the essential matrix constraint method. After removing outliers, the camera pose estimate corresponding to the new panoramic video is obtained. The camera pose estimation is transformed from the local coordinate system to the northeast-sky local rectangular coordinate system, and the grid position information of the new panoramic video in the two-dimensional spatial grid is calculated, so that the camera pose corresponding to the new panoramic video is associated as a leaf node with the corresponding panoramic video source in the four-level tree hierarchy.

9. An electronic device, characterized in that, include: The memory stores execution instructions; as well as A processor that executes the execution instructions stored in the memory, causing the processor to perform the physical region three-dimensional spatial reconstruction method based on panoramic video according to any one of claims 1 to 8.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the physical region three-dimensional spatial reconstruction method based on panoramic video as described in any one of claims 1 to 8.