Rapid temple landscape reconstruction method based on direct voxel grid optimization

Through direct voxel grid optimization technology, combined with multi-view image acquisition and adaptive resolution strategy, the problem of efficient, low-cost and high-precision reconstruction of temple buildings was solved, and high-quality 3D models were quickly generated, which are suitable for cultural heritage protection and virtual reality applications.

CN120612435AInactive Publication Date: 2025-09-09NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510765492.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing 3D reconstruction methods are inefficient and inaccurate when dealing with complex buildings such as temples, especially when dealing with complex details and changing lighting conditions, making it difficult to achieve fast and high-precision reconstruction.

Method used

A method based on direct voxel grid optimization is adopted, combined with a portable platform to collect multi-view image data. The density and color properties of the voxel grid are optimized through adaptive voxel density allocation and gradient descent. Combined with a multi-resolution optimization strategy, a high-precision three-dimensional model is generated, and material rendering and texture mapping are performed.

Benefits of technology

It achieves fast, efficient and low-cost 3D reconstruction, significantly improves the reconstruction speed and accuracy, can capture the detailed features of temple architecture, reduces the demand for computing resources, and the generated model has good compatibility and visual realism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612435A_ABST
    Figure CN120612435A_ABST
Patent Text Reader

Abstract

The invention discloses a temple landscape rapid reconstruction method based on direct voxel grid optimization, and belongs to the technical field of computer graphics and cultural heritage digital protection. According to the method, efficient three-dimensional reconstruction of a temple building is realized through multi-view image acquisition, preprocessing, camera pose estimation, direct voxel grid optimization and model post-processing. The method specifically comprises the following steps: acquiring a multi-angle high-resolution image by using an unmanned aerial vehicle and portable equipment; the image quality is improved through denoising, color correction and resolution optimization; estimating a camera pose based on a structural motion technology; constructing a voxel grid with adaptive resolution, optimizing voxel density and color attributes through gradient descent, and capturing geometric and material features of the temple; and finally, a high-precision three-dimensional model is generated and applied to Web3D and virtual reality platforms. The method obviously improves the efficiency while maintaining the reconstruction precision, is suitable for digital protection and display of complex temple buildings, and has the technical advantages of low cost and high compatibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and 3D reconstruction, specifically a method for rapid reconstruction of temple landscapes based on direct voxel grid optimization. This method falls within the interdisciplinary application of 3D reconstruction, digital cultural heritage preservation, and virtual reality technology. This method is particularly suitable for rapid, low-cost, and high-precision digital reconstruction of complex and detailed structures such as temples, pagodas, and religious statues. It has broad application in a variety of scenarios, including cultural heritage preservation, virtual tourism, digital museums, and educational displays. Background Art

[0002] With the rapid development of digital technology and growing awareness of cultural heritage preservation, the 3D digital reconstruction of historical buildings such as temples is gaining increasing attention. Traditional architectural surveying and 3D modeling methods face challenges such as low efficiency, low accuracy, and high costs when processing these structures. This is particularly true when capturing intricate details such as carvings, roof curves, and Buddha statues.

[0003] Currently, 3D reconstruction methods based on laser radar (LiDAR) scanning and traditional photogrammetry are widely used for digital preservation of buildings. While these methods can obtain relatively accurate geometric information, they are expensive, complex, and time-consuming when processing large-scale temple complexes. Furthermore, while traditional 3D reconstruction techniques based on multi-view stereo (MVS) reduce equipment costs, they often struggle to achieve satisfactory reconstruction results when dealing with complex architectural structures, varying materials, and lighting conditions. This is particularly true for the golden decorations, intricate carvings, and unusual surface materials commonly found in temple architecture.

[0004] In recent years, Neural Radiance Field (NeRF), a novel 3D reconstruction method, has achieved more accurate 3D scene reconstruction through deep learning technology. However, NeRF methods have high computational complexity, long training time, and demanding computing resources when dealing with large-scale buildings, making it difficult to meet the needs of fast reconstruction.

[0005] Recent research has shown that direct voxel grid optimization methods (such as PlenOxels) can significantly improve the efficiency of 3D reconstruction while maintaining high reconstruction quality by directly optimizing the voxel grid. This method avoids the complexity of neural network training and directly optimizes the density and color properties of voxels, greatly reducing computational time and resource consumption. However, when applied to buildings with complex structures and rich details such as temples, these methods still face challenges such as how to efficiently capture details, handle variable lighting conditions, and optimize large-scale data.

[0006] Therefore, combining modern image acquisition technology with direct voxel grid optimization methods to develop a fast, high-precision 3D reconstruction solution specifically for temple landscapes has become a key research topic in the field of digital cultural heritage preservation. This paper addresses this need by proposing a fast temple landscape reconstruction method based on direct voxel grid optimization, aiming to provide an efficient, low-cost, and high-precision solution for the digital reconstruction of temple architecture. Summary of the Invention

[0007] The present invention aims to provide a method for rapid reconstruction of temple landscapes based on direct voxel grid optimization, addressing the low efficiency and accuracy of existing 3D reconstruction methods for complex structures such as temples. This method combines multi-view image data collected by portable platforms (such as mobile phones and drones) with direct voxel grid optimization technology for 3D reconstruction, rapidly generating high-precision 3D models of temple landscapes. Compared to traditional reconstruction techniques, this method offers significant speed and cost advantages while preserving the detailed features of temple architecture, ensuring the authenticity and accuracy of the reconstruction results.

[0008] In order to solve the above technical problems, the present invention provides a method for rapid reconstruction of temple landscape based on direct voxel grid optimization, comprising the following steps:

[0009] In the first step, a drone equipped with a high-resolution camera captured panoramic images of the temple from various altitudes and angles, ensuring coverage of various lighting and climatic conditions. Simultaneously, a portable device was used to capture detailed images of specific architectural elements of the temple, such as statues, carvings, and roofs. GPS coordinates and flight parameters were recorded during image acquisition to support subsequent camera pose estimation.

[0010] The second step is to improve image quality by denoising, color correcting, and optimizing resolution. Image processing techniques are used to standardize the lighting conditions of images from different viewpoints, ensuring consistency during the voxel optimization process and providing high-quality input data for subsequent reconstruction.

[0011] The third step is to construct a preliminary voxel grid based on the preprocessed image, dividing the 3D space into regular voxel units. An adaptive voxel density allocation strategy is used to use higher-density voxel representations in complex areas of the temple architecture (such as carvings and roof curves) to capture detailed features.

[0012] In the fourth step, the processed image and camera pose information are fed into a direct voxel optimization model. Using algorithms like gradient descent, the density and color properties of the voxel grid are directly optimized to achieve high-precision modeling of the temple's geometry and material properties. A multi-resolution optimization strategy is employed, initially optimizing a coarse voxel grid and then gradually refining it, improving reconstruction efficiency and quality.

[0013] The fifth step is to convert the optimized voxel model into a standard 3D mesh model, and perform texture mapping and material optimization to enhance the model's visual quality. Specialized material rendering techniques are used to enhance the realism of the reconstructed model, specifically targeting the unique textures of the temple architecture (such as gold decorations and stone surfaces).

[0014] The sixth step is to integrate the generated 3D model with data from other sources (such as digital elevation models and historical surveying data) to further enhance the comprehensiveness and accuracy of the temple modeling. This multi-source data integration will provide a more complete digital resource for cultural heritage preservation, virtual tourism, and educational presentations.

[0015] The method for rapid reconstruction of temple landscapes based on direct voxel grid optimization adopted in the present invention is an innovative application in the field of digital reconstruction. The core idea of ​​the algorithm is derived from the direct voxel optimization technology in the field of computer vision, combined with the actual needs of temple building reconstruction, especially when dealing with large-scale, detailed and complex architectural objects (such as temple bodies, pagodas, sculptures, etc.). Direct voxel grid optimization can efficiently capture the geometric and material information of the building, thereby generating a high-precision three-dimensional model. Compared with traditional reconstruction methods and neural network-based methods, the direct voxel optimization method proposed in the present invention can achieve faster and more efficient reconstruction under multi-view image input, breaking through the limitations of traditional reconstruction methods in accuracy, speed and computing resource consumption.

[0016] Furthermore, considering the shortcomings of existing reconstruction methods when dealing with complex temple architecture, this paper introduces an adaptive resolution strategy based on direct voxel meshing, which solves the problem of accurately reconstructing detailed areas of temple architecture (such as carvings and roof curves). Traditional methods often require extensive computing resources or long training times when dealing with complex architecture, resulting in low reconstruction efficiency and high costs. To this end, this paper achieves a more efficient 3D reconstruction process through optimized voxel representation, direct parameter optimization, and adaptive resolution allocation, significantly reducing computing resource requirements and reconstruction time.

[0017] The specific technical solutions are as follows:

[0018] A method for rapid reconstruction of temple landscapes based on direct voxel grid optimization includes the following steps:

[0019] (1) Using a drone equipped with a high-resolution camera to capture panoramic images of the temple from multiple angles, and using a portable device to capture detailed images of specific architectural elements, recording GPS coordinates and flight parameters;

[0020] (2) De-noising, color correction, resolution optimization, and illumination balancing are performed on the collected images, and high-quality images are selected as input data;

[0021] (3) Estimate the camera pose based on structure-from-motion technology, generate sparse point clouds and establish a camera projection model;

[0022] (4) Construct the initial voxel grid and use the gradient descent algorithm to optimize the voxel density and color attributes to minimize the color reconstruction loss function:

[0023]

[0024] Among them, C(r) is the real image pixel color, is the voxel prediction value; r represents the light from the camera center passing through the image pixel, and each light corresponds to a pixel point in the image;

[0025] (5) Dynamically adjust voxel resolution based on regional complexity, using higher-density voxels in sculptured and roof curve areas and lower-density voxels in flat areas;

[0026] (6) Extracting a triangular mesh model from the optimized voxel grid using the Marching Cubes algorithm to generate the surface structure;

[0027] (7) Texture mapping and material optimization are performed on the extracted triangular mesh to generate the final three-dimensional model, which is then exported to a format compatible with Web3D and virtual reality platforms.

[0028] Preferably, the denoising process in step (2) uses a Gaussian filter, and the filtering formula is:

[0029]

[0030] Among them, I filtered (x, y) represents the pixel value at the coordinate (x, y) of the image after Gaussian filtering; x, y represent the coordinate position of the target pixel in the image; Indicates the summation operation of all pixel positions within the Gaussian kernel window; G(i,j) represents the weight value of the Gaussian kernel function at position (i,j), which is defined as G(i,j) = (1 / 2πσ 2 )e-(i 2 +j 2 ) / 2σ 2 ; I(xi,yj) represents the pixel value of the original image at the coordinate (xi,yj); i,j represent the offset relative to the center pixel, indicating the relative position within the Gaussian kernel window; σ represents the standard deviation parameter of the Gaussian kernel, which controls the smoothness of the filter; the standard deviation of the Gaussian kernel is 0.5, and the resolution is uniformly adjusted to 2048×1536 pixels.

[0031] Preferably, a multi-resolution optimization strategy is adopted in step (4), including:

[0032] (4a) The coarse optimization stage uses a 128 × 128 × 128 voxel grid;

[0033] (4b) The medium optimization stage is increased to 256×256×256;

[0034] (4c) The fine optimization stage uses 512×512×512 resolution in the key areas.

[0035] Preferably, the regional complexity measurement formula in step (5) is:

[0036] σ adaptive (x)=σ base (x)·(1+α·D(x))

[0037] Where D(x) is the regional complexity and α is the adjustment parameter.

[0038] Preferably, the reprojection error of the camera pose estimation in step (3) is controlled within 1.2 pixels.

[0039] Preferably, the Marching Cubes algorithm in step (6) includes:

[0040] (6a) Based on the voxel grid optimized in step (4), an isosurface is determined according to a density threshold;

[0041] (6b) Calculating the vertex positions of the isosurface by interpolation and generating triangular facets;

[0042] (6c) Topology optimization of triangular facets is performed to reduce the number of redundant facets and generate a lightweight surface mesh;

[0043] (6d) Output the optimized surface mesh to step (7) for texture mapping and material rendering.

[0044] Preferably, the topology optimization in step (c) includes edge collapsing and vertex merging operations, which simplifies the number of model faces from 5 million to less than 500,000.

[0045] Preferably, the material optimization in step (7) includes:

[0046] (7a) Use a specular reflection map for the gold decoration;

[0047] (7b) Use normal maps to enhance details on stone surfaces.

[0048] The main innovations of the present invention are embodied in the following three aspects:

[0049] 1. Image Data Acquisition and Preprocessing: A specialized image quality optimization strategy addresses the complex lighting and diverse material characteristics of temple architecture, addressing lighting inconsistencies and material variations between images from different viewpoints. The algorithm utilizes techniques such as image denoising, lighting equalization, and resolution optimization to ensure high input data quality, providing a reliable foundation for subsequent voxel optimization. Targeted processing methods are particularly effective for temple-specific materials such as gold decorations and reflective surfaces, effectively improving reconstruction quality.

[0050] 2. Direct Voxel Grid Optimization: Compared to indirect representation methods that rely on neural networks, this method directly optimizes voxel grid parameters, bypassing the complexity and high computational cost of neural network training. By directly optimizing the density and color properties of voxels through gradient descent, the reconstruction process is more efficient. This significantly improves reconstruction speed while maintaining high-quality results, particularly for large-scale temple complexes. Experiments have shown that compared to NeRF-based methods, this method can reduce reconstruction time by 5-10 times under the same hardware conditions.

[0051] 3. Adaptive Resolution and Multi-Scale Optimization: To effectively capture the detailed features of the temple architecture, this paper proposes an adaptive voxel resolution strategy that dynamically allocates computing resources based on the complexity of the region. This strategy uses higher-density voxels in detail-rich areas (such as carvings and Buddha faces), while using lower-density voxels in flat areas, achieving efficient utilization of computing resources. Furthermore, a multi-scale optimization strategy is employed to gradually refine the voxel grid from coarse to fine, further improving reconstruction efficiency and quality.

[0052] The method presented in this paper is highly capable of capturing details in temple architectural reconstruction, rapidly generating high-precision 3D models while reducing computing resource requirements. All steps can be performed on a standard computer, eliminating the need for specialized high-performance computing equipment, significantly reducing the technical barriers and costs of digital temple preservation. Furthermore, the 3D models generated by this method possess excellent compatibility and can be easily integrated into various virtual reality platforms and web applications, providing strong technical support for the digital preservation and display of cultural heritage. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The present invention will be further described below with reference to the accompanying drawings.

[0054] Figure 1 is the original input image.

[0055] Figure 2 The generated white model (left) and the final model (right).

[0056] Figure 3 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0057] To verify the feasibility and effectiveness of our proposed method for rapid temple landscape reconstruction based on direct voxel mesh optimization, we conducted a complete 3D reconstruction of the Guangjiao Zen Temple on Langshan Mountain in Nantong, China. As a typical temple building, Guangjiao Zen Temple, with its complex structure, rich carvings, and diverse materials, holds significant cultural and historical value, making it an ideal location for temple landscape reconstruction.

[0058] Step 1: Image data acquisition

[0059] A DJI Air 3 drone was used to capture multi-view images of Guangjiao Temple, ensuring comprehensive coverage. Handheld devices were used to capture close-up, high-resolution images of the temple's distinctive architectural elements, such as the main hall, pagoda, and stone carvings. Each image's pose data (such as GPS coordinates, pitch, heading, and tilt) was also recorded for camera pose estimation. A total of approximately 500 images of the temple were captured from various angles and lighting conditions, at a resolution of 4K (3840 × 2160 pixels), ensuring sufficient image detail and coverage.

[0060] Step 2: Image Preprocessing

[0061] For the collected images, a series of preprocessing steps will be performed to improve the modeling quality:

[0062] 1. Noise removal. Use a Gaussian filter to smooth the image to remove the noise generated during the shooting process. The processing formula is as follows:

[0063]

[0064] Where G(i,j) is the Gaussian kernel function. The Gaussian filter with the standard deviation set to 0.5 effectively removes noise while preserving details.

[0065] 2. Color Correction and Brightness Balance: Images are color corrected and brightness balanced to ensure consistent hue and brightness across images taken at different times and under different lighting conditions. Specifically, golden decorations and red architecture in temples are enhanced to preserve their unique visual characteristics.

[0066] Resolution optimization: The resolution of all images is uniformly adjusted to 2048×1536 pixels to balance computational efficiency and image details.

[0067] Image screening: Based on image quality, coverage, and angular distribution, 350 high-quality images were screened from 500 original images for subsequent reconstruction.

[0068] Step 3: Camera pose estimation

[0069] Based on the preprocessed image, the camera pose information is estimated using the Structure from Motion (SfM) technique:

[0070] 1. Use the SIFT (Scale Invariant Feature Transform) algorithm to extract feature points in the image and establish feature matching between different images.

[0071] 2. Estimate the camera's intrinsic parameter matrix and extrinsic parameter (position and attitude) information through feature matching relationships. The camera projection model is mainly constructed through the Structure-from-Motion (SfM) method: for each image, the three-dimensional to two-dimensional projection relationship is modeled based on the camera parameters, and then the camera intrinsic and extrinsic parameters corresponding to each image are estimated using matching feature points. Among them, the camera intrinsic parameters include focal length and principal point position, and the extrinsic parameters include the camera's rotation matrix and translation vector. This parameter estimation process is based on the detection, matching and geometric verification of image feature points, and is inferred through multi-view geometric relationships. Let the three-dimensional point P = [XYZ] T The homogeneous coordinates of are expressed as:

[0072]

[0073] The homogeneous pixel coordinates on the corresponding image plane are:

[0074]

[0075] Where: K∈R 3×3 is the camera intrinsic parameter matrix:

[0076]

[0077] [R|t]∈R 3×4 is the external parameter matrix, R is the rotation matrix, and t is the camera's translation vector, which represents the position offset of the camera in the world coordinate system. R is the camera's rotation matrix, which transforms the world coordinate system to the camera coordinate system. x f y is the focal length of the camera in the horizontal and vertical directions, in pixels. x c y The position of the camera's principal point (optical center) in the image coordinate system, in pixels. is the camera's two-dimensional homogeneous image coordinate point.

[0078]

[0079] Step 4: Voxel Mesh Construction and Optimization

[0080] Based on the estimated camera pose and image data, we construct and optimize a voxel grid representation of the temple:

[0081] Initialization: Divide the space where the temple is located into a 256×256×256 uniform voxel grid, where each voxel contains density and RGB color attributes.

[0082] Direct voxel grid optimization: Construct the input: image data + camera pose. Initialize a voxel grid in 3D space, where each voxel has density σ and color c attributes. For each pixel ray, calculate the color value according to the voxel grid. The optimization goal is to minimize the following color reconstruction loss function:

[0083]

[0084] Among them, C(r) is the real image pixel color, Predict values ​​for the voxel grid.

[0085] The light transport equation in a voxel grid can be expressed as:

[0086]

[0087] Among them, T i =exp(-σ i δ i )) is the cumulative transparency, σ i is the voxel density, δ i is the sampling interval, c i is the voxel color.

[0088] Adaptive resolution allocation: Based on the initial reconstruction results, complex areas in the temple (such as carvings and roof curves) are identified and higher-resolution voxel representations (up to 512×512×512) are used in these areas, while relatively flat areas are represented at a lower resolution to optimize the allocation of computing resources.

[0089] Model optimization and detail enhancement: Adaptive voxel resolution strategy is adopted to use higher resolution voxels in complex areas of the temple:

[0090] σ adaptive (x)=σ base (x)·(1+α·D(x))

[0091] Where D(x) is the regional complexity measure and α is the adjustment parameter.

[0092] The total loss function combines geometric accuracy and texture fidelity:

[0093] L total =L color +λ1L smooth +λ2Ldetail

[0094] Where: L smooth : voxel grid smoothness loss; L detail : detail preservation loss based on image gradient; λ1,λ2: are balance parameters.

[0095] The optimization process is divided into three stages:

[0096] Coarse optimization stage: Use a lower resolution voxel grid (128×128×128) to quickly obtain the general structure of the scene.

[0097] Medium optimization stage: Increase the resolution to 256×256×256 and further refine the scene representation.

[0098] Fine-tuning stage: High-resolution voxels of 512×512×512 are used in key areas to capture the detailed features of the temple.

[0099] Loss function design: Taking into account the three aspects of color reconstruction error, voxel smoothness and detail preservation, the following loss function is designed:

[0100] L total =L color +0.1×L smooth +0.5×L detail

[0101] Among them, L color is the color reconstruction loss, L smooth is the voxel smoothness loss, L detail A detail-preserving loss based on image gradients.

[0102] The optimization process was carried out on a workstation equipped with an NVIDIA RTX 4090 graphics card and took a total of about 2 hours, which is significantly more efficient than the traditional NeRF method (which usually takes 10-20 hours).

[0103] Step 5: Model post-processing and detail enhancement

[0104] After voxel optimization, we post-processed the model to improve visual quality and usability:

[0105] Mesh Extraction: Extracts triangular mesh models from the optimized voxel grid using an improved Marching Cubes algorithm, generating surface models that can be rendered and interacted with.

[0106] Texture Mapping: Mapping optimized voxel color attributes onto triangular meshes to generate high-quality texture maps. For the temple's unique materials (such as the golden roof and stone carvings), specialized material rendering technology is used to enhance the visual effect.

[0107] Detail Enhancement: For important architectural details (such as Buddha statues and sculptures), normal mapping technology is used to enhance the performance of surface details and improve the visual quality of the model.

[0108] Model simplification: To meet the needs of different application scenarios, multiple simplified model versions with different numbers of faces are generated, ranging from a high-precision version (approximately 5 million faces) to a lightweight version (approximately 500,000 faces), meeting the performance requirements of different platforms.

[0109] Step 6: Model Validation and Evaluation

[0110] To evaluate the quality of the reconstructed model, we performed the following validation steps:

[0111] Output and Verification: Convert the optimized voxel mesh into a triangular mesh model (Mesh) and output it in a standard format (.obj / .ply, etc.) for web display or VR platform deployment. Use reference data (such as laser scanning point cloud) to evaluate reconstruction accuracy, and calculate the error using point-to-surface distance (P2S):

[0112]

[0113] Among them, d(p i ,S) is point p i The shortest distance to the reference surface S.

[0114] For specific architectural elements of the temple (such as Buddha statues and carvings), the structural similarity index (SSIM) can also be used to evaluate the visual quality of the reconstruction:

[0115]

[0116] Where x and y are the rendered image and the reference image respectively.

[0117] Geometric Accuracy Assessment: Using point cloud data acquired through laser scanning as a reference, the geometric error between the reconstructed model and the reference data was calculated. The results showed that the average error was within 3 cm in the main building area and within 5 mm in detailed areas (such as engravings), meeting the expected accuracy requirements.

[0118] Visual Quality Assessment: The visual quality of the reconstruction was evaluated by rendering the model from a perspective never used for reconstruction and comparing it with the actual photos taken. Using the Structural Similarity Index (SSIM) as the evaluation metric, the average SSIM value reached 0.85, indicating that the reconstructed model has good visual fidelity.

[0119] User Experience Testing: Cultural heritage preservation experts and regular users were invited to evaluate the reconstructed model, rating it based on details, overall realism, and interactive fluency. The test results showed that both experts and users gave positive feedback on the overall quality of the model, particularly expressing satisfaction with the preservation of details and the rendering of materials.

[0120] Step 7: Application Display

[0121] Finally, we applied the reconstructed temple model to the following scene:

[0122] Web3D Display: The model was converted to glTF format, and a web-based 3D display platform was developed using Three.js, allowing users to interactively explore various parts of the temple through a browser. The platform supports features such as detailed zooming, free roaming, and information annotation, providing an immersive digital tour experience.

[0123] Virtual Reality Experience: The model was imported into the Unity engine, and a VR version of the temple tour app was developed, which supports an immersive experience through VR headsets (such as the Oculus Quest 2). Users can move freely in the virtual environment and observe the temple's architectural details and artistic features up close.

[0124] Educational Applications: Based on the reconstructed model, an interactive application for cultural heritage education has been developed, including analysis of temple architectural structures, displays of historical changes, and introductions to cultural background, providing a new digital means for cultural heritage.

[0125] Implementation effect and application value

[0126] The 3D reconstruction of Guangjiao Temple using this method successfully achieved high-precision and efficient digital preservation of the temple. Compared to traditional methods, this method increases reconstruction speed by 5-10 times while maintaining high reconstruction quality. The reconstructed 3D model not only accurately reflects the temple's architectural structure and artistic features, but also provides an important digital resource for the preservation, research, and dissemination of cultural heritage.

[0127] The implementation effect of this method shows that the rapid reconstruction method of temple landscape based on direct voxel grid optimization has significant technical advantages and application value. It can be widely used in cultural heritage protection, virtual tourism, digital museums and educational displays, etc., providing strong support for promoting cultural inheritance and innovation.

Claims

1. A method for rapid reconstruction of temple landscapes based on direct voxel grid optimization, characterized in that: The following steps are involved: (1) Using a drone equipped with a high-resolution camera to capture panoramic images of the temple from multiple angles, and using a portable device to capture detailed images of specific architectural elements, recording GPS coordinates and flight parameters; (2) De-noising, color correction, resolution optimization, and illumination balancing are performed on the collected images, and high-quality images are selected as input data; (3) Estimate the camera pose based on structure-from-motion technology, generate sparse point clouds and establish a camera projection model; (4) Construct the initial voxel grid and use the gradient descent algorithm to optimize the voxel density and color attributes to minimize the color reconstruction loss function: Among them, C(r) is the real image pixel color, is the voxel prediction value; r represents the light from the camera center passing through the image pixel, and each light corresponds to a pixel point in the image; (5) Dynamically adjust voxel resolution based on regional complexity, using higher-density voxels in sculptured and roof curve areas and lower-density voxels in flat areas; (6) Extracting a triangular mesh model from the optimized voxel grid using the Marching Cubes algorithm to generate the surface structure; (7) Texture mapping and material optimization are performed on the extracted triangular mesh to generate the final three-dimensional model, which is then exported to a format compatible with Web3D and virtual reality platforms.

2. The method according to claim 1, characterized in that The denoising process in step (2) uses a Gaussian filter, and the filtering formula is: Among them, I filtered (x, y) represents the pixel value at the coordinate (x, y) of the image after Gaussian filtering; x, y represent the coordinate position of the target pixel in the image; Indicates the summation operation of all pixel positions within the Gaussian kernel window; G(i,j) represents the weight value of the Gaussian kernel function at position (i,j), which is defined as G(i,j) = (1 / 2πσ 2 )e-(i 2 +j 2 ) / 2σ 2 ; I(xi,yj) represents the pixel value of the original image at the coordinate (xi,yj); i,j represent the offset relative to the center pixel, indicating the relative position within the Gaussian kernel window; σ represents the standard deviation parameter of the Gaussian kernel, which controls the smoothness of the filter; the standard deviation of the Gaussian kernel is 0.5, and the resolution is uniformly adjusted to 2048×1536 pixels.

3. The method according to claim 1, characterized in that In step (4), a multi-resolution optimization strategy is adopted, including: (4a) The coarse optimization stage uses a 128 × 128 × 128 voxel grid; (4b) The medium optimization stage is increased to 256×256×256; (4c) The fine optimization stage uses 512×512×512 resolution in the key areas.

4. The method according to claim 1, wherein The regional complexity measurement formula in step (5) is: s adaptive (x)=σ base (x)·(1+α·D(x)) Where D(x) is the regional complexity and α is the adjustment parameter.

5. The method according to claim 1, wherein The reprojection error of the camera pose estimation in step (3) is controlled within 1.2 pixels.

6. The method according to claim 1, characterized in that The Marching Cubes algorithm in step (6) includes: (6a) Based on the voxel grid optimized in step (4), an isosurface is determined according to a density threshold; (6b) Calculating the vertex positions of the isosurface by interpolation and generating triangular facets; (6c) Topology optimization of triangular facets is performed to reduce the number of redundant facets and generate a lightweight surface mesh; (6d) Output the optimized surface mesh to step (7) for texture mapping and material rendering.

7. The method according to claim 6, characterized in that The topology optimization in step (c) includes edge collapsing and vertex merging operations, which simplifies the number of model faces from 5 million to less than 500,000.

8. The method according to claim 1, characterized in that The material optimization in step (7) includes: (7a) Use a specular reflection map for the gold decoration; (7b) Use normal maps to enhance details on stone surfaces.