3D Gaussian Splitting-based single-view historic building rapid reconstruction method
Through the rapid reconstruction method of single-view ancient buildings based on 3D Gaussian Splatting, the 3D features of a single image are extracted using the U-Net network, the Gaussian point cloud is optimized and the Poisson surface reconstruction algorithm is combined, and the problems of inefficiency and low quality in the existing technology are solved, and the rapid reconstruction of high-quality ancient building models is achieved.
Patent Information
- Application Number
- CN202510281346.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The existing three-dimensional reconstruction methods require multiple images, resulting in complex and inefficient feature extraction, and are prone to distortion, artifacts and blur when dealing with real ancient architectural scenes, resulting in loss of details and poor quality.
A single-view ancient building rapid reconstruction method based on 3D Gaussian Splatting is adopted. A single two-dimensional image is extracted through an image translation network based on U-Net to generate a 3D Gaussian point cloud, and the Gaussian point distribution is optimized through regularization terms, and a triangle grid is generated combined with the Poisson surface reconstruction algorithm to bind Gaussian points to retain texture information.
It realizes the rapid reconstruction of high-quality ancient building models in a single image, avoids distortion, artifacts and blur, improves rendering quality and reconstruction efficiency, and is suitable for digital protection and research of ancient buildings.
Smart Images

Figure CN120219619A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to 3D reconstruction, and specifically to a fast single-view ancient building reconstruction method based on 3D Gaussian Splatting. Background Art
[0002] As an important part of cultural heritage, ancient buildings have extremely high historical, artistic and scientific values. However, due to factors such as the passage of time, natural disasters and human damage, many ancient buildings have been damaged or even disappeared. For example, the heavy rain in October 2021 caused very serious damage to many historical sites in Shanxi. According to incomplete statistics, more than 1,700 out of more than 28,000 ancient buildings in Shanxi Province were affected by heavy rain and floods. Therefore, the protection of above-ground ancient buildings is extremely urgent.
[0003] In order to protect and inherit these precious ancient building cultural heritages, digital reconstruction technology has emerged. Through digital reconstruction, the detailed information of ancient buildings can be permanently preserved, avoiding irreversible losses caused by natural disasters. It not only helps to protect cultural heritages, but also provides important reference for the subsequent repair and research of ancient buildings. At the same time, the digital ancient building model can be used for academic research and education to help historians, archaeologists and architects deeply analyze and understand ancient buildings. In addition, it can also be used for museum exhibitions and VR virtual reality experiences to enhance the public's awareness and interest in cultural heritages.
[0004] However, the existing reconstruction methods not only require multiple images, making the feature extraction process need to process a large amount of data and complex calculations, resulting in low reconstruction efficiency, but also when dealing with real ancient building scenes, the reconstructed models often have problems such as distortion, artifacts and blurring, resulting in the loss of details of ancient buildings and unable to accurately present their exquisite structures, leading to low quality of ancient building reconstruction models. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a fast single-view ancient building reconstruction method based on 3D Gaussian Splatting, which can quickly reconstruct a high-quality ancient building reconstruction model with only a single image.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions to implement:
[0007] A fast single-view ancient building reconstruction method based on 3D Gaussian Splatting, comprising the following steps:
[0008] Step 1: Input a single two-dimensional RGB ancient building image into the U-Net-based image translation network. After the encoder extracts the high-level features of the ancient building image, the decoder generates the mapped splatter image. Each pixel of the splatter image includes the parameters of a 3D Gaussian point;
[0009] Step 2: Update the parameters of the 3D Gaussian. The specific process is as follows:
[0010] Step 2.1: Set the two-dimensional coordinates of the pixel in the image plane as (μ1, μ2), and combine the depth value d of each pixel to calculate the central position μ of each Gaussian point in three-dimensional space, expressed as:
[0011]
[0012] In the formula, Δx, Δy, and Δz are offsets;
[0013] Step 2.2: Map the transparency S σ to the range of [0, 1] through the Sigmoid activation function, expressed as:
[0014] α = Sigmoid(S σ )
[0015] Step 2.3: Use the scale matrix and rotation matrix to calculate the covariance matrix Σ for describing the size and shape of the Gaussian point, expressed as:
[0016] Σ = R(q) · diag(e s ) 2 · R(q)T
[0017] In the formula, R(q) is the rotation matrix, indicating the direction of the Gaussian point in space; diag(e s ) represents the scale matrix, indicating the size of the Gaussian point and controlling the expansion degree of the Gaussian point in each direction;
[0018] Step 3: Input the updated parameters of the 3D Gaussian point into 3D Gaussian Splatting to generate a 3D Gaussian point cloud, and add a regularization term to the average reconstruction loss of 3D Gaussian Splatting to align the Gaussian point distribution with the scene surface, obtaining an optimized 3D Gaussian point cloud;
[0019] Step 4: Use the horizontal focus and normal vector of the optimized 3D Gaussian point cloud as the input of the Poisson surface reconstruction algorithm, and use the Poisson surface reconstruction algorithm to extract a triangular network from the optimized 3D Gaussian point cloud;
[0020] Step 5: First, calculate the central position of the Gaussian points through the barycentric coordinate interpolation of the triangle, so that the Gaussian points are evenly distributed on the triangle surface. Then, initialize the shape of the Gaussian points to a "flattened" Gaussian shape, which is expressed as:
[0021]
[0022] In the formula, σ1, σ2, and σ3 are all expansion coefficients, and both σ1 and σ2 are less than σ3. σ1 and σ2 are used to control the flattening in the plane direction, and σ3 is used to control the distribution range in the normal direction. R represents the regularization term;
[0023] Then, initialize the color and transparency of the Gaussian points to the color and transparency of the triangle where they are located to obtain new Gaussian points;
[0024] Step 6: Assign a set of new Gaussian points to each triangle of the triangular mesh, and bind the new Gaussian points to the surface of the triangular mesh to obtain a reconstructed ancient building model;
[0025] Step 7: Use the total loss function to perform consistency constraints on the image of the reconstructed ancient building and the real image.
[0026] Furthermore, the parameters of the 3D Gaussian points in Step 1 include: the transparency S for controlling the visibility of the Gaussian points σ , the central position μ of the Gaussian points in three-dimensional space, the covariance matrix Σ for describing the size and shape of the Gaussian points, and the color c of each Gaussian point.
[0027] Furthermore, the process of adding a regularization term to the average reconstruction loss of 3D Gaussian Splatting in Step 3 is as follows:
[0028] Step 3.1: Define a density function d(p) for representing the density at the spatial point p, which is expressed as:
[0029]
[0030] In the formula, μ g is the central position of the Gaussian point g, Σg is the covariance matrix of the Gaussian point g, and α g is the transparency of the Gaussian point g;
[0031] Step 3.2: Assume that the Gaussian points are evenly distributed on the surface of the scene. By calculating the distance from the spatial point p to the central position μ g and the normal vector n g , obtain the ideal density function which is expressed as:
[0032]
[0033] Wherein, s represents the scale factor of the Gaussian point, which is used to control the degree of expansion of the Gaussian distribution in the normal direction;
[0034] Step 3.3: The smaller the regularization term R is, the higher the alignment degree between the Gaussian point cloud and the scene surface. The regularization term R is expressed as:
[0035]
[0036] Wherein, P is the set of spatial points p;
[0037] Step 3.4: Add the regularization term R to the average reconstruction loss of 3D Gaussian Splatting.
[0038] Furthermore, the average reconstruction loss function is expressed as:
[0039]
[0040] Wherein, I is the input image, J is the ground truth image, π represents the camera pose, R'() represents the image rendered by the 3D Gaussian point cloud and π, and D is the public dataset ShapeNet.
[0041] Furthermore, the normal vector in Step 4 is expressed as:
[0042]
[0043] Wherein, represents the gradient of the density function, which is expressed as:
[0044]
[0045] Furthermore, the total loss function is expressed as:
[0046] L = L s + L LPIPS + R
[0047] Wherein, L S represents the average reconstruction loss, L LPIPS represents the LPIPS loss, and R represents the regularization term;
[0048] The LPIPS loss is used to optimize the visual quality of the image and is expressed as:
[0049] L LPIPS = ∑w l d l (I1, I2)
[0050] Wherein, w l represents the weight of the l-th layer, d lIndicates the L2 norm distance between images I1 and I2 in the feature space of the l-th layer.
[0051] Compared with the prior art, the present invention has the following technical effects:
[0052] 1) First, the present invention uses an image translation network based on U-Net to extract the parameter information required in the three-dimensional space of pixel points from the input two-dimensional ancient building image. It uses two-dimensional convolution instead of three-dimensional convolution to speed up the processing speed. Then, it generates Gaussian point clouds through 3D Gaussian Splatting and optimizes the Gaussian point clouds by adding a regularization term to align the Gaussian point distribution with the scene surface, preventing the disordered arrangement of Gaussian points, effectively preventing problems such as distortion, artifacts, and blurring, and improving the rendering quality. Next, the horizontal focus and normal vector of the optimized 3D Gaussian point clouds are used as the input of the Poisson surface reconstruction algorithm. The Poisson surface reconstruction algorithm generates a triangular mesh, and then by binding new Gaussian points to the mesh surface, more texture information can be retained, preventing the loss of details and further improving the rendering quality. Finally, a high-quality and realistic ancient building model is obtained. Therefore, the present invention has important significance and broad application prospects in protecting cultural heritage, promoting academic research, and driving technological innovation.
[0053] 2) The present invention only needs to input an ancient building image with any view to quickly generate a three-dimensional ancient building model, reducing the computational amount and complexity in the feature extraction process, improving the reconstruction efficiency, and being very suitable for the research on the digitalization of ancient buildings. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 : Flow chart of the present invention for reconstructing ancient buildings;
[0055] Figure 2 : Schematic diagram of 3D Gaussian Splatting of the present invention
[0056] Figure 3 : Pictures of the real ancient building image and the reconstructed model of the present invention from different perspectives. DETAILED DESCRIPTION OF THE INVENTION
[0057] The following further elaborates on the specific content of the present invention in detail in combination with embodiments.
[0058] As Figure 1 shown, a single-view ancient building rapid reconstruction method based on 3D Gaussian Splatting includes the following steps:
[0059] Step 1. As Figure 3A two-dimensional RGB ancient building image shown in (a) is input into an image-to-image network based on U-Net. After the encoder extracts the high-level features of the ancient building image, the decoder generates a mapped splatter image. Each pixel of this image includes the parameters of a 3D Gaussian point.
[0060] The two-dimensional RGB ancient building image is a picture of any perspective of the ancient building. The output image splatter image of the image-to-image network based on U-Net is a set of 3D Gaussian points. Each pixel of the output image represents the parameters of a Gaussian point. By rendering these Gaussian points, an ancient building model is reconstructed.
[0061] The parameters of the 3D Gaussian point include the transparency S for controlling the visibility of the Gaussian point σ , the central position μ of the Gaussian point in three-dimensional space, the covariance matrix Σ for describing the size and shape of the Gaussian point, and the color c of each Gaussian point.
[0062] Step 2: Update the parameters of the 3D Gaussian. The specific process is as follows:
[0063] Step 2.1: Set the two-dimensional coordinates of the pixel in the image plane as (μ1, μ2). Combining with the depth value d of each pixel point, calculate the central position μ of each Gaussian point in three-dimensional space, expressed as:
[0064]
[0065] In the formula, Δx, Δy, and Δz are offsets. By fine-tuning the three-dimensional position of the pixel point, the position of the object in three-dimensional space is made more accurate; the depth value d is equivalent to the distance from the camera to the pixel point, predicted by the image-to-image network based on U-Net. Each pixel point corresponds to a depth value d.
[0066] Step 2.2: Map the transparency S through the Sigmoid activation function σ to the range of [0, 1], expressed as:
[0067] α = Sigmoid(S σ )
[0068] Step 2.3: Use the scale matrix and rotation matrix to calculate the covariance matrix Σ for describing the size and shape of the Gaussian point, expressed as:
[0069] Σ = R(q) · diag(e s ) 2 · R(q)T
[0070] In the formula, R(q) is the rotation matrix, representing the direction of the Gaussian point in space; diag(es ) represents the scale matrix, which represents the size of the Gaussian points and controls the degree of expansion of the Gaussian points in each direction;
[0071] Step 3: First, input the parameters of the updated 3D Gaussian points into 3D Gaussian Splatting to generate a 3D Gaussian point cloud. The principle of 3D Gaussian Splatting is as Figure 2 shown. Then, add a regularization term to the average reconstruction loss of 3D Gaussian Splatting to align the Gaussian point distribution with the scene surface, prevent the Gaussian points from being disorderly arranged, and make the size and shape of the Gaussian points generated by the grid reasonable, obtaining an optimized 3D Gaussian point cloud. Among them, the process of adding a regularization term to the average reconstruction loss of 3D Gaussian Splatting is as follows:
[0072] Step 3.1: Define a density function d(p) representing the density at the spatial point p, expressed as:
[0073]
[0074] In the formula, μ g is the central position of the Gaussian point g, Σg is the covariance matrix of the Gaussian point g, and α g is the transparency of the Gaussian point g;
[0075] Step 3.2: Assume that the Gaussian points are evenly distributed on the surface of the scene. By calculating the distance from the spatial point p to the central position μ g and the normal vector n g , obtain the ideal density function expressed as:
[0076]
[0077] In the formula, s represents the scale factor of the Gaussian point, which is used to control the degree of expansion of the Gaussian distribution in the normal direction;
[0078] Step 3.3: In order to align the Gaussian point cloud with the scene surface, minimize the difference between the actual density and the ideal density. That is, the smaller the regularization term R, the higher the alignment degree of the Gaussian point cloud with the scene surface. The regularization term R is expressed as:
[0079]
[0080] In the formula, P is the set of spatial points p;
[0081] Step 3.4: Add the regularization term R to the average reconstruction loss of 3D Gaussian Splatting;
[0082] The average reconstruction loss function is expressed as:
[0083]
[0084] where I is the input image, J is the real image, π represents the camera pose, R'() represents the image rendered from the 3D Gaussian point cloud and π, and D is the public dataset ShapeNet;
[0085] Step 4: Use the horizontal focus and normal vector of the optimized 3D Gaussian point cloud as the input of the Poisson surface reconstruction algorithm, and extract a triangular network from the optimized 3D Gaussian point cloud by using the Poisson surface reconstruction algorithm. The normal vector is expressed as:
[0086]
[0087] where represents the gradient of the density function and is expressed as:
[0088]
[0089] Step 5: After the triangular network is generated, first calculate the central position of the Gaussian points by barycentric coordinate interpolation of the triangles, so that the Gaussian points are evenly distributed on the triangular surface. To adapt to the mesh surface, the Gaussian points are "flattened", that is, their covariance matrices are initialized to have two smaller normal directions and one larger normal direction, and the shape of the Gaussian points is initialized to a "flattened" Gaussian shape, which is expressed as:
[0090]
[0091] where σ1, σ2 and σ3 are all expansion coefficients, and both σ1 and σ2 are less than σ3. σ1 and σ2 are used to control the flattening in the plane direction, and σ3 is used to control the distribution range in the normal direction. R represents the regularization term;
[0092] Then initialize the color and transparency of the Gaussian points to the color and transparency of the triangles they are in to obtain new Gaussian points. By updating the shape, color and transparency, the Gaussian points can better match the input image and texture information;
[0093] Step 6: Assign a new set of Gaussian points to each triangle of the triangular mesh, bind the new Gaussian points to the triangular mesh surface, improve the rendering speed and rendering quality, and obtain the reconstructed ancient building model as shown in Figure 3 (b) to (d);
[0094] Step 7: Use the total loss function to perform consistency constraints on the image of the reconstructed ancient building and the real image. The total loss function is expressed as:
[0095] L = L s + L LPIPS + R
[0096] In the formula, L S represents the average reconstruction loss, which is used to minimize the pixel difference between the image generated by the network and the target image, ensuring that the reconstructed image is as similar as possible to the real image; L LPIPS represents the LPIPS loss, which is used to evaluate the perceptual similarity of the image, thereby optimizing the visual quality of the image and making the image more in line with human perception; R represents the regularization term, which is used to prevent the generation of unreasonable Gaussian points and prevent the model from overfitting;
[0097] The LPIPS loss is used and is expressed as:
[0098] L LPIPS = Σw l d l (I1, I2)
[0099] In the formula, w l represents the weight of the l-th layer, and d l represents the L2 norm distance between the images I1 and I2 in the feature space of the l-th layer.
Claims
1. A method for rapid reconstruction of ancient buildings from a single view based on 3D Gaussian Splatting, characterized in that: The steps include: Step 1: Input a single two-dimensional RGB ancient building image into the U-Net-based image translation network. After the encoder extracts the high-level features of the ancient building image, the decoder generates a mapped image splatter image. Each pixel point of the image splatter image includes the parameters of a 3D Gaussian point. Step 2: Update the parameters of the 3D Gaussian points. The specific process is as follows: Step 2.1, set the two-dimensional coordinates of the pixel in the image plane to (μ1, μ2), combine the depth value d of each pixel point, and calculate the center position μ of each Gaussian point in three-dimensional space, expressed as: Where Δx, Δy and Δz are the offsets; Step 2.2: Use the Sigmoid activation function to convert the transparency S σ Mapped to the range [0,1], expressed as: a=Simgoid(S σ ) Step 2.3, using the scale matrix and rotation matrix, calculate the covariance matrix Σ used to describe the size and shape of the Gaussian points, expressed as: Σ=R(q)·diag(e s ) 2 ·R(q)T Where R(q) is the rotation matrix, which indicates the direction of the Gaussian point in space; diag(e s ) represents the scale matrix, which indicates the size of the Gaussian points and controls the extent of expansion of the Gaussian points in all directions; Step 3: Input the updated parameters of the 3D Gaussian points into 3D Gaussian Splatting to generate a 3D Gaussian point cloud, and add a regularization term to the average reconstruction loss of 3D Gaussian Splatting to align the Gaussian point distribution with the scene surface to obtain an optimized 3D Gaussian point cloud; Step 4: Use the horizontal focus and normal vector of the optimized 3D Gaussian point cloud as input of the Poisson surface reconstruction algorithm, and use the Poisson surface reconstruction algorithm to extract a triangle network from the optimized 3D Gaussian point cloud; Step 5: First, calculate the center position of the Gaussian point by interpolating the barycentric coordinates of the triangle, so that the Gaussian points are evenly distributed on the surface of the triangle, and then initialize the shape of the Gaussian point to a "planarized" Gaussian shape, expressed as: Wherein, σ1, σ2 and σ3 are all expansion coefficients, and σ1 and σ2 are both smaller than σ3. σ1 and σ2 are used to control the flattening in the plane direction, σ3 is used to control the distribution range in the normal direction, and R represents the regularization term. Then the color and transparency of the Gaussian point are initialized to the color and transparency of the triangle where it is located, and a new Gaussian point is obtained; Step 6: assign a set of new Gaussian points to each triangle of the triangular mesh, bind the new Gaussian points to the surface of the triangular mesh, and obtain a reconstructed ancient building model; Step 7: Use the total loss function to constrain the consistency between the reconstructed image of the ancient building and the real image.
2. The single-view ancient building rapid reconstruction method based on 3D Gaussian Splatting according to claim 1 is characterized in that: The parameters of the 3D Gaussian points in step 1 include: transparency S for controlling the visibility of the Gaussian points σ , the central position μ of the Gaussian point in three-dimensional space, the covariance matrix Σ used to describe the size and shape of the Gaussian point, and the color c of each Gaussian point.
3. The single-view ancient building rapid reconstruction method based on 3D Gaussian Splatting according to claim 1 is characterized in that: The process of adding a regularization term to the average reconstruction loss of 3D Gaussian Splatting in step 3 is: Step 3.1, define a density function d(p) used to represent the density at a spatial point p, expressed as: In the formula, μ g is the center position of the Gaussian point g, Σg is the covariance matrix of the Gaussian point g, α g is the transparency of the Gaussian point g; Step 3.2: Assume that the Gaussian points are evenly distributed on the surface of the scene, and calculate the distance from the spatial point p to the center position μ g The distance and normal n g , we get the ideal density function It is expressed as: Where s represents the scale factor of the Gaussian point, which is used to control the expansion degree of the Gaussian distribution in the normal direction; Step 3.3: The smaller the regularization term R is, the higher the degree of alignment between the Gaussian point cloud and the scene surface is. The regularization term R is expressed as: Where P is the set of spatial points p; Step 3.
4. Add the regularization term R to the average reconstruction loss of 3D Gaussian Splatting.
4. The single-view ancient building rapid reconstruction method based on 3D Gaussian Splatting according to claim 1 or 3, characterized in that: The average reconstruction loss function is expressed as: Where I is the input image, J is the real image, π represents the camera pose, R'() represents the image rendered by 3D Gaussian point cloud and π, and D is the public dataset ShapeNet.
5. The single-view ancient building rapid reconstruction method based on 3D Gaussian Splatting according to claim 1 is characterized in that: The normal vector in step 4 is expressed as: In the formula, represents the gradient of the density function, expressed as:
6. The single-view ancient building rapid reconstruction method based on 3D Gaussian Splatting according to claim 4 is characterized in that: The total loss function is expressed as: L=L s +L LPIPS +R Where, L S represents the average reconstruction loss, L LPIPS represents LPIPS loss, R represents the regularization term; The LPIPS loss is used to optimize the visual quality of the image and is expressed as: L LPIPS =∑w l d l (I1,I2) In the formula, w l represents the weight of the lth layer, d l Represents the L2 norm distance between images I1 and I2 in the feature space of layer l.