Three-dimensional reconstruction method for high-precision urban updated ancient building

By improving the 3D Gaussian splashing method, increasing the deployment of Gaussian ellipsoids and refining the network iterative optimization, the existing technology cannot meet the requirements of high-precision three-dimensional reconstruction, and achieving high-precision three-dimensional reconstruction and efficiency improvement in complex areas.

CN120070759AInactive Publication Date: 2025-05-30BEIJING GUANGAN LIGHTING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510150728.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing 3D Gaussian splashing method cannot meet the requirements of high-precision three-dimensional reconstruction, especially in areas with complex geometric structures.

Method used

Through two shear and segmentation optimizations, the 3D Gaussian splattering method is improved, the number of Gaussian ellipsoids deployed in complex areas of geometric structures is increased, and the Gaussian ellipsoids are optimized through two refinement network iterations to generate a more accurate three-dimensional model.

Benefits of technology

High-precision three-dimensional reconstruction in geometric complex areas is realized, redundant Gaussian ellipsoids are reduced, reconstruction efficiency and accuracy are improved, and a three-dimensional reconstruction method with higher accuracy and speed is provided for urban renewal of ancient buildings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070759A_ABST
    Figure CN120070759A_ABST
Patent Text Reader

Abstract

The invention discloses a high-precision three-dimensional reconstruction method for an urban updated ancient building, and the method comprises the steps: obtaining image data and corresponding camera parameters of the urban ancient building, inputting the image data as an input image and the corresponding camera parameters into a reconstruction model for three-dimensional reconstruction and rendering, and obtaining a three-dimensional model of the urban ancient building; wherein the reconstruction model comprises a feature extraction module, a first refinement network and a second refinement network, the feature extraction module is used for carrying out feature extraction according to the image data and the corresponding camera parameters to obtain a first feature map, carrying out depth estimation on the first feature map to obtain a first Gaussian ellipsoid, and carrying out depth estimation on the second feature map to obtain a second Gaussian ellipsoid; and performing iterative refinement adjustment on the first Gaussian ellipsoid through the first refinement network to obtain a second Gaussian ellipsoid, performing iterative refinement adjustment on the second Gaussian ellipsoid through the second refinement network to obtain a third Gaussian ellipsoid, and rendering the third Gaussian ellipsoid to obtain the three-dimensional model of the urban ancient building.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a three-dimensional reconstruction method for ancient buildings in urban renewal with high precision. Background Art

[0002] In the field of applying ancient buildings in urban renewal, on the one hand, the number of ancient buildings is huge, and it is necessary to significantly reduce the modeling cost. Using general equipment to take photos as input is an optional method. On the other hand, these buildings are the crystallization of ancient Chinese wisdom, and subsequent digital models need to be repaired and protected. Therefore, a three-dimensional reconstruction method with high precision and high reliability is required.

[0003] The method of 3D Gaussian Splitting for rendering only through a small number of input images has good rendering effects and rendering efficiency. Compared with the neural radiance field method, 3D Gaussian Splitting is faster and can achieve real-time rendering of at least 30 FPS. However, the current 3D Gaussian Splitting still cannot meet the requirements for high-precision three-dimensional reconstruction. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a three-dimensional reconstruction method for ancient buildings in urban renewal with high precision to solve the problems existing in the above prior art.

[0005] To achieve the above object, the present invention provides a three-dimensional reconstruction method for ancient buildings in urban renewal with high precision, including:

[0006] Obtaining image data of ancient buildings in the city and corresponding camera parameters,

[0007] Inputting the image data as input images and the corresponding camera parameters into a reconstruction model for three-dimensional reconstruction and rendering to obtain a three-dimensional model of the ancient buildings in the city;

[0008] The reconstruction model includes a feature extraction module, a first refinement network, and a second refinement network. The feature extraction module is used to extract features according to the image data and the corresponding camera parameters, obtain a first feature map, estimate the depth of the first feature map to obtain a first Gaussian ellipsoid, iteratively refine and adjust the first Gaussian ellipsoid through the first refinement network to obtain a second Gaussian ellipsoid, iteratively refine and adjust the second Gaussian ellipsoid through the second refinement network to obtain a third Gaussian ellipsoid, and render the third Gaussian ellipsoid to obtain a three-dimensional model of the ancient buildings in the city.

[0009] Optionally, the reconstruction model further includes an input module, and the input module inputs the image data and the camera parameters into the reconstruction model in sequence according to the camera parameters.

[0010] Optionally, the reconstruction model further includes an output module, where the output module is used to rasterize and render the third Gaussian ellipsoid to obtain a three-dimensional model of the ancient urban architecture.

[0011] Optionally, the feature extraction module includes a parallel structure, a first MLP, and a multi-view depth estimation network connected in sequence, where the parallel structure includes a residual network and a cross-attention network connected in parallel. The image data is processed and fused through the parallel structure to obtain a fused feature map. The fused feature map is further fused through the first MLP to obtain a first feature map. The first feature map and camera parameters are used for depth estimation through the multi-view depth estimation network, and the depth estimation result is back-projected to obtain a first Gaussian ellipsoid, where the first Gaussian ellipsoid includes the center point position of the Gaussian ellipsoid.

[0012] Optionally, the first Gaussian ellipsoid is iteratively refined and adjusted several times through a first refinement network. In each iteration process, the Gaussian ellipsoid generated in the previous iteration process is used as the input data in the current iteration process. The input data is refined and adjusted through the first refinement network to obtain the Gaussian ellipsoid generated in the current iteration process. The first Gaussian ellipsoid is used as the input data in the first iteration process, and the Gaussian ellipsoid generated in the last iteration process is used as the second Gaussian ellipsoid.

[0013] Optionally, in the first refinement network, the first Gaussian ellipsoid is refined and adjusted through first threshold data, where the first threshold data includes a first high score threshold, a first low score threshold, a first opacity threshold, and a first gradient threshold. The process of the first threshold data includes:

[0014]

[0015] Among them, MLP(·) represents the processing of a multi-layer perceptron, DA(·) represents the processing of an attention mechanism, τ h and τ l respectively represent the first high score threshold and the first low score threshold, τ α represents the first opacity threshold, represents the first gradient threshold, ξ i represents the contribution factor, represents the score matrix, represents the saliency map, U(·) represents the projection operation, p sample represents the variable center point randomly sampled from the Gaussian ellipsoid, C i represents the input camera pose parameters corresponding to the saliency map, i represents the input image serial number, and N I represents the total number of input images;

[0016] Among them, the saliency map is obtained by processing the first feature map through a multi-layer perceptron and a softmax function, and the center point of the first Gaussian ellipsoid is projected onto the saliency map to obtain a score matrix.

[0017] Optionally, the process of refining and adjusting the first Gaussian ellipsoid includes:

[0018] Calculating the first Gaussian ellipsoid through a visual spatial position gradient algorithm to obtain a gradient value;

[0019] If the gradient value is less than the first gradient threshold, no refinement adjustment is performed; otherwise,

[0020] Judging whether the center point of the first Gaussian ellipsoid is greater than the first high score threshold. If so, the first Gaussian ellipsoid is split; otherwise,

[0021] Judging whether the center point of the first Gaussian ellipsoid is less than the first low score threshold and the opacity of the first Gaussian ellipsoid is greater than the first opacity threshold. If so, the first Gaussian ellipsoid is split; otherwise,

[0022] Judging whether the center point of the first Gaussian ellipsoid is less than the first low score threshold and the opacity of the first Gaussian ellipsoid is less than or equal to the first opacity threshold. If so, the first Gaussian ellipsoid is deleted; otherwise, no refinement adjustment is performed.

[0023] Optionally, in the first refinement network, the second Gaussian ellipsoid is refined and adjusted through second threshold data, where the second threshold data includes a second high score threshold, a second low score threshold, a second opacity threshold, and a second gradient threshold. The process of the second threshold data includes:

[0024]

[0025] Among them, represents the average score matrix in the i-th iteration process, DA(·) represents the attention mechanism processing, represents the average score matrix output in the previous iteration process of the i-th time, represents the feature map input in the i-th iteration, represents the feature map output in the previous time of the i-th time, MLP represents the processing by MLP, ξ 2i represents the contribution factor, U1(·) represents the projection operation, p represents the center point of the second Gaussian ellipsoid, C i is the pose parameter of the i-th input camera, τ h2 ,τ l2 ,τ α2 , They represent the second highest score threshold, the second lowest score threshold, the second opacity threshold, and the second gradient threshold respectively, wherein the average score matrix in the first iteration process is obtained by weighted average calculation of the score matrix, and the feature map in the first iteration process is the first feature map.

[0026] Optionally, the attention mechanism adopts a deformable attention mechanism.

[0027] Optionally, the loss function of the reconstruction network is:

[0028]

[0029] in, represents the loss value, represents L1 loss, represents the deformed version of structural similarity D-SSIM loss, and λ represents the weight factor.

[0030] Compared with the prior art, the present invention has the following advantages and technical effects:

[0031] This method improves 3D Gaussian splashing. Through two shearing and segmentation optimizations, more Gaussian ellipsoids can be deployed in areas with complex geometric structures, thereby enabling high-precision three-dimensional restoration and reconstruction, and reducing redundant Gaussian ellipsoids, providing a more accurate and faster three-dimensional reconstruction method for the application field of urban renewal of ancient buildings. The above technical solution of the present invention can achieve high-precision three-dimensional reconstruction: more 3D Gaussian ellipsoids are placed in geometrically complex areas to express more precise details; a small number of Gaussian ellipsoids are placed in geometrically simple areas to improve efficiency. The present invention improves the reconstruction accuracy through two detail optimization strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0033] Figure 1 A structural diagram of reconstruction model data processing according to an embodiment of the present invention;

[0034] Figure 2 A first detailed network structure diagram of an embodiment of the present invention;

[0035] Figure 3 This is a network structure diagram of the second detailed network according to an embodiment of the present invention. DETAILED DESCRIPTION

[0036] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0037] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0038] As Figure 1 shown, in this embodiment, a three-dimensional reconstruction method for high-precision urban renewal of ancient buildings is provided, including:

[0039] Obtain the input image and camera parameters, and process the input image and camera pose through the reconstruction model to obtain a three-dimensional model of the urban ancient building, where the input image and the camera pose correspond one by one, and the input image is an image of different urban ancient buildings taken by the camera.

[0040] The reconstruction model includes an input module, a feature extraction module, a first refinement network, a second refinement network, and an output module connected in sequence. The input module is used to input input data, where the input data includes the input image and the camera pose. The feature extraction module extracts features from the input data, and the first refinement network and the second refinement network refine the extracted features. The output module is used to output the final three-dimensional model according to the refined features.

[0041] For the above technical solution, the present invention makes the following specific and detailed description:

[0042] I. Reconstruction Model

[0043] The reconstruction model includes the following structures: an input module, a feature extraction module, a first refinement network, a second refinement network, and an output module.

[0044] 1. Input Module and Output Module:

[0045] (1) Input Module: The input parameters, that is, the input data, include a series of input images and a corresponding series of "input cameras". Among them, the "input camera" is the pose parameters of the corresponding camera that takes the input image, including the internal parameters and external parameters of the camera. The input camera and the input image are in one-to-one correspondence and are paired. In addition, according to the pose parameters of the input camera, they can be sorted in a certain order, for example, sorted by the translation vector or sorted by the focal length size, etc. As a sequence, the input data is input into the network model in the input order.

[0046] (2) Output Module: The output of the reconstruction model is a refined and cropped three-dimensional distribution of a high-precision Gaussian ellipsoid. Then, through rasterization rendering for the three-dimensional distribution of the Gaussian ellipsoid, different perspective images can be rendered onto the display screen to present the final three-dimensional model.

[0047] It should be noted that the Gaussian ellipsoid is represented as a three-dimensional Gaussian distribution. Among them, the two-dimensional Gaussian distribution is a normal distribution, and the three-dimensional Gaussian distribution is obtained by rotating the normal distribution one week to form an ellipsoidal shape, which is represented as a Gaussian ellipsoid.

[0048] (3) For the reconstruction model, the entire input-output process is expressed by the formula:

[0049]

[0050] Among them, Model represents the entire model system, I i represents the i-th input image, C i represents the camera pose parameters corresponding to the i-th input image, N I represents the number of input images, N O represents the number of output Gaussian ellipsoids, p j represents the position of the center point of the j-th Gaussian ellipsoid, s j represents the scaling coefficient of the j-th Gaussian ellipsoid, r j represents the rotation matrix of the j-th Gaussian ellipsoid, α j represents the opacity of the j-th Gaussian ellipsoid, h j represents the spherical harmonic function parameters of the j-th Gaussian ellipsoid, which are used to represent the color of the Gaussian ellipsoid.

[0051] 2. Feature extraction module

[0052] (1) The feature extraction module includes a residual network and a parallel cross-attention network, which are respectively used to extract the features of the input image, generate feature maps, then merge the generated feature maps, and then input them into the first MLP to encode the merged feature maps to generate feature vectors. Then, the encoded feature vectors are input into the multi-view depth estimation network. After accurately estimating the depth of the feature vectors, an initial point cloud is obtained through the method of back-projection;

[0053] Then, each point in the initial point cloud is used as the center point of the Gaussian ellipsoid. At this stage, what we are most concerned about is the position of the center point of the Gaussian ellipsoid. Other parameters of the Gaussian ellipsoid, such as scaling, rotation, opacity, spherical harmonic function parameters, etc., can be replaced by random numbers (because these parameters are learnable or will be adjusted in the subsequent refinement network. Therefore, during the network iteration process, these parameters will converge from random values to real values).

[0054] It should be noted that the multi-view depth estimation network has been trained and is used as a trained model. The training content of the multi-view depth estimation network is not involved in the present invention.

[0055] (2) The residual network can use a common ResNet-50, which is only used for feature extraction of each input image.

[0056] (3) Cross-attention network:

[0057] 1) According to the input order of the input images, cross-attention operations are performed on every two adjacent images. The input order has been determined according to the pose parameters of the camera. In this order, there is a correlation between the input images of the vectors. Based on this correlation, the cross-attention network is used to process the images.

[0058] 2) Two adjacent images are embedded into three matrices Q, K, and V through trilinear projection. Among them, the generation methods of the three matrices Q, K, and V are the same as those of the QKV matrix in Transformer, and the key matrix query matrix and value matrix and the key matrix query matrix and value matrix are obtained respectively. Here, the subscript cur represents the current image, and the subscript next represents the next image. is the size of the feature map, that is, width * height, c is the number of channels of the feature map, and R represents the real number field.

[0059] 3) The cross-attention formula is as follows:

[0060]

[0061] Among them, d represents the dimension of the Q, K, and V vectors. The dimensions of the Q, K, and V vectors are the same, all d, which is used to scale the dot product result to prevent the softmax function from saturating due to too large a numerical value, thus affecting the accuracy of the attention distribution.

[0062] 4) The other parts of the cross-attention network are the same as those of Transformer, including the multi-head attention mechanism, the position feed-forward network, the residual connection layer, and the normalization layer, etc., which will not be elaborated here.

[0063] 5) Finally, add a fusion layer. The fusion layer can choose Transformer or MLP. Preferably, Transformer is used to enhance the cross-embedding ability and integrate the two cross-attentions together.

[0064] (4) There is a plus sign, i.e., ⊕, in the reconstructed model architecture diagram, which is used to represent the fusion of the feature map generated by the cross-attention network and the feature map generated by the residual network. The fusion method can use Transformer or MLP for fusion. Preferably, Transformer is used to enhance the ability of cross-embedding.

[0065] It should be noted that: there are N I input images. After passing through the residual network and the cross-attention network respectively, N I feature maps are obtained. When performing operations, preferably Transformer is used. In Transformer, the above-mentioned images are processed through the above cross-attention mechanism and then fused with the current feature map in ResNet. After all fusions, N I fused feature maps are generated.

[0066] (5) The fused feature maps are input into the first MLP for further feature fusion and extraction.

[0067] (6) For the pre-trained multi-view depth estimation network, the specific model of this method is not limited. Common ones such as MAMo (Memory and Attention for Monocular Video Depth Estimation), FSRE-Depth, CVDE (Consistent Video Depth Estimation), DVW (Depth from Videos in the Wild), ACNNs (Atrous Convolutional Neural Networks), etc. By inputting the feature maps output from the first MLP and the pose parameters of the input camera, a depth map is output.

[0068] (7) Then, the input images are back-projected according to the output depth map to obtain the initial point cloud. Then, each point is regarded as the position of the center point of the Gaussian ellipsoid, and the initial Gaussian ellipsoid is constructed to generate the first Gaussian ellipsoid. Except for the position of the center point of each Gaussian ellipsoid in the first Gaussian ellipsoid, other parameters of the Gaussian ellipsoid, such as scaling, rotation, opacity, spherical harmonic function parameters, etc., can be replaced by random numbers. It is expressed by the formula:

[0069] p = U -1 (Model Depth (Features), C)#(4)

[0070] where p represents the position of the center point of a certain Gaussian ellipsoid, and U -1Denotes the back-projection operation, that is, the operation of mapping from a two-dimensional image to three-dimensional space coordinates, Model Depth Denotes a pre-trained multi-view depth estimation network. C denotes the camera pose parameters corresponding to the Gaussian ellipsoid, that is, each input image corresponds to a certain camera. After the input image undergoes the above processing, a corresponding feature map is generated, and the camera pose parameters correspond to the camera pose parameters of the input image corresponding to the generated feature map. Feature denotes the feature map output from the first MLP.

[0071] 3. First refinement network: The first Gaussian ellipsoid undergoes several iterations, such as 3 times, through the first refinement network to generate a second Gaussian ellipsoid.

[0072] It should be noted that during each iteration, the Gaussian ellipsoid generated in the previous iteration process is used as the input data of the first refinement network again for reprocessing to obtain the Gaussian ellipsoid generated in the current iteration process. In the first iteration process, the first Gaussian ellipsoid is used as the input data of the first refinement network. After the iteration stops, the Gaussian ellipsoid generated in the last iteration process is used as the second Gaussian ellipsoid.

[0073] Taking the 3-iteration process as an example, specifically, in the first iteration, the first Gaussian ellipsoid is input into the first refinement network to generate a Gaussian ellipsoid; in the second iteration, the Gaussian ellipsoid is used as the input again and input into the first refinement network again to generate a Gaussian ellipsoid; in the third iteration, is used as the input of the first refinement network again and input into the first refinement network again to generate the second Gaussian ellipsoid.

[0074] 4. Second refinement network: The second Gaussian ellipsoid is input into the second refinement network. The second Gaussian ellipsoid undergoes several iterations, such as 3 times, through the first refinement network, and finally generates a third Gaussian ellipsoid. Finally, the third Gaussian ellipsoid is rasterized and rendered to render images from different perspectives onto the display screen.

[0075] The iteration process of the second refinement network is similar to that of the first refinement network and will not be elaborated here.

[0076] II. First refinement network

[0077] As Figure 2 shown, 1. The input of the second MLP is the feature map output from the first MLP, and the purpose is to output a saliency map, which is expressed by the formula:

[0078]

[0079] Among them, denotes the i-th saliency map, and the total number is NI Zhang, softmax(·) is the activation function, MLP(·) represents the multi-layer perceptron network, and F i represents the i-th feature map output by the first MLP. The number of feature maps is the same as the number of input images, and the total number is N I ; a i represents the contribution degree of each feature map. β i As a learnable parameter, it also needs to be input into the second MLP together. Finally, the true contribution value of each feature map is learned, and then through the weighted average method, the sum of all a i is 1. F sig represents the overall saliency map, and its number is 1.

[0080] It should be noted that whether it is F sig or is a grayscale image and has only one channel, storing values from 0 to 255, which is used to represent the saliency intensity. 0 is the minimum and 255 is the maximum, and the size of the saliency map is the same as the size of the input image.

[0081] 2. Calculation and acquisition process of the average score matrix:

[0082] (1) Project the center points of all Gaussian ellipsoids in the first Gaussian ellipsoid through the projection P j onto N I saliency maps , and N I saliency projection maps will be obtained, that is, N I score matrices, denoted as where represents one of the score matrices, and the projection P j is implemented by the linear projection method.

[0083] (2) It should be noted that not all the center points of the Gaussian ellipsoids can be projected onto the saliency map. If there are occluded parts in the saliency map, since the saliency map is a two-dimensional image and the Gaussian ellipsoid is a three-dimensional space, the center points of these occluded Gaussian ellipsoids cannot be projected onto the saliency map. Therefore, the data that can be projected from the center point of the Gaussian ellipsoid to the saliency map will obtain a saliency value, and the grayscale value, that is, the saliency value, is stored at different positions in the saliency map; the data that cannot be projected onto the saliency map is set to 0; therefore, after projecting the center points of all Gaussian ellipsoids, a matrix composed of N G numbers, that is, the score matrix, will be obtained. Among them, N G is the number of all Gaussian ellipsoids in the first Gaussian ellipsoid. One Gaussian ellipsoid has only one center point.

[0084] (3) Perform weighted averaging on all fractional matrices to obtain an average fractional matrix, which stores the fractional values corresponding to the centers of all Gaussian ellipsoids in all the first Gaussian ellipsoids, that is, the saliency values. Denote the average fractional matrix as where N G is the number of centers of Gaussian ellipsoids and also the total number of parameters in the average fractional matrix, represents the fractional value corresponding to the center of the i-th Gaussian ellipsoid, that is, the saliency value.

[0085] 3. View space position gradient: For the specific algorithm, refer to Section A "DETAILS OF GRADIENT COMPUTATION" in Appendix A of the paper "3D Gaussian Splatting for Real-Time Radiance Field Rendering". By calculating the view space position gradient, a gradient value will be obtained, denoted as

[0086] 4. Generation process formula of the first refinement threshold:

[0087]

[0088] where MLP(·) in the formula represents the third MLP processing, and DA(·) represents the processing by the first attention mechanism, that is, the deformable attention mechanism. τ h and τ l respectively represent the high and low fractional thresholds output by the third MLP, and τ α represents the opacity threshold output by the third MLP, represents the gradient threshold output by the third MLP, and ξ i represents the contribution factor, which is calculated from the learnable parameter η i , that is, from N I learnable η i are also input into the third MLP as input parameters. The purpose is to make all ξ i add up to 1.

[0089] DA(·) represents the processing by the first attention mechanism. The first attention mechanism uses the deformable attention mechanism. Among them, the deformable attention mechanism is an improved attention mechanism, which comes from deformable ConvNets and can perform deformable operations on the Transformer architecture. This attention mechanism allows the model to focus on key regions with irregular shapes or non-uniform distributions in the input feature map; the input data of the first attention mechanism is the fractional matrix saliency map and the first projection map, represents the i-th fractional matrix, denote the i-th saliency map;

[0090] U(·) represents the projection, and U(p sample ,C i ) represents the first projection map, and p sample denotes the variable center point randomly sampled from the Gaussian ellipsoid. Since the center point of the Gaussian ellipsoid needs to be adjusted, sampling is performed. This is a learnable parameter. Through random sampling, a suitable position is learned, and C i denotes the input camera pose parameter corresponding to the saliency map. Each saliency map is obtained by feature extraction from the input image. Therefore, each saliency map corresponds to the camera pose parameter of the input image, and this pose parameter is consistent with the camera pose parameter of the input image.

[0091] In addition, N I denotes the number of projection maps. Since there are N I camera pose parameters, there are N I first projection maps. That is, each first projection map, the score matrix, and the saliency map need to pass through the first attention and then be merged together and input into the third MLP. Among them, the first projection map, the score matrix, and the saliency map all correspond to an input image.

[0092] 5. The Jd operation in the first refinement network, that is, Figure 2 Jd in

[0093] (1) The input parameters of the Jd operation include: the first Gaussian ellipsoid, the average score matrix the visual space position gradient the first refinement threshold τ h and τ l .

[0094] (2) When no redrawing operation is performed, so that judgment can be made quickly, denotes the gradient threshold output by the third MLP.

[0095] (3) When , all scores in all average score matrices are judged one by one:

[0096] 1) If a MLP network is adopted to split the i-th Gaussian ellipsoid into two independent Gaussian ellipsoids;

[0097] 2) If and α i >τ α , the opacity and scaling ratio of the i-th Gaussian ellipsoid are reduced to half of the original. Among them, α iRepresents the opacity of the i-th Gaussian ellipsoid, and its value range is between 0 and 1. τ α Represents the opacity threshold output by the third MLP.

[0098] 3) If and α i ≤τ α At this time, the i-th Gaussian ellipsoid will be directly deleted, that is, it is regarded as a redundant Gaussian ellipsoid;

[0099] 4) If Then no redrawing operation is performed. If no redrawing operation is performed, these Gaussian ellipsoids are considered appropriate, indicating that the Gaussian ellipsoid meets the requirements of high-precision reconstruction and no further relevant adjustments are required, and it is directly output to the next stage.

[0100] 6. The entire first refinement network iterates 3 times

[0101] (1) The first Gaussian ellipsoid is input into the first refinement network to generate the second Gaussian ellipsoid, which is the 1st iteration;

[0102] (2) Then, the second Gaussian ellipsoid is used as the input of the first refinement network to generate the second Gaussian ellipsoid x2, which is the 2nd iteration;

[0103] (3) Finally, the second Gaussian ellipsoid x2 is used as the input of the first refinement network again to generate the second Gaussian ellipsoid x3, which is the 3rd iteration.

[0104] (4) After 3 iterations, the effect will be better. 3 times is used as the selectable and interpretable number of times. For better effect, several iterations can be performed, and there is no limit here.

[0105] III. Second Refinement Network

[0106] As Figure 3 shown, 1. The second refinement network is similar in function and structure to the first refinement network. The purpose is to further improve the accuracy of 3D reconstruction and enable it to adapt to more complex reconstruction tasks.

[0107] 2. The system architecture diagram of the second refinement network, that is, Figure 3 the module in the dashed box, is called the iterative feature extraction module;

[0108] (1) The input includes 3 parameters: the feature map output by the first MLP in the reconstruction model, the average score matrix output by the first refinement network, and the second projection map obtained by projecting all the center points of the second Gaussian ellipsoid onto multiple corresponding input images. The size of the second projection map is the same as the size of the input image, and the number is the same as the number of input images.

[0109] (2) There are two outputs: one is the iterative average score matrix, and the other is the feature map output by the feature map output head.

[0110] (3) The second attention has the same structure as the first attention, both being deformable attention mechanisms.

[0111] (4) The fourth MLP is used for further feature extraction.

[0112] (5) The two output heads, namely the feature map output head and the score matrix output head, are used to convert the feature map into the final output parameters.

[0113] (6) Iterative mechanism: First, the iterative average score matrix will replace the average score matrix as the input of the second attention; second, the feature map output by the feature map output head will replace the feature map output by the first MLP as the input of the second attention.

[0114] Note that during the iteration process, the second projection map does not change. After several iterations, such as 3 iterations, the feature map output by the feature map output head will be input into the fifth MLP for further feature extraction. And the output iterative average score matrix will be directly used as the input parameter for the next operation.

[0115] (7) The formula for this process is:

[0116]

[0117] Among them, represents the average score matrix input into the second attention in the i-th iteration, represents the average score matrix output in the previous iteration of the i-th iteration, represents the feature map input in the i-th iteration, represents the feature map output in the previous iteration of the i-th iteration, is the feature map output by the first MLP. MLP in the first formula represents being processed by the fourth MLP, and MLP in the second formula represents being processed by the fifth MLP. N I represents the number of second projection maps corresponding to N I camera pose parameters. That is, each second projection map needs to go through the second attention and then be merged together and input into the fourth MLP.

[0118] ξ 2i represents the contribution factor, calculated from the learnable parameter η 2i DA(·) represents the second attention, that is, the deformable attention mechanism.

[0119] U1(·) represents a projection operation, p represents the center point of the second Gaussian ellipsoid, and C i is the pose parameter of the i-th input camera. τ h2 , τ l2 , τ α2 , is similar to the previous τ h , τ l , τ α , is defined, that is, τ h2 and τ l2 respectively represent the high and low score thresholds of the output of the fifth MLP, and τ α2 represents the opacity threshold of the output of the fifth MLP, represents the gradient threshold of the output of the fifth MLP.

[0120] (8) The Jd operation of the second refinement network, that is, Figure 3 in is consistent with the Jd operation of the first refinement network, both of which perform parameter comparison, except that the average score matrix in the first refinement network is replaced by an iterative average score matrix. Specifically:

[0121] (1) The input parameters of the Jd operation include: the second Gaussian ellipsoid, the average score matrix the gradient of the visual space position the first refinement threshold τ h2 and τ l2 .

[0122] (2) When no redrawing operation is performed, so that a quick judgment can be made, represents the gradient threshold of the output of the fifth MLP.

[0123] (3) When , all scores in all iterative average score matrices are judged one by one:

[0124] 1) If a MLP network is adopted to split the j-th Gaussian ellipsoid into two independent Gaussian ellipsoids;

[0125] 2) If and α j > τ α2 , the opacity and scaling ratio of the j-th Gaussian ellipsoid are reduced to half of the original, where α j represents the opacity of the j-th Gaussian ellipsoid, and its value range is between 0 and 1. τ α represents the opacity threshold of the output of the fifth MLP.

[0126] 3) If and αj ≤ τ α2 When it is, the j-th Gaussian ellipsoid will be directly deleted, that is, it is regarded as a redundant Gaussian ellipsoid;

[0127] 4) If Then no redrawing operation is performed. If no redrawing operation is performed, these Gaussian ellipsoids are considered appropriate, indicating that the Gaussian ellipsoid meets the requirements of high-precision reconstruction and no further relevant adjustments are required, and it is directly output to the next stage.

[0128] IV. Loss Function

[0129] The entire network model is trained uniformly using the L1 loss and the D-SSIM loss. The formula is:

[0130]

[0131] Among them, represents the L1 loss, which is used to measure the absolute difference in pixel values between the predicted image I and the target image . Specifically, are multiple input images, and I is the image rendered from the third Gaussian ellipsoid through the rasterization rendering pipeline according to the input camera pose parameters corresponding to the input image I.

[0132] represents the deformed version D-SSIM loss of structural similarity, which is used to evaluate the structural similarity between two images. The loss function here is consistent with the loss function in the paper "3D Gaussian Splatting for Real-Time RadianceField Rendering". If simplified, it can be replaced by .

[0133] λ is a weight factor, such as λ = 0.2, which is used to balance the contributions between and .

[0134] The above is only a preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A high-precision three-dimensional reconstruction method for ancient buildings in urban renewal, characterized in that: include: Obtain image data of ancient urban buildings and corresponding camera parameters, The image data is used as an input image and corresponding camera parameters are input into a reconstruction model for three-dimensional reconstruction and rendering to obtain a three-dimensional model of the ancient urban building; The reconstruction model includes a feature extraction module, a first refinement network and a second refinement network. The feature extraction module is used to extract features according to the image data and the corresponding camera parameters, obtain a first feature map, perform depth estimation on the first feature map, obtain a first Gaussian ellipsoid, iteratively refine and adjust the first Gaussian ellipsoid through a first refinement network to obtain a second Gaussian ellipsoid, iteratively refine and adjust the second Gaussian ellipsoid through a second refinement network to obtain a third Gaussian ellipsoid, render the third Gaussian ellipsoid, and obtain a three-dimensional model of ancient urban buildings.

2. The method according to claim 1, characterized in that: The reconstruction model further comprises an input module, wherein the input module sequentially inputs the image data and the camera parameters into the reconstruction model according to the camera parameters.

3. The method according to claim 1, characterized in that The reconstructed model also includes an output module, wherein the output module is used to perform raster rendering on the third Gaussian ellipsoid to obtain a three-dimensional model of the ancient urban buildings.

4. The method according to claim 1, characterized in that: The feature extraction module includes a parallel structure, a first MLP and a multi-view depth estimation network connected in sequence, wherein the parallel structure includes a residual network and a cross-attention network connected in parallel, the image data is processed and fused through the parallel structure to obtain a fused feature map, the fused feature map is further fused through the first MLP to obtain a first feature map, the first feature map and camera parameters are depth estimated through the multi-view depth estimation network, and the depth estimation result is back-projected to obtain a first Gaussian ellipsoid, wherein the first Gaussian ellipsoid includes the center point position of the Gaussian ellipsoid.

5. The method according to claim 1, characterized in that The first Gaussian ellipsoid is iteratively refined and adjusted several times through the first refinement network, wherein, in each iteration process, the Gaussian ellipsoid generated by the previous iteration process is used as input data in the current iteration process, the input data is refined and adjusted through the first refinement network to obtain the Gaussian ellipsoid generated by the current iteration process, the first Gaussian ellipsoid is used as input data in the first iteration process, and the Gaussian ellipsoid generated by the last iteration process is used as the second Gaussian ellipsoid.

6. The method according to claim 5, characterized in that In the first refinement network, the first Gaussian ellipsoid is refined and adjusted by first threshold data, wherein the first threshold data includes a first high score threshold, a first low score threshold, a first opacity threshold and a first gradient threshold, wherein the process of the first threshold data includes: Among them, MLP(·) represents multi-layer perceptron processing, DA(·) represents attention mechanism processing, τ h and τ l Respectively represent the first high score threshold and the first low score threshold, τ α represents the first opacity threshold, represents the first gradient threshold, ξ i represents the contribution factor, represents the score matrix, represents the saliency map, U(·) represents the projection operation, and p sample represents the variable center point randomly sampled from the Gaussian ellipsoid, C i represents the input camera pose parameter corresponding to the saliency map, i represents the input image number, N I Represents the total number of input images; The first feature map is processed by a multi-layer perceptron and a softmax function to obtain a saliency map, and the center point of the first Gaussian ellipsoid is projected onto the saliency map to obtain a score matrix.

7. The method according to claim 6, characterized in that The process of refining the first Gaussian ellipsoid includes: Calculating the first Gaussian ellipsoid by using a visual space position gradient algorithm to obtain a gradient value; If the gradient value is less than the first gradient threshold, no refinement adjustment is performed; otherwise, Determine whether the center point of the first Gaussian ellipsoid is greater than the first high score threshold. If so, split the first Gaussian ellipsoid; otherwise, Determine whether the center point of the first Gaussian ellipsoid is less than the first low score threshold and the opacity of the first Gaussian ellipsoid is greater than the first opacity threshold. If so, split the first Gaussian ellipsoid; otherwise, It is determined whether the center point of the first Gaussian ellipsoid is smaller than the first low score threshold and the opacity of the first Gaussian ellipsoid is less than or equal to the first opacity threshold. If so, the first Gaussian ellipsoid is deleted, otherwise no refinement adjustment is performed.

8. The method according to claim 7, characterized in that In the first refinement network, the second Gaussian ellipsoid is refined and adjusted by second threshold data, wherein the second threshold data includes a second high score threshold, a second low score threshold, a second opacity threshold and a second gradient threshold, wherein the process of the second threshold data includes: in, represents the average score matrix during the i-th iteration, DA(·) represents the attention mechanism processing, Represents the average score matrix output by the last iteration process of the i-th time, represents the feature map of the input of the i-th iteration, represents the feature map of the last output of the i-th time, MLP represents the processing by MLP, ξ 2i represents the contribution factor, U1(·) represents the projection operation, p represents the center point of the second Gaussian ellipsoid, C i is the pose parameter of the i-th input camera, τ h2 ,τ l2 ,τ α2 , They represent the second highest score threshold, the second lowest score threshold, the second opacity threshold, and the second gradient threshold respectively, wherein the average score matrix in the first iteration process is obtained by weighted average calculation of the score matrix, and the feature map in the first iteration process is the first feature map.

9. The method according to claim 8, characterized in that The attention mechanism adopts a deformable attention mechanism.

10. The method according to claim 1, characterized in that The loss function of the reconstruction network is: in, represents the loss value, represents L1 loss, represents the deformed version of structural similarity D-SSIM loss, and λ represents the weight factor.