A large-scale scene reconstruction method, system, electronic device and computer-readable storage medium based on a 3D Gaussian algorithm

By segmenting the large scene into multiple grid cells and applying a 3D Gaussian sputtering model for processing, the problem of excessive number of three-dimensional Gaussians in large scene reconstruction is solved, and high-quality three-dimensional reconstruction and rapid rendering are achieved.

CN119741449BActive Publication Date: 2025-06-27BEIJING DATA INTELLIGENCE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510240459.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-27
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The existing 3D Gaussian sputtering technology is difficult to effectively deal with a large number of three-dimensional Gaussian numbers in large-scenario reconstruction, resulting in low reconstruction quality or insufficient memory.

Method used

By segmenting the large scene into multiple grid cells, the sparse point clouds and original images of each cell are processed separately, the 3D Gaussian sputtering model is used for point cloud distribution transformation, rasterization and appearance modeling optimization, and finally iterative updates are performed through the comprehensive loss function to improve the reconstruction quality.

Benefits of technology

High-quality three-dimensional reconstruction in large scenarios is realized, the reconstruction rate is improved, and more scene details are retained. The generated rendered images are visually more realistic and consistent in structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741449B_ABST
    Figure CN119741449B_ABST
Patent Text Reader

Abstract

The present invention provides a large-scale scene reconstruction method based on a 3D Gaussian algorithm, including: S1 obtaining a first co-visual original image set; S2 generating a first sparse point cloud set by using the SfM algorithm and dividing the large-scale scene into multiple first grid cells; S3 expanding the boundaries of the first grid cells to obtain multiple second grid cells; S4 adding new cameras to the original camera set to obtain an optimized camera set; S5 optimizing each second sparse point cloud subset to obtain a corresponding third sparse point cloud subset; S6 using a 3D Gaussian sputtering model to perform 3D Gaussian point cloud distribution conversion, rasterization, and appearance modeling optimization processing on each third sparse point cloud subset to obtain multiple primary rendered image sub-blocks; S7 constructing a comprehensive loss function to iterate each primary rendered image sub-block to obtain multiple high-level rendered image sub-blocks; S8 using the multiple high-level rendered image sub-blocks to reconstruct a dense sub-scene, and seamlessly merging the dense sub-scenes to obtain the reconstructed large-scale scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer vision and deep learning, and particularly relates to a large-scale scene reconstruction method, system, electronic device, and computer-readable storage medium based on a 3D Gaussian algorithm. Background Art

[0002] Traditional 3D reconstruction methods mainly achieve 3D scene reconstruction through methods such as point clouds and meshes. Although the rendering speed is fast, they are slightly insufficient in presenting scene details. The full implicit 3D reconstruction method has excellent scene presentation capabilities, but it takes a large amount of time for training and rendering. To combine the advantages of both, recently, deep learning technology has made significant progress in scene reconstruction. In particular, 3D Gaussian Splatting (3DGS) is particularly outstanding in terms of the clarity of the reconstruction results and the fast rendering speed. The 3DGS technology uses Gaussian functions and sparse point cloud data to model the scene, converts 3D information into the form of Gaussian functions for large-scale scene model reconstruction, and realizes real-time rendering. Currently, the traditional 3DG runs well in small-scale scenes and object-centered scenes, but it encounters difficulties in large-scale scene reconstruction. First, the number of 3D Gaussian numbers is limited by a given video memory, and a large amount of 3D Gaussian numbers are required for the rich details of a large-scale scene. Directly applying 3DGS to a large-scale scene will result in problems such as low-quality reconstruction or insufficient memory. Therefore, how to use 3DGS for high-quality 3D reconstruction of large-scale scenes has become one of the research hotspots. Summary of the Invention

[0003] To solve the above technical problems, the first object of the present invention is to provide a large-scale scene reconstruction method based on a 3D Gaussian algorithm, the method comprising:

[0004] S1 Obtain multiple original images covering the target large-scale scene area, and screen the original images that meet the co-visibility condition from the multiple original images to form a first co-visibility original image set; the first co-visibility original image set includes multiple original images with different perspectives;

[0005] S2 Based on the first co-visibility original image set, use the SfM algorithm to generate a first sparse point cloud set, and calculate the point cloud boundary of the first sparse point cloud set; then divide the target large-scale scene area into multiple first grid units according to the point cloud boundary and the positions of the original camera set projected on the plane of the first co-visibility original image set;

[0006] S3 Expand the boundaries of the first grid units to obtain multiple second grid units; each second grid unit contains a second co-visibility original image subset and a second sparse point cloud subset;

[0007] S4 adds new cameras to the original camera sets corresponding to each second grid cell to obtain an optimized camera set corresponding to each second grid cell;

[0008] S5 optimizes each second sparse point cloud subset in turn according to the optimized camera set to obtain a corresponding third sparse point cloud subset;

[0009] S6 adopts a 3D Gaussian sputtering model to perform 3D Gaussian point cloud distribution conversion, rasterization, and appearance modeling optimization processing on each third sparse point cloud subset in turn to obtain multiple primary rendered image sub-blocks;

[0010] S7 constructs a comprehensive loss function and iteratively updates each primary rendered image sub-block in turn based on the comprehensive loss function until the loss function converges to obtain multiple high-level rendered image sub-blocks;

[0011] S8 reconstructs a dense sub-scene using the geometric information of multiple high-level rendered image sub-blocks and seamlessly merges the dense sub-scenes to obtain the reconstructed target large scene.

[0012] Specifically, the 3D Gaussian sputtering model described in step S6 includes a Gaussian point cloud distribution conversion unit, a differentiable Gaussian rasterization unit, and a decoupled appearance modeling unit.

[0013] Specifically, step S6 further includes:

[0014] S61 uses the Gaussian point cloud distribution conversion unit to convert each third sparse point cloud subset corresponding to each second grid cell into a set of primary 3D Gaussian point cloud distributions in turn; and then converts each set of primary 3D Gaussian point cloud distributions into each primary 3D Gaussian image sub-block;

[0015] S62 uses the differentiable Gaussian rasterization unit to project and render each primary 3D Gaussian image sub-block onto the 2D image plane corresponding to each camera pose to obtain multiple primary predicted image sub-blocks;

[0016] S63 based on multiple primary predicted image sub-blocks, uses the decoupled appearance modeling unit to generate multiple primary transformation maps, and performs appearance optimization on each corresponding primary predicted image sub-block in turn according to each primary transformation map to obtain multiple primary rendered image sub-blocks.

[0017] Specifically, step S7 further includes:

[0018] S71 constructs a comprehensive loss function, and the comprehensive loss function includes a first loss function and a second loss function; wherein the first loss function includes a first difference loss function and a D-SSIM loss function, and the second loss function includes a second difference loss function;

[0019] S72 iteratively updates the parameters of each primary 3D Gaussian image sub-block based on multiple primary rendered image sub-blocks through a first loss function until the first loss function converges, obtaining multiple high-level 3D Gaussian image sub-blocks;

[0020] S73 sequentially renders each high-level 3D Gaussian image sub-block onto a 2D image plane using a differentiable Gaussian rasterization unit, obtaining multiple high-level predicted image sub-blocks;

[0021] S74 generates multiple corresponding high-level transformation mappings based on multiple high-level predicted image sub-blocks using a decoupled appearance modeling unit, and iteratively optimizes the appearance of each high-level predicted image sub-block according to each high-level transformation mapping and a second loss function until the second loss function converges, obtaining multiple high-level rendered image sub-blocks.

[0022] Specifically, the first loss function in step S7 is:

[0023]

[0024] where, is the first difference loss function, used to calculate the absolute pixel value difference between the predicted image sub-block and the original image; is the D-SSIM loss function, used to calculate the structural similarity difference between the predicted image sub-block and the original image; The weight of the D-SSIM loss function

[0025] Specifically, the parameters of the 3D Gaussian point cloud distribution include the center point, covariance matrix, color information, and opacity.

[0026] Specifically, the co-visibility conditions in step S1 include: the satellite zenith angle of each original image is less than 30°, the angle between every two original images is less than 5°, and the shadow area ratio of each original image is less than 5%.

[0027] The second object of the present invention is to provide a large-scale scene reconstruction system based on a 3D Gaussian algorithm, and the system includes:

[0028] An image module, configured to obtain multiple original images covering the target large-scale scene area, screen the original images that meet the co-visibility conditions from the multiple original images, and form a first co-visible original image set;

[0029] A sparse point cloud generation module, configured to generate a first sparse point cloud set based on the first co-visible original image set using the SfM algorithm, and calculate the point cloud boundary of the first sparse point cloud set;

[0030] A large-scale scene segmentation module, and then based on the point cloud boundary, divides the target large-scale scene area into multiple first grid units according to the original camera positions projected on the plane;

[0031] A boundary expansion module, configured to expand the boundary of an initial first grid cell to obtain a plurality of second grid cells; each of the second grid cells contains a second co-visible original image subset and a second sparse point cloud subset;

[0032] A camera optimization module, configured to add new cameras to the original camera set corresponding to each second grid cell based on the airspace-aware visibility standard to obtain an optimized camera set corresponding to each second grid cell;

[0033] A sparse point cloud optimization module, configured to optimize the second sparse point cloud subset in each second grid cell in sequence according to the optimized camera set to obtain a corresponding third sparse point cloud subset;

[0034] A 3D Gaussian conversion module, configured to adopt a 3D Gaussian sputtering model to perform 3D Gaussian point cloud distribution conversion, rasterization, and appearance modeling optimization processing on the third sparse point cloud subset corresponding to each second grid cell in sequence to obtain a plurality of primary rendered image sub-blocks;

[0035] A rendered image optimization module, constructs a comprehensive loss function, and iteratively updates each primary rendered image sub-block based on the comprehensive loss function until the loss function converges to obtain a plurality of high-level rendered image sub-blocks;

[0036] A merging and reconstruction module, configured to reconstruct a dense sub-scene by using the geometric information of a plurality of high-level rendered image sub-blocks, and seamlessly merge the dense sub-scenes to obtain a reconstructed target large scene.

[0037] The third object of the present invention is to provide an electronic device, which includes: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the method described above.

[0038] The fourth object of the present invention is to provide a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor executes the method described above.

[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0040] (1) The present invention allocates the original images and sparse point clouds in a large scene to different grid cells, and sequentially performs 3D Gaussian point cloud distribution conversion, rasterization processing, and appearance optimization on each independent grid cell to realize the reconstruction in a large scene and improve the reconstruction rate;

[0041] (2) The expansion of the boundaries of the grid cells in the present invention helps to include more visual information, thereby obtaining better results in local optimization, getting a more accurate 3D structure, providing richer geometric details for the subsequent conversion of 3D Gaussian point cloud distribution, and retaining more scene details.

[0042] (3) The present invention introduces decoupled appearance modeling to optimize each predicted image sub-block independently in terms of appearance, making the finally generated rendered image sub-blocks more visually realistic while maintaining structural consistency; and introduces a comprehensive loss function for iterative optimization to reduce rendering errors. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0044] Figure 1 It is a technical flowchart of a large-scale scene reconstruction method based on 3D Gaussian algorithm in an embodiment of the present invention;

[0045] Figure 2 It is a technical framework diagram of a large-scale scene reconstruction method based on 3D Gaussian algorithm in an embodiment of the present invention;

[0046] Figure 3 It is a structural schematic diagram of a large-scale scene reconstruction system based on 3D Gaussian algorithm in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0048] It should be noted that the terms used here are only for describing specific embodiments, rather than intending to limit the exemplary embodiments according to the present invention. As used here, unless otherwise clearly specified in the context, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0049] Please refer to Figure 1 and Figure 2 , Figure 1This is a flow chart of a large scene reconstruction method based on a 3D Gaussian algorithm in an embodiment of the present invention; Figure 1 This is a technical flow chart of a large scene reconstruction method based on a 3D Gaussian algorithm in an embodiment of the present invention; the method comprises the following steps:

[0050] S1 acquires a plurality of original images covering a large target scene area, and selects original images satisfying a common-view condition from the plurality of original images to form a first common-view original image set; the first common-view original image set includes a plurality of original images from different perspectives;

[0051] In the embodiment of the present invention, the target large scene area is a large scene in the real world, including roads, buildings, ruins, woodlands, elevated roads and bridges, etc. The above scene data records multiple original images taken from different positions and angles in the large scene. These images include not only the front view, but also the side, top and other non-traditional view, which can capture the overall view and complex geometric structure of the scene. The common view conditions for shooting the original images include: the satellite zenith angle of each original image is less than 30°, the angle between each two original images is less than 5°, and the shadow area of ​​each original image accounts for less than 5%. 80 original images from different perspectives are obtained to form the first common view original image set.

[0052] S2 generates a first sparse point cloud set based on the first common view original image set using the SfM algorithm, and calculates the point cloud boundary of the first sparse point cloud set; then divides the target large scene area into a plurality of first grid units according to the point cloud boundary and the position of the original camera set projected on the plane by the first common view original image set;

[0053] S3 expands the boundary of the first grid unit to obtain a plurality of second grid units; each of the second grid units contains a second common view original image subset and a second sparse point cloud subset.

[0054] In an embodiment of the present invention, first, a Structure from Motion (SfM) algorithm is used to estimate and generate a first sparse point cloud set from a first co-visual original image set. The target large-scene area is segmented according to the point cloud boundary of the first sparse point cloud set and the original camera position set projected on the plane of the first co-visual original image set. The ground plane is divided into m parts along the x-axis, and then the m parts are sequentially divided into n segments along the y-axis. Each segment contains approximately 80 / (m×n) original images with different perspectives. A segment is regarded as a first grid cell, and each first grid cell contains a similar number of original images with different perspectives. The boundaries of each first grid cell are extended by a range of 10% to ensure the overlap of the original images between adjacent cells, resulting in a plurality of second grid cells. Each of the second grid cells contains a second co-visual original image subset and a second sparse point cloud subset; wherein, each second co-visual original image subset is composed of original images with different perspectives in each first grid cell and the original images within the 10% extended range; the size of each original image is W×H, where W is 1920 and H is 1080.

[0055] S4 Add new cameras to the original camera set corresponding to each second grid cell to obtain an optimized camera set corresponding to each second grid cell;

[0056] S5 Optimize each second sparse point cloud subset in sequence according to the optimized camera set to obtain a corresponding third sparse point cloud subset.

[0057] In an embodiment of the present invention, for each second grid cell q i , an Axis-Aligned Bounding Box (AABB) is constructed based on its corresponding second sparse point cloud subset. The height direction of the axis-aligned bounding box extends to the distance between the highest point of the target large-scene area and the ground plane, covering all possible airspaces within the second grid cell; project the bounding box onto the image plane of the candidate camera E i to form a convex hull area, and calculate the area of the convex hull area ; The visibility ratio refers to the ratio of the visible area of a certain camera to a specific grid cell to the total area of the camera's image. Through airspace perception calculation, it can cover the complete airspace within the cell, avoiding missing important perspectives due to sparse surface points, thereby improving the reconstruction quality and reducing floating artifacts. Its calculation formula is: , where is the total area of the image of the candidate camera E i . If the visibility ratio of the candidate camera E i to the current grid cell exceeds the predefined threshold, i.e., 25%, then add it to the optimized camera set V i。And expand the grid cells by 10% according to the boundary, and initially allocate the newly added cameras. Traverse the remaining unoptimized original camera set, and evaluate the visibility ratio one by one according to the above steps. Add the cameras that meet the predefined threshold, i.e., 25%, to V i 。Finally, the optimized camera set V for each cell i contains two types of cameras: the original camera set within 10% of the expanded boundary and the newly added cameras screened by the visibility of airspace perception. Add the point clouds observed by the newly added cameras (even if they are outside the second grid cell of the current cell) to the second sparse point cloud subset of the second grid cell to obtain the corresponding third sparse point cloud subset of the second grid cell.

[0058] S6 adopts a 3D Gaussian sputtering model, and sequentially performs 3D Gaussian point cloud distribution conversion, rasterization, and appearance modeling optimization processing on each third sparse point cloud subset to obtain multiple primary rendered image sub-blocks.

[0059] In the embodiment of the present invention, the 3D Gaussian sputtering model includes a Gaussian point cloud distribution conversion unit, a differentiable Gaussian rasterization unit, and a decoupled appearance modeling unit.

[0060] In the embodiment of the present invention, step S6 further includes:

[0061] S61 Use the Gaussian point cloud distribution conversion unit to sequentially convert the corresponding third sparse point cloud subset of each second grid cell q i into a set of primary 3D Gaussian point cloud distributions G i ; then convert each set of primary 3D Gaussian point cloud distributions G i into each primary 3D Gaussian image sub-block.

[0062] In the embodiment of the present invention, each third sparse point cloud subset is sequentially mapped to the 3D Gaussian point cloud distribution space to generate multiple sets of primary 3D Gaussian point cloud distributions G i 。The parameters of the primary 3D Gaussian point cloud distribution include the center point coordinates μ, the covariance matrix ∑, the color value Y, and the opacity A. The primary 3D Gaussian point cloud is represented by the three-dimensional covariance matrix ∑ and the center point coordinates μ in three-dimensional space:

[0063]

[0064] where x is the three-dimensional space point coordinate; the three-dimensional covariance matrix , where is the rotation matrix and S is the scaling matrix.

[0065] S62 Use the differentiable Gaussian rasterization unit to project and render each primary 3D Gaussian image sub-block onto the 2D image plane corresponding to each camera pose to obtain multiple primary predicted image sub-blocks.

[0066] In the embodiments of the present invention, each primary 3D Gaussian image sub-block is respectively projected onto the 2D image planes corresponding to each camera pose to obtain 2D Gaussian point clouds corresponding to different camera poses. Specifically, for a camera pose N, the projection transformation matrix from world coordinates to camera coordinates is M. The primary 3D Gaussian point cloud distribution corresponding to the primary 3D Gaussian image sub-block is projected onto the 2D image plane corresponding to the camera pose N, and the coordinates in the camera coordinates after projection are calculated and the covariance matrix are:

[0067]

[0068]

[0069] where J is the Jacobian matrix of the projection transformation, which is used to convert a 3D vector into a 2D vector, thereby realizing the projection transformation from the world coordinate system to the camera coordinate system. The purpose of this transformation is to transform the primary 3D Gaussian image sub-block in the world coordinate system to the 2D image plane space in the camera coordinate system. Then, differentiable Gaussian rasterization processing is performed on the 2D Gaussian point cloud. In the differentiable rasterization processing, the 2D image plane is divided into multiple pixel regions. Each pixel region is responsible for processing the 2D Gaussian point cloud projected into its range, and spatial hashing or grid indexing is used to quickly determine the subset of the 2D Gaussian point cloud corresponding to each pixel point in the pixel region. Within each pixel region, the 2D Gaussian point clouds are sorted from far to near according to the depth value in the camera coordinate system to ensure the correct subsequent blending order. For each pixel point, the color values and opacities of all 2D Gaussian point clouds covering the pixel point are blended in depth order using Alpha blending to obtain the corresponding primary predicted image sub-block with a size of σW×βH . The eigenvalue f of each pixel point in the primary predicted image sub-block w,h is calculated through AlphaBlending, and the calculation formula is:

[0070]

[0071] where, ; is the subset in the 2D Gaussian point cloud corresponding to this pixel; is the eigenvalue of the i-th 2D Gaussian point cloud in ; is the opacity of the i-th 2D Gaussian point cloud in ; is the center point coordinate of the i-th 2D Gaussian point cloud in

[0072] The differentiable Gaussian rasterization method can quickly render a 3D Gaussian distribution onto a 2D image plane and effectively sort Gaussians according to depth information. This process involves calculating the contribution of each Gaussian to the pixel color and accumulating these contributions to form the eigenvalue of the final pixel. Differentiable Gaussian rasterization combines the capabilities of efficient rendering and neural network optimization, allowing for high performance in complex image tasks.

[0073] S63 is based on multiple primary prediction image patches, and uses a decoupled appearance modeling unit to generate multiple primary transformation maps. Then, according to each primary transformation map, appearance optimization is sequentially performed on each corresponding primary prediction image patch to obtain multiple primary rendered image patches.

[0074] In the embodiments of the present invention, the appearance of the primary prediction image patches is optimized by the decoupled appearance modeling unit, aiming to eliminate appearance inconsistencies (such as brightness differences and color offsets) caused by lighting changes, camera exposure differences, etc., so as to generate visually consistent primary rendered image patches.

[0075] First, each primary prediction image patch is sequentially downsampled to a low resolution to obtain a downsampled image; the size of the downsampled image is and an appearance embedding of length t is introduced , and then the appearance embedding is stitched to the 3-channel downsampled image in a pixel-by-pixel manner to obtain an embedded image with 3 + t channels ;

[0076]

[0077] Then, each embedded image is sequentially input into a convolutional neural network CNN, and the CNN network gradually performs upsampling to generate multiple primary transformation maps with the same resolution as . Finally, each corresponding primary prediction image patch is sequentially optimized in appearance using , to obtain multiple primary rendered image patches after appearance optimization by performing appearance optimization on each corresponding primary prediction image patch to obtain multiple primary rendered image patches. .

[0078] S7 constructs a comprehensive loss function and sequentially performs iterative updates on each primary rendered image patch based on the comprehensive loss function until the loss function converges, obtaining multiple high-level rendered image patches.

[0079] In the embodiments of the present invention, step S7 further includes:

[0080] S71 Construct a comprehensive loss function, which includes a first loss function and a second loss function; wherein the first loss function includes a first difference loss function and a D-SSIM loss function, and the second loss function includes a second difference loss function;

[0081] S72 Based on multiple primary rendered image sub-blocks, iteratively update the parameters of each primary 3D Gaussian image sub-block through the first loss function until the first loss function converges, obtaining multiple advanced 3D Gaussian image sub-blocks;

[0082] S73 Use a differentiable Gaussian rasterization unit to render each advanced 3D Gaussian image sub-block onto a 2D image plane in turn, obtaining multiple advanced predicted image sub-blocks;

[0083] S74 Based on multiple advanced predicted image sub-blocks, then use a decoupled appearance modeling unit to generate multiple corresponding advanced transformation mappings, and iteratively optimize the appearance of each advanced predicted image sub-block according to each advanced transformation mapping and the second loss function until the second loss function converges, obtaining multiple advanced rendered image sub-blocks.

[0084] In the embodiment of the present invention, the first loss function in step S7 is:

[0085]

[0086] Wherein, is the first difference loss function, which is used to calculate the absolute pixel value difference between the predicted image sub-block and the original image; is the D-SSIM loss function, which is used to calculate the structural similarity difference between the predicted image sub-block and the original image; is the weight of the D-SSIM loss function. The calculation formula of is:

[0087]

[0088] Where is the total number of predicted image sub-blocks, is the i-th predicted image sub-block, is the corresponding original image. Ensure the color and brightness alignment between the predicted image sub-block and the original image.

[0089] The calculation formula of is:

[0090]

[0091] Measure the consistency between the predicted image and the real image in texture, contrast and structure through the structural similarity index (SSIM), and avoid the loss of geometric details.

[0092] In an embodiment of the present invention, the second loss function in step S7 includes a second difference loss function: is the second difference loss function, which is used to calculate the absolute pixel value difference between the rendered image sub-block and the predicted image sub-block. The calculation formula of is:

[0093]

[0094] where is the total number of primary predicted image sub-blocks, is the i-th predicted image sub-block, is the corresponding rendered image sub-block. The consistency between the rendered image sub-block and the predicted image sub-block is constrained to prevent deviation during the optimization process.

[0095] In an embodiment of the present invention, by constructing a comprehensive loss function, the pixel accuracy, structural consistency, and rendering stability are balanced to avoid a single loss dominating the optimization direction. A strategy of first optimizing geometric parameters and then optimizing appearance parameters is adopted to prevent parameter coupling.

[0096] S8 uses the geometric information of multiple high-level rendered image sub-blocks to reconstruct a dense sub-scene, and seamlessly merges the dense sub-scenes to obtain the reconstructed target large scene.

[0097] In an embodiment of the present invention, dense point clouds are sequentially extracted from the geometric information of each high-level rendered image sub-block to construct local dense sub-scenes; each dense sub-scene corresponds to a triangular mesh; and the dense sub-scenes are sequentially aligned, smoothed, and optimized to generate a complete reconstructed target large scene.

[0098] Please refer to Figure 3 , Figure 3 is a schematic structural diagram of a large scene reconstruction system based on a 3D Gaussian algorithm in an embodiment of the present invention; an embodiment of the present invention also provides a large scene reconstruction system 100 based on a 3D Gaussian algorithm, and the system includes:

[0099] An image module 101, configured to obtain multiple original images covering the target large scene area, and screen the original images that meet the co-visibility condition from the multiple original images to form a first co-visibility original image set;

[0100] A sparse point cloud generation module 102, configured to generate a first sparse point cloud set based on the first co-visibility original image set by using the SfM algorithm, and calculate the point cloud boundary of the first sparse point cloud set;

[0101] A large scene segmentation module 103, based on the point cloud boundary, divides the target large scene area into multiple first grid units according to the original camera positions projected on the plane;

[0102] The boundary expansion module 104 is configured to expand the boundary of the initial first grid cell to obtain a plurality of second grid cells; each of the second grid cells contains a second co-visible original image subset and a second sparse point cloud subset;

[0103] The camera optimization module 105 is configured to add new cameras to the original camera set corresponding to each second grid cell based on the airspace-aware visibility standard to obtain an optimized camera set corresponding to each second grid cell;

[0104] The sparse point cloud optimization module 106 is configured to optimize the second sparse point cloud subset in each second grid cell in sequence according to the optimized camera set to obtain a corresponding third sparse point cloud subset;

[0105] The 3D Gaussian conversion module 107 is configured to adopt a 3D Gaussian sputtering model to perform 3D Gaussian point cloud distribution conversion, rasterization, and appearance modeling optimization processing on the third sparse point cloud subset corresponding to each second grid cell in sequence to obtain a plurality of primary rendered image sub-blocks;

[0106] The rendered image optimization module 108 constructs a comprehensive loss function and iteratively updates each primary rendered image sub-block based on the comprehensive loss function until the loss function converges to obtain a plurality of high-level rendered image sub-blocks;

[0107] The merging and reconstruction module 109 is configured to reconstruct a dense sub-scene using the geometric information of a plurality of high-level rendered image sub-blocks and seamlessly merge the dense sub-scenes to obtain the reconstructed target large scene.

[0108] An embodiment of the present invention also provides an electronic device, which includes: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the method described above.

[0109] An embodiment of the present invention also provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor executes the method described above.

[0110] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A large scene reconstruction method based on 3D Gaussian algorithm, characterized in that: The method comprises: S1 acquires a plurality of original images covering a large target scene area, and selects original images satisfying a common-view condition from the plurality of original images to form a first common-view original image set; the first common-view original image set includes a plurality of original images from different perspectives; S2 generates a first sparse point cloud set based on the first common view original image set using the SfM algorithm, and calculates the point cloud boundary of the first sparse point cloud set; then divides the target large scene area into a plurality of first grid units according to the point cloud boundary and the position of the original camera set projected on the plane by the first common view original image set; S3 expands the boundary of the first grid unit to obtain a plurality of second grid units; each second grid unit contains a second common view original image subset and a second sparse point cloud subset; S4 adds a new camera to the original camera set corresponding to each second grid unit to obtain an optimized camera set corresponding to each second grid unit; S5 optimizes each second sparse point cloud subset in turn according to the optimized camera set to obtain a corresponding third sparse point cloud subset; S6 adopts a 3D Gaussian sputtering model to perform 3D Gaussian point cloud distribution conversion, rasterization and appearance modeling optimization processing on each third sparse point cloud subset in turn to obtain multiple primary rendering image sub-blocks; S7 constructs a comprehensive loss function, and iteratively updates each primary rendered image sub-block in turn based on the comprehensive loss function until the loss function tends to be stable, thereby obtaining multiple advanced rendered image sub-blocks; S8 uses the geometric information of multiple high-level rendered image sub-blocks to reconstruct dense sub-scenes, and seamlessly merges the dense sub-scenes to obtain the reconstructed target large scene.

2. The method according to claim 1, characterized in that The 3D Gaussian sputtering model in step S6 includes a Gaussian point cloud distribution conversion unit, a differentiable Gaussian rasterization unit and a decoupled appearance modeling unit.

3. The method according to claim 2, characterized in that Step S6 further comprises: S61: using a Gaussian point cloud distribution conversion unit to sequentially convert the third sparse point cloud subset corresponding to each second grid unit into a group of primary 3D Gaussian point cloud distributions; and then convert each group of primary 3D Gaussian point cloud distributions into each primary 3D Gaussian image sub-block; S62 projects and renders each primary 3D Gaussian image sub-block onto a 2D image plane corresponding to each camera pose using a differentiable Gaussian rasterization unit to obtain a plurality of primary predicted image sub-blocks; S63 generates multiple primary transformation maps based on the multiple primary prediction image sub-blocks using a decoupled appearance modeling unit, and optimizes the appearance of each corresponding primary prediction image sub-block in turn according to each primary transformation map to obtain multiple primary rendering image sub-blocks.

4. The method according to claim 3, characterized in that Step S7 further comprises: S71 constructs a comprehensive loss function, wherein the comprehensive loss function includes a first loss function and a second loss function; wherein the first loss function includes a first difference loss function and a D-SSIM loss function, and the second loss function includes a second difference loss function; S72, based on the multiple primary rendered image sub-blocks, iteratively updating the parameters of each primary 3D Gaussian image sub-block by using the first loss function until the first loss function converges, to obtain multiple advanced 3D Gaussian image sub-blocks; S73 uses a differentiable Gaussian rasterization unit to render each high-level 3D Gaussian image sub-block onto a 2D image plane in sequence to obtain a plurality of high-level prediction image sub-blocks; S74 generates multiple corresponding high-level transformation maps based on multiple high-level prediction image sub-blocks using a decoupled appearance modeling unit, and performs iterative appearance optimization on each high-level prediction image sub-block in turn according to each high-level transformation map and the second loss function until the second loss function converges to obtain multiple high-level rendering image sub-blocks.

5. The method according to claim 4, characterized in that The first loss function in step S7 is: L=(1-β)L1+βL D-SSIM Wherein, L1 is the first difference loss function, which is used to calculate the absolute difference in pixel values ​​between the predicted image sub-block and the original image; L D-SSIM is the D-SSIM loss function, which is used to calculate the structural similarity difference between the predicted image sub-block and the original image; β is the weight of the D-SSIM loss function.

6. The method according to claim 4, characterized in that The parameters of the 3D Gaussian point cloud distribution include a center point, a covariance matrix, color information, and opacity.

7. The method according to claim 1, characterized in that The common viewing conditions in step S1 include: the satellite zenith angle of each original image is less than 30°, the angle between every two original images is less than 5°, and the shadow area of ​​each original image accounts for less than 5%.

8. A large scene reconstruction system based on 3D Gaussian algorithm, characterized in that: include: An image acquisition module is configured to acquire a plurality of original images covering a target large scene area, and select original images satisfying a common viewing condition from the plurality of original images to form a first common viewing original image set; A sparse point cloud generation module is configured to generate a first sparse point cloud set based on the first common view original image set by using the SfM algorithm, and calculate a point cloud boundary of the first sparse point cloud set; The large scene segmentation module further segments the target large scene area into a plurality of first grid units based on the point cloud boundary and the original camera position projected on the plane; A boundary expansion module is configured to expand the boundary of the initial first grid unit to obtain a plurality of second grid units; each second grid unit contains a second common view original image subset and a second sparse point cloud subset; A camera optimization module is configured to add a new camera to an original camera set corresponding to each second grid unit based on an airspace perception visibility standard to obtain an optimized camera set corresponding to each second grid unit; a sparse point cloud optimization module, configured to optimize the second sparse point cloud subset in each second grid unit in turn according to the optimized camera set to obtain a corresponding third sparse point cloud subset; A 3D Gaussian transformation module is configured to use a 3D Gaussian sputtering model to sequentially perform 3D Gaussian point cloud distribution transformation, rasterization, and appearance modeling optimization processing on the third sparse point cloud subset corresponding to each second grid unit to obtain a plurality of primary rendering image sub-blocks; The rendering image optimization module constructs a comprehensive loss function and iteratively updates each primary rendering image sub-block in turn based on the comprehensive loss function until the loss function tends to be stable, thereby obtaining multiple advanced rendering image sub-blocks; The merging and reconstruction module is configured to reconstruct a dense sub-scene using geometric information of multiple high-level rendered image sub-blocks, and seamlessly merge the dense sub-scenes to obtain a reconstructed target large scene.

9. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are enabled to perform the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Executable instructions are stored thereon, and when the instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • 3D modeling reconstruction system, method and device based on point cloud information and Gaussian cloud cluster

    CN118196306A

  • 3D scene reconstruction method and device, electronic equipment and computer readable medium

    CN118298000A