A 3D Gaussian reconstruction method for large-scene UAV images

Through the 3D Gaussian reconstruction method of chunked training and depth information matching of large-scene drone images, the problems of huge data volume, high computing complexity, insufficient video memory and poor visual effects in large-scene image processing are solved, and high-quality three-dimensional reconstruction and good visual effects are achieved.

CN119313828BActive Publication Date: 2025-06-06BAOLUE TECH (ZHEJIANG) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411854270.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-06-06
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

When processing large-scene drone images, the existing technology faces problems such as huge data volume, high computing complexity, insufficient video memory and poor visual effects, which makes it difficult to complete high-quality three-dimensional reconstruction.

Method used

A 3D Gaussian reconstruction method of large-scene drone images is adopted. By chunking large scenes and using depth information to match the required images for training, the chunking small scene training and fusion of high-pixel images is achieved, and the model rendering speed and detailed information are improved.

Benefits of technology

Through block training and depth information matching, this method significantly improves the model rendering speed and detailed information, ensures the visual effect of large scenes after 3D Gaussian reconstruction, and avoids memory overflow, achieving high-quality 3D reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119313828B_ABST
    Figure CN119313828B_ABST
Patent Text Reader

Abstract

The present invention relates to a 3D Gaussian reconstruction method for large-scene UAV images, including: using a UAV to shoot a large scene to obtain a number of UAV images; constructing a panoramic 3D Gaussian prior model; using the trained panoramic 3D Gaussian prior model to output the true depth value of each pixel in each UAV image; shrinking all Gaussian center points into a fixed area; evenly dividing the obtained fixed area into blocks; training each small area in turn; fusing all trained small areas to obtain a complete large scene. The method trains large scenes in blocks and uses depth information to match images required for training, thereby using high-pixel images to train small scenes in blocks and fusing the trained small scenes, greatly improving the model rendering speed and detail information, ensuring the visual effect of the large scene after three-dimensional Gaussian reconstruction, and avoiding memory overflow, thereby completing high-quality three-dimensional reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and deep learning technology, and in particular to a 3D Gaussian reconstruction method for large-scene drone images. Background Art

[0002] 3D Gaussian Splatting (3DGS) is a cutting-edge AI algorithm that can build high-quality 3D models from images. This technology uses a series of Gaussian functions to represent objects in three-dimensional space, and can efficiently render new perspective images with extremely high quality.

[0003] Compared with traditional photogrammetry technology, 3DGS technology has significant advantages in processing transparent or reflective surfaces, and can reduce the occurrence of visual defects such as holes even when image data is limited. Traditional photogrammetry methods often require manual repair of these defects, resulting in inefficient generation of 3D models and poor immersion. 3DGS technology can effectively capture the details of transparent and reflective areas in the scene through the transparency of Gaussian functions and the color simulation of spherical harmonics.

[0004] At present, 3DGS technology is widely used in many fields. For example, in medical image processing, 3DGS is used to process CT, MRI and other scanning data for three-dimensional reconstruction, achieving high-precision visualization of organs such as the liver and heart, allowing doctors to observe tissue structures more intuitively and determine treatment plans; in virtual reality and game development, 3DGS is used to generate high-quality natural environment effects, such as realistic rendering of translucent bodies such as smoke and fire clouds, to improve the sense of reality and immersion and increase user experience; in the digitization of ancient monuments and sites, drones or ground laser scanners are used to collect point cloud data at the site, and the 3DGS algorithm is used to convert these discrete point clouds into continuous, smooth three-dimensional surfaces, accurately restoring the details of ancient buildings and sculptures.

[0005] Although the 3D Gaussian reconstruction method surpasses the traditional 3D modeling method in terms of training speed and visual effect, it often faces the following challenges when processing large-scene drone images: (1) Huge data volume and high computational complexity: The number of images collected by drones is huge, and each image contains a large amount of pixel information, which places extremely high demands on computing resources. Conventional algorithms are difficult to complete high-quality 3D reconstruction while ensuring timeliness; (2) Video memory problem: The video memory of mainstream graphics cards is mostly 24GB, and the number of Gaussian point parameters is too large, which will lead to insufficient video memory during training; (3) Poor visual effect: Due to the lack of spatial information, overfitting is prone to occur during the training process, resulting in some blurred Gaussian points generated in front of the camera, creating floating objects in the scene, and resulting in poor visual effect of the rendered scene. Summary of the invention

[0006] In view of the problems existing in the prior art, the present invention provides a 3D Gaussian reconstruction method for large-scene UAV images, which is used to solve the technical problems that the prior art cannot complete high-quality three-dimensional reconstruction when faced with huge data volumes, insufficient video memory during overall training of large-scene images, and poor visual effects of large scenes after three-dimensional Gaussian reconstruction.

[0007] The technical solution adopted by the present invention is a 3D Gaussian reconstruction method for large-scene UAV images, which includes the following steps:

[0008] S1. Using a drone to photograph a large scene to obtain a number of drone images, wherein the drone images are drone images containing GPS information;

[0009] S2. Process each drone image to generate a corresponding pixel point cloud and camera parameters of a corresponding virtual camera, wherein the camera parameters include a camera coordinate position;

[0010] S3, constructing a panoramic 3D Gaussian prior model;

[0011] S4, compressing the size of the drone image obtained in step S1, and then inputting it into the panoramic 3D Gaussian prior model for training to obtain a trained panoramic 3D Gaussian prior model, wherein the trained panoramic 3D Gaussian prior model outputs a true depth value of each pixel in each drone image;

[0012] S5, reading all Gaussian center points of the trained panoramic 3D Gaussian prior model, and shrinking all Gaussian center points into a fixed area according to the camera coordinate position of the virtual camera corresponding to each drone image;

[0013] S6, evenly divide the fixed area obtained in step S5 into N×M small areas containing Gaussian center points;

[0014] S7, according to the true depth value of each pixel point in each drone image and the position of each pixel point obtained in step S4, project the coordinates of each pixel point in each drone image to the world coordinates to obtain a projection area image of each drone image on the world coordinates;

[0015] S8, for each small area obtained in step S6, check each drone image in turn, if part or all of the projection area of ​​the projection area image of the drone image on the world coordinates appears in the small area, then the part or all of the projection area appearing in the small area is taken as the visible area part, and a mask is generated for the remaining part except the visible area part, and for each small area, a projection area image of each drone image after the mask is generated is obtained;

[0016] S9, training the N×M small areas obtained in step S6 in sequence: for each small area, using the projection area image after the mask generation obtained in step S8 to train the small area, to obtain a trained small area;

[0017] S10, removing the overlapping parts of all the trained small areas obtained in step S9 and then fusing them to obtain a complete large scene.

[0018] The beneficial effect of the present invention is that the above-mentioned 3D Gaussian reconstruction method for large-scene drone images is adopted. The method trains the large scene in blocks and uses depth information to match the images required for training, thereby using high-pixel images to train small scenes in blocks and fusing the trained small scenes, greatly improving the model rendering speed and detail information, ensuring the visual effect of the large scene after three-dimensional Gaussian reconstruction, and avoiding memory overflow, thereby completing high-quality three-dimensional reconstruction.

[0019] Preferably, in step S2, the specific process of processing each drone image is: obtaining the camera parameters of the virtual camera corresponding to each drone image through metashape software, and aligning all drone images through image matching technology to generate their respective corresponding pixel point clouds; the camera parameters include camera coordinate position, camera shooting angle, camera viewing field and camera radial distortion.

[0020] Preferably, in step S4, The actual depth value of a pixel is specifically expressed as:

[0021] ;

[0022] in, Represents the z-axis depth value of the i-th 3D Gaussian center point in the camera coordinate space; , Indicates the transparency of the 3D Gaussian. represents the correlation function of 3D Gaussian projection to 2D pixel space; j represents an integer smaller than i, , Indicates the transparency of the 3D Gaussian.

[0023] Preferably, in step S4, during the training of the panoramic 3D Gaussian prior model, a loss function is used to calculate the training loss of the panoramic 3D Gaussian prior model; and the panoramic 3D Gaussian prior model is trained by gradient back propagation to obtain a trained panoramic 3D Gaussian prior model; the loss function is specifically expressed as:

[0024] ;

[0025] = ;

[0026] ;

[0027] = ;

[0028] ;

[0029] in, represents the total number of raw pixels in one of the drone images; represents the mean value of all original pixels of the drone image; represents the mean value of all rendered pixels of the drone image; represents the variance of all raw pixels of the drone image; represents the variance of all rendered pixels of the drone image; Represents the correlation between the original pixels of the drone image and the rendered pixels; and All represent constants;

[0030] ; Indicates the estimated depth value of each pixel estimated by the Depth Anything V2 algorithm; Represents the true depth value of each pixel output by the panoramic 3D Gaussian prior model in each round of training; represents opacity entropy regularization, Indicates the number of Gaussian centers visible after rendering. represents the sigmoid function, ; H represents the entropy function, ; o indicates opacity; , , as well as All represent coefficients.

[0031] Preferably, the specific process of step S5 includes the following steps:

[0032] S5.1. Read all Gaussian center points of the trained panoramic 3D Gaussian prior model, and project the Gaussian center points of the panoramic 3D Gaussian prior model and the camera coordinate positions corresponding to all drone images onto the XY plane;

[0033] S5.2. Get the median of the center points of the virtual cameras corresponding to all drone images , and then obtain the straight lines formed by all the center points through Hough line detection;

[0034] S5.3. The angle between the straight line and the x-axis constitutes a rotation angle, and a rotation matrix R is obtained according to the rotation angle. The rotation matrix R is specifically expressed as: ;in, Indicates the rotation angle;

[0035] S5.4. Rotate the Gaussian center point of the panoramic 3D Gaussian model according to the rotation matrix R to obtain a rotated Gaussian center point. The rotated Gaussian center point is specifically expressed as: ;in, Indicates Gaussian center points; Represents the center of the center points of all virtual cameras;

[0036] S5.5. Rotate the center point of the virtual camera corresponding to all drone images according to the rotation matrix R to obtain the center point of the rotated virtual camera, which is specifically expressed as: ;in, Indicates The center point of a virtual camera;

[0037] S5.6. According to the center point of the rotated virtual camera obtained in step S5.5, find its maximum and minimum values ​​on the X coordinate and the maximum and minimum values ​​on the Y coordinate, so as to determine a rectangular boundary, and all the center points of the rotated virtual cameras are located within the rectangular boundary. The rectangular boundary is specifically expressed as:

[0038] ;

[0039] ;

[0040] in, Represents the upper right coordinate of the rectangle boundary, Represents the lower left coordinate of the rectangle boundary;

[0041] S5.7, according to the rectangular boundary obtained in step S5.6, shrink all the rotated Gaussian center points obtained in step S5.4 to a fixed area The specific expression is:

[0042] ;

[0043] in, ; Indicates Gaussian center points, Represents a constant.

[0044] Preferably, in step S6, the (i, j)th small area is specifically expressed as:

[0045] ;

[0046] ;

[0047] Among them, N and M represent the number of times the x-axis direction and the y-axis direction are divided, respectively. Represents the overlap coefficient.

[0048] Preferably, in step S9, the process of training the small area further includes: removing redundant Gaussian center points; the specific process is: according to the depth value of the pixel point after rendering each drone image, all the Gaussian center points in each small area are converted into pixel coordinates, and according to the depth value of the Gaussian center point in the visible range, it is determined whether to delete the Gaussian center point; if the depth value of the Gaussian center point in the visible range satisfies:

[0049] , then delete the Gaussian center point; if it does not meet the requirement, then do not delete it; where, Indicates the depth value of the Gaussian center point; represents a constant; Indicates the pixel depth value after the drone image is rendered. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 This is a flow chart of a 3D Gaussian reconstruction method for large-scene UAV images according to the present invention;

[0051] Figure 2 It is a schematic diagram of evenly dividing a large scene into blocks as described in the present invention; wherein, Figure 2 (a) is a schematic diagram showing that the center points of all rotated virtual cameras are within the rectangular boundary; Figure 2 (b) is a schematic diagram of uniformly dividing a large scene into blocks;

[0052] Figure 3 Schematic diagram of the projection area image after the mask is generated in the present invention; wherein, Figure 3 (a) is a schematic diagram of the projection area image of the drone image on the world coordinates; Figure 3 (b) is a schematic diagram of a mask image generated based on a fixed area; Figure 3 (c) Based on Figure 3 (b) Schematic diagram of the projection area image after the generated mask;

[0053] Figure 4 This is a schematic diagram of rendering a segmented small scene in the present invention;

[0054] Figure 5 Rendering schematic diagram before and after removing floating objects in the present invention;

[0055] Figure 6 It is a schematic diagram of the overall rendering after the small scenes are fused in the present invention. DETAILED DESCRIPTION

[0056] The invention will be further described below with reference to the accompanying drawings and in combination with specific implementations, so that those skilled in the art can implement the invention with reference to the description. The protection scope of the invention is not limited to the specific implementations.

[0057] The technical solution adopted by the present invention is a 3D Gaussian reconstruction method for large-scene drone images, such as Figure 1 As shown, the method comprises the following steps:

[0058] Step 1: Get drone images of large scenes

[0059] Using a drone to shoot a large scene to obtain a number of drone images, wherein the drone images are drone images containing GPS information and are high-pixel images;

[0060] Furthermore, the specific process of obtaining a number of drone images is: using a drone to photograph a large scene, pre-planning the flight path of the drone, automatically generating a flight route plan through the drone software, adjusting the drone to automatic flight mode to perform flight missions according to the planned route, and at the same time changing the altitude of the drone, photographing the large scene again at a pre-set level and altitude, to obtain a number of drone images at different altitude levels, each drone image is a high-pixel image and contains GPS information.

[0061] Step 2: UAV image preprocessing

[0062] Processing each drone image to generate a corresponding pixel point cloud and camera parameters of a corresponding virtual camera, wherein the camera parameters include camera coordinate positions;

[0063] Furthermore, the specific process of processing each drone image is as follows: obtaining the camera parameters of the virtual camera corresponding to each drone image through the metashape software, that is, each drone image corresponds to a virtual camera, and the virtual camera has corresponding camera parameters, including camera coordinate position, camera shooting angle, camera field of view and camera radial distortion; and aligning all drone images through image matching technology to generate their corresponding pixel point clouds.

[0064] Step 3: Construct a panoramic 3D Gaussian prior model

[0065] 3D Gaussian Splatting is a method that uses a three-dimensional Gaussian function to transform discrete information points (pixel point cloud) into a smooth and continuous three-dimensional model, where the three-dimensional Gaussian function is defined as:

[0066] ;

[0067] in, represents the center point of the three-dimensional Gaussian, represents the 3D covariance matrix, , R represents the rotation matrix, S represents the scaling matrix. Each three-dimensional Gaussian has a series of spherical harmonic coefficients, collectively referred to as SH. When , 2D pixel The corresponding colors are:

[0068] ;

[0069] Among them, i represents the 3D Gaussian center point that contributes to the pixel point. Represents pixel Color, , Indicates the transparency of the 3D Gaussian. Represents the correlation function of 3D Gaussian projection to 2D pixel space. If the coordinates of the Gaussian center projected to the pixel space are The farther The smaller it is; SH represents spherical harmonics.

[0070] Step 4: Estimate the true depth value of each pixel in the drone image

[0071] In order to reduce the resolution of the image, the size of the drone image obtained in step 1 is downsampled, and the pixel size of the obtained drone image is about 1 / 4 of the original image. Then the compressed image is input into the panoramic 3D Gaussian prior model for training to obtain a trained panoramic 3D Gaussian prior model. The trained panoramic 3D Gaussian prior model is used to output the true depth value of each pixel in each drone image.

[0072] Further, in step 4, the first The actual depth value of a pixel is specifically expressed as:

[0073] ;

[0074] ;

[0075] in, Represents the z-axis depth value of the i-th 3D Gaussian center point in the camera coordinate space; , Indicates the transparency of the 3D Gaussian. Represents the correlation function of 3D Gaussian projection to 2D pixel space.

[0076] Furthermore, in the process of training the panoramic 3D Gaussian prior model, a loss function is used to calculate the training loss of the panoramic 3D Gaussian prior model; and the panoramic 3D Gaussian prior model is trained by gradient back propagation to obtain a trained panoramic 3D Gaussian prior model; the loss function is specifically expressed as:

[0077] ;

[0078] = ;

[0079] ;

[0080] = ;

[0081] ;

[0082] in, represents the total number of raw pixels in one of the drone images; represents the mean value of all original pixels of the drone image; represents the mean value of all rendered pixels of the drone image; represents the variance of all raw pixels of the drone image; represents the variance of all rendered pixels of the drone image; Represents the correlation between the original pixels of the drone image and the rendered pixels; and All represent constants;

[0083] ; Indicates the estimated depth value of each pixel estimated by the Depth Anything V2 algorithm; Represents the true depth value of each pixel output by the panoramic 3D Gaussian prior model in each round of training; represents opacity entropy regularization, Indicates the number of Gaussian centers visible after rendering. represents the sigmoid function, ; H represents the entropy function, ; o indicates opacity; , , as well as All represent coefficients;

[0084] Among them, Depth Anything V2 is a depth estimation model that can infer the three-dimensional structure of a scene from a single two-dimensional image. Therefore, a monocular depth estimation algorithm such as Depth Anything V2 is used to generate an estimated depth value in the real world for each image.

[0085] In the process of training the panoramic 3D Gaussian prior model, in order to prevent the problem of memory overflow, opacity entropy regularization is introduced. ;

[0086] The opacity entropy regularization loss function makes the transparency tend to 1 or 0. Combined with the algorithm's own process of deleting Gaussian functions with too small transparency, the model's memory consumption can be further reduced, thereby realizing the training of a one-time panoramic 3D Gaussian prior model.

[0087] The use of downsampling and opacity regularization loss functions will cause the loss of details of the reconstructed 3D model, but at this stage, the main purpose is to obtain Gaussian distribution and reliable depth information, without too many requirements on the details and texture of the model.

[0088] Step 5: Divide the large scene into blocks evenly

[0089] Read all Gaussian center points of the trained panoramic 3D Gaussian prior model, and shrink all Gaussian center points into a fixed area according to the camera coordinate position of the virtual camera corresponding to each drone image;

[0090] The obtained fixed area is evenly divided into N×M small areas containing Gaussian center points;

[0091] Furthermore, the specific process of shrinking all Gaussian center points into a fixed area includes:

[0092] Since the flying height does not change much during drone collection, the z-axis can be removed first during block processing; all Gaussian center points of the trained panoramic 3D Gaussian prior model are read, and the Gaussian center points of the panoramic 3D Gaussian prior model and the camera coordinate positions corresponding to all drone images are projected onto the XY plane;

[0093] Get the median of the center points of the virtual cameras corresponding to all drone images , and then use Hough line detection to get the straight line formed by all the center points;

[0094] The angle between the straight line and the x-axis constitutes a rotation angle, and a rotation matrix R is obtained according to the rotation angle. The rotation matrix R is specifically expressed as: ;in, Indicates the rotation angle;

[0095] The Gaussian center point of the panoramic 3D Gaussian model is rotated according to the rotation matrix R to obtain the rotated Gaussian center point, and the rotated Gaussian center point is specifically expressed as: ;in, Indicates Gaussian center points; Represents the center of the center points of all virtual cameras;

[0096] According to the rotation matrix R, the center point of the virtual camera corresponding to all drone images is rotated to obtain the center point of the rotated virtual camera, which is specifically expressed as: ;in, Indicates The center point of a virtual camera;

[0097] According to the center point of the rotated virtual camera, find its maximum and minimum values ​​on the X coordinate and the maximum and minimum values ​​on the Y coordinate, so as to determine a rectangular boundary, such as Figure 2 As shown in (a), the center points of all the rotated virtual cameras are located within the rectangular boundary, and the rectangular boundary is specifically represented as:

[0098] ;

[0099] ;

[0100] in, Represents the upper right coordinate of the rectangle boundary, Represents the lower left coordinate of the rectangle boundary;

[0101] Figure 2 In the figure, the center point of the rotated virtual camera is the red point in the figure, which is located within the rectangular boundary, and the center point of the rotated Gaussian is the yellow point in the figure;

[0102] According to the obtained rectangular boundary, all the obtained rotated Gaussian center points are shrunk to a fixed area The specific expression is:

[0103] ;

[0104] in, ; Indicates Gaussian center points, represents a constant;

[0105] For a fixed region containing a Gaussian center point Evenly divide the blocks into N×M blocks of small areas containing Gaussian center points, that is, N×M blocks of small scenes are obtained, such as Figure 2 (b) shows that the (i,j)th small area is represented as:

[0106] ;

[0107] ;

[0108] Among them, N and M represent the number of times the x-axis direction and the y-axis direction are divided, respectively. represents the overlap coefficient, Generally, 10% is used, that is, 10% more range is taken in the x-axis and y-axis directions each time the blocks are divided. The introduction of the overlap coefficient facilitates the processing of the boundary area during subsequent fusion, so that there is no sense of fragmentation in the visual effect of the fused large scene.

[0109] Step 6: Get the visible area of ​​each small scene

[0110] According to the true depth value of each pixel point in each drone image and the position of each pixel point obtained in step 4, the coordinates of each pixel point in each drone image are projected to the world coordinates to obtain the projection area image of each drone image on the world coordinates;

[0111] For each small area obtained, each of the drone images is checked in turn. If part or all of the projection area of ​​the projection area image of the drone image on the world coordinates appears in the small area, the part or all of the projection area appearing in the small area is taken as the visible area part, and a mask is generated for the remaining part except the visible area part. Then, for each small area, a projection area image after mask generation of each drone image is obtained;

[0112] like Figure 3 As shown, Figure 3 (a) is the projection area image of the drone image in the world coordinates; Figure 3 (b) is a mask image generated based on a fixed area, where the white part is within the fixed area and the black part is outside the fixed area. Figure 3 (c) Based on Figure 3 (b) The generated masked projection area image, where the visible part is trained and the black part is not trained.

[0113] Step 7: Perform distribution training on each small scene by generating the masked projection area image

[0114] The N×M blocks of small areas obtained in the step are trained: for each of the small areas, the corresponding projection area image after the mask generation is used to train the small area of ​​the block, so as to obtain the trained small area;

[0115] Since fewer images are needed for training each small area, the UAV image can be directly used to perform distribution training on each small area to achieve better visual effects. The specific rendering effect is as follows: Figure 4 As shown, Figure 4 In the middle, there is no obvious sense of separation and splicing marks near the dividing line, and each area is trained in parallel to reduce the time required for training;

[0116] Furthermore, in the small area training process, it also includes: removing redundant Gaussian center points, that is, removing redundant floating objects; the specific process is: according to the pixel depth value after each drone image is rendered, all Gaussian center points in each small area are converted into pixel coordinates, and according to the depth value of the Gaussian center point in the visible range, it is determined whether to delete the Gaussian center point; if the depth value of the Gaussian center point in the visible range satisfies:

[0117] , then delete the Gaussian center point to achieve the purpose of removing floating objects. For details, please refer to Figure 5 ; If not satisfied, it will not be deleted; Among them, Indicates the depth value of the Gaussian center point; represents a constant; Indicates the pixel depth value after the drone image is rendered.

[0118] Step 7: Fuse the trained small scenes to get a complete large scene

[0119] After each small area is trained, 10% of the overlapping parts are removed and these areas are merged; the fused large scene retains more detail information, and there are no obvious segmentation marks in the large scene, which has a higher visual effect and meets the needs of actual production. The details are as follows Figure 6 shown.

[0120] The following experiments are used to illustrate the superiority of the 3D Gaussian reconstruction method for large-scene UAV images of the present invention.

[0121] Three large-scale 3D scene reconstruction datasets are used to conduct experiments and present the experimental results, namely Rubble and Building of the Mill-19 dataset, and Residence of the UrbanScene3D dataset.

[0122] Mill-19: This dataset consists of two scenes around a former industrial complex. The Mill-19 building consists of shots taken in a grid pattern across a large The area around the industrial building. The Rubble in Mill-19 contains 1657 original images with a resolution of 4608×3456, and the Building contains 1920 original images with a resolution of 4608×3456.

[0123] UrbanScene3D: This dataset contains more than 128,000 high-resolution images covering 16 scenes, including large real urban scenes and virtual urban scenes with a total area of ​​136 square kilometers. It also contains high-precision LiDAR scans and hundreds of image sets with different observation modes. Residence in UrbanScene3D contains 2582 original images with a resolution of 5480×3648.

[0124] Evaluation indicators: SSIM, PSNR, LPIPS and the number of Gaussian points are used as evaluation indicators of the 3DGS model on the Mill-19 and UrbanScene3D datasets. Among them, SSIM refers to the structural similarity index used to measure the similarity between two images. The larger the SSIM value, the higher the similarity between the two images and the better the reconstruction effect. PSNR refers to the peak signal-to-noise ratio, which is an indicator of image quality. The larger the PSNR value, the better the image quality. LPIPS is the learning perception of image block similarity, also known as perceptual loss, which is used to measure the difference between two images. The lower the LPIPS value, the more similar the two images are and the better the reconstruction effect. In addition, the number of Gaussian points in the model is also recorded.

[0125] Training details: The experiments were conducted on a Linux system equipped with an RTX4090D GPU. The panoramic 3D Gaussian prior model was trained for 40,000 iterations using images with a width of 1,000 pixels after downsampling. Densification was performed from iteration 1,000 to 30,000, with a densification interval of 200, and scene floats were removed every 7,000 iterations. After that, each small scene after segmentation was trained for 30,000 iterations, and densification was performed from iteration 1,000 to 15,000, with a densification interval of 100, and scene floats were removed every 7,000 iterations.

[0126] During training, the loss function is: ;

[0127] right , as well as The value of is: , , ;

[0128] Through experiments, the experimental evaluation index results in Table 1 are obtained. Table 1 shows the experimental evaluation index results of different 3D reconstruction models in four large scene 3D reconstruction datasets: Building, Rubble, Residence, and Sci-Art. Among them, SSIM is the structural similarity index. The larger the value, the closer the reconstructed 3D scene is to the real image and the better the visual effect. PSNR is the peak signal-to-noise ratio. The larger the value, the higher the image quality. LPIPS is a method to measure image similarity. The smaller the value, the more realistic the reconstructed 3D scene is. It can be seen from Table 1 that the method of the present invention performs well under different evaluation indicators in multiple different datasets, especially for Building and Rubble datasets. The reason for the slightly lower scores in Residence and Sci-Art is that the original images are overexposed or too dark, resulting in poor evaluation indicators.

[0129] Table 1 Experimental evaluation index table of different 3D reconstruction models for different scenes

[0130] .

Claims

1. A 3D Gaussian reconstruction method for large-scene drone images, characterized by: The method comprises the following steps: S1. Using a drone to photograph a large scene to obtain a number of drone images, wherein the drone images are drone images containing GPS information; S2. Process each drone image to generate a corresponding pixel point cloud and camera parameters of a corresponding virtual camera, wherein the camera parameters include a camera coordinate position; S3, constructing a panoramic 3D Gaussian prior model; S4, compressing the size of the drone image obtained in step S1, and then inputting it into the panoramic 3D Gaussian prior model for training to obtain a trained panoramic 3D Gaussian prior model, wherein the trained panoramic 3D Gaussian prior model outputs a true depth value of each pixel in each drone image; S5, reading all Gaussian center points of the trained panoramic 3D Gaussian prior model, and shrinking all Gaussian center points into a fixed area according to the camera coordinate position of the virtual camera corresponding to each drone image; S6, evenly divide the fixed area obtained in step S5 into N×M small areas containing Gaussian center points; S7, according to the true depth value of each pixel point in each drone image and the position of each pixel point obtained in step S4, project the coordinates of each pixel point in each drone image to the world coordinates to obtain a projection area image of each drone image on the world coordinates; S8, for each small area obtained in step S6, check each drone image in turn, if part or all of the projection area of ​​the projection area image of the drone image on the world coordinates appears in the small area, then the part or all of the projection area appearing in the small area is taken as the visible area part, and a mask is generated for the remaining part except the visible area part, and for each small area, a projection area image of each drone image after the mask is generated is obtained; S9, training the N×M small areas obtained in step S6 in sequence: for each small area, using the projection area image after the mask generation obtained in step S8 to train the small area, to obtain a trained small area; S10, removing the overlapping parts of all the trained small areas obtained in step S9 and then fusing them to obtain a complete large scene.

2. The 3D Gaussian reconstruction method of large-scene UAV images according to claim 1 is characterized by: In step S2, the specific process of processing each drone image is as follows: obtaining the camera parameters of the virtual camera corresponding to each drone image through the metashape software, and aligning all the drone images through image matching technology to generate their respective corresponding pixel point clouds; the camera parameters include camera coordinate position, camera shooting angle, camera viewing field and camera radial distortion.

3. The 3D Gaussian reconstruction method of large-scene UAV images according to claim 1 or 2, characterized in that: In step S4, The actual depth value of a pixel is specifically expressed as: ; ; in, Represents the z-axis depth value of the i-th 3D Gaussian center point in the camera coordinate space; , Indicates the transparency of the 3D Gaussian. represents the correlation function of 3D Gaussian projection to 2D pixel space; j represents an integer smaller than i, , Indicates the transparency of the 3D Gaussian.

4. The 3D Gaussian reconstruction method of large-scene UAV images according to claim 3 is characterized by: In step S4, during the training of the panoramic 3D Gaussian prior model, a loss function is used to calculate the training loss of the panoramic 3D Gaussian prior model; and the panoramic 3D Gaussian prior model is trained by gradient back propagation to obtain a trained panoramic 3D Gaussian prior model; the loss function is specifically expressed as: ; = ; ; = ; ; in, represents the total number of raw pixels in one of the drone images; represents the mean value of all original pixels of the drone image; represents the mean value of all rendered pixels of the drone image; represents the variance of all raw pixels of the drone image; represents the variance of all rendered pixels of the drone image; Represents the correlation between the original pixels of the drone image and the rendered pixels; and All represent constants; ; Indicates the estimated depth value of each pixel estimated by the Depth Anything V2 algorithm; Represents the true depth value of each pixel output by the panoramic 3D Gaussian prior model in each round of training; represents opacity entropy regularization, Indicates the number of Gaussian centers visible after rendering. represents the sigmoid function, ; H represents the entropy function ; o indicates opacity; , , as well as All represent coefficients.

5. The 3D Gaussian reconstruction method of large-scene UAV images according to claim 4 is characterized in that: The specific process of step S5 includes the following steps: S5.

1. Read all Gaussian center points of the trained panoramic 3D Gaussian prior model, and project the Gaussian center points of the panoramic 3D Gaussian prior model and the camera coordinate positions corresponding to all drone images onto the XY plane; S5.

2. Get the median of the center points of the virtual cameras corresponding to all drone images , and then obtain the straight lines formed by all the center points through Hough line detection; S5.

3. The angle between the straight line and the x-axis constitutes a rotation angle, and a rotation matrix R is obtained according to the rotation angle. The rotation matrix R is specifically expressed as: ;in, Indicates the rotation angle; S5.

4. Rotate the Gaussian center point of the panoramic 3D Gaussian model according to the rotation matrix R to obtain a rotated Gaussian center point. The rotated Gaussian center point is specifically expressed as: ;in, Indicates Gaussian center points; Represents the center of the center points of all virtual cameras; S5.

5. Rotate the center point of the virtual camera corresponding to all drone images according to the rotation matrix R to obtain the center point of the rotated virtual camera, which is specifically expressed as: ;in, Indicates The center point of a virtual camera; S5.

6. According to the center point of the rotated virtual camera obtained in step S5.5, find its maximum and minimum values ​​on the X coordinate and the maximum and minimum values ​​on the Y coordinate, so as to determine a rectangular boundary, and all the center points of the rotated virtual cameras are located within the rectangular boundary. The rectangular boundary is specifically expressed as: ; ; in, Represents the upper right coordinate of the rectangle boundary, Represents the lower left coordinate of the rectangle boundary; S5.7, according to the rectangular boundary obtained in step S5.6, shrink all the rotated Gaussian center points obtained in step S5.4 to a fixed area The specific expression is: ; in, ; Indicates Gaussian center points, Represents a constant.

6. The 3D Gaussian reconstruction method of large-scene UAV images according to claim 5 is characterized by: In step S6, the (i, j)th small area is specifically expressed as: ; ; Among them, N and M represent the number of times the a-axis direction and the y-axis direction are divided, respectively. Represents the overlap coefficient.

7. The 3D Gaussian reconstruction method of large-scene UAV images according to claim 6 is characterized by: In step S9, the process of training the small area also includes: removing redundant Gaussian center points; the specific process is: according to the depth value of the pixel point after rendering each drone image, all Gaussian center points in each small area are converted into pixel coordinates, and according to the depth value of the Gaussian center point in the visible range, it is determined whether to delete the Gaussian center point; if the depth value of the Gaussian center point in the visible range satisfies: , then delete the Gaussian center point; if it does not meet the requirement, then do not delete it; where, Indicates the depth value of the Gaussian center point; represents a constant; Indicates the pixel depth value after the drone image is rendered.

Citation Information

Patent Citations

  • Three-dimensional scene arbitrary angle model segmentation method based on 3D Gaussian Splitting

    CN118351277A

  • Three-dimensional model reconstruction method based on 3D Gaussian Splitting

    CN119091051A