Method and system for optimizing unmanned aerial vehicle aerial photography scene reconstruction based on GAN model
By using a multi-source data fusion and optimization mechanism based on the GAN model, the problems of insufficient parameter mining and lack of guidance mechanism in 3DGS large scene reconstruction are solved, and high-fidelity reconstruction of drone aerial photography scenes is achieved.
Patent Information
- Application Number
- CN202511425066.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2025-12-23
AI Technical Summary
Existing 3DGS large-scene reconstruction methods suffer from insufficient parameter mining and a lack of effective guidance mechanisms, resulting in insufficient realism of the reconstructed models and difficulty in robustly coping with complex flight attitudes and environmental interference.
A GAN-based approach is adopted to fuse satellite imagery, point cloud data, and UAV aerial images into multi-source data. An optimization mechanism is formed through the GAN model to gradually guide and optimize the 3D Gaussian sputtering model, thereby improving the scene reconstruction capability.
By fusing multi-source data and optimizing the GAN model, the realism and reconstruction quality of drone aerial photography scenes were significantly improved, the generalization ability of the generator and discriminator was enhanced, and the application scope was expanded.
Smart Images

Figure CN121190677A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of scene reconstruction, and relates to a technology of scene reconstruction by using unmanned aerial vehicle aerial images, in particular to a method and system for optimizing unmanned aerial vehicle aerial scene reconstruction based on a GAN model. BACKGROUND
[0002] Nowadays, the application of unmanned aerial vehicles is more and more extensive. Among them, the implementation of 3D Gaussian sputtering for large scene reconstruction by using unmanned aerial vehicles is an important application of unmanned aerial vehicles. It is to collect multi-view sequence images of a target area by using an unmanned aerial vehicle, and to perform three-dimensional reconstruction on a scene by using computer vision and deep learning algorithms. This technology has important application value in the fields of city planning, disaster emergency assessment, geographic information system updating, etc.
[0003] The traditional 3DGS (3D Gaussian sputtering) large scene reconstruction method often causes insufficient fidelity of the reconstructed model and a significant gap with the real scene when processing image data collected by unmanned aerial vehicles due to factors such as complex flight posture, environmental interference and data noise. This is mainly due to two bottlenecks: first, insufficient parameter mining, the existing methods mostly rely on dense design rules and simple loss functions (such as L1 loss function, SSIM loss function) designed by hand, and the mining ability of the deep relationship between Gaussian parameters (position, color, covariance, etc.) and complex scene features is limited, which restricts the restoration ability of the model; second, there is a lack of effective guidance mechanism for optimization. The traditional gradient descent relies on a pre-defined mathematical objective, and lacks a mechanism that can dynamically evaluate and feedback the "visual reality" to adaptively guide the model optimization, making it difficult to robustly cope with data defects and scene diversity. SUMMARY
[0004] In view of the technical problems of the existing 3DGS large scene reconstruction method described in the background technology, due to insufficient parameter mining in the 3DGS algorithm and lack of effective guidance mechanism for optimization, the present application proposes a method and system for optimizing unmanned aerial vehicle aerial scene reconstruction based on a GAN model to solve the technical problems.
[0005] The present application improves the restoration ability of the 3D Gaussian sputtering model constructed by the 3DGS algorithm by performing multi-source data fusion on satellite image data I sat , point cloud data P lidar , unmanned aerial vehicle aerial images I photo and multi-view rendered images I rend ; and forms an optimization mechanism by using a GAN model to continuously guide and optimize the 3D Gaussian sputtering model, thereby improving the scene reconstruction ability of the 3D Gaussian sputtering model and the fidelity of the unmanned aerial vehicle aerial scene.
[0006] To solve the above technical problems, the application adopts the following technical solutions: A method for optimizing unmanned aerial vehicle aerial scene reconstruction based on a GAN model, comprising the following steps: Feature vector acquisition: acquiring satellite image data I sat , point cloud data P lidar , and unmanned aerial vehicle aerial image I photo based on the unmanned aerial vehicle aerial image I photo ; rend ; The unmanned aerial vehicle aerial image I photo , the multi-view rendering image I rend , the satellite image data I sat , and the point cloud data P lidar are respectively subjected to feature extraction and fusion to obtain generated data fusion feature vector f fake =[f sat ;f lidar ;f rend ] and real data fusion feature vector f real =[f sat ;f lidar ;f photo ]; Discriminator update: based on the update condition of the discriminator, the discriminator in the GAN model is updated through the generated data fusion feature vector f fake =[f sat ;f lidar ;f rend ] and the real data fusion feature vector f real =[f sat ;f lidar ;f photo ], and the updated discriminator generates a discrimination result; Generator update: extracting the model parameters of the 3D Gaussian sputtering model; generating joint encoding features, taking the joint encoding features as the input of the generator in the GAN model, generating parameter increments through the generator, and updating the model parameters of the 3D Gaussian sputtering model according to the parameter increments using the LOD dynamic update mechanism; Iteration: repeating the feature vector acquisition, discriminator update, and generator update steps until the iteration is completed; Scene reconstruction: updating the 3D Gaussian sputtering model using the last updated model parameters as the optimal 3D Gaussian sputtering model, and performing unmanned aerial vehicle aerial scene reconstruction based on the optimal 3D Gaussian sputtering model.
[0007] Further limitation, in the feature vector acquisition step, the unmanned aerial vehicle aerial image I photo, multi-view rendered image I rend , satellite image data I sat , and point cloud data P lidar , respectively, to obtain generated data fusion feature vector f fake =[f sat ;f lidar ;f rend ] and real data fusion feature vector f real =[f sat ;f lidar ;f photo ], specifically: The feature vectors of the unmanned aerial vehicle aerial image I photo , multi-view rendered image I rend , satellite image data I sat , and point cloud data P lidar are extracted respectively, and the corresponding feature vectors f photo , feature vector f rend , feature vector f sat , and feature vector f lidar are obtained; The feature vector f photo and the feature vector f rend are fused with the feature vector f sat and the feature vector f lidar respectively to form generated data fusion feature vector f fake =[f sat ;f lidar ;f rend ] and real data fusion feature vector f real =[f sat ;f lidar ;f photo ].
[0008] Further limitation, the model parameters include Gaussian center coordinates μ i , covariance matrix ∑ i , opacity α i and appearance feature c i , i represents the serial number of Gaussian distribution in 3D Gaussian splashing model.
[0009] Further limitation, in the generator updating step, the generator generates parameter increment specifically includes: The Gaussian center coordinates μ i and the covariance matrix ∑ i are combined to form an input vector V i , and the input vector V i is extracted by using a multilayer perception algorithm to output geometric feature encoding E geo (V i ). Appearance features c are processed using a one-dimensional convolutional network. i Feature extraction is performed to obtain the appearance feature code E. app (c i ); Encode geometric features E geo (V i ) and appearance feature code E app (c i The concatenation process yields the joint code h. i =[E geo (V i E app (c i )], will jointly encode h i As input to the generator in the GAN model, the parameter increments are generated by the generator.
[0010] Further specifying, in the generator update step, the generation process of the LOD dynamic update mechanism is as follows: Calculate geometric error ε i ; Using the Gaussian center coordinates μ i The sum of the covariance matrix ∑ i The coarse optimization stage LOD0 is formed by combining the parameter increment ΔG; Using the Gaussian center coordinates μ i , covariance matrix ∑ i and opacity α i The LOD1 in the optimization stage is formed by combining the parameter increment ΔG; Using the Gaussian center coordinates μ i , covariance matrix ∑ i Opacity α i and appearance features c i The LOD2 fine optimization stage is formed by combining the parameter increment ΔG; Based on the coarse optimization stage LOD0, the intermediate optimization stage LOD1, and the fine optimization stage LOD2, the geometric error ε is... i As a condition for judgment, a dynamic LOD update mechanism is generated.
[0011] Further specifying, in the feature vector acquisition step, a multi-view rendering image I is generated through differentiable rasterization rendering. rend Specifically, it includes: The virtual camera uniformly acquires K virtual camera poses in the scene space. The scene space is the space reconstructed using a 3D Gaussian sputtering model or an optimized 3D Gaussian sputtering model, and m is the virtual camera pose index. Using differentiable rendering technology Generate the pose corresponding to each virtual camera a rendered image, i.e., a multi-view rendered image I rend , wherein, , , and are updated model parameters, is a projection function; is a projection transformation parameter of a virtual camera; M is the total number of Gaussian cells in the 3D Gaussian Splatting model; j is a Gaussian cell index.
[0012] Further limited, the generator and the discriminator are both formed by optimizing the total loss of the adversarial loss, the perceptual loss and the geometric consistency loss.
[0013] Further limited, in the iteration step, the condition for completing the iteration is to reach the maximum iteration number or the 3D Gaussian Splatting model has converged.
[0014] The system for optimizing the UAV aerial scene reconstruction based on the GAN model formed by the above method comprises: a feature vector acquisition module: used for acquiring satellite image data I sat , point cloud data P lidar and UAV aerial image I photo in the same scene, based on the UAV aerial image I photo , a 3D Gaussian Splatting model is constructed by using a 3DGS algorithm, and a multi-view rendered image I rend is generated by differentiable rasterization rendering; the UAV aerial image I photo , the multi-view rendered image I rend , the satellite image data I sat and the point cloud data P lidar are respectively subjected to feature extraction and fusion, so as to obtain a generated data fusion feature vector f fake =[f sat ;f lidar ;f rend ] and a real data fusion feature vector f real =[f sat ;f lidar ;f photo ]; a discriminator updating module: used for updating the discriminator based on the updating condition of the discriminator, by using the generated data fusion feature vector f fake =[f sat ;f lidar ;f rend ] and the real data fusion feature vector f real =[f sat ;f lidar ;f photoThe discriminator in the GAN model is updated, and a discrimination result is generated through the updated discriminator; The generator updating module is used for extracting model parameters of the 3D Gaussian sputtering model, generating joint coding features, taking the joint coding features as inputs of a generator in the GAN model, generating a parameter increment through the generator, and updating the model parameters of the 3D Gaussian sputtering model according to the parameter increment by using the LOD dynamic updating mechanism; The iteration module is used for repeating the feature vector obtaining module, the discriminator updating module and the generator updating module until iteration is completed. The scene reconstruction module is used for updating the 3D Gaussian sputtering model by using the last updated model parameters as an optimal 3D Gaussian sputtering model, and reconstructing a scene of a UAV aerial photograph based on the optimal 3D Gaussian sputtering model.
[0015] A memory stores a program file, and the program file is executed to realize the program instructions of the method for optimizing the scene reconstruction of the UAV aerial photograph based on the GAN model.
[0016] Compared with the prior art, the method has the advantages that: 1. The method for optimizing the scene reconstruction of the UAV aerial photograph based on the GAN model is provided, and the steps of the method include feature vector obtaining, discriminator updating, generator updating, iteration and scene reconstruction. sat The satellite image data I lidar , the point cloud data P photo , the UAV aerial image I rend and the multi-view rendering image I i are fused to improve the restoration ability of the 3D Gaussian sputtering model constructed by the 3DGS algorithm. The 3D Gaussian sputtering model is continuously guided and optimized by the optimization mechanism formed by the GAN model, and the scene reconstruction ability of the 3D Gaussian sputtering model is improved, so that the fidelity of the UAV aerial scene is improved.
[0017] 2. The method is based on the coarse optimization stage LOD0, the medium optimization stage LOD1 and the fine optimization stage LOD2, and the geometric error ε i is taken as a judgment condition to generate the LOD dynamic updating mechanism. Based on the LOD dynamic updating mechanism, the key optimization region can be adaptively adjusted according to the geometric error of different regions in the scene. From the macro scene layout to the micro detail description, the 3D Gaussian sputtering model is gradually improved, and the scene reconstructed by the 3D Gaussian sputtering model can highly match the real scene at each level.
[0018] 3、The application utilizes differentiable rendering technology, LOD dynamic updating mechanism and dynamic loss to evaluate the difference between the optimized 3D Gaussian sputtering model reconstructed scene and the real scene, accurately guides the 3D Gaussian sputtering model to generate more realistic results, and significantly improves the geometric accuracy, appearance details and overall visual effect of the generated 3D Gaussian sputtering model.
[0019] 4、The application updates the generator and the discriminator based on the deep joint training and dynamic parameter adjustment mechanism of the GAN, so that the generator and the discriminator can better adapt to different types of real scene data, enhance the generalization ability of the generator and the discriminator, effectively optimize the 3D Gaussian sputtering model, improve the quality and realism of the 3D Gaussian sputtering model reconstructed scene, and expand the application range of the 3D Gaussian sputtering model. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 A schematic diagram of the method for optimizing the unmanned aerial vehicle aerial scene reconstruction based on the GAN model of the application; Figure 2 A schematic diagram of the system for optimizing the unmanned aerial vehicle aerial scene reconstruction based on the GAN model of the application. DETAILED DESCRIPTION
[0021] The technical solutions of the application will be further explained and described below in combination with the drawings and embodiments, but the application is not limited to the following described embodiments.
[0022] Reference Figure 1 , the method for optimizing the unmanned aerial vehicle aerial scene reconstruction based on the GAN model of the application comprises the following steps: Feature vector acquisition: acquiring satellite image data I sat , point cloud data P lidar and unmanned aerial vehicle aerial image I photo based on the unmanned aerial vehicle aerial image I photo , constructing a 3D Gaussian sputtering model by using a 3DGS algorithm, the 3D Gaussian sputtering model generating multi-view rendering images I rend through differentiable rasterization rendering; performing feature extraction and fusion on the unmanned aerial vehicle aerial image I photo , the multi-view rendering images I rend , the satellite image data I sat and the point cloud data P lidar respectively to obtain generated data fusion feature vector f fake =[f sat ;f lidar ;f rend ] and real data fusion feature vector f real =[f sat ;f lidar ;f photo]; wherein the model parameters include Gaussian center coordinates μ i , covariance matrix ∑ i , opacity a i , and appearance feature c i , i represents the serial number of the Gaussian distribution in the 3D Gaussian splatting model. The Gaussian center coordinates μ i represent the position of the i-th Gaussian distribution in the three-dimensional space, μ i = (μ ix , μ iy , μ iz ). The covariance matrix ∑ i is composed of a rotation and scaling matrix, which is used to control the shape and direction of the Gaussian ellipsoid, and has 6 independent parameters, ∑ i = , wherein the diagonal elements (σ ) represent the variance (scaling factor) in the direction of each coordinate axis, describing the stretching degree of the Gaussian distribution in the x, y, z axis; the non-diagonal elements (σ ) represent the covariance between different coordinate axes, reflecting the rotation characteristics of the Gaussian distribution. The opacity a i is used to determine the visibility of the Gaussian distribution in the scene, and each Gaussian distribution corresponds to an opacity value. The appearance feature c i is usually controlled by spherical harmonic coefficients, which is used to represent the appearance attributes of each Gaussian distribution. The present application can provide a macroscopic global bird's eye view, large area scene information, high precision three-dimensional coordinate information and local detail information by obtaining satellite image data I sat , point cloud data P lidar and unmanned aerial vehicle aerial image I photo , which provides a reference for the overall layout and geographical features of the 3D Gaussian splatting model, helps to accurately restore the terrain and object shape of the scene, and improves the restoration ability of the 3D Gaussian splatting model, making the 3D Gaussian splatting model more realistic. The satellite image data I sat is to uniformly adjust the resolution of the collected original satellite image to 512x512 using the bilinear interpolation algorithm. Specifically, for each target pixel in the original satellite image, its coordinates are mapped to the spatial position of the original image, and its pixel value is calculated based on the weighted average of the four adjacent pixels, so as to obtain satellite image data with consistent resolution, which provides standardized data input for feature extraction and 3D Gaussian splatting model optimization.
[0023] Discriminator warm-up: the generated data fusion feature vector f fake =[f sat ;f lidar ;f rend ] and the real data fusion feature vector f real =[f sat ;flidar ;f photo ]Input discriminator preheating is performed on the discriminator, and when the accuracy of the discrimination result generated by the discriminator is ≥85% or the update round is ≥10, the discriminator preheating stage ends, and the discriminator update and generator update cycle stage is entered.
[0024] Discriminator update: based on the update condition of the discriminator, the generated data fusion feature vector f fake =[f sat ;f lidar ;f rend ] and the real data fusion feature vector f real =[f sat ;f lidar ;f photo ] are updated to the discriminator in the GAN model, and the discrimination result is generated through the updated discriminator; Generator update: the model parameters of the 3D Gaussian sputtering model are extracted; the joint coding feature is generated, and the joint coding feature is used as the input of the generator in the GAN model; the parameter increment is generated through the generator, and the updated model parameters are generated according to the parameter increment using the LOD dynamic update mechanism; Iteration: repeat the feature vector acquisition, discriminator update and generator update steps until the iteration is completed; Scene reconstruction: update the 3D Gaussian sputtering model using the last updated model parameters as the optimal 3D Gaussian sputtering model, and perform unmanned aerial vehicle aerial scene reconstruction based on the optimal 3D Gaussian sputtering model.
[0025] In the above feature vector acquisition step, the unmanned aerial vehicle aerial image I photo , the multi-view rendering image I rend , the satellite image data I sat and the point cloud data P lidar are respectively extracted and fused to obtain the generated data fusion feature vector f fake =[f sat ;f lidar ;f rend ] and the real data fusion feature vector f real =[f sat ;f lidar ;f photo ]Specifically: The feature vectors of the unmanned aerial vehicle aerial image I photo , the multi-view rendering image I rend , the satellite image data I sat and the point cloud data P lidar are extracted respectively, and the feature vectors f photo , the feature vector f rend , the feature vector f satand eigenvector f lidar ; the feature vector f photo and eigenvector f rend respectively with the eigenvector f sat and eigenvector f lidar The data is fused to form a fused feature vector f. fake =[f sat ;f lidar ;f rend ] and the feature vector f fused with real data real =[f sat ;f lidar ;f photo Among them, drone aerial images I photo and multi-view rendering image I rend It extracts feature vectors based on the convolutional neural network in PatchGAN. and eigenvectors Satellite imagery data I sat Feature vector extraction is performed using ResNet50. Point cloud data P lidar Feature vector extraction is performed using PointNet++. , It is an identity matrix.
[0026] The feature vector f photo and eigenvector f rend respectively with the eigenvector f sat and eigenvector f lidar The data is fused to form a fused feature vector f. fake =[f sat ;f lidar ;f rend ] and the feature vector f fused with real data real =[f sat ;f lidar ;f photo Specifically, in the discriminator, f needs to be... fake and f real Perform linear transformation, batch normalization, and dimensionality reduction respectively to obtain and ,Will and Input to the fully connected layer via Output true probability value In the formula, bias , for or .
[0027] In the generator update steps described above, the generator parameter increments specifically include: The coordinates of the Gaussian center are μ. i The sum of the covariance matrix ∑ i Merge to form the input vector V i , The multilayer perceptron algorithm is used to process the input vector V. i Perform feature extraction and output geometric feature code E geo (V i ); where the input dimension of the multilayer perceptron algorithm is... , No. layer The weight matrix is , The number of layers in the multilayer perceptron algorithm; For each layer, the output dimension is... , .
[0028] Appearance features c are processed using a one-dimensional convolutional network. i Feature extraction is performed to obtain the appearance feature code E. app (c i ) Among them, appearance feature c i The spherical harmonic coefficients are represented as a 16×3 matrix, where 16 represents the number of spherical harmonic bases and 3 represents the RGB channels.
[0029] Encode geometric features E geo (V i ) and appearance feature code E app (c i The concatenation process yields the joint code h. i =[E geo (V i E app (c i )], will jointly encode h i As input to the generator in the GAN model, the generator generates parameter increments, and the updated model parameters are generated using the LOD dynamic update mechanism based on the parameter increments. This can capture the characteristics of the 3D Gaussian sputtering model from different perspectives and achieve the fusion expression of geometric features and appearance features.
[0030] In the generator update step, the generation process of the LOD dynamic update mechanism is as follows: Voxelization of point cloud data: The boundaries of the voxel mesh are determined based on the maximum and minimum values of the point cloud data in the 3D coordinate system. Using a 2m×2m×2m cube as the voxel unit, the point cloud space formed by the point cloud data is divided into a regular voxel mesh. For point cloud data P...lidar For each point P in the array, the row, column, and layer index of the voxel to which it belongs is determined by calculating the offset of its spatial coordinates from the grid origin and dividing it by the voxel size. This voxel index is then recorded in the corresponding voxel grid. This allows for quick location of points within the voxel to approximate the minimum distance when calculating geometric errors, reducing the search range and improving computational efficiency.
[0031] According to point cloud data P lidar Determine the geometric error ε i Specifically: for each Gaussian center coordinate μ i Within the voxelized mesh structure, the point cloud data of the voxel and its adjacent voxels are searched, and the coordinates of the Gaussian center of each voxel are calculated. The Euclidean distance to each point P in the neighborhood of the voxel is calculated, and the minimum value among them is taken as the geometric error. In this context, a voxel neighborhood is a spatial region concept, and a voxel index is a tool for quickly accessing and locating these regions. Through the voxel index, the voxel neighborhood that needs to be searched can be found.
[0032] The generator predicts parameter increments ΔG=[Δμ;Δ∑;Δα;Δc], which are used to update the model parameters. The model parameters are optimized through staged decoupling to avoid gradient conflicts and training instability caused by optimizing all parameters simultaneously. Specifically: Using the Gaussian center coordinates μ i The sum of the covariance matrix ∑ i The coarse optimization stage LOD0 is formed by combining the parameter increment ΔG, and the corresponding calculation formula is: In the formula, This represents the number of iterations. This represents the parameter increment for the Gaussian center coordinates; This represents the parameter increment of the covariance matrix; The weight corresponding to the parameter increment of the Gaussian center coordinates is 0.08 in this stage; This represents the weight corresponding to the parameter increment of the covariance matrix, and its value is 0.03 in this stage.
[0033] In the coarse optimization stage, LOD0 focuses on adjusting the global geometry of the scene to make it macroscopically approximate the geometric layout of the real scene.
[0034] Using the Gaussian center coordinates μ i , covariance matrix ∑ i and opacity α i The formula for calculating LOD1 during the optimization phase, based on the parameter increment ΔG, is as follows: In the formula, For the Sigmoid function; This represents the number of iterations. This represents the parameter increment for the Gaussian center coordinates; This represents the parameter increment of the covariance matrix; This is the increment of the opacity parameter; The weight corresponding to the parameter increment of the Gaussian center coordinates is 0.05 in this stage; This represents the weight corresponding to the parameter increment of the covariance matrix, and its value is 0.02 in this stage. This is the weight corresponding to the opacity parameter increment, and its value is 0.008 at this stage.
[0035] The focus of LOD1 in the optimization stage is to control the density of the Gaussian distribution to make the 3D Gaussian sputtering model more visually reasonable and enhance its sense of layering and three-dimensionality.
[0036] Using the Gaussian center coordinates μ i , covariance matrix ∑ i Opacity α i and appearance features c i The LOD2 fine optimization stage is formed by combining the parameter increment ΔG; The corresponding calculation formula is: In the formula, This represents the number of iterations. This represents the parameter increment for the Gaussian center coordinates; This represents the parameter increment of the covariance matrix; For the Sigmoid function; This is the increment of the opacity parameter; For the parameter increments of appearance features; The weight corresponding to the parameter increment of the Gaussian center coordinates is 0.01 in this stage; This represents the weight corresponding to the parameter increment of the covariance matrix, and its value is 0.005 in this stage. This is the weight corresponding to the opacity parameter increment, and its value is 0.005 at this stage; The weight corresponding to the parameter increment of the appearance feature is 0.015 in this stage.
[0037] The focus of LOD2's fine optimization phase is on adjusting parameters such as spherical harmonic coefficients or RGB colors, finely adjusting local details of the model, such as textures and lighting effects, to significantly improve the model's realism and lifelikeness.
[0038] Based on the coarse optimization stage LOD0, the intermediate optimization stage LOD1, and the fine optimization stage LOD2, the geometric error ε is... i As a condition for judgment, a dynamic LOD update mechanism is generated, wherein the dynamic LOD update mechanism is as follows: Set the lower limit of the geometric error threshold 0.3, upper limit of geometric error threshold A value of 0.7 is used to control the switching timing between the coarse optimization stage (LOD0), the intermediate optimization stage (LOD1), and the fine optimization stage (LOD2). At this point, we are in the LOD0 stage, and only the coordinates of the Gaussian center μ are considered. i , covariance matrix ∑ i Optimization was performed, with a focus on adjusting the global geometry of the 3D Gaussian sputtering model. At this point, it enters the LOD1 stage, which adds an adjustment to the opacity α based on the LOD0 stage. i Optimization was performed to adjust the density of the Gaussian distribution and enhance the sense of hierarchy in the 3D Gaussian sputtering model. At this point, upon reaching LOD2 stage, additional steps are taken to enhance the appearance feature c. i Optimize and improve the detail of the 3D Gaussian sputtering model.
[0039] In this invention, the conditions for updating the discriminator are divided into two cases: The first scenario is event-driven updates. When the coarse optimization stage LOD0, the intermediate optimization stage LOD1, and the fine optimization stage LOD2 switch, the generator parameters are frozen to ensure that the generator parameters remain unchanged when the discriminator is updated. Then, the discriminator is updated by minimizing the total loss function.
[0040] The second scenario involves periodic updates, where the discriminator is forcibly updated after every three generator optimizations. Periodic updates ensure that the discriminator and generator optimizations remain synchronized, preventing the discriminator's performance from lagging behind the generator and causing the generator's generated fake data to be ineffectively detected, thus affecting the overall model optimization performance.
[0041] In the above feature vector acquisition steps, a multi-view rendering image I is generated through differentiable rasterization rendering. rend Specifically, it includes: The virtual camera uniformly acquires K virtual camera poses in the scene space. The scene space is the space reconstructed using a 3D Gaussian sputtering model or an optimized 3D Gaussian sputtering model, and m is the virtual camera pose index. Using differentiable rendering technology Generate the pose corresponding to each virtual camera The rendered image is the multi-view rendered image I. rend In the formula, , , and These are all updated model parameters. For projection functions; represents the projection transformation parameters of the virtual camera; M represents the total number of Gaussian elements in the 3D Gaussian sputtering model; j represents the index of the Gaussian element.
[0042] In this invention, both the generator and the discriminator are optimized by using the total loss obtained from adversarial loss, perceptual loss and geometric consistency loss.
[0043] Adversarial loss comprises a discriminator loss function and a generator loss function. The discriminator aims to maximize its ability to distinguish between real and generated scenes, while the generator aims to deceive the discriminator, making it difficult for it to differentiate between real and generated scenes. Both the discriminator and generator loss functions are generated using WGAN-GP (Wasserstein Generative Adversarial Network - Gradient Penalty).
[0044] The perceptual loss is calculated by adaptively adjusting the LOD dynamic update mechanism and the resolution of the multi-view rendered images. When Or the resolution of the multi-view rendered image is At that time, the perceptual loss is calculated by rendering images from multiple perspectives. With drone aerial images The feature differences are determined in the ReLU2-2 layer of the VGG-19 network. At this stage, the 3D Gaussian sputtering model focuses on the coarse structure or global framework, and shallow features are sufficient to capture the fundamental differences. This is achieved when LOD=2 and the resolution of the multi-view rendered image is... At that time, the perceptual loss is calculated by rendering images from multiple perspectives. With drone aerial images The feature differences at ReLU3-3 layers are used to determine the perceptual loss. At this stage, the 3D Gaussian sputtering model needs to optimize local details; deeper features are more suitable for characterizing complex semantic differences because the 3D Gaussian sputtering model has not yet reached the stage where perceptual loss needs to be calculated, and is more focused on the coarse structure or global framework. Perceptual loss function. The definition of is: in, This represents the output feature map of the ReLU3-3 or ReLU2-2 layer of the VGG-19 network. The perceptual loss balances the structural optimization at the low level of detail with the perceptual realism enhancement at the high level of detail by dynamically selecting the feature layer (shallow / deep) and the joint constraints of the LOD stage, ensuring that the generated rendered image is perceptually similar to a real photo and improving the visual quality of the image generated by the model.
[0045] Geometric consistency loss is used to ensure that the geometry of the optimized 3D Gaussian sputtering model matches the point cloud data P. lidar To ensure consistency with the true geometric information, the Symmetric Chamfer Distance (SCD) is used to measure geometric consistency. The mathematical expression for the geometric consistency loss is: In the formula, B is from point cloud data. A subset of random samples from the middle It is the optimized set of Gaussian center coordinates. It is point cloud data The coordinates of a single point in subset B, yes The coordinates of a single point, and B and respectively The total number of subsets. Geometric consistency loss ensures that the geometric structure of the 3D Gaussian sputtering model remains consistent with the true geometric information represented by the point cloud data by comparing the distance between the projected Gaussian center line and the point cloud data, thereby improving the geometric accuracy of the 3D Gaussian sputtering model.
[0046] The total loss is the sum of the adversarial loss, the perceptual loss, and the geometric consistency loss, and its expression is: In the formula, Total loss; To combat the losses; To perceive loss; This is the geometric consistency loss; Weights to counteract losses; Weights for perceived loss; The weights for the geometric consistency loss are at LOD0 in the coarse optimization stage. It is 0.7. It is 0.0. The value is 0.3, focusing on stabilizing the global geometry; during the intermediate optimization stage at LOD1, It is 0.5. It is 0.2. The initial value is 0.3, focusing on adjusting model density and basic appearance; during the fine optimization stage at LOD2, It is 0.3. It is 0.5. The value is 0.2, focusing on generating a refined appearance and details. When optimizing the generator, The loss calculated for the generator loss function; when optimizing the discriminator. The loss calculated for the discriminator loss function.
[0047] The invention also includes adaptively optimizing the Gaussian density of the 3D Gaussian sputtering model using total loss before each iteration: By calculating the parameter gradient contribution of each Gaussian distribution ,in Indicates total loss For the coordinates of the Gaussian center μ i The gradient norm is used to quantify the importance of the current Gaussian distribution in the optimized 3D Gaussian sputtering model.
[0048] Based on gradient contribution and opacity α i The adaptive adjustment rule is designed as follows: when and At this point, the Gaussian distribution contributes little to the optimized 3D Gaussian sputtering model and has low opacity, constituting redundant information. Therefore, it is removed from the 3D Gaussian sputtering model to reduce its complexity and improve computational efficiency. and This indicates that the Gaussian distribution corresponds to the detailed region in the scene, playing a crucial role in the accuracy and visual realism of the 3D Gaussian sputtering model. Applying the Gaussian center coordinates μ to this detailed region... i and opacity α i To improve the accuracy and realism of 3D Gaussian sputtering models.
[0049] In the above iterative steps, the iteration is completed when the maximum number of iterations is reached or the 3D Gaussian sputtering model converges. Generally, the maximum number of iterations is 500-1000 epochs. This avoids overtraining of the 3D Gaussian sputtering model, preventing overfitting, while ensuring optimization is completed within reasonable time and computational resources. The 3D Gaussian sputtering model converges when the change in total loss occurs over 20 consecutive epochs. hour, The difference in moving average loss indicates that the 3D Gaussian sputtering model has converged and the iteration has terminated.
[0050] See Figure 2The present invention also proposes a system for optimizing drone aerial scene reconstruction based on GAN model, which is formed by the above-mentioned method for optimizing drone aerial scene reconstruction based on GAN model, including a feature vector acquisition module, a discriminator preheating module, a discriminator update module, a generator update module, an iteration module, and a scene reconstruction module. Feature vector acquisition module: used to acquire satellite imagery data I of the same scene. sat Point cloud data P lidar and drone aerial images I photo Based on drone aerial images I photo A 3D Gaussian sputtering model was constructed using the 3DGS algorithm. This model was then used to generate multi-view rendered images (I) through differentiable rasterization rendering. rend ; to capture aerial images from drones photo Multi-view rendering image I rend Satellite imagery data I sat Point cloud data P lidar Feature extraction and fusion are performed separately to obtain the generated data fusion feature vector f. fake =[f sat ;f lidar ;f rend ] and the feature vector f fused with real data real =[f sat ;f lidar ;f photo ]; Discriminator warm-up module: fused the generated data into a feature vector f fake =[f sat ;f lidar ;f rend ] and the feature vector f fused with real data real =[f sat ;f lidar ;f photo The discriminator is preheated in the input discriminator. When the accuracy of the discriminator in generating the discrimination result is ≥85% or the number of update rounds is ≥10, the discriminator preheating stage ends and the discriminator update and generator update cycle stage begins.
[0051] Discriminator update module: Based on the discriminator's update conditions, it generates a data fusion feature vector f. fake =[f sat ;f lidar ;f rend ] and the feature vector f fused with real data real =[f sat ;f lidar ;f photo The discriminator in the GAN model is updated, and the discrimination result is generated using the updated discriminator. Generator update module: used to extract model parameters of 3D Gaussian sputtering model; generate joint encoded features, use joint encoded features as input to generator in GAN model, generate parameter increments through generator, and generate updated model parameters based on parameter increments using LOD dynamic update mechanism; Iteration module: Used to repeat the feature vector acquisition module, discriminator update module, and generator update module until the iteration is complete; Scene Reconstruction Module: This module is used to update the 3D Gaussian sputtering model using the last updated model parameters as the optimal 3D Gaussian sputtering model, and then reconstruct the drone aerial photography scene based on the optimal 3D Gaussian sputtering model.
[0052] The system for optimizing UAV aerial scene reconstruction based on GAN model of the present invention corresponds to the method for optimizing UAV aerial scene reconstruction based on GAN model described above. The specific contents of the feature vector acquisition module, discriminator preheating module, discriminator update module, generator update module, iteration module and scene reconstruction module are described in the description of the method for optimizing UAV aerial scene reconstruction based on GAN model described above, and will not be repeated here.
[0053] This invention also provides a memory storing a program file, which is executed to implement the program instructions formed by the above-described method for optimizing UAV aerial scene reconstruction based on a GAN model. Specifically, please refer to the description above for the method for optimizing UAV aerial scene reconstruction based on a GAN model.
[0054] The memory in this invention may specifically include random access memory (RAM), main memory, read-only memory (ROM), programmable ROM, erasable programmable ROM, registers, hard disk, removable disk, or CD-ROM. It should be noted that those skilled in the art can choose the form and type of storage medium according to actual usage needs, and this invention does not impose further specific limitations.
[0055] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of this application.
Claims
1. A method for optimizing UAV aerial scene reconstruction based on a GAN model, characterized in that, Includes the following steps: Feature vector acquisition: Obtain satellite imagery data I from the same scene sat Point cloud data P lidar and drone aerial images I photo Based on drone aerial images I photo A 3D Gaussian sputtering model was constructed using the 3DGS algorithm, and multi-view rendered images were generated using differentiable rasterization rendering. rend ; drone aerial images I photo Multi-view rendering image I rend Satellite imagery data I sat Point cloud data P lidar Feature extraction and fusion are performed separately to obtain the generated data fusion feature vector f. fake =[f sat ;f lidar ;f rend ] and the feature vector f fused with real data real =[f sat ;f lidar ;f photo ]; Discriminator Update: Based on the discriminator's update conditions, a data fusion feature vector f is generated. fake =[f sat ;f lidar ;f rend ] and the feature vector f fused with real data real =[f sat ;f lidar ;f photo The discriminator in the GAN model is updated, and the discrimination result is generated using the updated discriminator. Generator update: Extract model parameters of the 3D Gaussian sputtering model; generate joint encoded features, use the joint encoded features as input to the generator in the GAN model, generate parameter increments through the generator, and update the model parameters of the 3D Gaussian sputtering model using the LOD dynamic update mechanism based on the parameter increments; Iteration: Repeat the steps of feature vector acquisition, discriminator update, and generator update until the iteration is complete; Scene reconstruction: The 3D Gaussian sputtering model is updated using the model parameters from the last update and used as the optimal 3D Gaussian sputtering model. The drone aerial photography scene is then reconstructed based on the optimal 3D Gaussian sputtering model.
2. The method for optimizing UAV aerial scene reconstruction using a GAN model according to claim 1, characterized in that, In the feature vector acquisition step, the UAV aerial image I photo Multi-view rendering image I rend Satellite imagery data I sat Point cloud data P lidar Feature extraction and fusion are performed separately to obtain the generated data fusion feature vector f. fake =[f sat ;f lidar ;f rend ] and the feature vector f fused with real data real =[f sat ;f lidar ;f photo Specifically: Extract drone aerial images I photo Multi-view rendering image I rend Satellite imagery data I sat Point cloud data P lidar The eigenvectors of are obtained, and the corresponding eigenvectors f are obtained. photo eigenvector f rend eigenvector f sat and eigenvector f lidar ; The feature vector f photo and eigenvector f rend respectively with the eigenvector f sat and eigenvector f lidar The data is fused to form a fused feature vector f. fake =[f sat ;f lidar ;f rend ] and the feature vector f fused with real data real =[f sat ;f lidar ;f photo ].
3. The method for optimizing UAV aerial scene reconstruction using a GAN model according to claim 1, characterized in that, The model parameters include the Gaussian center coordinates μ. i , covariance matrix ∑ i Opacity α i and appearance features c i , where i represents the index of the Gaussian distribution in the 3D Gaussian sputtering model.
4. The method for optimizing UAV aerial scene reconstruction using a GAN model according to claim 3, characterized in that, The generator update step specifically includes the following generator parameter increments: The coordinates of the Gaussian center are μ. i The sum of the covariance matrix ∑ i Merge to form the input vector V i The multilayer perceptron algorithm is used to process the input vector V. i Perform feature extraction and output geometric feature code E geo (V i ); Appearance features c are processed using a one-dimensional convolutional network. i Feature extraction is performed to obtain the appearance feature code E. app (c i ); Encode geometric features E geo (V i ) and appearance feature code E app (c i The concatenation process yields the joint code h. i =[E geo (V i E app (c i )], will jointly encode h i As input to the generator in the GAN model, the parameter increments are generated by the generator.
5. The method for optimizing UAV aerial scene reconstruction using a GAN model according to claim 3, characterized in that, In the generator update step, the generation process of the LOD dynamic update mechanism is as follows: Calculate geometric error ε i ; Using the Gaussian center coordinates μ i The sum of the covariance matrix ∑ i The coarse optimization stage LOD0 is formed by combining the parameter increment ΔG; Using the Gaussian center coordinates μ i , covariance matrix ∑ i and opacity α i The LOD1 in the optimization stage is formed by combining the parameter increment ΔG; Using the Gaussian center coordinates μ i , covariance matrix ∑ i Opacity α i and appearance features c i The LOD2 fine optimization stage is formed by combining the parameter increment ΔG; Based on the coarse optimization stage LOD0, the intermediate optimization stage LOD1, and the fine optimization stage LOD2, the geometric error ε is... i As a condition for judgment, a dynamic LOD update mechanism is generated.
6. The method for optimizing UAV aerial scene reconstruction using a GAN model according to claim 5, characterized in that, In the feature vector acquisition step, a multi-view rendering image I is generated through differentiable rasterization rendering. rend Specifically, it includes: The virtual camera uniformly acquires K virtual camera poses in the scene space. The scene space is the space reconstructed using a 3D Gaussian sputtering model or an optimized 3D Gaussian sputtering model, and m is the virtual camera pose index. Using differentiable rendering technology Generate the pose corresponding to each virtual camera The rendered image is the multi-view rendered image I. rend In the formula, , , and These are all updated model parameters. For projection functions; represents the projection transformation parameters of the virtual camera; M represents the total number of Gaussian elements in the 3D Gaussian sputtering model; j represents the index of the Gaussian element.
7. The method for optimizing UAV aerial scene reconstruction using a GAN model according to claim 1, characterized in that, Both the generator and the discriminator are optimized using the total loss obtained from adversarial loss, perceptual loss, and geometric consistency loss.
8. The method for optimizing UAV aerial scene reconstruction using a GAN model according to claim 1, characterized in that, In the iterative steps, the iteration is completed when the maximum number of iterations is reached or the 3D Gaussian sputtering model has converged.
9. A system for optimizing UAV aerial scene reconstruction based on a GAN model, formed using the method for optimizing UAV aerial scene reconstruction based on a GAN model as described in claim 1, is characterized in that... include: Feature vector acquisition module: used to acquire satellite imagery data I of the same scene. sat Point cloud data P lidar and drone aerial images I photo Based on drone aerial images I photo A 3D Gaussian sputtering model was constructed using the 3DGS algorithm, and multi-view rendered images were generated using differentiable rasterization rendering. rend ; to capture aerial images from drones photo Multi-view rendering image I rend Satellite imagery data I sat Point cloud data P lidar Feature extraction and fusion are performed separately to obtain the generated data fusion feature vector f. fake =[f sat ;f lidar ;f rend ] and the feature vector f fused with real data real =[f sat ;f lidar ;f photo ]; Discriminator update module: Used to generate a data fusion feature vector f based on the discriminator's update conditions. fake =[f sat ;f lidar ;f rend ] and the feature vector f fused with real data real =[f sat ;f lidar ;f photo The discriminator in the GAN model is updated, and the discrimination result is generated using the updated discriminator. Generator update module: used to extract model parameters of 3D Gaussian sputtering model; generate joint encoded features, use joint encoded features as input to generator in GAN model, generate parameter increments through generator, and update model parameters of 3D Gaussian sputtering model according to parameter increments using LOD dynamic update mechanism; Iteration module: Repeat the feature vector acquisition module, discriminator update module, and generator update module until the iteration is complete; And the scene reconstruction module: used to update the 3D Gaussian sputtering model with the model parameters of the last update, as the optimal 3D Gaussian sputtering model, and to reconstruct the drone aerial photography scene based on the optimal 3D Gaussian sputtering model.
10. A memory, characterized in that, The system stores a program file, which is executed to implement the program instructions formed by the method for optimizing UAV aerial scene reconstruction based on the GAN model as described in any one of claims 1-8.