Unbounded scene three-dimensional reconstruction method based on unmanned aerial vehicle aerial photography
Multi-view images are acquired through drone aerial photography, combined with 3D Gaussian sputtering optimization and geometric consistency verification, the problems of high consumption of three-dimensional modeling resources and insufficient geometric accuracy in large-scale scenarios are solved, and efficient and accurate three-dimensional reconstruction is achieved.
Patent Information
- Application Number
- CN202510482448.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
In large-scale scenarios, traditional drone three-dimensional modeling methods have problems such as high resource consumption and insufficient geometric accuracy.
A three-dimensional reconstruction method for unbounded scenes based on drone aerial photography is proposed. By acquiring multi-view image sets, building a spatial topological association network, extracting image features, generating scene partitions, performing 3D Gaussian sputtering optimization, dynamically combining underlying assets, performing geometric consistency verification, and finally generating an unbounded three-dimensional scene.
This method can effectively reduce the amount of calculation and improve modeling efficiency. The generated three-dimensional scenes have high precision and flexibility, and are suitable for different application scenarios.
Smart Images

Figure CN120014178A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a method for three-dimensional reconstruction of an unbounded scene based on drone aerial photography. Background Art
[0002] With the rapid development of science and technology, 3D reconstruction technology has become one of the core technologies in many fields. Especially in application scenarios such as urban planning, topographic mapping, disaster assessment, and environmental monitoring, the 3D reconstruction technology of scenes is extremely valuable. By accurately modeling the scenes in 3D, it can provide decision makers with more accurate geospatial information, improve various work efficiencies, reduce manual intervention, and save time and costs.
[0003] However, traditional UAV 3D modeling methods usually rely on image matching and point cloud data processing, and the modeling accuracy is poor. Especially when the scene details are complex or there are occlusions, the high accuracy of the 3D model cannot be guaranteed, and a lot of manual intervention is required for data optimization and correction. At the same time, traditional 3D reconstruction methods require a lot of computing time, especially in large-scale scenes. The reconstruction process often takes a long time to calculate, which is unrealistic for practical applications.
[0004] In order to overcome the above problems, image-based 3D reconstruction methods have gradually become a research hotspot in the field of UAV modeling. Through the efficient and convenient image acquisition of UAVs, combined with deep learning and computer vision technology, more accurate 3D reconstruction can be achieved. Deep learning methods can automatically extract feature information from images through training of large-scale data sets, optimize the accuracy of the model, reduce manual intervention, and improve the degree of automation of modeling. At the same time, when processing certain scenes, image-based 3D reconstruction methods can greatly reduce the dependence on point cloud data in traditional modeling methods, thereby reducing computing time and memory consumption and improving modeling efficiency. However, as the scale of the scene continues to increase, the computational cost and memory space requirements of 3D reconstruction also increase.
[0005] Therefore, how to avoid high consumption of 3D modeling resources and insufficient geometric accuracy in large-scale scenarios remains a technical challenge that needs to be solved urgently. Summary of the invention
[0006] The main purpose of the present invention is to provide a method for 3D reconstruction of unbounded scenes based on drone aerial photography, aiming to solve the problems of high consumption of 3D modeling resources and insufficient geometric accuracy in large-scale scenes.
[0007] In order to achieve the above object, the present invention proposes a method for 3D reconstruction of an unbounded scene based on drone aerial photography, comprising the following steps: Obtain a multi-view image set of the target area and construct a spatial topological association network; Image features are extracted based on the Hessian matrix and SIFT operator, and scene partitions are generated through Euclidean distance matching; Perform 3D Gaussian sputtering optimization on each partition and simultaneously generate a reusable 3D base asset, which includes structural components and texture features; Based on the spatial transformation rules defined by programmatic code, the underlying assets are dynamically combined to generate a hierarchical three-dimensional model; Geometric consistency verification is performed in latent space by jointly optimizing a diffusion model with a differentiable renderer. The optimized hierarchical models are assembled into an unbounded three-dimensional scene according to the spatial layout parameters.
[0008] In one embodiment of the present application, the 3D Gaussian sputtering optimization includes: Initialize programmatic code for each partition to define layout constraints of the underlying assets, including spatial boundaries, rotation parameters, and scaling; During 3D Gaussian sputtering optimization, the Gaussian distribution range is limited by bounding box adaptive constraints, including: (a) Gaussian center Beyond soft bounding box , forcing its position to align to the boundary; (b) If the Gaussian center exceeds the boundary, the threshold is extended according to the boundary Update its location; The formula is as follows:
[0009] Synchronously update the Gaussian parameters of repeated underlying assets and achieve parameter sharing through gradient back propagation; represents the adjusted Gaussian center position, represents the original x-axis lower bound of partition i; Represents the original x-axis upper bound of partition i; Indicates the boundary extension alignment threshold; Soft boundary relaxation amount; represents the original position of the Gaussian center; Represents the center coordinate of the kth Gaussian distribution on the x-axis.
[0010] In one embodiment of the present application, the programmatic code generation includes: Segment repetitive structural components from multi-view images and reconstruct their geometric and texture features through 3D Gaussian sputtering optimization; Encode the transformation rules of components into a programmed instruction set, including layer expansion, axial arrangement, and size adaptive adjustment; Automatically generate asset distribution rules based on spatial layout parameters; When the unbounded scene is expanded, scene continuity is achieved through block loading and real-time stitching.
[0011] In one embodiment of the present application, the diffusion model is combined with a differentiable renderer for optimization, including: In the latent space of the diffusion model, the base asset texture is enhanced by multi-scale denoising; Project the optimized asset to multiple views through a differentiable renderer and calculate the geometric consistency loss:
[0012] in, represents the geometric consistency loss; is the number of partitions, Reconstruct the depth truth for the partition, To render the depth map; support users to modify the programmed code parameters in real time and dynamically trigger the re-optimization of local Gaussian parameters.
[0013] In one embodiment of the present application, geometric consistency verification is performed in the latent space, including the following steps: The result of the partitioned 3D reconstruction is used as input, and the rendering diffusion model performs a diffusion process in a low-resolution latent space. Given a camera pose obtained by a spherical linear interpolation method, a differentiable renderer generates an initial image according to the camera pose as a condition of the diffusion process, wherein the low-resolution latent space is a partial area or sub-area of low resolution, and the differentiable renderer is a renderer whose output can differentiate the input parameters; The initial image is encoded in a low-resolution latent space using a variational autoencoder, and feature reverse sampling and extraction is performed in the latent space. The encoding is a process of iteratively adding Gaussian noise to the initial image, and the feature reverse sampling and extraction is to extract features similar to the original ordered image from the chaotic and disordered image through a denoising operation to generate a denoised image; Decoding the denoised image in pixel space using a decoder in a variational autoencoder, and converting the decoded denoised image into a pixel space image, wherein the pixel space is the original representation space of the image, in which each pixel has its specific color value and grayscale value; Pixel loss is used to maintain the three-dimensional consistency between the generated pixel space image and the input partitioned three-dimensional reconstruction result, and the three-dimensional network reconstruction optimization result is output.
[0014] In one embodiment of the present application, assembling the optimized hierarchical model according to the spatial layout parameters includes: Input the boundary point clouds of adjacent partitions into the graph convolutional network and calculate the feature similarity of the overlapping areas; When the overlap exceeds a threshold, partitions are automatically merged using programmatic rules, including: (a) Align the partition coordinate system to eliminate the stitching gap; (b) Smooth interpolation of Gaussian parameters of the merged area; Generate seamless global point cloud maps, supporting dynamic loading and memory compression.
[0015] In one embodiment of the present application, the feature similarity of the overlapping area is calculated using a graph attention mechanism: ; in, , E is the edge set of the overlapping area, For point and The characteristic vector of Represents two point clouds whose similarity is to be calculated; Represents a pair of adjacent points (nodes) in an edge set; Represents the feature vector The cosine similarity of Attention weight; and Represents learnable query and key transformation matrices; The dimension of the feature vector; The combined Gaussian interpolation formula is:
[0016] in and The Gaussian parameters to be merged, the weight w is calculated based on the attenuation of the distance to the overlapping center; The parameters of the merged Gaussian; and represents the weight coefficient; Represents Gaussian distribution The center coordinates of Indicates the center coordinates of the overlapping area.
[0017] In one embodiment of the present application, the SIFT operator extracts image features, specifically including the following steps: The SIFT operator is used to perform multi-scale Gaussian difference filtering on the input image, and the extreme points in the image are detected as key points; the key points are located in the scale space and the local neighborhood, and the main direction is assigned to them according to the gradient direction; the neighborhood of the key points is sampled through the graph convolutional neural network, and the local features of the neighborhood of the key points are aggregated to complete the feature extraction; For each pixel point in the target image, the pixel coordinates of the image features in the adjacent images in the two partitions are calculated and subtracted to obtain the disparity value, thereby constructing the corresponding disparity map between the two images; The depth value corresponding to each pixel in the target image is calculated through the camera parameters and the disparity value in the disparity map, so as to construct a depth map of the target image, where the pixel value of the pixel corresponds to the depth value of the pixel; Through the depth value in the depth map, each pixel is converted into a three-dimensional coordinate in the real world to obtain a point cloud; The above steps are repeated to obtain the point cloud data containing texture feature information in each partition.
[0018] In one embodiment of the present application, the calculation formula using the SIFT operator is:
[0019]
[0020] in, represents the accuracy of the gradient vector, represents the direction of the gradient vector, represents the coordinates of each sampling point, , Represents the horizontal and vertical coordinates of the coordinate point.
[0021] In one embodiment of the present application, the calculation formula for calculating the Hessian matrix of each image in the image set is as follows: ; in = , represents the Hessian matrix, Represents each sampling point, express, express, express, They all represent the values of the second-order derivatives, , Represents the horizontal and vertical coordinates of the coordinate point.
[0022] By adopting the above technical solution, multi-view images are obtained through drone aerial photography, and combined with advanced image processing and 3D reconstruction technology, the 3D structure of the unbounded scene can be accurately restored. Through efficient image feature matching and partition generation, the amount of calculation can be reduced and the modeling efficiency can be improved. The reusable 3D basic assets generated simultaneously further improve the flexibility and scalability of the model, so that the final generated 3D scene can be flexibly adjusted in different application scenarios. It solves the problem of high consumption of 3D modeling resources and insufficient geometric accuracy in large-scale scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The present invention is described in detail below with reference to specific embodiments and accompanying drawings, wherein: Figure 1 This is a schematic diagram of the process structure of the first embodiment of the present invention. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention is described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the following specific embodiments are only used to explain the present invention and do not constitute a limitation to the present invention.
[0025] like Figure 1 As shown, in order to achieve the above purpose, the present invention proposes a method for 3D reconstruction of an unbounded scene based on drone aerial photography, comprising the following steps: Obtain a multi-view image set of the target area and construct a spatial topological association network; Image features are extracted based on the Hessian matrix and SIFT operator, and scene partitions are generated through Euclidean distance matching; Perform 3D Gaussian sputtering optimization on each partition and simultaneously generate a reusable 3D base asset, which includes structural components and texture features; Based on the spatial transformation rules defined by programmatic code, the underlying assets are dynamically combined to generate a hierarchical three-dimensional model; Geometric consistency verification is performed in latent space by jointly optimizing a diffusion model with a differentiable renderer. The optimized hierarchical models are assembled into an unbounded three-dimensional scene according to the spatial layout parameters.
[0026] Specifically, it is necessary to first use a drone to take aerial photos of the target area, and obtain a multi-view image set by shooting at multiple different angles and viewing distances. The image set should cover all key locations and features of the target area to ensure that the reconstructed scene has sufficient detail information. Then, the spatial relationship information between images is extracted through image processing algorithms to construct a spatial topological association network. The spatial topological association network can describe the geometric relationship between images at different perspectives.
[0027] The Hessian matrix is used to detect and describe the feature points of the acquired image. The Hessian matrix can help detect structural features such as edges and corners in the image. Then, the SIFT (Scale Invariant Feature Transform) operator is combined to further extract the local feature points of the image and generate the corresponding feature descriptors. The extracted feature descriptors are used to match images in the image set. The correlation between images is determined by calculating the Euclidean distance between image features. Images with a high degree of correlation indicate that they have a strong similarity or spatial overlap.
[0028] According to the image correlation, the image collection is partitioned using the image partitioning algorithm. Each partition represents a sub-area or perspective of the target area, and can classify interrelated images into the same group.
[0029] For each partition, 3D reconstruction optimization is performed through the 3D Gaussian sputtering optimization algorithm, and the 3D structure of the scene is restored using the geometric information in the multi-view images. In this process, reusable 3D basic assets are also generated synchronously. These basic assets include structural components and texture features. Structural components refer to the fixed geometric parts of the scene, while texture features refer to the surface details in the scene, such as the ground, walls, vegetation, etc. The basic assets can be reused to reduce repeated calculations and improve reconstruction efficiency.
[0030] Use rules to adjust the spatial parameters of each basic asset, such as position, rotation angle, and scale, to ensure its reasonable layout and efficient combination in the overall 3D model. In this way, 3D models of different levels can be automatically generated, making the entire 3D scene more expressive and diverse. Hierarchical 3D models can clearly show the structure and details of the scene, and models of different levels can be detailed or simplified as needed.
[0031] The generated 3D model is further optimized by using the joint optimization technology of diffusion model and differentiable renderer. The diffusion model is used to explore various geometric shapes in the latent space, and the differentiable renderer is used to make detail corrections to the geometric model to ensure that the final generated 3D model maintains consistency in visual effects. Geometric consistency verification can ensure that the various parts of the 3D model match each other in space, thereby reducing geometric errors and ensuring the authenticity and accuracy of the scene.
[0032] By parameterizing the spatial layout of the optimized hierarchical model, the three-dimensional components are assembled into an unbounded three-dimensional scene according to specific spatial relationships based on the actual needs of the scene. An unbounded scene means that the boundaries of the scene are not restricted and can be freely expanded or scaled as needed, which is suitable for a variety of application scenarios such as virtual reality and augmented reality. Through precise spatial layout and reasonable model assembly, the final generated three-dimensional scene can provide a flexible, efficient and realistic virtual experience in different application scenarios.
[0033] By adopting the above technical solution, multi-view images are obtained through drone aerial photography, and combined with advanced image processing and 3D reconstruction technology, the 3D structure of the unbounded scene can be accurately restored. Through efficient image feature matching and partition generation, the amount of calculation can be reduced and the modeling efficiency can be improved. The reusable 3D basic assets generated simultaneously further improve the flexibility and scalability of the model, so that the final generated 3D scene can be flexibly adjusted in different application scenarios. It solves the problem of high consumption of 3D modeling resources and insufficient geometric accuracy in large-scale scenes.
[0034] In one embodiment of the present application, the 3D Gaussian sputtering optimization includes: Initialize programmatic code for each partition to define layout constraints of the underlying assets, including spatial boundaries, rotation parameters, and scaling; During 3D Gaussian sputtering optimization, the Gaussian distribution range is limited by bounding box adaptive constraints, including: (a) Gaussian center Beyond soft bounding box , forcing its position to align to the boundary; (b) If the Gaussian center exceeds the boundary, the threshold is extended according to the boundary Update its location; The formula is as follows:
[0035] Synchronously update the Gaussian parameters of repeated underlying assets and achieve parameter sharing through gradient back propagation; represents the adjusted Gaussian center position, represents the original x-axis lower bound of partition i; Represents the original x-axis upper bound of partition i; Indicates the boundary extension alignment threshold; Soft boundary relaxation amount; represents the original position of the Gaussian center; Represents the center coordinate of the kth Gaussian distribution on the x-axis.
[0036] Specifically, the basic assets in each partition will set the spatial boundaries, rotation parameters and scaling ratio according to the requirements of the target scene. The spatial boundaries define the area where the asset can be placed in the three-dimensional space, the rotation parameters determine the direction of the asset, and the scaling ratio defines the size of the asset. These constraints ensure that the layout of the assets meets the predetermined design requirements and ensures the rationality and consistency of the scene during the reconstruction process.
[0037] In 3D Gaussian sputtering optimization, the bounding box adaptive constraint is used to limit the range of the Gaussian distribution to avoid the distribution points exceeding the target area or producing unreasonable distribution. The bounding box defines the spatial range of each partition. Through this constraint, the center point of the Gaussian distribution can be limited ( ) within a specific range. Specific constraint operations include: (a) If the center point of the Gaussian distribution ( ) exceeds the defined soft bounding box (i.e. ), its position is forced to be aligned to the bounding box. This means that if the Gaussian center point exceeds the soft boundary in the x-axis direction, the position will be adjusted to the inside of the bounding box to keep the point within a reasonable range.
[0038] (b) If the Gaussian center point exceeds the range of the hard bounding box, it will be extended according to the boundary threshold ( ) adjusts its position so that it can return to the valid area. This operation ensures that even in extreme cases, the center point of the Gaussian distribution remains within an acceptable range, thus avoiding uncontrollable distribution deviation.
[0039] The formula is as follows:
[0040] in, represents the adjusted Gaussian center position, represents the original x-axis lower bound of partition i; Represents the original x-axis upper bound of partition i; Indicates the boundary extension alignment threshold; Soft boundary relaxation amount; represents the original position of the Gaussian center; Represents the center coordinate of the kth Gaussian distribution on the x-axis; The Gaussian parameters of the underlying assets repeated in the same partition are updated synchronously by means of gradient back propagation. For multiple underlying assets of the same type, the parameters of their Gaussian distribution, such as the center position, are shared and adjusted through the optimization algorithm. This parameter sharing method can effectively reduce redundant calculations, improve reconstruction efficiency, and ensure that the generated 3D scene is more consistent and coordinated.
[0041] The above technical solution, through Gaussian sputtering optimization technology and combined with bounding box adaptive constraints, can effectively limit the distribution range of basic assets in three-dimensional space, avoiding the situation where assets exceed the predetermined boundaries of the scene. The sharing of Gaussian parameters realizes efficient update and consistency control of basic assets through gradient back propagation, thereby ensuring the consistency of the three-dimensional scene in details and overall layout. This solution can greatly improve the accuracy and efficiency of three-dimensional reconstruction.
[0042] In one embodiment of the present application, the programmatic code generation includes: Segment repetitive structural components from multi-view images and reconstruct their geometric and texture features through 3D Gaussian sputtering optimization; Encode the transformation rules of components into a programmed instruction set, including layer expansion, axial arrangement, and size adaptive adjustment; Automatically generate asset distribution rules based on spatial layout parameters; When the unbounded scene is expanded, scene continuity is achieved through block loading and real-time stitching.
[0043] Specifically, repetitive structural components are segmented from multi-view images. Through image analysis and segmentation algorithms, repetitive parts in the target scene, such as buildings, roads, vegetation, etc., are identified. Then, these repetitive components are reconstructed in three dimensions using a 3D Gaussian sputtering optimization algorithm. In this process, not only the geometric structure is reconstructed, but also the texture features need to be extracted and reconstructed to ensure that the reconstruction result is visually consistent with the original image. 3D Gaussian sputtering optimization can effectively repair and refine the geometric and texture features of repetitive components, improving the realism and detail restoration of the scene.
[0044] The transformation rules for the repeating structural components are encoded to generate a procedural instruction set. The procedural code defines how to operate on these components and includes three main transformation rules: Layer expansion: According to the scene requirements, adjust the number of layers of the component to increase or decrease the vertical height or depth of the component so that it can adapt to the spatial layout of different levels.
[0045] Axial arrangement: Determine the axial arrangement of components in space, such as arranging components along the x-axis, y-axis, or z-axis. By transforming the rules, ensure the reasonable arrangement of components in the scene to meet the needs of different scene structures.
[0046] Adaptive size adjustment: According to the density and spatial layout of the target area, the size of the component is adjusted to adapt it to the proportion of the surrounding environment, ensuring that the scene meets the design requirements both visually and spatially.
[0047] The programmatic code generates asset distribution rules based on spatial layout parameters, including topological boundaries and regional density.
[0048] Topological boundaries: By defining the spatial boundaries of the scene, determine the area where each asset should be placed and its location range, and ensure that each component is laid out in the scene according to the set boundaries to avoid duplication and redundancy.
[0049] Regional density: Determine the density of each region based on the overall needs and design of the scene, that is, how many assets to place in a specific area. By calculating the density within the region, adjust the quantity and arrangement of each asset to achieve a reasonable distribution of assets.
[0050] When the scene is expanded, it is divided into multiple small blocks, and the content of each small block is loaded separately and spliced in real time. The scene is divided into multiple small blocks, and only the part that needs to be displayed at the moment is loaded into the memory, thereby reducing the consumption of computing resources. When loading a new scene block, the new block is seamlessly connected to the existing scene data through a real-time splicing algorithm to ensure the continuity and stability of the scene when expanding. In this way, seamless and smooth scene expansion can be achieved without boundary restrictions, which is suitable for application scenarios such as virtual reality and augmented reality.
[0051] By adopting the above technical solution, by segmenting and reconstructing repeated structural components from multi-view images, and generating transformation rules and asset distribution rules through programmatic code, a 3D scene with consistency and high restoration can be efficiently constructed. Especially when the scene is expanded, the application of block loading and real-time splicing technology ensures the continuity and smoothness of the scene. This technical solution improves the efficiency of large-scale 3D scene reconstruction, and ensures the stability and scalability of the scene, and can adapt to the needs of dynamic environment and real-time update.
[0052] In one embodiment of the present application, the diffusion model is combined with a differentiable renderer for optimization, including: In the latent space of the diffusion model, the base asset texture is enhanced by multi-scale denoising; Project the optimized asset to multiple views through a differentiable renderer and calculate the geometric consistency loss:
[0053] in, represents the geometric consistency loss; is the number of partitions, Reconstruct the depth truth for the partition, To render the depth map; support users to modify the programmed code parameters in real time and dynamically trigger the re-optimization of local Gaussian parameters.
[0054] Specifically, the diffusion model is used to perform multi-scale denoising and enhancement on the texture of the underlying asset. The core idea of the diffusion model is to gradually remove noise and enhance texture details by processing the texture at multiple levels and scales. The texture information is mapped to the latent space of the diffusion model, and the denoising operation is applied in this space. The denoising process processes the texture through multi-scale convolution or diffusion processes to ensure that the final output texture can maintain details at different scales and the noise is effectively suppressed.
[0055] The optimized assets are rendered using a differentiable renderer, which projects them into a depth map from multiple perspectives. The differentiable renderer enables depth map-based rendering to obtain depth information consistent with the real scene. The rendered results are compared with the true depth value reconstructed by the partitions to calculate the geometric consistency loss. Geometric consistency loss The difference between the rendered depth map and the reconstructed depth truth is measured. The specific calculation formula is as follows:
[0056] in, represents the geometric consistency loss, is the number of partitions, Reconstruct the depth truth for the partition, To render the depth map; support users to modify the programmed code parameters in real time and dynamically trigger the re-optimization of local Gaussian parameters.
[0057] The system provides a user interaction interface that allows users to modify the parameters of the programmatic code in real time. By dynamically modifying parameters, such as spatial layout, component transformation rules, etc., users can adjust the details and structure of the scene according to their needs. When the user modifies the parameters, the system automatically triggers the re-optimization of local Gaussian parameters. When the programmatic code is executed, the Gaussian distribution parameters of the underlying assets related to the user's modifications are recalculated and updated to ensure that these modifications are correctly reflected in the 3D model. Through this dynamic optimization mechanism, users can control the scene generation process in real time to ensure that the final generated 3D scene meets the user's actual needs and design intent.
[0058] The above technical solution can significantly improve the texture quality and geometric accuracy of the basic assets in the three-dimensional scene through the multi-scale denoising enhancement of the diffusion model and the geometric consistency optimization of the differentiable renderer. The calculation and optimization of the geometric consistency loss ensures a high degree of consistency between the rendering results and the depth information of the real world, thereby improving the realism and usability of the model. In addition, the function of supporting users to modify the programmatic code parameters in real time and triggering the re-optimization of local Gaussian parameters allows users to flexibly adjust the details of the three-dimensional scene, further enhancing the applicability of this technical solution under dynamic scene generation and personalized needs.
[0059] In one embodiment of the present application, geometric consistency verification is performed in the latent space, including the following steps: The result of the partitioned 3D reconstruction is used as input, and the rendering diffusion model performs a diffusion process in a low-resolution latent space. Given a camera pose obtained by a spherical linear interpolation method, a differentiable renderer generates an initial image according to the camera pose as a condition of the diffusion process, wherein the low-resolution latent space is a partial area or sub-area of low resolution, and the differentiable renderer is a renderer whose output can differentiate the input parameters; The initial image is encoded in a low-resolution latent space using a variational autoencoder, and feature reverse sampling and extraction is performed in the latent space. The encoding is a process of iteratively adding Gaussian noise to the initial image, and the feature reverse sampling and extraction is to extract features similar to the original ordered image from the chaotic and disordered image through a denoising operation to generate a denoised image; Decoding the denoised image in pixel space using a decoder in a variational autoencoder, and converting the decoded denoised image into a pixel space image, wherein the pixel space is the original representation space of the image, in which each pixel has its specific color value and grayscale value; Pixel loss is used to maintain the three-dimensional consistency between the generated pixel space image and the input partitioned three-dimensional reconstruction result, and the three-dimensional network reconstruction optimization result is output.
[0060] Specifically, the results of the partitioned 3D reconstruction are used as input and rendered to perform the diffusion process. In order to verify the geometric consistency, the diffusion model first operates in a low-resolution latent space. The low-resolution latent space refers to a simplified part or sub-region of the 3D reconstruction result, which contains the most critical geometric information but is presented at a low resolution. In this process, a reasonable camera pose is calculated by the spherical linear interpolation method. The camera pose determines the viewing angle of the 3D scene and can affect the effect of image rendering. Then, a differentiable renderer (i.e., a renderer that can output and differentiate the input parameters) is used to generate an initial image based on the camera pose. This initial image is a condition for the diffusion process and provides a basis for subsequent denoising operations.
[0061] The initial image in the low-resolution latent space is encoded using a variational autoencoder. The encoding process is achieved by iteratively adding Gaussian noise to the initial image, gradually disturbing the details of the image and introducing noise. This process helps the model learn how to recover important geometric and texture information from the noisy image. By performing feature reverse sampling extraction in the latent space, the variational autoencoder can extract features similar to the original ordered image from the chaotic and disordered image during the denoising operation, and finally generate a denoised image. The core of this denoising process is to remove the noise in the image and restore its true structure and detail information.
[0062] The denoised image is decoded using the decoder in the variational autoencoder. The decoder converts the denoised latent space image into a pixel space image. In the pixel space, each image pixel has its specific color value and grayscale value, which represents the original representation of the image. The decoding process restores clear and realistic images by mapping the image from the latent space to the pixel space. These images can provide more intuitive visual effects and can be further used for geometric consistency verification.
[0063] Pixel loss is used to ensure the 3D consistency between the generated pixel space image and the input partitioned 3D reconstruction result. Specifically, the pixel loss is performed by comparing the pixel differences between the denoised image and the partitioned result of the 3D reconstruction. By minimizing this loss function, it is ensured that the generated image can accurately reflect the geometric structure of the partitioned 3D reconstruction result. In this process, the pixel loss ensures high consistency between the image details and the 3D model, thereby optimizing the results of the 3D network reconstruction.
[0064] By adopting the above technical solution, the reconstruction accuracy and consistency of the 3D model can be effectively improved by performing geometric consistency verification in the latent space. By combining the diffusion model with the variational autoencoder, not only can denoising be performed in the low-resolution latent space, but the difference between the generated image and the actual reconstruction result can also be accurately compared through pixel loss, thereby optimizing the 3D reconstruction process. Ultimately, the output 3D network reconstruction optimization result can better preserve the geometric consistency of the 3D structure and ensure the realism and accuracy of the scene.
[0065] In one embodiment of the present application, assembling the optimized hierarchical model according to the spatial layout parameters includes: Input the boundary point clouds of adjacent partitions into the graph convolutional network and calculate the feature similarity of the overlapping areas; When the overlap exceeds a threshold, partitions are automatically merged using programmatic rules, including: (a) Align the partition coordinate system to eliminate the stitching gap; (b) Smooth interpolation of Gaussian parameters of the merged area; Generate seamless global point cloud maps, supporting dynamic loading and memory compression.
[0066] In one embodiment of the present application, the feature similarity of the overlapping area is calculated using a graph attention mechanism: ; in, , E is the edge set of the overlapping area, For point and The characteristic vector of Represents two point clouds whose similarity is to be calculated; Represents a pair of adjacent points (nodes) in an edge set; Represents the feature vector The cosine similarity of Attention weight; and Represents learnable query and key transformation matrices; The dimension of the feature vector; The combined Gaussian interpolation formula is:
[0067] in and The Gaussian parameters to be merged, the weight w is calculated based on the attenuation of the distance to the overlapping center; The parameters of the merged Gaussian; and represents the weight coefficient; Represents Gaussian distribution The center coordinates of Indicates the center coordinates of the overlapping area.
[0068] Specifically, the point cloud to be processed is calculated through the graph attention mechanism and The graph attention mechanism can assign different attention weights to each pair of adjacent nodes through the connection relationship between nodes (i.e., edge set E), thereby accurately measuring the similarity of different parts in the point cloud. The specific calculation process is:
[0069] Through the graph attention mechanism, the system can automatically adjust the feature similarity of each pair of adjacent nodes to obtain accurate overlapping area similarity. This mechanism can effectively process the local feature relationship in point cloud data and help optimize data matching and fusion in the reconstruction process.
[0070] When merging overlapping areas in point clouds, the Gaussian interpolation method is used to weight the Gaussian parameters to be merged. This process helps to smooth the transition and ensure the geometric consistency of the merged area. Specifically, the two Gaussian distributions to be merged are and The parameters of are combined by the following formula:
[0071] are the parameters of the merged Gaussian distribution, which achieves a smooth transition between the two Gaussian distributions during the Gaussian interpolation process, ensuring the natural fusion of the geometric and texture features of the point cloud in the overlapping area.
[0072] By using the above technical solution and calculating the feature similarity of overlapping areas in the point cloud through the graph attention mechanism, the similarity between different point clouds can be evaluated more accurately, and the point cloud alignment and matching process can be optimized. The merging method based on Gaussian interpolation ensures the geometric consistency of the point cloud in the overlapping area, and smoothes the transition between different distributions through the distance attenuation strategy. Finally, the merged Gaussian parameters can maintain high accuracy and naturalness during the point cloud synthesis process, thereby improving the quality of the 3D reconstruction model.
[0073] In one embodiment of the present application, the SIFT operator extracts image features, specifically including the following steps: The SIFT operator is used to perform multi-scale Gaussian difference filtering on the input image, and the extreme points in the image are detected as key points; the key points are located in the scale space and the local neighborhood, and the main direction is assigned to them according to the gradient direction; the neighborhood of the key points is sampled through the graph convolutional neural network, and the local features of the neighborhood of the key points are aggregated to complete the feature extraction; For each pixel point in the target image, the pixel coordinates of the image features in the adjacent images in the two partitions are calculated and subtracted to obtain the disparity value, thereby constructing the corresponding disparity map between the two images; The depth value corresponding to each pixel in the target image is calculated through the camera parameters and the disparity value in the disparity map, so as to construct a depth map of the target image, where the pixel value of the pixel corresponds to the depth value of the pixel; Through the depth value in the depth map, each pixel is converted into a three-dimensional coordinate in the real world to obtain a point cloud; The above steps are repeated to obtain point cloud data containing texture feature information in each partition.
[0074] Specifically, the image is processed into multiple scale levels, and each scale is smoothed using a Gaussian filter. Then, the local feature changes of the image are calculated on images of different scales through the Gaussian difference method, so as to detect the extreme points in the image at different scales. Extreme points refer to points where the image intensity value is locally maximum or minimum in the current scale and the adjacent scales. These points usually correspond to key feature points in the image (such as corners, edges, etc.).
[0075] Locate the detected extreme points. First, it is necessary to accurately locate the position of the key points in the scale space, and to more finely locate each key point in its local neighborhood. By analyzing the gradient direction of the local area where each key point is located, the main direction of the key point is calculated. The main direction refers to the mainstream direction of the gradient direction of the pixel in the local area, which is used to ensure scale invariance under different rotation angles. Therefore, after assigning the main direction, the stability and consistency of the feature points under rotation transformation can be improved.
[0076] The neighborhood of each key point is sampled using a graph convolutional neural network (GCN). A graph convolutional neural network is a deep learning model that can process graph-structured data. By sampling the image information of the neighborhood around the key point, it can aggregate the local features of the neighborhood. This process effectively extracts the texture features around the key point and obtains the detailed information of the local area in the image. These extracted local features will provide the necessary feature support for subsequent matching, recognition, and reconstruction.
[0077] For each pixel in the target image, the disparity value is obtained by calculating the pixel coordinates in the adjacent images and subtracting them. The disparity value represents the spatial offset of the same scene point in the two images under the change of viewing angle. By matching the feature points in the images, the difference in their pixel positions under different viewing angles is calculated to construct a disparity map. The disparity map provides a disparity value for each pixel, which is used to infer the depth information of the object in stereo vision.
[0078] Based on the obtained disparity map, the depth value of each pixel is calculated using the internal and external parameters of the camera (such as focal length, camera baseline length, etc.). The depth value represents the physical distance between the pixel and the camera. Based on the geometric model of the camera, the depth of each pixel can be calculated using the disparity value to generate a complete depth map. The depth map is an important form of three-dimensional information representation, in which the value of each pixel corresponds to the depth information of a point in the scene, and is used to represent the position of an object in three-dimensional space.
[0079] Using the depth information in the depth map, the pixel position coordinates (i.e. x, y in the image coordinate system) are combined with the depth value, and back-projected through the internal and external parameters of the camera to calculate the actual position of each pixel in three-dimensional space. Finally, the three-dimensional coordinates of all pixels are combined to obtain a complete point cloud.
[0080] Repeat the above steps for each partition in the target image. In each partition, the features extracted by the SIFT operator, the calculated disparity value, depth value and other information will help generate point cloud data containing texture features in the partition. By looping through each partition, the point cloud data of the entire scene can be obtained. These point clouds contain not only spatial information, but also texture features in the image, which can more realistically reflect the structure and details of the scene.
[0081] The above technical solution uses the SIFT operator to extract multi-scale features in the image, and combines it with the deep learning graph convolutional neural network for feature extraction, which can accurately capture the local features in the image and effectively aggregate them, thereby providing stable and reliable texture information for point cloud generation. The construction of the disparity map and depth map enables the three-dimensional spatial position of each pixel to be accurately calculated and high-quality point cloud data to be generated.
[0082] In one embodiment of the present application, the calculation formula using the SIFT operator is:
[0083]
[0084] in, represents the accuracy of the gradient vector, represents the direction of the gradient vector, represents the coordinates of each sampling point, , Represents the horizontal and vertical coordinates of the coordinate point.
[0085] Using the above technical solution, by calculating the accuracy of the gradient vector and direction , the SIFT operator can accurately capture local features in an image and provide a basis for subsequent feature description. The gradient magnitude reflects the intensity of detail changes in an image, and the gradient direction describes the direction of the change. These two quantities are the core of feature point extraction, matching, and rotation invariance. In different areas of an image, this information helps identify stable and unique local features, providing high-quality feature data for applications such as 3D reconstruction and target recognition.
[0086] In one embodiment of the present application, the calculation formula for calculating the Hessian matrix of each image in the image set is as follows: ; in = , represents the Hessian matrix, Represents each sampling point, express, express, express, They all represent the values of the second-order derivatives, , Represents the horizontal and vertical coordinates of the coordinate point.
[0087] With the above technical solution, the Hessian matrix can accurately capture the areas with dramatic changes in the image, especially for the detection of structural features such as edges and corners. In an image, edges and corners are usually the places with the most dramatic changes, and the second-order derivative can effectively reflect these changes. The second-order derivative has good robustness to scale changes and noise. This allows the Hessian matrix to maintain a relatively stable performance in multi-scale image analysis.
[0088] The above description is only a preferred embodiment of the present invention, and does not limit the patent scope of the present invention. All equivalent structural changes made by using the contents of the present invention specification and drawings under the inventive concept of the present invention, or directly / indirectly applied in other related technical fields are included in the patent protection scope of the present invention.
Claims
1. A method for 3D reconstruction of unbounded scenes based on drone aerial photography, characterized in that: The following steps are involved: Obtain a multi-view image set of the target area and construct a spatial topological association network; Image features are extracted based on the Hessian matrix and SIFT operator, and scene partitions are generated through Euclidean distance matching; Perform 3D Gaussian sputtering optimization on each partition and simultaneously generate a reusable 3D base asset, which includes structural components and texture features; Based on the spatial transformation rules defined by programmatic code, the underlying assets are dynamically combined to generate a hierarchical three-dimensional model; Geometric consistency verification is performed in latent space by jointly optimizing a diffusion model with a differentiable renderer. The optimized hierarchical models are assembled into an unbounded three-dimensional scene according to the spatial layout parameters.
2. The method for 3D reconstruction of unbounded scenes based on drone aerial photography according to claim 1, characterized in that: The 3D Gaussian sputtering optimization includes: Initialize programmatic code for each partition to define layout constraints of the underlying assets, including spatial boundaries, rotation parameters, and scaling; During 3D Gaussian sputtering optimization, the Gaussian distribution range is limited by bounding box adaptive constraints, including: (a) Gaussian center Beyond soft bounding box , forcing its position to align to the boundary; (b) If the Gaussian center exceeds the boundary, the threshold is extended according to the boundary Update its location; The formula is as follows: Synchronously update the Gaussian parameters of repeated underlying assets and achieve parameter sharing through gradient back propagation; represents the adjusted Gaussian center position, represents the original x-axis lower bound of partition i; Represents the original x-axis upper bound of partition i; Indicates the boundary extension alignment threshold; Soft boundary relaxation amount; represents the original position of the Gaussian center; Represents the center coordinate of the kth Gaussian distribution on the x-axis.
3. The method for 3D reconstruction of unbounded scenes based on drone aerial photography as claimed in claim 2, characterized in that: The programmatic code generation includes: Segment repetitive structural components from multi-view images and reconstruct their geometric and texture features through 3D Gaussian sputtering optimization; Encode the transformation rules of components into a programmed instruction set, including layer expansion, axial arrangement, and size adaptive adjustment; Automatically generate asset distribution rules based on spatial layout parameters; When the unbounded scene is expanded, scene continuity is achieved through block loading and real-time stitching.
4. The method for 3D reconstruction of unbounded scenes based on drone aerial photography according to claim 1, characterized in that: Joint optimization through diffusion model and differentiable renderer, including: In the latent space of the diffusion model, the base asset texture is enhanced by multi-scale denoising; Project the optimized asset to multiple views through a differentiable renderer and calculate the geometric consistency loss: in, represents the geometric consistency loss; is the number of partitions, Reconstruct the depth truth for the partition, To render the depth map; support users to modify the programmed code parameters in real time and dynamically trigger the re-optimization of local Gaussian parameters.
5. The method for 3D reconstruction of unbounded scenes based on drone aerial photography as claimed in claim 4, characterized in that: Performing geometric consistency verification in latent space includes the following steps: The result of the partitioned 3D reconstruction is used as input, and the rendering diffusion model performs a diffusion process in a low-resolution latent space. Given a camera pose obtained by a spherical linear interpolation method, a differentiable renderer generates an initial image according to the camera pose as a condition of the diffusion process, wherein the low-resolution latent space is a partial area or sub-area of low resolution, and the differentiable renderer is a renderer whose output can differentiate the input parameters; The initial image is encoded in a low-resolution latent space using a variational autoencoder, and feature reverse sampling and extraction is performed in the latent space. The encoding is a process of iteratively adding Gaussian noise to the initial image, and the feature reverse sampling and extraction is to extract features similar to the original ordered image from the chaotic and disordered image through a denoising operation to generate a denoised image; Decoding the denoised image in pixel space using a decoder in a variational autoencoder, and converting the decoded denoised image into a pixel space image, wherein the pixel space is the original representation space of the image, in which each pixel has its specific color value and grayscale value; Pixel loss is used to maintain the three-dimensional consistency between the generated pixel space image and the input partitioned three-dimensional reconstruction result, and the three-dimensional network reconstruction optimization result is output.
6. The method for 3D reconstruction of unbounded scenes based on drone aerial photography according to claim 1, characterized in that: The step of assembling the optimized hierarchical model according to the spatial layout parameters includes: Input the boundary point clouds of adjacent partitions into the graph convolutional network and calculate the feature similarity of the overlapping areas; When the overlap exceeds a threshold, partitions are automatically merged using programmatic rules, including: (a) Align the partition coordinate system to eliminate the stitching gap; (b) Smooth interpolation of Gaussian parameters of the merged area; Generate seamless global point cloud maps, supporting dynamic loading and memory compression.
7. The method for 3D reconstruction of unbounded scenes based on drone aerial photography according to claim 1, characterized in that: The graph attention mechanism is used to calculate the feature similarity of the overlapping area: ; in, , E is the edge set of the overlapping area, For point and The characteristic vector of Represents two point clouds whose similarity is to be calculated; Represents a pair of adjacent points (nodes) in an edge set; Represents the feature vector The cosine similarity of Attention weight; and Represents learnable query and key transformation matrices; The dimension of the feature vector; The combined Gaussian interpolation formula is: in and The Gaussian parameters to be merged, the weight w is calculated based on the attenuation of the distance to the overlapping center; The parameters of the merged Gaussian; and represents the weight coefficient; Represents Gaussian distribution The center coordinates of Indicates the center coordinates of the overlapping area.
8. The method for 3D reconstruction of unbounded scenes based on drone aerial photography according to claim 1, characterized in that: The SIFT operator extracts image features, which specifically includes the following steps: The SIFT operator is used to perform multi-scale Gaussian difference filtering on the input image, and the extreme points in the image are detected as key points; the key points are located in the scale space and the local neighborhood, and the main direction is assigned to them according to the gradient direction; the neighborhood of the key points is sampled through the graph convolutional neural network, and the local features of the neighborhood of the key points are aggregated to complete the feature extraction; For each pixel point in the target image, the pixel coordinates of the image features in the adjacent images in the two partitions are calculated and subtracted to obtain the disparity value, thereby constructing the corresponding disparity map between the two images; The depth value corresponding to each pixel in the target image is calculated through the camera parameters and the disparity value in the disparity map, so as to construct a depth map of the target image, where the pixel value of the pixel corresponds to the depth value of the pixel; Through the depth value in the depth map, each pixel is converted into a three-dimensional coordinate in the real world to obtain a point cloud; The above steps are repeated to obtain the point cloud data containing texture feature information in each partition.
9. The method for 3D reconstruction of unbounded scenes based on drone aerial photography according to claim 1, characterized in that: The calculation formula using the SIFT operator is: in, represents the accuracy of the gradient vector, represents the direction of the gradient vector, represents the coordinates of each sampling point, , Represents the horizontal and vertical coordinates of the coordinate point.
10. The method for 3D reconstruction of unbounded scenes based on drone aerial photography according to claim 1, characterized in that: The calculation formula for calculating the Hessian matrix of each image in the image set is as follows: ; in = , represents the Hessian matrix, Represents each sampling point, express, express, express, They all represent the values of the second-order derivatives, , Represents the horizontal and vertical coordinates of the coordinate point.
Citation Information
Patent Citations
Unmanned aerial vehicle aerial photography target detection and identification method and system
CN114842365A
Sparse visual angle three-dimensional reconstruction method based on depth prior information
CN118657888A
Dense RGB-D SLAM method based on multistage 3D Gaussian
CN118781189A
Image reconstruction method and system based on aerial photography of unmanned aerial vehicle
CN119027590A
Laser enhanced vision three-dimensional reconstruction method and system based on Gaussian splashing
CN119180908A
Cited By
Three-dimensional model adjusting method and system and medium
CN120852636A
Cross-source data three-dimensional reconstruction method and system based on improved Gaussian sputtering
CN120976449A
A method and system for three-dimensional reconstruction of cross-source data based on improved Gaussian sputtering
CN120976449B
Three-dimensional visualization region model construction method for surveying and mapping data
CN121213770A
Three-dimensional modeling method and device based on neural implicit surface reconstruction and storage medium
CN121366256A