3D Reconstruction and Visualization Method of News Scenes Based on Multi-Source Remote Sensing Data
By combining multi-source remote sensing data processing technology and deep learning methods, hierarchical feature descriptors are extracted and data enhancement is performed, and the problems of low three-dimensional reconstruction accuracy and low visualization efficiency of news scenes are solved, achieving high-precision and efficient three-dimensional model generation and visualization.
Patent Information
- Application Number
- CN202510399528.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-01
AI Technical Summary
When handling complex scenarios, especially news scenarios, the prior art has problems such as difficulty in fusion of multi-source remote sensing data, low reconstruction accuracy, and low scene visualization efficiency.
Using iterative closest point algorithm based on spatial position weighting, tensor projection algorithm, bidirectional attention deep neural network, variational autoencoder network and generation adversarial network, hierarchical feature descriptors are extracted, data augmentation and depth estimation are performed, and multi-view projection consistency constraints and grid optimization algorithms are combined to generate and optimized three-dimensional scene models, and semantic segmentation and annotation are performed through a dual-branch semantic segmentation network.
The accuracy and visualization efficiency of three-dimensional reconstruction of news scenes are improved, and the generated three-dimensional model is more realistic and detailed, which can effectively highlight the location of news events and provide richer news information.
Smart Images

Figure CN119904592B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional reconstruction, and particularly relates to a method for three-dimensional reconstruction and visualization of news scenes based on multi-source remote sensing data. Background Art
[0002] The three-dimensional reconstruction technology aims to restore the three-dimensional structure of the real world from two-dimensional images or other data, and has been widely applied in fields such as urban planning, disaster assessment, virtual reality, etc. With the rapid development of remote sensing technology, the three-dimensional reconstruction based on multi-source remote sensing data has become a research hotspot, which can comprehensively utilize the information obtained by different sensors, such as the spectral information of optical images, the penetration ability of radar, and the high-precision elevation information of lidar, so as to reconstruct the three-dimensional scene more completely and accurately;
[0003] Traditional remote sensing three-dimensional reconstruction methods mainly rely on technologies such as stereophotogrammetry or structured light. However, when dealing with complex scenes, especially news scenes with dynamic changes and rich information, these methods have certain limitations, and there are still problems such as difficulty in fusing multi-source remote sensing data, low reconstruction accuracy of complex scenes, and low scene visualization efficiency;
[0004] Therefore, there is an urgent need for a solution to solve the problems existing in the prior art. Summary of the Invention
[0005] An embodiment of the present invention provides a method for three-dimensional reconstruction and visualization of news scenes based on multi-source remote sensing data, which can at least solve some of the problems existing in the prior art.
[0006] In the first aspect of the embodiment of the present invention, a method for three-dimensional reconstruction and visualization of news scenes based on multi-source remote sensing data is provided, including:
[0007] Collect multi-source remote sensing data corresponding to the target news scene through a remote sensing satellite, extract hierarchical feature descriptors based on the iterative closest point algorithm and tensor projection algorithm weighted by spatial position and a bidirectional attention deep neural network, and screen through a random sampling algorithm based on local geometric consistency constraints, eliminate feature point pairs with projection errors greater than a preset projection threshold and unify them to the global geographic coordinates to obtain registered multi-source remote sensing data, and perform data enhancement on the registered multi-source remote sensing data through a variational autoencoder network and a noise-aware loss function, in combination with a generative adversarial network based on pixel-level adversarial discrimination to obtain enhanced remote sensing data;
[0008] Add the enhanced remote sensing data to the feature extraction network to obtain the spatial attention feature map and the channel attention feature map. Add the spatial attention feature map to the monocular depth estimation network based on dense connections to generate depth information, fuse it with the dual-polarization radar remote sensing data and the multi-echo laser point cloud data to obtain the sparse three-dimensional point cloud, generate the initial three-dimensional mesh model through the Poisson reconstruction algorithm, add the channel attention feature map to the residual graph convolutional neural network for texture feature transfer, and combine the multi-view projection consistency constraint and the mesh optimization algorithm to perform topological optimization to obtain the optimized three-dimensional scene model;
[0009] Calculate the curvature change value and the normal vector change value of each grid in the optimized three-dimensional scene model, construct a spatial division evaluation function to divide the three-dimensional scene into octree spatial blocks of different sizes, calculate the quadratic error matrix of the grid vertices in each spatial block and construct an edge collapse cost function to merge and simplify the grid vertices to obtain four precision level models, construct a quadtree frustum to divide it into three depth regions and load the corresponding precision level models to obtain the visualization scene, extract local features and global features through the dual-branch semantic segmentation network, generate the scene semantic segmentation map through fusion upsampling and label the news event location, collect the interactive location information and calculate the viewpoint transformation matrix to update the visualization scene for scene roaming and perform scene roaming.
[0010] In an alternative embodiment,
[0011] Collect the multi-source remote sensing data corresponding to the target news scene through the remote sensing satellite, extract the hierarchical feature descriptors through the iterative closest point algorithm based on spatial position weighting, the tensor projection algorithm, and the bidirectional attention depth neural network, and screen them through the random sampling algorithm based on the local geometric consistency constraint, remove the feature point pairs with projection errors greater than the preset projection threshold and unify them to the global geographic coordinates to obtain the registered multi-source remote sensing data. Through the variational autoencoder network and the noise-aware loss function, combined with the generative adversarial network based on pixel-level adversarial discrimination, data augmentation is performed to obtain the enhanced remote sensing data including:
[0012] Collect the multi-source remote sensing data corresponding to the target news scene through the remote sensing satellite, and the multi-source remote sensing data includes optical remote sensing image data, dual-polarization radar remote sensing data, and multi-echo laser point cloud data;
[0013] Add the multi-source remote sensing data to the iterative closest point algorithm weighted by spatial position to construct a spatial position weight matrix. Construct a feature distance function based on the distance from the data point to the scene center, calculate the weight value, perform a nearest neighbor search on each data point through an octree spatial indexing structure to establish a set of nearest neighbor points, calculate the position deviation vector between the data point and the set of nearest neighbor points, multiply the position deviation vector by the spatial position weight matrix to obtain a weighted position deviation, repeat the iteration until the weighted position deviation is less than a preset threshold, and output the optimized set of data points;
[0014] Add the multi-source remote sensing data to the tensor projection algorithm, construct a feature covariance matrix through principal component analysis, determine the optimal projection direction by calculating eigenvalues and eigenvectors, perform a tensor projection operation in the optimal projection direction, construct a local curvature calculation function and a normal vector change function for the projection result, and extract multi-modal data feature points through a curvature threshold segmentation and region growing algorithm;
[0015] Add the multi-modal data feature points to a bidirectional attention deep neural network. The bidirectional attention deep neural network extracts multi-scale features through multi-layer convolution and pooling operations of the encoder. The multi-layer convolution includes point convolution layers, edge convolution layers, and surface convolution layers, and restores the features through deconvolution and upsampling operations of the decoder. A bidirectional attention module is set between the encoder and the decoder. Among them, the bidirectional attention module includes a self-attention mechanism for calculating the internal spatial correlation degree of feature points and a cross-attention mechanism for calculating the topological correlation degree between different feature points, and fuses the feature mapping results of the self-attention mechanism and the cross-attention mechanism to obtain a hierarchical feature descriptor;
[0016] Screen the hierarchical feature descriptor through a random sampling algorithm based on local geometric consistency constraints, determine candidate feature point pairs through feature similarity calculation, randomly select a feature point pair as a seed match, calculate an initial geometric transformation relationship including a rotation matrix and a translation vector based on the seed match, perform a local geometric consistency constraint test on the remaining feature point pairs to obtain a projection error, remove the feature point pairs with a projection error greater than a preset projection threshold, repeat the random sampling process multiple times to select the geometric transformation relationship with the most inlier support, and unify the multi-source remote sensing data to the global geographic coordinates based on the selected geometric transformation relationship to generate registered multi-source remote sensing data;
[0017] Input the registered multi-source remote sensing data into a variational autoencoder network. The encoder of the variational autoencoder network maps the registered multi-source remote sensing data to a latent feature space through multi-layer residual convolution, and the decoder reconstructs the latent features into the form of the original data through deconvolution and skip connections. Random noise perturbations are introduced into the latent feature space to construct a noise-aware loss function to constrain the reconstruction result;
[0018] Construct a generative adversarial network based on pixel-level adversarial discrimination. The generative adversarial network adopts a densely connected fully convolutional structure to output a pixel-level authenticity score map. Use the decoder of the variational autoencoder network as the generator, construct an adversarial loss function, and optimize the parameters of the generator and the generative adversarial network through a min-max game until reaching the Nash equilibrium to obtain enhanced remote sensing data.
[0019] In an alternative implementation,
[0020] Screen the hierarchical feature descriptors through a random sampling algorithm based on local geometric consistency constraints, determine candidate feature point pairs through feature similarity calculation, randomly select a feature point pair as a seed match, calculate an initial geometric transformation relationship including a rotation matrix and a translation vector based on the seed match, perform a local geometric consistency constraint test on the remaining feature point pairs to obtain a projection error, remove the feature point pairs with a projection error greater than a preset projection threshold, repeat the random sampling process multiple times to select the geometric transformation relationship with the most inlier support, and unify the multi-source remote sensing data to the global geographic coordinates based on the selected geometric transformation relationship. The generated registered multi-source remote sensing data includes:
[0021] Construct a local geometric feature extraction unit and a global semantic feature extraction unit. Among them, the local geometric feature extraction unit extracts a local shape descriptor by constructing a feature point neighborhood and analyzing the distribution characteristics of points in the neighborhood. The global semantic feature extraction unit extracts hierarchical semantic information from bottom to top through a multi-scale feature pyramid, and concatenates the local shape descriptor and the hierarchical semantic information to form a feature vector;
[0022] Construct a feature similarity calculation module. The feature similarity calculation module is based on the cosine similarity metric method, introduces a spatial distance weighting term to make feature points with closer distances have a higher matching priority, and pairs the feature descriptors with a similarity higher than a preset threshold to form a set of candidate feature point pairs;
[0023] Randomly select a feature point pair from the set of candidate feature point pairs as a seed match, construct a geometric transformation estimation module based on the seed match, establish a mapping relationship between corresponding points through the geometric transformation estimation module to construct a three-dimensional rigid body transformation model, decompose the rotation matrix in the transformation process into three basic rotation matrices and solve them through an iterative optimization algorithm, and solve the translation vector by minimizing the Euclidean distance between corresponding points;
[0024] Construct a geometric consistency checking module, which includes a feature transformation unit and an error calculation unit. The feature transformation unit transforms the source feature points to the target coordinate system through the three-dimensional rigid body transformation model. The error calculation unit calculates the projection error between the transformed position and the actual target position, dynamically adjusts the projection error threshold according to the data distribution characteristics, and eliminates the feature point pairs with projection errors greater than the projection error threshold;
[0025] Construct a random sample consensus (RANSAC) optimization module. The RANSAC optimization module randomly selects different seed matches to execute the geometric transformation estimation module and the geometric consistency checking module, counts the number of inliers that meet the constraint conditions, scores by comprehensively considering the number of inliers and the stability of the geometric transformation, and selects the geometric transformation relationship with the highest score as the optimal geometric transformation relationship;
[0026] Based on the optimal geometric transformation relationship, construct a coordinate transformation module. Based on the coordinate transformation module, use the optical remote sensing image data as the reference data, construct an affine transformation model considering the imaging geometric characteristics for the dual-polarization radar remote sensing data, solve the transformation parameters through polynomial fitting method to achieve geometric correction, construct a three-dimensional space rigid body transformation model for the multi-echo laser point cloud data, solve the transformation parameters through iterative optimization method, and uniformly transform the optical remote sensing image data, the dual-polarization radar remote sensing data and the multi-echo laser point cloud data to the global geographic coordinate system.
[0027] In an alternative embodiment,
[0028] Add the enhanced remote sensing data to the feature extraction network to obtain a spatial attention feature map and a channel attention feature map. Add the spatial attention feature map to the monocular depth estimation network based on dense connections to generate depth information, fuse it with the dual-polarization radar remote sensing data and the multi-echo laser point cloud data to obtain a sparse three-dimensional point cloud, generate an initial three-dimensional mesh model through Poisson reconstruction algorithm, add the channel attention feature map to the residual graph convolutional neural network for texture feature transfer, and perform topological optimization through multi-view projection consistency constraint and mesh optimization algorithm to obtain an optimized three-dimensional scene model, including:
[0029] Input the enhanced remote sensing data into the feature extraction network, extract multi-scale spatial features through a multi-level cascaded convolutional structure and a batch normalization layer, use an adaptive weight calculation unit to assign spatial attention weights to each position, aggregate information through global pooling to obtain a channel descriptor, perform multi-layer transformation on the channel descriptor to learn the dependencies between channels to generate channel attention weights, and generate a spatial attention feature map and a channel attention feature map through adaptive fusion of the spatial attention weights and the channel attention weights;
[0030] Input the spatial attention feature map into a monocular depth estimation network based on dense connections. In the monocular depth estimation network, establish multi-level cascade connections in a jumping manner, establish a direct mapping relationship between shallow features and deep features, gradually restore the feature resolution through a multi-level deconvolution structure, and use a multi-scale feature pyramid structure to adaptively fuse the weights of features at different levels to generate depth information;
[0031] Sample the depth information and establish a mapping relationship from pixel coordinates to three-dimensional space points. Calculate the confidence weight based on the signal-to-noise ratio and spatial resolution for the pre-acquired dual-polarization radar remote sensing data and multi-echo laser point cloud data. Construct an adaptive weighted fusion strategy according to the spatial distance and feature similarity, and perform point cloud fusion on the overlapping area to generate a sparse three-dimensional point cloud;
[0032] Input the sparse three-dimensional point cloud into the Poisson reconstruction algorithm. Calculate the normal vector of each point through the feature decomposition method, dynamically adjust the spatial subdivision granularity according to the point cloud density distribution, obtain the implicit surface expression by iteratively solving the Poisson equation, extract the grid structure in combination with an adaptive threshold and perform Laplacian smoothing to generate an initial three-dimensional grid model. Input the channel attention feature map into the residual graph convolutional neural network, transfer node information through local neighborhood aggregation, add the input feature and the transformed feature to construct a residual connection, adaptively capture texture features through the graph attention mechanism, and use projection transformation to map the texture features to the surface of the initial three-dimensional grid model for texture feature transfer;
[0033] Uniformly sample multiple viewing perspectives in the spherical space, project the initial three-dimensional grid model onto each viewing perspective plane to generate multi-view projection images, calculate the similarity between the multi-view projection images and the original images through feature descriptor matching to establish a multi-view projection consistency constraint, and optimize the grid vertex positions and connection relationships based on the multi-view projection consistency constraint while maintaining the geometric details of the model to generate an optimized three-dimensional scene model.
[0034] In an alternative embodiment,
[0035] Input the sparse three-dimensional point cloud into the Poisson reconstruction algorithm. Calculate the normal vector of each point through the feature decomposition method, dynamically adjust the spatial subdivision granularity according to the point cloud density distribution, obtain the implicit surface expression by iteratively solving the Poisson equation, extract the grid structure in combination with an adaptive threshold and perform Laplacian smoothing to generate an initial three-dimensional grid model. Input the channel attention feature map into the residual graph convolutional neural network, transfer node information through local neighborhood aggregation, add the input feature and the transformed feature to construct a residual connection, adaptively capture texture features through the graph attention mechanism, and use projection transformation to map the texture features to the surface of the initial three-dimensional grid model for texture feature transfer includes:
[0036] Input the sparse three-dimensional point cloud into the Poisson reconstruction algorithm, construct a point cloud nearest neighbor search tree structure, use the spherical neighborhood search method to determine the local neighborhood point set corresponding to each point in the sparse three-dimensional point cloud, construct a covariance matrix based on the local neighborhood point set, calculate the normal vector of each point through the eigenvalue decomposition method, construct a minimum spanning tree to establish the connection relationship between points, select the normal vector at the center position of the point cloud as the reference direction, and recursively traverse the minimum spanning tree to adjust the pointing directions of all points' normal vectors to obtain a set of normal vectors;
[0037] Calculate the outer bounding box of the point cloud as the initial spatial range, dynamically adjust the spatial subdivision granularity according to the point cloud density distribution, count the number of point clouds contained in each octree node and calculate the point density distribution, set an adaptive subdivision threshold based on the point density, divide the nodes with point density greater than the adaptive subdivision threshold into eight child nodes and record the depth information and the contained point set of each leaf node, perform a merging operation on the divided leaf nodes, and merge adjacent nodes with similar point densities to obtain an optimized octree structure;
[0038] Define a three-dimensional scalar field function on the optimized octree structure, convert the set of normal vectors into a gradient field constraint, obtain an implicit surface representation by iteratively solving the Poisson equation, construct a solution hierarchy using the multigrid strategy, convert the Poisson equation into a linear system of equations by constructing a sparse matrix to represent the Laplacian operator, and solve the linear system of equations using an iterative optimization method and update the scalar values of the grid nodes to obtain a three-dimensional scalar field;
[0039] Extract the grid structure by combining an adaptive threshold, extract an isosurface on the optimized octree structure based on the marching cubes algorithm to generate an initial grid, perform a topological check on the initial grid and repair non-manifold edges and singular points, perform Laplacian smoothing to iteratively optimize the vertex positions, and simplify the grid based on the edge collapse criterion to obtain an initial three-dimensional grid model;
[0040] Input the channel attention feature map into a residual graph convolutional neural network, transfer node information through local neighborhood aggregation, define an adjacency matrix to describe the connection relationship between nodes, perform message aggregation based on the adjacency matrix to weight-combine the node features and neighborhood features, and add the input features and the transformed features to construct a residual connection to obtain enhanced features;
[0041] Adaptively capture texture features through the graph attention mechanism, obtain feature statistical information through global pooling operation on the enhanced features and perform a non-linear transformation, learn the correlation weights between different channels to generate channel attention weights, calculate the importance scores of different positions on the feature map to generate spatial attention weights, and apply the channel attention weights and the spatial attention weights to the feature map to obtain texture features;
[0042] The texture features are mapped to the surface of the initial three-dimensional mesh model by projective transformation for texture feature transfer, the mapping relationship between the vertices of the initial three-dimensional mesh model and the image pixels is established, the vertices are projected onto the image plane according to the camera projection matrix, and barycentric coordinate interpolation is used to calculate the texture coordinates inside the mesh patches, and multiple views are constructed Figure 1 Consistency constraints are used to evaluate the texture mapping quality from different perspectives, the optimal texture source is selected based on the graph cut optimization algorithm, and the texture fineness is dynamically adjusted according to the viewing distance to obtain a textured three-dimensional scene model.
[0043] In an alternative embodiment,
[0044] Calculate the curvature change value and the normal vector change value of each mesh in the optimized three-dimensional scene model, construct a spatial partition evaluation function to divide the three-dimensional scene into octree spatial blocks of different sizes, calculate the quadratic error matrix of the mesh vertices in each spatial block and construct an edge collapse cost function to merge and simplify the mesh vertices, obtain four precision level models, construct a quadtree frustum divided into three depth regions to load the corresponding precision level models to obtain a visualization scene, extract local features and global features through a double-branch semantic segmentation network, generate a scene semantic segmentation map by fusion upsampling and label the news event positions, collect interaction position information and calculate the viewpoint transformation matrix to update the visualization scene for scene roaming, and the scene roaming includes:
[0045] For each mesh patch in the optimized three-dimensional scene model, project the normal vector of each mesh patch onto the vertex local coordinate system and calculate the weighted average to obtain the vertex normal vector, construct a local surface fitting equation of the vertex to calculate the principal curvature to obtain the curvature change value, traverse adjacent vertices to calculate the normal vector angle change rate, combine the curvature change value and the normal vector angle change rate to form a vertex feature description, and establish an adjacency relationship graph recording the vertex feature description;
[0046] Input the vertex feature description, the vertex density distribution and the feature point distribution ratio in the adjacency relationship graph into the spatial partition evaluation function, recursively construct an octree with the scene bounding box as the initial space, calculate the evaluation function value of each node to be divided, dynamically adjust the segmentation threshold according to the level of the node and the evaluation value of the parent node, and record the spatial range, the vertex feature description and the mesh topology relationship of each leaf node to form a spatial index structure;
[0047] For each leaf node in the spatial index structure, construct a quadratic error matrix by the sum of the squares of the distances from the vertex to the adjacent patches, combine the quadratic error matrix, the normal vector change rate and the texture coordinate change amount to construct an edge collapse cost function, traverse the mesh edges in the node to calculate the optimal collapse position and store it in a priority queue, and iteratively perform the edge collapse operation according to the magnitude of the value of the collapse cost function, and update the error matrix of the affected vertices and the cost values of the relevant edges;
[0048] Control the number of collapse iterations based on the edge collapse cost function, maintain the continuity of the vertex feature description, generate four precision-level models that retain 90%, 70%, 50%, and 30% of the original number of vertices, construct a quadtree structure corresponding to the frustum, divide the frustum into a near-view region, a mid-view region, and a far-view region, set an overlapping transition zone between adjacent regions, load the two models with the highest level of detail in the four precision-level models into the near-view region, load the model with the third-highest level of detail into the mid-view region, and load the model with the lowest level of detail into the far-view region;
[0049] Calculate vertex position interpolation based on the vertex feature description within the overlapping transition zone, perform smooth switching between different precision models to obtain a visualization scene, input the visualization scene into a dual-branch semantic segmentation network, extract local features through multi-layer convolution and downsampling, and extract global features through multi-level pooling;
[0050] Fuse the local features and the global features through skip connections, perform deconvolution processing on the fused features to restore the spatial resolution to generate a semantic segmentation map, analyze the semantic segmentation map to obtain the distribution of different category regions in the scene, establish an association between the input news event description and the corresponding category regions, and convert the annotation position to the scene space coordinate system through line-of-sight projection;
[0051] Obtain the position coordinates and operation instructions of the interaction device, convert the screen space coordinates to the normalized device coordinate system, construct an interaction ray to calculate the intersection position with the visualization scene, calculate the rotation matrix, translation vector, and scaling factor of the viewpoint based on the intersection position, generate a viewpoint transformation path, and dynamically update the precision configuration of the models within the visible range and the display state of the annotation symbols according to the viewpoint position.
[0052] In an alternative embodiment,
[0053] For each leaf node in the spatial index structure, construct a quadratic error matrix from the sum of the squares of the distances from the vertices to the adjacent patches, combine the quadratic error matrix, the normal vector change rate, and the texture coordinate change amount to construct an edge collapse cost function, traverse the mesh edges within the node to calculate the optimal collapse position and store it in a priority queue, and iteratively perform the edge collapse operation according to the magnitude of the values of the collapse cost function, updating the error matrix of the affected vertices and the cost values of the relevant edges, including:
[0054] Obtain the grid information of each leaf node in the spatial index structure, establish an association table between vertices and adjacent patches, obtain the patch set and topological relationship to which each vertex belongs, calculate the patch normal vector and plane equation based on the patch set, calculate the sum of the squares of the distances from the vertex to the adjacent patches to construct a quadratic error matrix, where the quadratic error matrix characterizes the influence degree of the change of the vertex position on the adjacent patches;
[0055] Combine the quadratic error matrix, the normal vector change rate, and the texture coordinate change amount to construct an edge collapse cost function, calculate the shape deformation value based on the quadratic error matrices of the two endpoints of the edge, obtain the normal vector change rate by calculating the normal vectors of the two endpoints of the edge and the adjacent vertices, analyze the vertex texture coordinates to obtain the coordinate change amount, and perform linear combination after normalizing the shape deformation value, the normal vector change rate, and the texture coordinate change amount to obtain the edge collapse cost function value;
[0056] Traverse the grid edges in the node to calculate the optimal collapse position, obtain the patches adjacent to the edge to construct a local grid area, select multiple candidate collapse points on the straight line where the edge is located, calculate the grid error corresponding to each candidate point, and select the position with the smallest grid error and maintaining topological validity as the optimal collapse position, and store the edge information and the optimal collapse position into a priority queue sorted according to the cost function value;
[0057] Extract the edge with the smallest cost function value from the priority queue, obtain the optimal collapse position corresponding to the current edge after checking the validity of the edge, merge the two endpoints of the edge into a new vertex, update the coordinates, normal vector, and texture coordinates of the new vertex according to the optimal collapse position, update the vertex indices of the relevant patches, delete the degenerate patches and adjust the grid connection relationship;
[0058] Update the error matrix of the affected vertices, integrate the geometric relationship between the new vertex and the adjacent patches into the error matrix, recalculate the cost function value and the optimal collapse position for the edges connected to the new vertex, and re-store the updated edge information into the priority queue to maintain the optimal simplification order.
[0059] In the second aspect of the embodiments of the present invention, a three-dimensional reconstruction and visualization system for news scenes based on multi-source remote sensing data is provided, including:
[0060] The first unit is used to collect multi-source remote sensing data corresponding to the target news scene through remote sensing satellites, extract hierarchical feature descriptors based on the iterative closest point algorithm and tensor projection algorithm with spatial position weighting, and a bidirectional attention deep neural network, and screen them through a random sampling algorithm based on local geometric consistency constraints, remove feature point pairs with projection errors greater than a preset projection threshold, and unify them to global geographic coordinates to obtain registered multi-source remote sensing data. Through a variational autoencoder network and a noise-aware loss function, combined with a generative adversarial network based on pixel-level adversarial discrimination, data enhancement is performed to obtain enhanced remote sensing data;
[0061] The second unit is used to add the enhanced remote sensing data to a feature extraction network to obtain a spatial attention feature map and a channel attention feature map, add the spatial attention feature map to a monocular depth estimation network based on dense connections to generate depth information, fuse it with dual-polarization radar remote sensing data and multi-echo laser point cloud data to obtain sparse three-dimensional point clouds, generate an initial three-dimensional mesh model through a Poisson reconstruction algorithm, add the channel attention feature map to a residual graph convolutional neural network for texture feature transfer, and combine multi-view projection consistency constraints and a mesh optimization algorithm for topological optimization to obtain an optimized three-dimensional scene model;
[0062] The third unit is used to calculate the curvature change value and normal vector change value of each grid in the optimized three-dimensional scene model, construct a spatial division evaluation function to divide the three-dimensional scene into octree spatial blocks of different sizes, calculate the quadratic error matrix of the grid vertices in each spatial block, and construct an edge collapse cost function to merge and simplify the grid vertices to obtain four precision-level models. Construct a quadtree frustum to divide it into three depth regions and load the corresponding precision-level models to obtain a visualization scene. Extract local features and global features through a dual-branch semantic segmentation network, generate a scene semantic segmentation map through fusion upsampling and label the news event positions, collect interactive position information, calculate a viewpoint transformation matrix to update the visualization scene for scene roaming and perform scene roaming.
[0063] In the third aspect of the embodiments of the present invention,
[0064] A kind of electronic device is provided, including:
[0065] A processor;
[0066] A memory for storing instructions executable by the processor;
[0067] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0068] In the fourth aspect of the embodiments of the present invention,
[0069] Provided is a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, the foregoing method is implemented.
[0070] In the present invention, by using multi-source remote sensing data (including satellite images, radar, and lidar point clouds), and combining deep learning techniques for feature extraction, data augmentation, and depth estimation, a more refined and realistic 3D model of the news scene can be generated. The texture feature transfer and mesh optimization algorithms further enhance the visual effect of the model. By adopting octree space partitioning and quadtree frustum techniques, models with different precisions are dynamically loaded according to the viewing distance, achieving efficient scene rendering and roaming, and avoiding the computational burden caused by loading models with too high precision. By using semantic segmentation technology to annotate the scene and combining interactive operations, the location of the news event can be highlighted, facilitating users to understand the scene and providing richer news information. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 is a schematic flow chart of a method for 3D reconstruction and visualization of a news scene based on multi-source remote sensing data according to an embodiment of the present invention;
[0072] Figure 2 is a schematic structural diagram of a system for 3D reconstruction and visualization of a news scene based on multi-source remote sensing data according to an embodiment of the present invention;
[0073] Figure 3 is a comparison diagram of the simulation effects of 3D model reconstruction of a method for 3D reconstruction and visualization of a news scene based on multi-source remote sensing data according to an embodiment of the present invention;
[0074] Figure 4 is a comparison diagram of the simplified effects of texture models of a method for 3D reconstruction and visualization of a news scene based on multi-source remote sensing data according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0075] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0076] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0077] Figure 1The flowchart of the 3D reconstruction and visualization method for news scenes based on multi-source remote sensing data according to the embodiments of the present invention is shown as Figure 1 follows. The method includes:
[0078] S1. Collect multi-source remote sensing data corresponding to the target news scene through remote sensing satellites. Extract hierarchical feature descriptors based on the spatially weighted iterative closest point algorithm, tensor projection algorithm, and bidirectional attention deep neural network. Screen through the random sampling algorithm based on local geometric consistency constraints, remove feature point pairs with projection errors greater than the preset projection threshold, and unify them to global geographic coordinates to obtain registered multi-source remote sensing data. Through the variational autoencoder network and noise-aware loss function, combined with the generative adversarial network based on pixel-level adversarial discrimination, perform data augmentation to obtain enhanced remote sensing data;
[0079] The spatially weighted iterative closest point algorithm is an improved point cloud registration method that assigns different weights to points in different regions, thereby improving the registration accuracy. It is particularly suitable for non-uniformly sampled or noisy point cloud data. The tensor projection algorithm is a high-dimensional data dimensionality reduction method that extracts main information by projecting high-order tensor data into a low-dimensional subspace. It can improve the computational efficiency while maintaining the data structure information and is widely used in fields such as image processing and pattern recognition. The hierarchical feature descriptor is a multi-level feature expression method that extracts key features of data at different scales and levels, enabling it to better adapt to data matching tasks with different resolutions and complexities, and improving the robustness and generalization ability of the matching. The random sampling algorithm based on local geometric consistency constraints is an optimized feature matching method that combines local geometric consistency constraints during the random sampling process to exclude incorrect matching point pairs and improve the accuracy of the algorithm. It is commonly used in 3D point cloud registration and feature matching tasks. The noise-aware loss function is a loss function design for data noise robustness. By adaptively adjusting the penalty degree for noise points, it reduces the impact of noise on model training, improves the stability and generalization ability of the model, and is widely used in image denoising, object detection, and the deep learning training process.
[0080] In an alternative embodiment,
[0081] Collecting multi-source remote sensing data corresponding to the target news scene through remote sensing satellites, extracting hierarchical feature descriptors based on the spatially weighted iterative closest point algorithm, tensor projection algorithm, and bidirectional attention deep neural network, screening through the random sampling algorithm based on local geometric consistency constraints, removing feature point pairs with projection errors greater than the preset projection threshold, and unifying them to global geographic coordinates to obtain registered multi-source remote sensing data, and performing data augmentation through the variational autoencoder network and noise-aware loss function, combined with the generative adversarial network based on pixel-level adversarial discrimination, includes:
[0082] Collect multi-source remote sensing data corresponding to the target news scene through a remote sensing satellite. The multi-source remote sensing data includes optical remote sensing image data, dual-polarization radar remote sensing data, and multi-echo laser point cloud data;
[0083] Add the multi-source remote sensing data to the iterative closest point algorithm weighted by spatial position, construct a spatial position weight matrix, construct a feature distance function based on the distance from the data point to the scene center, calculate the weight value, perform a nearest neighbor search on each data point through an octree spatial index structure to establish a nearest neighbor point set, calculate the position deviation vector between the data point and the nearest neighbor point set, multiply the position deviation vector by the spatial position weight matrix to obtain a weighted position deviation, and repeat the iteration until the weighted position deviation is less than a preset threshold, and output the optimized data point set;
[0084] Add the multi-source remote sensing data to the tensor projection algorithm, construct a feature covariance matrix through principal component analysis, determine the optimal projection direction by calculating eigenvalues and eigenvectors, perform a tensor projection operation in the optimal projection direction, construct a local curvature calculation function and a normal vector change function for the projection result, and extract multi-modal data feature points through a curvature threshold segmentation and region growing algorithm;
[0085] Add the multi-modal data feature points to a bidirectional attention deep neural network. The bidirectional attention deep neural network extracts multi-scale features through multi-layer convolution and pooling operations of the encoder. The multi-layer convolution includes point convolution layers, edge convolution layers, and surface convolution layers, restores the features through deconvolution and upsampling operations of the decoder, and a bidirectional attention module is set between the encoder and the decoder. Among them, the bidirectional attention module includes a self-attention mechanism for calculating the internal spatial correlation degree of feature points and a cross-attention mechanism for calculating the topological correlation degree between different feature points, and fuses the feature mapping results of the self-attention mechanism and the cross-attention mechanism to obtain a hierarchical feature descriptor;
[0086] Screen the hierarchical feature descriptor through a random sampling algorithm based on local geometric consistency constraints, determine candidate feature point pairs through feature similarity calculation, randomly select a feature point pair as a seed match, calculate an initial geometric transformation relationship including a rotation matrix and a translation vector based on the seed match, perform a local geometric consistency constraint test on the remaining feature point pairs to obtain a projection error, eliminate the feature point pairs with a projection error greater than a preset projection threshold, repeat the random sampling process multiple times to select the geometric transformation relationship with the most inlier support, and unify the multi-source remote sensing data to the global geographic coordinates based on the selected geometric transformation relationship to generate registered multi-source remote sensing data;
[0087] Input the registered multi-source remote sensing data into a variational autoencoder network. The encoder of the variational autoencoder network maps the registered multi-source remote sensing data to a latent feature space through multi-layer residual convolution, and the decoder reconstructs the latent features into the form of original data through deconvolution and skip connections. Random noise perturbations are introduced into the latent feature space to construct a noise-aware loss function to constrain the reconstruction results;
[0088] Construct a generative adversarial network based on pixel-level adversarial discrimination. The generative adversarial network uses a densely connected fully convolutional structure to output a pixel-level authenticity score map. Use the decoder of the variational autoencoder network as the generator, construct an adversarial loss function, and optimize the parameters of the generator and the generative adversarial network through a minimax game until reaching a Nash equilibrium to obtain enhanced remote sensing data.
[0089] The octree spatial index structure is a hierarchical data structure for three-dimensional space data management. It recursively divides the three-dimensional space into eight sub-regions, which helps to efficiently store, index, and query large-scale three-dimensional point cloud data and is commonly used in fields such as computer vision, geographic information systems, and robot navigation. The curvature threshold segmentation is an image or point cloud segmentation method based on geometric characteristics. By calculating curvature information and setting thresholds, regions with different curvature characteristics are distinguished, and it is applicable to tasks such as object surface analysis, medical image processing, and point cloud segmentation. The region growing algorithm is a segmentation method based on pixel or point neighborhood relationships. By selecting seed points and expanding to adjacent regions that meet specific similarity criteria, connected regions are formed, and it is widely used in image segmentation, point cloud clustering, and medical image analysis. The seed matching is a feature matching strategy. By selecting a set of highly confident matching points (seed points) in the data, and then using these matching points to guide the matching process of the entire dataset, the accuracy and stability of the matching are improved, and it is commonly used in three-dimensional reconstruction and target tracking. Inliers refer to data points that are considered to conform to the transformation model and contribute to the final matching result during feature matching or point cloud registration. Compared with the mis-matched outliers, inliers usually have a higher matching confidence.
[0090] Obtain multi-source remote sensing data corresponding to the target news scene through remote sensing satellites, including optical remote sensing image data, dual-polarization radar remote sensing data, and multi-echo lidar point cloud data. For example, it is possible to obtain the surface optical image, X-band dual-polarization radar data, and multi-echo lidar point cloud data of a certain area.
[0091] Input the obtained multi-source remote sensing data into the Iterative Closest Point algorithm based on spatial position weighting for preliminary registration. Construct a spatial position weight matrix, where the weight value is calculated according to the distance of the data point to the scene center, and the closer the distance, the greater the weight. Use the octree spatial index structure to perform nearest neighbor search for each data point and establish a set of nearest neighbor points. Calculate the position deviation vector between the data point and the set of nearest neighbor points, and multiply this vector by the spatial position weight matrix to obtain the weighted position deviation. Iteratively calculate the weighted position deviation until it is less than a preset threshold (e.g., 0.5 pixels), and output the optimized set of data points.
[0092] Input the multi-source remote sensing data into the tensor projection algorithm for feature extraction. Construct a feature covariance matrix through principal component analysis, and calculate the eigenvalues and eigenvectors to determine the optimal projection direction, and perform the tensor projection operation in the optimal projection direction. Construct a local curvature calculation function and a normal vector change function for the projection result. Extract the multi-modal data feature points by setting a curvature threshold (e.g., 0.01) and the region growing algorithm.
[0093] Input the extracted multi-modal data feature points into the bidirectional attention deep neural network. The encoder of this network extracts multi-scale features through multi-layer convolution and pooling operations (such as point convolution, edge convolution, and face convolution). The decoder restores the features through deconvolution and upsampling operations. Set a bidirectional attention module between the encoder and the decoder, and this module contains a self-attention mechanism and a cross-attention mechanism. The self-attention mechanism calculates the internal spatial correlation degree of the feature points, and the cross-attention mechanism calculates the topological correlation degree between different feature points. Fuse the feature mapping results of the two attention mechanisms to obtain a hierarchical feature descriptor.
[0094] Use a random sampling algorithm based on local geometric consistency constraints to screen the hierarchical feature descriptor. First, determine candidate feature point pairs through feature similarity calculation. Then randomly select a feature point pair as a seed match, and calculate the initial geometric transformation relationship including the rotation matrix and the translation vector based on the seed match. Perform a local geometric consistency constraint test on the remaining feature point pairs to obtain the projection error. Eliminate the feature point pairs with a projection error greater than the preset threshold (e.g., 1 pixel). Repeat the random sampling process multiple times and select the geometric transformation relationship with the most inlier support. Unify the multi-source remote sensing data to the global geographic coordinates based on the selected geometric transformation relationship to generate registered multi-source remote sensing data.
[0095] Input the registered multi-source remote sensing data into the variational autoencoder network for data augmentation. The encoder of the variational autoencoder maps the data to the latent feature space through multi-layer residual convolution, and the decoder reconstructs the latent features into the original data form through deconvolution and skip connections. Introduce random noise perturbation in the latent feature space and construct a noise-aware loss function to constrain the reconstruction result.
[0096] Construct a generative adversarial network based on pixel-level adversarial discrimination, and use a densely connected fully convolutional structure to output a pixel-level authenticity score map. Use the decoder of the variational autoencoder as the generator, and construct an adversarial loss function. Optimize the parameters of the generator and discriminator through a min-max game until reaching the Nash equilibrium to obtain enhanced remote sensing data.
[0097] In this embodiment, by fusing multi-source remote sensing data and deep learning technologies, feature points can be extracted more accurately and matching relationships can be established, thereby improving the accuracy of remote sensing data registration. Through the variational autoencoder and generative adversarial network, noise can be effectively removed, missing information can be filled, and details can be enhanced, thereby improving the quality of remote sensing data. By fusing multi-source remote sensing data and deep learning technologies, scene information can be obtained more comprehensively, thereby enhancing the ability to understand the scene.
[0098] In an alternative embodiment,
[0099] Screen the hierarchical feature descriptors through a random sampling algorithm based on local geometric consistency constraints, determine candidate feature point pairs through feature similarity calculation, randomly select a feature point pair as a seed match, calculate an initial geometric transformation relationship including a rotation matrix and a translation vector based on the seed match, perform a local geometric consistency constraint test on the remaining feature point pairs to obtain a projection error, and eliminate the feature point pairs with a projection error greater than a preset projection threshold. Repeat the random sampling process multiple times to select the geometric transformation relationship with the most inlier support, and unify the multi-source remote sensing data to the global geographic coordinates based on the selected geometric transformation relationship. The generated registered multi-source remote sensing data includes:
[0100] Construct a local geometric feature extraction unit and a global semantic feature extraction unit. Among them, the local geometric feature extraction unit extracts a local shape descriptor by constructing a feature point neighborhood and analyzing the distribution characteristics of points in the neighborhood. The global semantic feature extraction unit extracts hierarchical semantic information from bottom to top through a multi-scale feature pyramid, and cascades the local shape descriptor and the hierarchical semantic information to form a feature vector;
[0101] Construct a feature similarity calculation module. The feature similarity calculation module is based on the cosine similarity metric method, introduces a spatial distance weighting term to make feature points with closer distances have a higher matching priority, and pairs the feature descriptors with a similarity higher than a preset threshold to form a set of candidate feature point pairs;
[0102] Randomly select feature point pairs from the set of candidate feature point pairs as seed matches, construct a geometric transformation estimation module based on the seed matches, establish a mapping relationship between corresponding points through the geometric transformation estimation module to construct a three-dimensional rigid body transformation model, decompose the rotation matrix in the transformation process into three basic rotation matrices and solve it by combining an iterative optimization algorithm, and solve the translation vector by minimizing the Euclidean distance between corresponding points;
[0103] Construct a geometric consistency verification module. The geometric consistency verification module includes a feature transformation unit and an error calculation unit. The feature transformation unit transforms the source feature points to the target coordinate system through the three-dimensional rigid body transformation model. The error calculation unit calculates the projection error between the transformed position and the actual target position, dynamically adjusts the projection error threshold according to the data distribution characteristics, and eliminates the feature point pairs with projection errors greater than the projection error threshold;
[0104] Construct a random sampling iterative optimization module. The random sampling iterative optimization module randomly selects different seed matches to execute the geometric transformation estimation module and the geometric consistency verification module, counts the number of inliers that meet the constraint conditions, scores by comprehensively considering the number of inliers and the stability of the geometric transformation, and selects the geometric transformation relationship with the highest score as the optimal geometric transformation relationship;
[0105] Construct a coordinate transformation module based on the optimal geometric transformation relationship. Based on the coordinate transformation module, use the optical remote sensing image data as the reference data, construct an affine transformation model considering the imaging geometric characteristics for the dual-polarization radar remote sensing data, solve the transformation parameters through polynomial fitting to achieve geometric correction, construct a three-dimensional space rigid body transformation model for the multi-echo laser point cloud data, solve the transformation parameters through iterative optimization, and uniformly transform the optical remote sensing image data, the dual-polarization radar remote sensing data, and the multi-echo laser point cloud data to the global geographic coordinate system.
[0106] The three-dimensional rigid body transformation model is a mathematical model used to describe the rigid motion of three-dimensional objects, including translation and rotation transformations, but not involving scaling or deformation, and is widely used in the fields of computer vision, robot navigation, and medical imaging. The iterative optimization algorithm is a class of optimization methods that continuously update parameters and gradually approach the optimal solution, and is commonly used to solve non-linear problems. The affine transformation model considering the imaging geometric characteristics is a mathematical model used to represent spatial transformations in image processing and computer vision, which can describe geometric transformations such as rotation, scaling, translation, and shear, and is particularly suitable for remote sensing image correction and image registration tasks.
[0107] Construct a feature extraction unit, which includes a local geometric feature extraction part and a global semantic feature extraction part. The local geometric feature extraction part determines a feature point, selects a certain number of neighboring points around it, and analyzes the distribution of these neighboring points, such as the density, direction of the points, and the distance relationship between them, so as to extract features describing the local shape, such as divergence, curvature, histogram of oriented gradients, etc. The global semantic feature extraction part uses a multi-scale feature pyramid to gradually extract semantic information from the low-resolution levels of the image. Features at each scale are extracted, such as texture, color, edges, etc. Finally, the features of each layer are combined into a hierarchical semantic information descriptor. Finally, the local shape descriptor and the hierarchical semantic information descriptor are connected together to form a complete feature vector. For example, for a building corner, the local geometric features can describe the sharpness of the corner, and the global semantic features can describe the semantic category that the corner belongs to, i.e., the building.
[0108] Construct a feature similarity calculation module, which uses cosine similarity to measure the similarity between two feature vectors. In order to make feature points closer in space have a higher matching priority, a spatial distance weighting term is introduced. Calculate the spatial distance between two feature points and convert it into a weight value. The closer the distance, the greater the weight. Then multiply this weight value by the cosine similarity to obtain the final similarity score. Pair the feature descriptors with similarity scores higher than a preset threshold (e.g., 0.8) to form a set of candidate feature point pairs. For example, if the feature vectors of two building corners are highly similar and their spatial distance is close, they will be paired into a candidate feature point pair.
[0109] Randomly select a part of the feature point pairs from the set of candidate feature point pairs as seed matches. Based on these seed matches, construct a geometric transformation estimation module, establish a mapping relationship between corresponding points, and construct a three-dimensional rigid body transformation model, which includes two parts: rotation and translation. Through an iterative optimization algorithm, decompose the rotation matrix into three basic rotation matrices and solve the parameters of these three rotation matrices respectively. At the same time, solve the translation vector by minimizing the Euclidean distance between corresponding points. For example, assume that three seed match point pairs are selected. By minimizing the distance between the seed match point pairs, a rotation matrix and a translation vector can be estimated to describe the geometric transformation relationship between two data sets.
[0110] Construct a geometric consistency checking module to transform the feature points in the source data to the coordinate system of the target data through the estimated three-dimensional rigid body transformation model. Calculate the projection error between the transformed position and the actual target position. Dynamically adjust the projection error threshold according to the data distribution characteristics, for example, use statistical methods to calculate the distribution of the errors and set the threshold based on the distribution. Eliminate the feature point pairs with projection errors greater than the threshold. For example, if the distance between the position of a feature point after transformation and the actual target position exceeds the set threshold, it is considered that this point pair does not meet the geometric consistency and is eliminated.
[0111] Construct a random sample consensus (RANSAC) optimization module to repeat the above processes of random sampling, geometric transformation estimation, and geometric consistency checking multiple times. Each iteration will select different seed matches and calculate the number of inliers that meet the constraint conditions, that is, the number of feature point pairs with projection errors less than the threshold. Score by comprehensively considering the number of inliers and the stability of the geometric transformation. For example, use the proportion of the number of inliers to the total number of feature points as the scoring criterion. Select the geometric transformation relationship with the highest score as the optimal geometric transformation relationship.
[0112] Based on the obtained optimal geometric transformation relationship, construct a coordinate transformation module. Take the optical remote sensing image data as the reference data, construct an affine transformation model considering the imaging geometric characteristics for the dual-polarization radar remote sensing data, and solve the transformation parameters through polynomial fitting method to achieve geometric correction. Construct a three-dimensional space rigid body transformation model for the multi-echo laser point cloud data and solve the transformation parameters through iterative optimization method. Finally, uniformly transform the optical remote sensing image data, dual-polarization radar remote sensing data, and multi-echo laser point cloud data to the global geographic coordinate system.
[0113] In this embodiment, by combining local geometric features and global semantic features, and introducing spatial distance weighting and geometric consistency constraints, false matches can be effectively eliminated, the registration accuracy can be improved. The random sampling and iterative optimization strategy can effectively handle problems such as noise and data loss, enhance the robustness of the algorithm, and can automatically complete the registration of multi-source remote sensing data without manual intervention, improving the degree of automation.
[0114] S2. Add the enhanced remote sensing data to the feature extraction network to obtain a spatial attention feature map and a channel attention feature map. Add the spatial attention feature map to the monocular depth estimation network based on dense connections to generate depth information, fuse it with the dual-polarization radar remote sensing data and the multi-echo laser point cloud data to obtain a sparse three-dimensional point cloud, generate an initial three-dimensional mesh model through the Poisson reconstruction algorithm. Add the channel attention feature map to the residual graph convolutional neural network for texture feature transfer, and perform topology optimization through the multi-view projection consistency constraint and the mesh optimization algorithm to obtain an optimized three-dimensional scene model;
[0115] The monocular depth estimation network based on dense connection is a neural network that uses the dense connection strategy to improve the accuracy of monocular image depth estimation. By enhancing feature transmission and gradient flow, it alleviates the problem of gradient disappearance in depth networks, improves the robustness of depth prediction and the ability to restore details. The Poisson reconstruction algorithm is a three-dimensional reconstruction method based on the Poisson equation. By solving the Poisson equation with the normal information of the input point cloud, a smooth and coherent three-dimensional surface is constructed, which is widely used in fields such as computer vision, reverse engineering, and medical image processing. The texture feature transfer is a method used to preserve or enhance texture information in image or three-dimensional reconstruction tasks. By transferring texture features between different domains or scales, the target image or model is improved in terms of style consistency and detail richness. The multi-view projection consistency constraint is an optimization strategy for three-dimensional reconstruction and depth estimation. By ensuring the consistency of depth estimation or geometric features in different view projections, the accuracy and visual coherence of the reconstruction results are improved, and it is applicable to multi-view stereo matching and structured light scanning. The mesh optimization algorithm is a class of methods for optimizing three-dimensional mesh structures, including mesh simplification, topology optimization, and adaptive resampling, to improve the quality of three-dimensional models, reduce redundant data, and optimize rendering efficiency, and is widely used in computer graphics and simulation.
[0116] In an alternative embodiment,
[0117] Adding enhanced remote sensing data to the feature extraction network to obtain a spatial attention feature map and a channel attention feature map, adding the spatial attention feature map to the monocular depth estimation network based on dense connection to generate depth information, fusing it with dual-polarization radar remote sensing data and multi-echo laser point cloud data to obtain sparse three-dimensional point clouds, generating an initial three-dimensional mesh model through the Poisson reconstruction algorithm, adding the channel attention feature map to the residual graph convolutional neural network for texture feature transfer, and combining the multi-view projection consistency constraint and the mesh optimization algorithm for topology optimization to obtain an optimized three-dimensional scene model, including:
[0118] Input the enhanced remote sensing data into the feature extraction network, extract multi-scale spatial features through a multi-level cascaded convolutional structure and a batch normalization layer, use an adaptive weight calculation unit to assign spatial attention weights to each position, perform information aggregation through global pooling to obtain a channel descriptor, perform multi-level transformation on the channel descriptor to learn the dependencies between channels to generate channel attention weights, and generate a spatial attention feature map and a channel attention feature map through adaptive fusion of the spatial attention weights and the channel attention weights;
[0119] Input the spatial attention feature map into a monocular depth estimation network based on dense connections. In the monocular depth estimation network, establish multi-level cascade connections in a skip manner, establish a direct mapping relationship between shallow features and deep features, gradually restore the feature resolution through a multi-layer deconvolution structure, and use a multi-scale feature pyramid structure to adaptively fuse the weights of features at different levels to generate depth information;
[0120] Sample the depth information and establish a mapping relationship from pixel coordinates to three-dimensional space points. Calculate the confidence weights for pre-acquired dual-polarization radar remote sensing data and multi-echo laser point cloud data based on the signal-to-noise ratio and spatial resolution. Construct an adaptive weighted fusion strategy according to the spatial distance and feature similarity, and perform point cloud fusion on the overlapping area to generate a sparse three-dimensional point cloud;
[0121] Input the sparse three-dimensional point cloud into the Poisson reconstruction algorithm. Calculate the normal vector of each point through the eigen-decomposition method, dynamically adjust the spatial dissection granularity according to the point cloud density distribution, obtain the implicit surface expression by iteratively solving the Poisson equation, combine the adaptive threshold to extract the grid structure and perform Laplacian smoothing to generate an initial three-dimensional grid model. Input the channel attention feature map into the residual graph convolutional neural network, transfer node information through local neighborhood aggregation, add the input feature and the transformed feature to construct a residual connection, adaptively capture texture features through the graph attention mechanism, and use projective transformation to map the texture features to the surface of the initial three-dimensional grid model for texture feature transfer;
[0122] Uniformly sample multiple viewing perspectives in the spherical space, project the initial three-dimensional grid model onto each viewing perspective plane to generate multi-view projection images, calculate the similarity between the multi-view projection images and the original images through feature descriptor matching to establish a multi-view projection consistency constraint, and optimize the grid vertex positions and connection relationships based on the multi-view projection consistency constraint while maintaining the geometric details of the model to generate an optimized three-dimensional scene model.
[0123] The multi-layer transposed convolution structure is a deep learning architecture for image generation and feature reconstruction. It gradually restores low-resolution features to high-resolution through multiple levels of transposed convolution (or inverse convolution) layers, thereby generating more refined outputs, such as tasks like depth estimation, image super-resolution, and semantic segmentation. The point cloud density distribution refers to the distribution characteristics of point cloud data in three-dimensional space, which affects processes such as three-dimensional reconstruction, point cloud sampling, and filtering. The performance of point cloud processing algorithms can be optimized by estimating the local point cloud density. The implicit surface representation is a mathematical model for representing continuous three-dimensional surfaces. It defines an isosurface in space through an implicit function and can be used for smooth surface reconstruction, shape modeling, and surface fitting in computer graphics. Laplacian smoothing is a geometric optimization method based on the Laplacian operator. It iteratively adjusts vertex positions to reduce noise and optimize surface smoothness, and is commonly used for point cloud denoising, mesh optimization, and shape reconstruction.
[0124] Input the enhanced remote sensing data (for example, high-resolution satellite images after atmospheric correction, geometric correction, and image fusion processing, such as WorldView-4 image data, with a size of 1024x1024 pixels and including red, green, blue, and near-infrared bands) into the feature extraction network. Use a multi-level cascaded convolution structure and batch normalization layers to extract multi-scale spatial features. For example, use four convolutional layers with kernel sizes of 7x7, 5x5, 3x3, and 3x3 respectively. After each layer of convolution, perform batch normalization operations. Use an adaptive weight calculation unit to assign spatial attention weights to each position. For example, assign weights by calculating the similarity between the features at each position and the global average feature, and perform information aggregation through global pooling to obtain channel descriptors. Perform multi-layer transformations (for example, use two fully connected layers) on the channel descriptors to learn the dependencies between channels and generate channel attention weights. Adaptively fuse the spatial attention weights and channel attention weights (for example, through weighted summation) to generate a spatial attention feature map and a channel attention feature map, both with a size of 128x128.
[0125] Input the spatial attention feature map into a monocular depth estimation network based on dense connections. Establish multi-level cascaded connections in a skip manner to establish a direct mapping relationship between shallow features and deep features. For example, connect the output of the first convolutional layer to the output of the third convolutional layer. The network gradually restores the feature resolution through a multi-layer transposed convolution structure. For example, use two transposed convolution layers to restore the feature map size from 32x32 to 128x128. Use a multi-scale feature pyramid structure to adaptively fuse the weights of features at different levels to generate depth information. For example, perform weighted summation on the outputs of different transposed convolution layers to obtain the final depth information map with a size of 128x128.
[0126] Sample the depth information and establish the mapping relationship from pixel coordinates to three-dimensional space points. For example, convert pixel coordinates to three-dimensional coordinates according to the intrinsic matrix and extrinsic matrix of the camera. Calculate the confidence weight based on the signal-to-noise ratio and spatial resolution for pre-acquired dual-polarization radar remote sensing data (e.g., TerraSAR-X data, with the same coverage as satellite images) and multi-echo laser point cloud data (e.g., ALS data, with a point cloud density of 10 points per square meter). For example, the higher the signal-to-noise ratio and the higher the spatial resolution, the greater the weight. Construct an adaptive weighted fusion strategy based on spatial distance and feature similarity. For example, the closer the distance and the more similar the features, the greater the weight. Perform point cloud fusion on the overlapping area to generate a sparse three-dimensional point cloud. For example, perform weighted averaging on the point cloud data from different data sources according to the weight to obtain the fused sparse three-dimensional point cloud, which contains 10,000 points.
[0127] Input the sparse three-dimensional point cloud into the Poisson reconstruction algorithm. Calculate the normal vector of each point through the eigen decomposition method. For example, use the principal component analysis method to calculate the normal vector. Dynamically adjust the spatial subdivision granularity according to the point cloud density distribution. For example, the higher the density, the finer the subdivision granularity. Obtain the implicit surface representation by iteratively solving the Poisson equation. Combine the adaptive threshold to extract the grid structure and perform Laplacian smoothing to generate the initial three-dimensional grid model, which contains 5,000 triangular patches.
[0128] Input the channel attention feature map into the residual graph convolutional neural network. Transmit node information through local neighborhood aggregation. Add the input feature and the transformed feature to construct a residual connection. Adaptively capture texture features through the graph attention mechanism. Use projection transformation to map the texture features to the surface of the initial three-dimensional grid model for texture feature transfer.
[0129] Uniformly sample multiple viewing perspectives in the spherical space (e.g., sample 8 perspectives). Project the initial three-dimensional grid model onto each viewing perspective plane to generate multi-view projection images. Calculate the similarity between the multi-view projection images and the original images through feature descriptor matching (e.g., use SIFT feature descriptors) to establish the multi-view projection consistency constraint. Optimize the grid vertex positions and connection relationships based on the multi-view projection consistency constraint, while maintaining the geometric details of the model to generate an optimized three-dimensional scene model.
[0130] In this embodiment, by fusing multi-source remote sensing data, the accuracy of three-dimensional scene modeling can be effectively improved. Especially in complex scenes, through the attention mechanism and the residual graph convolutional neural network, the detail expression ability of the model can be effectively enhanced, making the generated model more realistic. Through adaptive weight calculation and multi-scale feature fusion, the modeling efficiency can be effectively improved and the calculation time can be reduced.
[0131] In an alternative embodiment,
[0132] Input the sparse three-dimensional point cloud into the Poisson reconstruction algorithm, calculate the normal vector of each point through the eigen - decomposition method, dynamically adjust the spatial subdivision granularity according to the point cloud density distribution, obtain the implicit surface representation by iteratively solving the Poisson equation, extract the mesh structure by combining an adaptive threshold and perform Laplacian smoothing to generate an initial three - dimensional mesh model. Input the channel attention feature map into the residual graph convolutional neural network, transfer node information through local neighborhood aggregation, add the input feature and the transformed feature to construct a residual connection, adaptively capture texture features through the graph attention mechanism, and use projective transformation to map the texture features to the surface of the initial three - dimensional mesh model for texture feature transfer, including:
[0133] Input the sparse three - dimensional point cloud into the Poisson reconstruction algorithm, construct a point cloud nearest - neighbor search tree structure, use the spherical neighborhood search method to determine the local neighborhood point set corresponding to each point in the sparse three - dimensional point cloud, construct a covariance matrix based on the local neighborhood point set, calculate the normal vector of each point through the eigen - decomposition method, construct a minimum spanning tree to establish the connection relationship between points, select the normal vector at the center position of the point cloud as the reference direction, and recursively traverse the minimum spanning tree to adjust the pointing directions of all points' normal vectors to obtain a set of normal vectors;
[0134] Calculate the outer bounding box of the point cloud as the initial spatial range, dynamically adjust the spatial subdivision granularity according to the point cloud density distribution, count the number of point clouds contained in each octree node and calculate the point density distribution, set an adaptive subdivision threshold based on the point density, divide the nodes with point density greater than the adaptive subdivision threshold into eight child nodes and record the depth information and the contained point set of each leaf node, perform a merging operation on the divided leaf nodes, and merge adjacent nodes with similar point densities to obtain an optimized octree structure;
[0135] Define a three - dimensional scalar field function on the optimized octree structure, convert the set of normal vectors into a gradient field constraint, obtain the implicit surface representation by iteratively solving the Poisson equation, construct a solution hierarchy using the multigrid strategy, convert the Poisson equation into a linear system of equations by constructing a sparse matrix to represent the Laplacian operator, and use an iterative optimization method to solve the linear system of equations and update the scalar values of the grid nodes to obtain a three - dimensional scalar field;
[0136] Extract the mesh structure by combining an adaptive threshold, generate an initial mesh by extracting the isosurface in the optimized octree structure based on the marching cubes algorithm, perform a topological check on the initial mesh and repair non - manifold edges and singular points, perform Laplacian smoothing to iteratively optimize the vertex positions, and simplify the mesh based on the edge collapse criterion to obtain an initial three - dimensional mesh model;
[0137] Input the channel attention feature map into the residual graph convolutional neural network, transfer node information through local neighborhood aggregation, define an adjacency matrix to describe the connection relationship between nodes, perform message aggregation based on the adjacency matrix to weight - combine the node features and neighborhood features, and add the input feature and the transformed feature to construct a residual connection to obtain an enhanced feature;
[0138] Adaptively capture texture features through the graph attention mechanism, obtain feature statistical information through global pooling operation on the enhanced feature and perform non - linear transformation, learn the correlation weights between different channels to generate channel attention weights, calculate the importance scores at different positions of the feature map to generate spatial attention weights, and apply the channel attention weights and the spatial attention weights to the feature map to obtain texture features;
[0139] Use projection transformation to map the texture features to the surface of the initial three - dimensional mesh model for texture feature transfer, establish the mapping relationship between the vertices of the initial three - dimensional mesh model and the image pixels, project the vertices to the image plane according to the camera projection matrix, calculate the texture coordinates by barycentric coordinate interpolation inside the mesh patches, and construct multi - view Figure 1 Consistency constraints evaluate the texture mapping quality under different perspectives, select the optimal texture source based on the graph cut optimization algorithm, and dynamically adjust the texture fineness according to the viewing distance to obtain a textured three - dimensional scene model.
[0140] The point - cloud nearest - neighbor search tree structure is a spatial indexing method for accelerating point - cloud data queries, which can be used for efficient nearest - neighbor search, point - cloud registration, and three - dimensional feature matching. The local neighborhood point set refers to a group of spatially adjacent points relative to a certain target point in three - dimensional point - cloud data, which is usually used for calculating local geometric features, normal estimation, and point - cloud filtering. The barycentric coordinate interpolation is an interpolation method based on triangular or polygonal meshes. By calculating the barycentric coordinate weights of a point at its neighborhood vertices, smooth interpolation is achieved, and it is widely used in computer graphics, finite - element analysis, and three - dimensional deformation modeling.
[0141] Preprocess the input sparse three-dimensional point cloud. Construct a tree structure, such as a KD tree or an octree, that can quickly search for neighboring points. With each point as the center, use spherical neighborhood search to find all neighboring points within a certain radius around the point, forming the local neighborhood point set of the point. Using these neighboring points, calculate the covariance matrix of the local area where the point is located. Perform eigenvalue decomposition on the covariance matrix, and the eigenvector corresponding to the smallest eigenvalue is the normal vector of the point. To ensure the consistency of the normal vector direction, construct the minimum spanning tree of the point cloud and connect the points. Select the normal vector of the center point of the point cloud as the reference direction, and along the minimum spanning tree, sequentially adjust the normal vector direction of each point to make it consistent with the reference direction, finally obtaining the set of normal vectors of all points. For example, assume that the local neighborhood point set of point A contains points B, C, and D. After eigenvalue decomposition of the calculated covariance matrix, the normal vector n is obtained.
[0142] Calculate the bounding box of the point cloud and use it as the initial spatial range. Dynamically adjust the granularity of spatial partitioning according to the density distribution of the point cloud. Divide the space into multiple small cube units, count the number of point clouds contained in each unit, and calculate the point density. Set an adaptive partitioning threshold, such as the average point density. For the units with point density higher than the threshold, continue to divide them into eight smaller sub-units, and record the depth information of each leaf node and the contained point set. To optimize the partitioning result, merge adjacent leaf nodes with similar point densities, finally obtaining an optimized octree structure. Assume that the range of the point cloud bounding box is from (0, 0, 0) to (100, 100, 100), the point density threshold is set to 10, and a unit contains 20 points, then it will be further divided.
[0143] Define a three-dimensional scalar field function on the optimized octree structure. Convert the previously calculated set of normal vectors into a gradient field constraint. Obtain the implicit surface representation by iteratively solving the Poisson equation. Adopt a multigrid strategy to construct the solution hierarchy. Convert the Poisson equation into a linear system of equations and use an iterative optimization method to solve the system of equations to update the scalar values of the grid nodes, finally obtaining the three-dimensional scalar field. For example, assume that the initial scalar value of a certain grid node is 0, and after one iteration, it is updated to 0.5.
[0144] Combined with an adaptive threshold, use the marching cubes algorithm to extract the isosurface in the optimized octree structure and generate the initial mesh. Conduct a topological check on the initial mesh to repair non-manifold edges and singular points. Use the Laplacian smoothing algorithm to iteratively optimize the vertex positions to make the mesh smoother. Simplify the mesh according to the edge collapse criterion, remove redundant triangular patches, and finally obtain the initial three-dimensional mesh model. For example, the three vertex coordinates of a triangular patch are (0, 0, 0), (1, 0, 0), and (0, 1, 0) respectively. After smoothing, they may become (0.1, 0.1, 0), (0.9, 0.1, 0), and (0.1, 0.9, 0).
[0145] Input the channel attention feature map into the residual graph convolutional neural network. Transmit node information through local neighborhood aggregation. Define the adjacency matrix to describe the connection relationship between nodes. Based on the adjacency matrix, perform message aggregation to weight-combine the node features with their neighborhood features. Add the input feature and the transformed feature to construct a residual connection to obtain the enhanced feature. Adaptively capture texture features through the graph attention mechanism. Conduct a global pooling operation on the enhanced feature to obtain feature statistical information and perform a non-linear transformation. Learn the correlation weights between different channels to generate the channel attention weights. Calculate the importance scores at different positions of the feature map to generate the spatial attention weights. Apply the channel attention weights and the spatial attention weights to the feature map to obtain the texture features. Use projective transformation to map the texture features onto the surface of the initial three-dimensional mesh model. Establish the mapping relationship between the vertices of the initial three-dimensional mesh model and the image pixels. Project the vertices onto the image plane according to the camera projection matrix. For the interior of the mesh patch, calculate the texture coordinates using barycentric coordinate interpolation. Construct multi-view Figure 1 consistency constraints to evaluate the texture mapping quality from different perspectives. Based on the graph cut optimization algorithm, select the optimal texture source. Dynamically adjust the texture fineness according to the viewing distance, and finally obtain the three-dimensional scene model with textures.
[0146] Figure 3This is a comparison chart of the simulation results of 3D model reconstruction for the news scene 3D reconstruction and visualization method based on multi-source remote sensing data in the embodiments of the present invention, showing the comparison of the simulation results of 3D model reconstruction based on the same input point cloud (with a density of 235 points / m²). The upper left figure is the original point cloud data, where it can be seen that the point cloud distribution is uneven and there is a certain amount of noise. The upper right figure is the reconstruction result of the technical solution of the present invention. The model surface is smooth and natural, the edge features are clear, the reconstruction accuracy reaches 0.93, and it only takes 18.5 seconds to complete the reconstruction. The lower left figure is the result of traditional Poisson reconstruction. Although the overall shape is correct, there are some uneven areas on the surface, the accuracy is 0.85, and it takes 29.7 seconds. The lower right figure is the result of the Marching Cubes method. Some edge details are lost, the surface smoothness is insufficient, and the accuracy is only 0.79, taking 24.3 seconds. It can be clearly seen from the reconstruction effects that the technical solution of the present invention not only retains the geometric details of complex structures but also achieves a higher surface smoothness. Especially at the edges and corners of objects, the technical solution of the present invention can more accurately restore the shape features, while other methods have varying degrees of detail loss or deformation.
[0147] In this embodiment, a high-precision 3D model is reconstructed from sparse point cloud data and rich geometric details are retained. Through adaptive spatial partitioning and Poisson reconstruction, the surface shape of the point cloud can be effectively captured to generate a smooth and realistic 3D mesh. By using the graph attention mechanism and residual graph convolutional neural network, texture features can be effectively extracted and transferred, enabling the reconstructed 3D model to have a realistic texture effect. Figure 1 The multi-view consistency constraint and adaptive texture fineness adjustment further improve the texture fitting degree and visual quality. By normal vector adjustment and topology repair, the noise and missing in the point cloud data can be effectively processed to generate a high-quality 3D model.
[0148] S3. Calculate the curvature change value and normal vector change value of each grid in the optimized 3D scene model, construct a spatial partitioning evaluation function to divide the 3D scene into octree spatial blocks of different sizes, calculate the quadratic error matrix of the grid vertices in each spatial block and construct an edge collapse cost function to merge and simplify the grid vertices to obtain four precision-level models. Construct a quadtree frustum to divide it into three depth regions to load the corresponding precision-level models to obtain a visualization scene. Extract local features and global features through a dual-branch semantic segmentation network, generate a scene semantic segmentation map through fusion upsampling and label the news event positions, collect interactive position information and calculate the view point transformation matrix to update the visualization scene for scene roaming and perform scene roaming.
[0149] The octree spatial block is a spatial data structure that organizes and stores data by recursively dividing a three-dimensional space into eight sub-regions (i.e., an octree), enabling efficient querying, accessing, and processing of point clouds, volume data, etc. in the three-dimensional space. The edge collapse cost function is an optimization function used in image or geometric mesh processing to measure the cost of edge collapse operations, usually considering the geometric accuracy, smoothness, and structural preservation degree after deformation to optimize processes such as mesh simplification, triangle subdivision, or surface reconstruction. The quadratic error matrix refers to a matrix obtained by calculating the sum of squared errors and is commonly used in optimization algorithms such as the least squares method, aiming to minimize the error of the objective function and is usually applied in fields such as image processing, data fitting, and regression analysis in machine learning. The accuracy level model is a method model for measuring and controlling calculation or measurement accuracy by defining multiple accuracy levels and standards. The quadtree frustum is a quadtree data structure for spatial partitioning and view management that recursively divides a two-dimensional plane into four regions. The dual-branch semantic segmentation network is a semantic segmentation model usually composed of two parallel branch structures that process different features or different scales of an image respectively to perform image segmentation tasks through a deep learning network. The view point transformation matrix is a mathematical matrix used to describe view point transformation that adjusts the viewing angle, projection direction, and camera parameters through coordinate transformation.
[0150] In an alternative embodiment,
[0151] Calculate the curvature change value and normal vector change value of each mesh in the optimized three-dimensional scene model, construct a spatial partitioning evaluation function to divide the three-dimensional scene into octree spatial blocks of different sizes, calculate the quadratic error matrix of the mesh vertices within each spatial block and construct an edge collapse cost function to merge and simplify the mesh vertices to obtain four accuracy level models, construct a quadtree frustum divided into three depth regions to load the corresponding accuracy level models to obtain a visualization scene, extract local features and global features through a dual-branch semantic segmentation network, generate a scene semantic segmentation map through fusion and upsampling and mark the news event locations, collect interactive location information and calculate the view point transformation matrix to update the visualization scene for scene roaming, and the scene roaming includes:
[0152] For each mesh patch in the optimized three-dimensional scene model, project the normal vector of each mesh patch onto the vertex local coordinate system and calculate the weighted average to obtain the vertex normal vector, construct a local surface fitting equation of the vertex to calculate the principal curvature to obtain the curvature change value, traverse adjacent vertices to calculate the normal vector angle change rate, combine the curvature change value and the normal vector angle change rate to form a vertex feature description, and establish an adjacency relationship graph recording the vertex feature description;
[0153] Input the vertex feature description, the vertex density distribution in the adjacency relationship graph, and the distribution ratio of feature points into the spatial division evaluation function. Recursively construct an octree with the scene bounding box as the initial space, calculate the evaluation function value of each node to be segmented, dynamically adjust the segmentation threshold based on the level of the node and the evaluation value of the parent node, and record the spatial range of each leaf node, the vertex feature description, and the mesh topology relationship to form a spatial index structure;
[0154] For each leaf node in the spatial index structure, construct a quadratic error matrix based on the sum of the squares of the distances from the vertices to the adjacent patches. Combine the quadratic error matrix, the normal vector change rate, and the texture coordinate change amount to construct an edge collapse cost function. Traverse the mesh edges within the node to calculate the optimal collapse position and store it in a priority queue. Iteratively perform the edge collapse operation according to the magnitude of the value of the collapse cost function, and update the error matrix of the affected vertices and the cost values of the relevant edges;
[0155] Control the number of collapse iterations based on the edge collapse cost function to maintain the continuity of the vertex feature description. Generate four precision-level models that retain 90%, 70%, 50%, and 30% of the original vertex numbers. Construct a quadtree structure corresponding to the frustum, divide the frustum into a near-view region, a middle-view region, and a far-view region, set an overlapping transition zone between adjacent regions, load the two models with the highest level of detail in the four precision-level models into the near-view region, load the model with the third highest level of detail into the middle-view region, and load the model with the lowest level of detail into the far-view region;
[0156] Calculate vertex position interpolation based on the vertex feature description within the overlapping transition zone to perform smooth switching between different precision models to obtain a visualization scene. Input the visualization scene into a dual-branch semantic segmentation network, extract local features through multi-layer convolution and downsampling, and extract global features through multi-level pooling;
[0157] Fuse the local features and the global features through skip connections, perform deconvolution processing on the fused features to restore the spatial resolution to generate a semantic segmentation map, analyze the semantic segmentation map to obtain the distribution of different category regions in the scene, establish an association between the input news event description and the corresponding category regions, and convert the annotation position to the scene space coordinate system through line-of-sight projection;
[0158] Obtain the position coordinates and operation instructions of the interaction device, convert the screen space coordinates to the normalized device coordinate system, construct an interaction ray to calculate the intersection position with the visualization scene, calculate the rotation matrix, translation vector, and scaling factor of the viewpoint based on the intersection position, generate a viewpoint transformation path, and dynamically update the precision configuration of the models within the visible range and the display state of the annotation symbols according to the viewpoint position.
[0159] The grid patch refers to a planar fragment obtained by grid division in 3D modeling, which is often used to represent the discretized grid of an object's surface and is applied in fields such as computer graphics, finite element analysis, and 3D modeling. The local surface fitting equation is a mathematical model for fitting a surface based on a local data point set, which obtains the equation of the local surface by minimizing the fitting error and is often used in tasks such as point cloud processing, image reconstruction, and 3D surface fitting. The scene bounding box is a minimum bounding box used to enclose a 3D object or scene, which is usually used for collision detection, object detection, and scene representation in computer graphics, and improves the calculation efficiency by simplifying the geometric shape. The view point transformation path refers to the path trajectory from one view point to another in an image or 3D scene, which is usually used for view angle conversion in computer graphics, path planning, and navigation control in virtual reality.
[0160] Preprocess the 3D scene model. For each grid patch in the model, convert the normal vector of each patch to the local coordinate system where the vertex is located, and calculate the weighted average to obtain the vertex normal vector. For example, if a vertex is shared by three patches, the normal vector of this vertex is the weighted average of the normal vectors of the three patches, and the weight can be set as the patch area. Then, based on the local coordinate system of the vertex and the position information of the surrounding vertices, fit a local surface equation. By analyzing this surface equation, calculate the principal curvature at this vertex, and then obtain the curvature change value. For example, a quadratic polynomial can be used to fit the local surface, and the principal curvature can be calculated through the eigenvalues of the Hessian matrix. At the same time, traverse the adjacent vertices of this vertex and calculate the change rate of the included angle between their normal vectors. Combine the calculated curvature change value and the change rate of the normal vector included angle to form the feature description of this vertex. Finally, establish a graph structure to record the feature descriptions of all vertices and the adjacency relationships between vertices.
[0161] Perform spatial division according to the vertex features. Input the vertex feature description, the vertex density distribution in the adjacency relationship graph, and the feature point distribution ratio into the spatial division evaluation function. Use the bounding box of the scene as the initial space and recursively construct an octree structure. When dividing each time, calculate the evaluation function value of each node to be divided. The division threshold is dynamically adjusted according to the level of the node and the evaluation value of the parent node. For example, for nodes with a deeper level or a higher evaluation value of the parent node, set a more stringent division threshold. After the division is completed, each leaf node records its spatial range, the included vertex feature description, and the mesh topology relationship, forming a spatial index structure. Suppose the size of a scene bounding box is 10x10x10 and it contains 1000 vertices. After octree division, leaf nodes of different sizes such as 5x5x5 and 2.5x2.5x2.5 may be obtained.
[0162] For each leaf node in the spatial index structure, calculate the sum of the squares of the distances from each vertex within the node to the adjacent patches, and construct a quadratic error matrix. Combine this error matrix, the normal vector change rate, and the texture coordinate change amount to construct an edge collapse cost function. Traverse all the mesh edges within the node, calculate the optimal position after collapsing each edge, and store the results in a priority queue. According to the value of the edge collapse cost function, take out the edge with the minimum cost from the priority queue for the collapse operation, and update the error matrix of the affected vertices and the cost values of the relevant edges. Repeat this process until the preset vertex number target is reached. For example, for a leaf node containing 100 vertices, the target vertex number can be set to 50, and the edge collapse operation is iteratively executed.
[0163] Control the number of collapse iterations according to the edge collapse cost function to generate four models with different precisions. To maintain the continuity of vertex feature description, during the edge collapse process, preferentially select the edges with less impact on vertex features for collapse. Finally, generate four precision level models that retain 90%, 70%, 50%, and 30% of the original vertex numbers. Then, construct a quadtree structure corresponding to the frustum, divide the frustum into three regions: near view, middle view, and far view, and set an overlapping transition zone between adjacent regions. Load the two models with the highest precision into the near view region, the third highest precision model into the middle view region, and the lowest precision model into the far view region.
[0164] Within the overlapping transition zone, calculate the interpolation of vertex positions based on vertex feature description to achieve smooth switching between models with different precisions, and finally obtain a visualization scene. Input this scene into a dual-branch semantic segmentation network, extract local features through multi-layer convolution and downsampling operations, and extract global features through multi-level pooling operations. Fuse the local features and global features through skip connections, and then perform deconvolution processing to restore the spatial resolution to generate a semantic segmentation map.
[0165] Establish an association between the input news event description and the corresponding category area. For example, associate the "fire" event with the area labeled as "building". Through line-of-sight projection, convert the labeled position into the scene space coordinate system.
[0166] Obtain the position coordinates and operation instructions of the interaction device, convert the screen space coordinates to the normalized device coordinate system, construct an interaction ray, and calculate its intersection position with the visualization scene. Based on this intersection position, calculate the rotation matrix, translation vector, and scaling factor of the viewpoint to generate a viewpoint transformation path. Dynamically update the precision configuration of the models within the visible range and the display status of the annotation symbols according to the viewpoint position. For example, when the viewpoint approaches a certain area, load the high-precision model of that area and display the relevant annotation symbols.
[0167] In this embodiment, by adaptively loading models with different precisions based on the viewing distance and scene complexity, the rendering computation amount is significantly reduced, the frame rate and smoothness of scene roaming are improved. By optimizing the models, important geometric features and texture details are retained, and smooth transitions are made between models with different precisions, making the scene more realistic and enhancing the user's sense of immersion. Combining with semantic segmentation technology, different objects in the scene can be accurately labeled, and news events can be associated with scene elements, facilitating users to quickly understand the location and context information where the events occur.
[0168] In an alternative embodiment,
[0169] For each leaf node in the spatial index structure, construct a quadratic error matrix with the sum of the squares of the distances from the vertex to adjacent patches, and combine the quadratic error matrix, the normal vector change rate, and the texture coordinate change amount to construct an edge collapse cost function. Traverse the mesh edges within the node to calculate the optimal collapse position and store it in a priority queue. Iteratively perform edge collapse operations according to the magnitude of the values of the collapse cost function. Updating the error matrix of the affected vertices and the cost values of the relevant edges includes:
[0170] Obtain the mesh information of each leaf node in the spatial index structure, establish an association table between vertices and adjacent patches, obtain the patch set and topological relationship to which each vertex belongs, calculate the patch normal vector and plane equation based on the patch set, and calculate the sum of the squares of the distances from the vertex to adjacent patches to construct a quadratic error matrix, where the quadratic error matrix characterizes the influence degree on adjacent patches when the vertex position changes;
[0171] Combine the quadratic error matrix, the normal vector change rate, and the texture coordinate change amount to construct an edge collapse cost function, calculate the shape deformation value based on the quadratic error matrices of the two ends of the edge, obtain the normal vectors of the two ends and adjacent vertices of the edge to calculate the normal vector change rate, analyze the vertex texture coordinates to obtain the coordinate change amount, and perform linear combination after normalizing the shape deformation value, the normal vector change rate, and the texture coordinate change amount to obtain the edge collapse cost function value;
[0172] Traverse the mesh edges within the node to calculate the optimal collapse position, obtain the patches adjacent to the edge to construct a local mesh area, select multiple candidate collapse points on the straight line where the edge is located, calculate the mesh error corresponding to each candidate point, and select the position with the smallest mesh error and maintaining topology validity as the optimal collapse position. Store the edge information and the optimal collapse position in a priority queue sorted according to the cost function value;
[0173] Extract the edge with the minimum cost function value from the priority queue. After verifying the validity of the edge, obtain the optimal collapse position corresponding to the current edge, merge the two endpoints of the edge into a new vertex, update the coordinates, normal vectors, and texture coordinates of the new vertex according to the optimal collapse position, update the vertex indices of the relevant patches, delete the degenerate patches, and adjust the mesh connection relationship;
[0174] Update the error matrix of the affected vertices, integrate the geometric relationships between the new vertex and the adjacent patches into the error matrix, recalculate the cost function value and the optimal collapse position for the edges connected to the new vertex, and re-store the updated edge information into the priority queue to maintain the optimal simplification order.
[0175] The candidate collapse point is a point selected for simplification or merging in geometric meshes or image processing. It is usually determined by calculating the surrounding environment, deformation cost, and geometric features of each point and is used for operations such as mesh simplification and point cloud compression. The topological validity means maintaining the validity and connectivity of the spatial data structure or graphical model during topological optimization and mesh processing, avoiding structures that do not conform to physical or geometric constraints, so that the optimization results meet the actual application requirements.
[0176] Load the 3D mesh model into memory and construct a spatial index structure, such as an octree or a KD - tree. The spatial index structure divides the model space into multiple leaf nodes, and each leaf node contains a part of the mesh data.
[0177] Traverse each leaf node in the spatial index structure. For each leaf node, obtain the mesh information it contains, including vertex coordinates, patch connection relationships, etc. Establish an association table between vertices and adjacent patches, recording the set of patches to which each vertex belongs and their topological relationships.
[0178] Based on the set of patches to which each vertex belongs, calculate the normal vector and plane equation of each patch, calculate the sum of the squared distances from the vertex to the adjacent patches, and construct a quadratic error matrix. This matrix is used to characterize the degree of influence of the change in vertex position on the adjacent patches. For example, if there are three patches around a vertex and the sum of the squared distances to these three patches are 1, 4, and 9 respectively, then the quadratic error matrix of this vertex can be expressed as [1, 4, 9].
[0179] Combine the quadratic error matrix, the normal vector change rate, and the texture coordinate change amount to construct an edge collapse cost function. For each edge, calculate the shape deformation value based on the quadratic error matrices of the vertices at both ends of the edge. Obtain the normal vectors of the two ends of the edge and the adjacent vertices, and calculate the normal vector change rate. Analyze the vertex texture coordinates to obtain the coordinate change amount. Linearly combine the shape deformation value, the normal vector change rate, and the texture coordinate change amount after normalization to obtain the edge collapse cost function value. For example, assume the shape deformation value is 0.8, the normal vector change rate is 0.5, the texture coordinate change amount is 0.2, and the weights are 0.5, 0.3, and 0.2 respectively. Then the edge collapse cost function value is 0.5 * 0.8 + 0.3 * 0.5 + 0.2 * 0.2 = 0.59.
[0180] Traverse all the mesh edges within the node and calculate the optimal collapse position. For each edge, obtain the adjacent faces and construct a local mesh region. Select multiple candidate collapse points on the line where the edge lies. For example, divide the edge into five equal parts and take these five equally divided points as candidate collapse points. Calculate the mesh error corresponding to each candidate point, and select the position with the minimum mesh error and maintaining topological validity as the optimal collapse position. Store the edge information and the optimal collapse position into a priority queue sorted by the cost function value.
[0181] Extract the edge with the minimum cost function value from the priority queue. Check the validity of the edge. For example, check whether collapsing this edge will generate degenerate triangles or flipped normals, etc. If the edge is valid, obtain the optimal collapse position corresponding to the current edge, merge the two endpoints of the edge into a new vertex, and update the coordinates, normal vectors, and texture coordinates of the new vertex according to the optimal collapse position. Update the vertex indices of the relevant faces, delete the degenerate faces, and adjust the mesh connection relationship.
[0182] Update the error matrix of the affected vertices. Integrate the geometric relationship between the new vertex and the adjacent faces into the error matrix. Recalculate the cost function value and the optimal collapse position for the edges connected to the new vertex, and store the updated edge information back into the priority queue to maintain the optimal simplification order. Repeat the above steps until the preset simplification goal is reached. For example, the number of faces is reduced to a specified ratio.
[0183] Figure 4This is a comparison chart of the texture model simplification effect of the 3D reconstruction and visualization method for news scenes based on multi-source remote sensing data in the embodiments of the present invention, showing the visual effect comparison of the texture model at a simplification rate of 80%. The upper left is the original model (538,124 patches), the upper right is the simplification result of the present technical solution (107,625 patches, texture distortion degree 0.275), the lower left is the simplification result of the QEM method (107,625 patches, texture distortion degree 0.463), and the lower right is the simplification result of the vertex clustering method (107,635 patches, texture distortion degree 0.798). It can be clearly seen from the visual effect that the present technical solution performs best in maintaining the texture mapping quality, with accurate texture boundary alignment and clear pattern details. Although the QEM method maintains the general texture distribution, obvious distortion and blurring occur in the texture-dense areas (such as the pattern boundary). The vertex clustering method leads to serious texture distortion, with many details completely lost and incorrect boundary alignment. From the data, the texture distortion degree of the present technical solution (0.275) is 40.6% lower than that of the QEM method (0.463) and 65.5% lower than that of the vertex clustering method (0.798). Especially in the high-frequency texture areas and texture boundaries on the model surface, the preservation effect of the present technical solution is more prominent, fully demonstrating the effectiveness of incorporating the texture coordinate change amount into the edge collapse cost function in the present technical solution, which can still maintain good texture mapping quality at a high simplification rate and is suitable for application scenarios that require high-quality texture performance.
[0184] In this embodiment, the grid model can be divided into multiple local regions by using a spatial index structure, and the edge collapse operation can be independently performed in each region, so as to achieve parallel processing and improve the simplification efficiency. By considering factors such as the quadratic error matrix, the normal vector change rate, and the texture coordinate change amount, the model deformation during the simplification process can be effectively controlled, and the visual quality of the model can be maintained.
[0185] In an alternative embodiment,
[0186] Taking the reconstruction and visualization of the production recovery situation of a certain city as an example. Collect multi-source remote sensing data: Use the Gaofen-2 to obtain 1-meter resolution multi-spectral images for ground object recognition; use the Gaofen-3 radar data to achieve all-weather observation; obtain stereo mapping data through the Gaofen-7 for elevation extraction; at the same time, obtain meteorological satellite data such as Himawari-8 for surface temperature inversion.
[0187] The data obtained by different sensors are registered by using the iterative closest point algorithm based on spatial position weighting, and the data are unified into the same coordinate system through the tensor projection algorithm. The multi-source data features are extracted by using a bidirectional attention deep neural network and a hierarchical feature descriptor is generated. The feature matching and screening are carried out by using a random sampling algorithm based on local geometric consistency constraints. The variational autoencoder network and the generative adversarial network based on pixel-level adversarial discrimination are used to enhance the data.
[0188] The enhanced data is used to generate a spatial attention feature map and a channel attention feature map through a feature extraction network. The depth information is generated by a monocular depth estimation network based on dense connections using the spatial attention feature map. The depth information is fused with the GF-3 radar data and the GF-7 point cloud data to obtain the sparse three-dimensional point cloud of urban buildings. The Poisson reconstruction algorithm is used to convert the point cloud into an initial three-dimensional mesh model, and then the texture transfer of the model is performed by a residual map convolutional neural network using the channel attention feature map. Topological optimization is carried out by combining the multi-view projection consistency constraint and the mesh optimization algorithm.
[0189] To achieve a hierarchical visualization effect, the curvature change value and the normal vector change value of the meshes in the optimized model are calculated, and a spatial division evaluation function is constructed to divide the scene into octree spatial blocks of different sizes. The quadratic error matrix of the mesh vertices in each spatial block is calculated and an edge collapse cost function is constructed, and the mesh vertices are merged and simplified according to the cost function to generate models with four different precision levels. A quadtree frustum is constructed and divided into three depth regions: far, middle, and near, and the models with corresponding precision levels are loaded in different regions.
[0190] The visualization scene is input into a dual-branch semantic segmentation network to extract local features and global features, and a scene semantic segmentation map is generated through feature fusion to label key regions such as industrial areas and transportation hubs. The system supports collecting user interaction location information and calculating the viewpoint transformation matrix to implement the scene roaming function.
[0191] Figure 2 FIG. is a schematic structural diagram of a three-dimensional reconstruction and visualization system for news scenes based on multi-source remote sensing data according to an embodiment of the present invention, as Figure 2 shown, the system includes:
[0192] A first unit, configured to collect multi-source remote sensing data corresponding to a target news scene through a remote sensing satellite, extract hierarchical feature descriptors through an iterative closest point algorithm based on spatial position weighting, a tensor projection algorithm, and a bidirectional attention depth neural network, and perform screening through a random sampling algorithm based on local geometric consistency constraints, eliminate feature point pairs with projection errors greater than a preset projection threshold and unify them to global geographic coordinates to obtain registered multi-source remote sensing data, and perform data enhancement through a variational autoencoder network and a noise-aware loss function, in combination with a generative adversarial network based on pixel-level adversarial discrimination to obtain enhanced remote sensing data;
[0193] The second unit is used to add enhanced remote sensing data to a feature extraction network to obtain a spatial attention feature map and a channel attention feature map, add the spatial attention feature map to a monocular depth estimation network based on dense connections to generate depth information, fuse it with dual-polarization radar remote sensing data and multi-echo laser point cloud data to obtain sparse three-dimensional point clouds, generate an initial three-dimensional mesh model through a Poisson reconstruction algorithm, add the channel attention feature map to a residual graph convolutional neural network for texture feature transfer, and perform topological optimization through a multi-view projection consistency constraint and a mesh optimization algorithm to obtain an optimized three-dimensional scene model;
[0194] The third unit is used to calculate the curvature change value and the normal vector change value of each grid in the optimized three-dimensional scene model, construct a spatial division evaluation function to divide the three-dimensional scene into octree spatial blocks of different sizes, calculate the quadratic error matrix of the grid vertices in each spatial block and construct an edge collapse cost function to merge and simplify the grid vertices to obtain four precision level models, construct a quadtree frustum to divide it into three depth regions and load the corresponding precision level models to obtain a visualization scene, extract local features and global features through a dual-branch semantic segmentation network, generate a scene semantic segmentation map through fusion upsampling and label the news event positions, collect interactive position information and calculate a viewpoint transformation matrix to update the visualization scene for scene roaming and perform scene roaming.
[0195] In the third aspect of the embodiments of the present invention,
[0196] A kind of electronic device is provided, including:
[0197] A processor;
[0198] A memory for storing instructions executable by the processor;
[0199] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0200] In the fourth aspect of the embodiments of the present invention,
[0201] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0202] The present invention can be a method, a device, a system and / or a computer program product. The computer program product can include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.
[0203] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A three-dimensional reconstruction and visualization method for news scenes based on multi-source remote sensing data, characterized in that: include: Multi-source remote sensing data corresponding to the target news scene is collected through remote sensing satellites. Hierarchical feature descriptors are extracted based on the iterative closest point algorithm and tensor projection algorithm based on spatial position weighting and bidirectional attention deep neural network. The data is screened through a random sampling algorithm based on local geometric consistency constraints, and feature point pairs with projection errors greater than a preset projection threshold are eliminated and unified to global geographic coordinates to obtain registered multi-source remote sensing data. Enhanced remote sensing data is obtained through data enhancement using a variational autoencoder network and a noise-aware loss function combined with a generative adversarial network based on pixel-level adversarial discrimination. The enhanced remote sensing data is added to the feature extraction network to obtain the spatial attention feature map and the channel attention feature map. The spatial attention feature map is added to the densely connected monocular depth estimation network to generate depth information. It is fused with the dual-polarization radar remote sensing data and the multi-echo laser point cloud data to obtain a sparse three-dimensional point cloud. The initial three-dimensional mesh model is generated by the Poisson reconstruction algorithm. The channel attention feature map is added to the residual graph convolutional neural network for texture feature migration. The topology optimization is performed by combining the multi-view projection consistency constraint and the mesh optimization algorithm to obtain an optimized three-dimensional scene model. The curvature change value and normal vector change value of each mesh in the three-dimensional scene model are calculated and optimized, and a space division evaluation function is constructed to divide the three-dimensional scene into octree space blocks of different sizes. The quadratic error matrix of the mesh vertices in each space block is calculated and an edge collapse cost function is constructed to merge and simplify the mesh vertices to obtain four precision level models. A quadtree viewing cone is constructed to divide it into three depth areas and load the corresponding precision level models to obtain a visualized scene. Local and global features are extracted through a dual-branch semantic segmentation network, and a scene semantic segmentation map is generated through fusion upsampling and the location of news events is marked. The interactive position information is collected and the viewpoint transformation matrix is calculated to update the visualized scene for scene roaming.
2. The method according to claim 1, characterized in that Multi-source remote sensing data corresponding to the target news scene is collected through remote sensing satellites. Hierarchical feature descriptors are extracted based on the iterative closest point algorithm and tensor projection algorithm based on spatial position weighting and bidirectional attention deep neural network. The feature point pairs with projection errors greater than the preset projection threshold are screened through a random sampling algorithm based on local geometric consistency constraints, and the feature points are unified to the global geographic coordinates to obtain the registered multi-source remote sensing data. The enhanced remote sensing data includes: Collect multi-source remote sensing data corresponding to the target news scene through remote sensing satellites, wherein the multi-source remote sensing data includes optical remote sensing image data, dual-polarization radar remote sensing data and multi-echo laser point cloud data; Add the multi-source remote sensing data to the iterative nearest point algorithm based on spatial position weighting, construct a spatial position weight matrix, construct a feature distance function according to the distance from the data point to the center of the scene, calculate the weight value, perform a nearest neighbor search on each data point through an octree spatial index structure to establish a nearest neighbor point set, calculate the position deviation vector between the data point and the nearest neighbor point set, multiply the position deviation vector by the spatial position weight matrix to obtain a weighted position deviation, repeat the iteration until the weighted position deviation is less than a preset threshold, and output the optimized data point set; Add the multi-source remote sensing data to the tensor projection algorithm, construct a characteristic covariance matrix through principal component analysis, determine the optimal projection direction by calculating eigenvalues and eigenvectors, perform a tensor projection operation in the optimal projection direction, construct a local curvature calculation function and a normal vector change function for the projection results, and extract multimodal data feature points through curvature threshold segmentation and region growing algorithm; Adding the multimodal data feature points to a bidirectional attention deep neural network, the bidirectional attention deep neural network extracts multi-scale features through multi-layer convolution and pooling operations of an encoder, the multi-layer convolution includes a point convolution layer, an edge convolution layer and a surface convolution layer, and restores the features through deconvolution and upsampling operations of a decoder, and a bidirectional attention module is set between the encoder and the decoder, wherein the bidirectional attention module includes a self-attention mechanism for calculating the internal spatial correlation of feature points and a cross-attention mechanism for calculating the topological correlation between different feature points, and the feature mapping results of the self-attention mechanism and the cross-attention mechanism are integrated to obtain a hierarchical feature descriptor; The hierarchical feature descriptors are screened by a random sampling algorithm based on local geometric consistency constraints, candidate feature point pairs are determined by feature similarity calculation, feature point pairs are randomly selected as seed matching, initial geometric transformation relationships including rotation matrices and translation vectors are calculated based on the seed matching, local geometric consistency constraint checks are performed on the remaining feature point pairs to obtain projection errors, feature point pairs whose projection errors are greater than a preset projection threshold are eliminated, the random sampling process is repeated multiple times to select the geometric transformation relationship with the most internal point support, the multi-source remote sensing data are unified to global geographic coordinates based on the selected geometric transformation relationship, and registered multi-source remote sensing data are generated; Inputting the registered multi-source remote sensing data into a variational autoencoder network, the encoder of the variational autoencoder network maps the registered multi-source remote sensing data to a latent feature space through multiple layers of residual convolution, the decoder reconstructs the latent features into the original data form through deconvolution and jump connection, introduces random noise disturbance into the latent feature space, and constructs a noise-aware loss function to constrain the reconstruction result; A generative adversarial network based on pixel-level adversarial discrimination is constructed. The generative adversarial network adopts a densely connected fully convolutional structure to output a pixel-level authenticity score map. The decoder of the variational autoencoder network is used as a generator, and an adversarial loss function is constructed. The parameters of the generator and the generative adversarial network are optimized through a minimum-maximum game until a Nash equilibrium is reached to obtain enhanced remote sensing data.
3. The method according to claim 2, characterized in that The hierarchical feature descriptor is screened by a random sampling algorithm based on local geometric consistency constraints, candidate feature point pairs are determined by feature similarity calculation, feature point pairs are randomly selected as seed matching, initial geometric transformation relationships including rotation matrices and translation vectors are calculated based on the seed matching, local geometric consistency constraint checks are performed on the remaining feature point pairs to obtain projection errors, feature point pairs whose projection errors are greater than a preset projection threshold are eliminated, the random sampling process is repeated multiple times to select the geometric transformation relationship with the most internal point support, and the multi-source remote sensing data is unified to global geographic coordinates based on the selected geometric transformation relationship, and the generation of registered multi-source remote sensing data includes: Constructing a local geometric feature extraction unit and a global semantic feature extraction unit, wherein the local geometric feature extraction unit extracts a local shape descriptor by constructing a feature point neighborhood and analyzing the distribution characteristics of points in the neighborhood, and the global semantic feature extraction unit extracts hierarchical semantic information from bottom to top through a multi-scale feature pyramid, and cascades the local shape descriptor and the hierarchical semantic information to form a feature vector; Constructing a feature similarity calculation module, which is based on the cosine similarity measurement method, introduces a spatial distance weighting term so that feature points with similar distances have a higher matching priority, and pairs feature descriptors with similarities higher than a preset threshold to form a set of candidate feature point pairs; Randomly select a feature point pair from the candidate feature point pair set as a seed match, build a geometric transformation estimation module based on the seed match, establish a mapping relationship between corresponding points through the geometric transformation estimation module to build a three-dimensional rigid body transformation model, combine an iterative optimization algorithm to decompose the rotation matrix in the transformation process into three basic rotation matrices and solve them, and solve the translation vector by minimizing the Euclidean distance between corresponding points; Constructing a geometric consistency verification module, the geometric consistency verification module includes a feature transformation unit and an error calculation unit, the feature transformation unit transforms the source feature point to the target coordinate system through the three-dimensional rigid body transformation model, the error calculation unit calculates the projection error between the transformed position and the actual target position, dynamically adjusts the projection error threshold according to the data distribution characteristics, and eliminates feature point pairs with projection errors greater than the projection error threshold; Constructing a random sampling iterative optimization module, wherein the random sampling iterative optimization module randomly selects different seeds to match and execute the geometric transformation estimation module and the geometric consistency verification module, counts the number of interior points that meet the constraint conditions, comprehensively considers the number of interior points and the stability of the geometric transformation to perform scoring, and selects the geometric transformation relationship with the highest score as the optimal geometric transformation relationship; A coordinate conversion module is constructed based on the optimal geometric transformation relationship. Based on the coordinate conversion module, the optical remote sensing image data is used as reference data. An affine transformation model considering imaging geometric characteristics is constructed for the dual-polarization radar remote sensing data, and transformation parameters are solved by a polynomial fitting method to achieve geometric correction. A three-dimensional rigid body transformation model is constructed for the multi-echo laser point cloud data, and transformation parameters are solved by an iterative optimization method. The optical remote sensing image data, the dual-polarization radar remote sensing data and the multi-echo laser point cloud data are uniformly converted to a global geographic coordinate system.
4. The method according to claim 1, characterized in that: The enhanced remote sensing data is added to the feature extraction network to obtain the spatial attention feature map and the channel attention feature map. The spatial attention feature map is added to the densely connected monocular depth estimation network to generate depth information. It is fused with the dual-polarization radar remote sensing data and the multi-echo laser point cloud data to obtain a sparse three-dimensional point cloud. The initial three-dimensional mesh model is generated by the Poisson reconstruction algorithm. The channel attention feature map is added to the residual graph convolutional neural network for texture feature migration. The optimized three-dimensional scene model is obtained by combining the multi-view projection consistency constraint and the mesh optimization algorithm for topology optimization. The following includes: Input the enhanced remote sensing data into the feature extraction network, extract multi-scale spatial features through a multi-layer cascade convolution structure and a batch normalization layer, use an adaptive weight calculation unit to assign spatial attention weights to each position, obtain channel descriptors through global pooling for information aggregation, perform multi-layer transformation on the channel descriptors to learn the inter-channel dependency relationship to generate channel attention weights, and adaptively fuse the spatial attention weights and the channel attention weights to generate a spatial attention feature map and a channel attention feature map; The spatial attention feature map is input into a densely connected monocular depth estimation network, a multi-layer cascade connection is established in the monocular depth estimation network by a jump method, a direct mapping relationship is established between shallow features and deep features, feature resolution is gradually restored through a multi-layer deconvolution structure, and a multi-scale feature pyramid structure is used to perform weight adaptive fusion of features at different levels to generate depth information; The depth information is sampled and a mapping relationship between pixel coordinates and three-dimensional space points is established. The confidence weight is calculated for the pre-acquired dual-polarization radar remote sensing data and multi-echo laser point cloud data based on the signal-to-noise ratio and spatial resolution. An adaptive weighted fusion strategy is constructed according to the spatial distance and feature similarity. The point clouds in the overlapping areas are fused to generate a sparse three-dimensional point cloud. The sparse three-dimensional point cloud is input into the Poisson reconstruction algorithm, the normal vector of each point is calculated by the feature decomposition method, the spatial subdivision granularity is dynamically adjusted according to the point cloud density distribution, the implicit surface expression is obtained by iteratively solving the Poisson equation, the grid structure is extracted by combining the adaptive threshold and Laplace smoothing is performed to generate the initial three-dimensional grid model, the channel attention feature map is input into the residual graph convolutional neural network, the node information is transferred by local neighborhood aggregation, the input feature and the transformed feature are added to construct the residual connection, the texture feature is adaptively captured by the graph attention mechanism, and the texture feature is mapped to the surface of the initial three-dimensional grid model by projection transformation to perform texture feature migration; Multiple observation perspectives are uniformly sampled in the spherical space, and the initial three-dimensional mesh model is projected onto each observation perspective plane to generate a multi-perspective projection image. The similarity between the multi-perspective projection image and the original image is calculated by feature descriptor matching to establish a multi-perspective projection consistency constraint. Based on the multi-perspective projection consistency constraint, the mesh vertex positions and connection relationships are optimized while maintaining the model geometric details to generate an optimized three-dimensional scene model.
5. The method according to claim 4, characterized in that The sparse three-dimensional point cloud is input into the Poisson reconstruction algorithm, the normal vector of each point is calculated by the feature decomposition method, the spatial subdivision granularity is dynamically adjusted according to the point cloud density distribution, the implicit surface expression is obtained by iteratively solving the Poisson equation, the grid structure is extracted by combining the adaptive threshold and Laplace smoothing is performed to generate the initial three-dimensional grid model, the channel attention feature map is input into the residual graph convolutional neural network, the node information is transferred by local neighborhood aggregation, the input feature and the transformation feature are added to construct the residual connection, the texture feature is adaptively captured by the graph attention mechanism, and the texture feature is mapped to the surface of the initial three-dimensional grid model by projection transformation to perform texture feature migration, including: Input the sparse three-dimensional point cloud into the Poisson reconstruction algorithm, construct a point cloud neighbor search tree structure, use a spherical neighborhood search method to determine the local neighborhood point set corresponding to each point in the sparse three-dimensional point cloud, construct a covariance matrix based on the local neighborhood point set, calculate the normal vector of each point by the eigendecomposition method, construct a minimum spanning tree to establish the connection relationship between points, select the normal vector at the center of the point cloud as a reference direction, recursively traverse the minimum spanning tree to adjust the normal vector directions of all points to obtain a normal vector set; Calculate the outer bounding box of the point cloud as the initial spatial range, dynamically adjust the spatial subdivision granularity according to the point cloud density distribution, count the number of point clouds contained in each octree node and calculate the point density distribution, set an adaptive subdivision threshold based on the point density, divide the node with a point density greater than the adaptive subdivision threshold into eight child nodes and record the depth information and point set contained in each leaf node, merge the divided leaf nodes, and merge adjacent nodes with similar point density to obtain an optimized octree structure; A three-dimensional scalar field function is defined on the optimized octree structure, the normal vector set is converted into a gradient field constraint, an implicit surface expression is obtained by iteratively solving the Poisson equation, a solution hierarchy is constructed by adopting a multi-grid strategy, the Poisson equation is converted into a linear equation system by constructing a sparse matrix to represent the Laplace operator, an iterative optimization method is adopted to solve the linear equation system and the scalar values of the grid nodes are updated to obtain a three-dimensional scalar field; Extracting the mesh structure in combination with the adaptive threshold, extracting the isosurface in the optimized octree structure based on the marching cube algorithm to generate the initial mesh, performing a topological check on the initial mesh and repairing the non-manifold edges and singular points, performing Laplace smoothing iteration to optimize the vertex positions, and simplifying the mesh based on the edge collapse criterion to obtain the initial three-dimensional mesh model; Input the channel attention feature map into the residual graph convolutional neural network, transfer node information through local neighborhood aggregation, define an adjacency matrix to describe the connection relationship between nodes, perform message aggregation based on the adjacency matrix, perform weighted combination of node features and neighborhood features, add the input features and the transformed features to construct a residual connection to obtain enhanced features; The texture features are adaptively captured through the graph attention mechanism, the feature statistics are obtained through the global pooling operation and nonlinear transformation is performed on the enhanced features, the correlation weights between different channels are learned to generate channel attention weights, the importance scores of different positions of the feature map are calculated to generate spatial attention weights, and the channel attention weights and the spatial attention weights are applied to the feature map to obtain the texture features; A projection transformation is used to map the texture features to the surface of the initial three-dimensional mesh model for texture feature migration, a mapping relationship between the vertices of the initial three-dimensional mesh model and the image pixels is established, the vertices are projected to the image plane according to the camera projection matrix, the texture coordinates are calculated inside the mesh facets by using barycentric coordinate interpolation, a multi-view consistency constraint is constructed to evaluate the texture mapping quality under different viewing angles, the optimal texture source is selected based on the graph cut optimization algorithm, and the texture fineness is dynamically adjusted according to the observation distance to obtain a textured three-dimensional scene model.
6. The method according to claim 1, characterized in that Calculate and optimize the curvature change value and normal vector change value of each grid in the three-dimensional scene model, construct a space division evaluation function to divide the three-dimensional scene into octree space blocks of different sizes, calculate the quadratic error matrix of the grid vertices in each space block and construct an edge collapse cost function to merge and simplify the grid vertices to obtain four precision level models, construct a quadtree view cone divided into three depth areas to load the corresponding precision level model to obtain a visualization scene, extract local features and global features through a dual-branch semantic segmentation network, generate a scene semantic segmentation map through fusion upsampling and mark the location of news events, collect interactive location information and calculate the viewpoint transformation matrix to update the visualization scene for scene roaming and perform scene roaming, including: For each mesh face in the optimized three-dimensional scene model, project the normal vector of each mesh face to the local coordinate system of the vertex and calculate the weighted average to obtain the vertex normal vector, construct the local surface fitting equation of the vertex to calculate the principal curvature to obtain the curvature change value, traverse the adjacent vertices to calculate the normal vector angle change rate, combine the curvature change value with the normal vector angle change rate to form a vertex feature description, and establish an adjacency relationship graph that records the vertex feature description; Input the vertex feature description, vertex density distribution in the adjacency relationship graph, and feature point distribution ratio into the space partition evaluation function, recursively construct an octree with the scene bounding box as the initial space, calculate the evaluation function value of each node to be segmented, dynamically adjust the segmentation threshold according to the node level and the parent node evaluation value, and record the spatial range of each leaf node, the vertex feature description, and the mesh topology relationship to form a spatial index structure; For each leaf node in the spatial index structure, the sum of the squares of the distances from the vertex to the adjacent facets is used to construct a quadratic error matrix, the normal vector change rate is calculated based on the normal vectors of the vertices at both ends of the mesh edge corresponding to each leaf node and the adjacent vertices, the quadratic error matrix, the normal vector change rate, and the texture coordinate change are combined to construct an edge collapse cost function, the mesh edges in the node are traversed to calculate the optimal collapse position and stored in a priority queue, the edge collapse operation is iteratively performed according to the value of the collapse cost function, and the error matrix of the affected vertex and the cost value of the related edge are updated; Controlling the number of collapse iterations based on the edge collapse cost function, maintaining the continuity of the vertex feature description, generating four precision level models that retain 90%, 70%, 50%, and 30% of the original vertex number, constructing a quadtree structure corresponding to the viewing frustum, dividing the viewing frustum into a near view area, a mid-view area, and a far view area, setting an overlapping transition zone between adjacent areas, loading the two models with the highest precision among the four precision level models into the near view area, loading the model with the third highest precision into the mid-view area, and loading the model with the lowest precision into the far view area; Calculating vertex position interpolation based on the vertex feature description within the overlapping transition zone, performing smooth switching between models of different precisions to obtain a visualization scene, inputting the visualization scene into a dual-branch semantic segmentation network, extracting local features through multi-layer convolution and downsampling, and extracting global features through multi-level pooling; The local features are fused with the global features through a skip connection, the fused features are deconvolved to restore the spatial resolution to generate a semantic segmentation map, the semantic segmentation map is parsed to obtain the distribution of different category areas in the scene, the input news event description is associated with the corresponding category area, and the marked position is converted to the scene space coordinate system through line of sight projection; Obtain the position coordinates and operation instructions of the interactive device, convert the screen space coordinates to the normalized device coordinate system, construct the intersection position of the interactive ray calculation and the visualization scene, calculate the rotation matrix, translation vector and scaling factor of the viewpoint based on the intersection position, generate the viewpoint transformation path, and dynamically update the accuracy configuration of the model within the visible range and the display status of the annotation symbols according to the viewpoint position.
7. The method according to claim 6, characterized in that For each leaf node in the spatial index structure, the sum of the squares of the distances from the vertex to the adjacent facets is used to construct a quadratic error matrix, the normal vector change rate is calculated based on the normal vectors of the vertices at both ends of the mesh edge corresponding to each leaf node and the adjacent vertices, the quadratic error matrix, the normal vector change rate, and the texture coordinate change are combined to construct an edge collapse cost function, the mesh edges in the node are traversed to calculate the optimal collapse position and stored in a priority queue, the edge collapse operation is iteratively performed according to the value of the collapse cost function, and the error matrix of the affected vertex and the cost value of the related edge are updated, including: Obtain the mesh information of each leaf node in the spatial index structure, establish an association table between vertices and adjacent facets, obtain the facet set and topological relationship to which each vertex belongs, calculate the facet normal vector and plane equation based on the facet set, calculate the sum of the squares of the distances from the vertex to the adjacent facets to construct a quadratic error matrix, wherein the quadratic error matrix represents the degree of influence on the adjacent facets when the vertex position changes; The quadratic error matrix, the normal vector change rate, and the texture coordinate change amount are combined to construct an edge collapse cost function, the shape deformation value is calculated based on the quadratic error matrix of the vertices at both ends of the edge, the normal vectors at both ends of the edge and the adjacent vertices are obtained to calculate the normal vector change rate, the vertex texture coordinates are analyzed to obtain the coordinate change amount, the shape deformation value, the normal vector change rate, and the texture coordinate change amount are normalized and then linearly combined to obtain the edge collapse cost function value; Traversing the mesh edges in the node to calculate the optimal collapse position, obtaining the facets adjacent to the edge to construct a local mesh area, selecting multiple candidate collapse points on the straight line where the edge is located, calculating the mesh error corresponding to each candidate point, selecting the position with the smallest mesh error and maintaining topological validity as the optimal collapse position, and storing the edge information and the optimal collapse position in a priority queue sorted by the cost function value; Extracting the edge with the smallest cost function value from the priority queue, obtaining the optimal collapse position corresponding to the current edge after verifying the validity of the edge, merging the two endpoints of the edge into a new vertex, updating the coordinates, normal vector and texture coordinates of the new vertex according to the optimal collapse position, updating the vertex index of the relevant facet, deleting the degenerate facet and adjusting the mesh connection relationship; Update the error matrix of the affected vertex, integrate the geometric relationship between the new vertex and the adjacent facets into the error matrix, recalculate the cost function value and the optimal collapse position for the edge connected to the new vertex, and store the updated edge information back into the priority queue to maintain the optimal simplification order.
8. A news scene three-dimensional reconstruction and visualization system based on multi-source remote sensing data, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to collect multi-source remote sensing data corresponding to the target news scene through remote sensing satellites, extract hierarchical feature descriptors based on the iterative closest point algorithm and tensor projection algorithm weighted by spatial position, and bidirectional attention deep neural network, screen them through the random sampling algorithm based on local geometric consistency constraints, remove feature point pairs with projection errors greater than the preset projection threshold, and unify them to global geographic coordinates to obtain registered multi-source remote sensing data, and enhance remote sensing data through the variational autoencoder network and noise-aware loss function combined with the generative adversarial network based on pixel-level adversarial discrimination; The second unit is used to add the enhanced remote sensing data to the feature extraction network to obtain the spatial attention feature map and the channel attention feature map, add the spatial attention feature map to the densely connected monocular depth estimation network to generate depth information, fuse it with the dual-polarization radar remote sensing data and the multi-echo laser point cloud data to obtain a sparse three-dimensional point cloud, generate an initial three-dimensional mesh model through the Poisson reconstruction algorithm, add the channel attention feature map to the residual graph convolutional neural network for texture feature migration, and combine the multi-view projection consistency constraint and the mesh optimization algorithm for topology optimization to obtain an optimized three-dimensional scene model; The third unit is used to calculate and optimize the curvature change value and normal vector change value of each mesh in the three-dimensional scene model, construct a space division evaluation function to divide the three-dimensional scene into octree space blocks of different sizes, calculate the quadratic error matrix of the mesh vertices in each space block and construct an edge collapse cost function to merge and simplify the mesh vertices to obtain four accuracy level models, construct a quadtree viewing cone divided into three depth areas, load the corresponding accuracy level model to obtain a visualized scene, extract local and global features through a dual-branch semantic segmentation network, generate a scene semantic segmentation map through fusion upsampling and mark the location of news events, collect interactive position information and calculate the viewpoint transformation matrix to update the visualized scene for scene roaming and perform scene roaming.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-view three-dimensional reconstruction method based on deep learning
CN116310095A
Real scene three-dimensional modeling method and system based on unmanned aerial vehicle oblique photogrammetry
CN118967980A