Three-dimensional structure recovery method for high-quality urban updated landscape building

By introducing deep learning technology into the SfM framework, the E2ESfM framework is proposed, which solves the computational complexity and inaccuracy problems of the traditional SfM framework when processing large-scale image collections, and realizes high-quality three-dimensional structure recovery.

CN119991952AInactive Publication Date: 2025-05-13BEIJING GUANGAN LIGHTING TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510075186.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional SfM framework is complex and time-consuming to process large-scale image collections, and is susceptible to noise and mismatch, resulting in inaccurate results, and fails to make full use of all image information, resulting in incompleteness or missing point clouds.

Method used

A deep learning-based end-to-end SfM framework E2ESfM is proposed to extract 2D trajectories from the input images, reconstruct the camera using image and track features, initialize the point cloud, and apply differentiable beam optimization layer for reconstruction refinement.

Benefits of technology

Realize global camera recovery and reliable pixel accurate trajectory extraction, reduces computational complexity, improves matching accuracy and robustness, and can better handle complex scenarios and large-scale data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991952A_ABST
    Figure CN119991952A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of artificial intelligence, computer vision and computer graphics, and discloses a three-dimensional structure recovery method for a high-quality urban updated landscape building, which comprises the following steps: providing a three-dimensional structure recovery method for a high-quality urban updated landscape building, and providing a new deep learning-based end-to-end SfM framework-E2ESfM. A 2D trajectory is extracted from an input image, a camera is reconstructed using the image and trajectory features, a point cloud is initialized based on these trajectories and camera parameters, and a beam optimization layer is applied for reconstruction refinement. The whole frame is all differentiable, that is, each component is differentiable and can be trained in an end-to-end mode, a reliable pixel accurate track can be extracted, all camera postures can be synchronously recovered, and the camera and triangularized 3D points can be optimized through a differentiable beam optimization layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of artificial intelligence, computer vision and computer graphics, and specifically relates to a method for restoring the three-dimensional structure of high-quality urban renewal landscape buildings. Background Art

[0002] Landscape architecture in urban renewal is an important topic in urban development in recent years. It aims to improve the image of the city, enhance urban functions, and promote harmonious coexistence with nature by transforming and renewing existing buildings. Urban renewal landscape architecture not only focuses on the beauty and practicality of the building itself, but also pays more attention to the integration and interaction with the surrounding environment. In the landscape architecture design of urban renewal, it also pays attention to the restoration of urban mixed functions at different levels. Through comprehensive consideration of planning, architecture, cultural protection and other levels, the renovated buildings not only meet the spatial form and functional requirements of modern cities, but also retain the inheritance of historical culture. At the same time, some cities also use historical context to enhance regional activity, turning the entire area into a living historical museum area, providing citizens with a rich cultural experience. In general, landscape architecture in urban renewal is an important part of urban development. It can not only enhance the image and function of the city, but also promote harmonious coexistence with nature and provide citizens with a better living environment. These landscape buildings are not only beautiful and practical, but also carry the historical culture and future vision of the city.

[0003] 3D structure recovery (SfM for short) is a technology that can automatically recover camera parameters and scene 3D structure from multiple images or video sequences. It has become a long-standing and important problem in computer vision.

[0004] Traditional SfM Framework

[0005] (1) First, we find image pairs with overlapping view cones by detecting and matching key points.

[0006] (2) These image pairs are then verified using the essential matrix or homography matrix of the two views.

[0007] (3) Next, a pair or a small group of images are selected for initialization, and new images are gradually registered. During the registration process, the perspective-n-point (PnP) problem is used to solve the camera pose, which is then triangulated to obtain 3D points and bundle optimized.

[0008] (4) This process is repeated iteratively until all frames are registered or discarded.

[0009] (5) The basis of the whole process is 2D correspondences (multi-view trajectories), however these trajectories are usually only constructed through chain pairing.

[0010] Deep learning-based framework

[0011] Recent SfM research efforts have focused on enhancing specific components of traditional SfM frameworks using deep learning techniques.

[0012] These research works are usually based on the original non-differentiable processes.

[0013] Traditional SfM has some problems:

[0014] Traditional SfM requires a pair-by-pair matching step, but this step has some problems:

[0015] (1) Pairwise matching requires a lot of computing resources and time, especially when processing large-scale image collections.

[0016] (2) Pairwise matching is susceptible to noise and false matches, leading to inaccurate results.

[0017] (3) Pairwise matching cannot guarantee the continuity of points and may cause breaks or omissions in point trajectories.

[0018] Traditional SfM requires step-by-step camera registration

[0019] (1) Progressive camera registration requires processing each image sequentially, which leads to increased computational complexity, especially when processing large-scale image collections.

[0020] (2) The progressive camera registration is susceptible to cumulative errors, which may lead to inaccuracies in camera pose and 3D structure.

[0021] (3) The step-by-step registration camera cannot fully utilize the information of all images, which may result in incomplete or missing point clouds.

[0022] Traditional SfM uses a non-differentiable solver

[0023] (1) The non-differentiable solver limits the trainability of the overall optimization process of SfM. This means that the entire SfM process cannot be trained end-to-end using optimization algorithms such as gradient descent, but requires manual design and adjustment of the parameters of each component. This increases the complexity of development and debugging, and may lead to inconsistencies between subcomponents.

[0024] (2) Non-differentiable solvers have difficulty handling complex loss functions. In SfM, it is usually necessary to minimize multiple loss functions of different types, such as reprojection error, photometric error, etc. Using non-differentiable solvers requires manually designing and implementing the gradient calculation of these loss functions, which can be very difficult and error-prone.

[0025] (3) Non-differentiable solvers usually require manual selection and adjustment of optimization algorithm hyperparameters, such as learning rate, number of iterations, etc. This requires the experience and trial and error of domain experts, increasing the difficulty of development and debugging.

[0026] In the 3D structural restoration of high-quality urban renewal landscape buildings, the non-differentiable SfM framework performs poorly in handling such large-scale and high-quality structural restoration applications, and the workload of manual adjustment is large, making the training process complex, costly, and ineffective. Summary of the invention

[0027] In order to solve the problems existing in the traditional SfM architecture and the problems of non-differentiability, poor effect and low performance of the current deep learning-based SfM framework, the present invention provides a high-quality 3D structure restoration method for urban renewal landscape architecture, and proposes a new end-to-end SfM framework based on deep learning, E2ESfM, which extracts 2D trajectories from the input image, reconstructs the camera using image and trajectory features, initializes the point cloud based on these trajectories and camera parameters, and applies a bundle optimization layer for reconstruction and refinement. The entire framework is fully differentiable, that is, each component is differentiable, can be trained in an end-to-end manner, can extract reliable pixel-accurate trajectories, can synchronously restore all camera poses, and can optimize the camera and triangulated 3D points through a differentiable bundle optimization layer.

[0028] To achieve the above object, the present invention provides the following solutions:

[0029] A method for restoring a three-dimensional structure of a high-quality urban renewal landscape building, the method comprising:

[0030] Build an end-to-end SfM framework E2ESfM based on deep learning;

[0031] Using E2ESfM, 2D trajectories are extracted from the input image;

[0032] The camera is reconstructed using image and trajectory features, the point cloud is initialized based on the 2D trajectory and camera parameters, and a bundle optimization layer is applied to optimize the camera and triangulated 3D points.

[0033] Preferably, E2ESfM uses a reconstruction function to represent SfM;

[0034] The reconstruction function is fully differentiable and its parameters are optimized by minimizing the training loss.

[0035] Preferably, the input image in Where H×W represents the resolution of the input image, 3 represents the three channels of RGB, and N I represents the number of input images, i I Represents the image index;

[0036] Camera projection matrix in Represents a 3×4 matrix consisting of a posture external parameter and a camera intrinsic Composition, among which represents the special Euclidean group, i.e., the rotation and translation of a rigid body in three-dimensional space; i P Represents the camera projection matrix index;

[0037] Scene point cloud Where N x represents the number of 3D points in the point cloud, Represents the spatial position coordinates of a 3D point.

[0038] Preferably, E2ESfM decomposes the reconstruction function into four stages: feature extraction and matching stage, camera pose estimation stage, 3D point cloud triangulation stage, and bundle optimization stage;

[0039] Feature extraction and matching stage: extract feature points from the input image sequence, perform feature matching, and establish the correspondence between images. The calculation formula is: The trajectory tracker T is based on the input image To estimate the 2D trajectory

[0040] Camera pose estimation stage: By analyzing the geometric relationship of feature points, the camera pose of each image is estimated, including the position and direction of the camera. The calculation formula is: Initialize the camera estimator According to the input image and 2D trajectory To estimate the initial camera projection parameters

[0041] 3D point cloud triangulation stage: Using the correspondence between multiple images, the position of points in 3D space is calculated by triangulation method. The calculation formula is: Converter According to the 2D trajectory and the initial camera projection parameters To estimate the initial point cloud

[0042] Bundle optimization stage: By optimizing the camera posture and the position of the 3D points, the matching relationship of the feature points is satisfied. The calculation formula is: The bundle optimizer BA is based on the 2D trajectory Initial camera projection parameters and the initial point cloud Optimize the camera and 3D points together to improve accuracy.

[0043] Preferably, the trajectory tracker T estimates the confidence of each predicted trajectory point including:

[0044] The arithmetic uncertainty model is used to estimate the confidence of trajectory point prediction, and the covariance matrix is ​​assumed to be a diagonal matrix, that is, the uncertainty in the horizontal and vertical directions is The arithmetic uncertainty model predicts each 2D trajectory point Variance With each 2D trajectory point together form a tightly clustered normal distribution in is the true value of each trajectory point. After training, the confidence metric Proportional to the inverse of the prediction variance.

[0045] Preferably, the trajectory tracking process is divided into two stages: a coarse tracking stage and a fine tracking stage;

[0046] In the coarse tracking stage, the approximate positions of corresponding points are located. In the coarse tracking stage, a deep learning model is used for point tracking. The model accepts a set of images as input and directly outputs reliable point trajectories in all images. Among them, the model uses the advanced technology of the nearest point tracking method to accurately estimate the trajectory of the point without the need for temporal continuity.

[0047] In the fine tracking stage, the initial prediction is further optimized; in the fine tracking stage, the initial prediction is processed using a shallow transformer; specifically, the initial predicted position and visibility are used as input and processed by several self-attention layers and multi-layer perceptrons to obtain more accurate point positions and confidences.

[0048] Preferably, a pair of Transformer networks are used to initialize the camera and point cloud including:

[0049] Camera parameter initialization, The camera estimator is initialized According to the input image and 2D trajectory To estimate the initial camera projection parameters Camera Estimator is a neural network module, represents the ResNet-50 neural network, Indicates that the i C The input image is fed into ResNet-50. is a descriptor;

[0050] Point cloud initialization, Converter According to the 2D trajectory and the initial camera projection parameters To estimate the initial point cloud Descriptors Includes a trajectory tracker feature, as well as the initial point cloud midpoint Position harmonic embedding, the initial point cloud is formed by 3D triangulation via closed-form multi-view direct linear transformation.

[0051] Preferably, applying the bundle optimization layer optimization comprises:

[0052] Get the initial camera pose and 3D point cloud estimate;

[0053] By calculating the error between the reprojected position of each feature point in the image and the actual observed position, a loss function is constructed: Reprojection loss if Points with low visibility or low confidence or projection error are filtered out, and the loss function is the sum of the reprojection errors of all feature points;

[0054] An optimization algorithm is used to iteratively adjust the camera pose and 3D point cloud estimates to minimize the loss function.

[0055] In each iteration, the direction and size of the next parameter update are determined by calculating the partial derivatives of the loss function with respect to the camera pose and the 3D point cloud.

[0056] Through multiple iterations, bundle optimization gradually optimizes the estimated values ​​of the camera pose and the 3D point cloud, and the final optimization result is used to generate a more accurate 3D reconstruction model.

[0057] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any one of the methods described above when executing the program.

[0058] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed, any one of the methods described above is implemented.

[0059] Compared with the prior art, the present invention has the following beneficial effects:

[0060] The invention is fully differentiable: Each component in E2ESfM is fully differentiable, including 2D point trajectories, camera predictors, and triangulators, and can be trained end-to-end. This allows the entire framework to be optimized through back-propagation without the need to manually design and tune each component.

[0061] Reliable pixel-accurate trajectories: E2ESfM uses deep learning techniques to extract reliable pixel-accurate trajectories, eliminating the need for chained pairing in traditional SfM. This improves the accuracy and robustness of matching.

[0062] Global Camera Restoration: Different from traditional methods that register cameras step by step, E2ESfM recovers all camera poses simultaneously based on image and trajectory features instead of registering cameras step by step. This global camera restoration method can better handle complex scenes and large-scale datasets.

[0063] Differentiable bundle optimization: E2ESfM uses a differentiable bundle optimization layer to optimize the camera and triangulated 3D points. Compared with traditional non-differentiable bundle optimization methods, differentiable bundle optimization can be better integrated into the entire deep learning framework and improve the accuracy of reconstruction.

[0064] State-of-the-art performance: By improving the “pairwise matching”, “progressive camera registration”, and “non-differentiable solver” of the SfM framework, E2ESfM achieves state-of-the-art performance on multiple popular datasets. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0066] Figure 1 A schematic diagram of a process flow of a high-quality three-dimensional structure restoration method for urban renewal landscape architecture according to an embodiment of the present invention;

[0067] Figure 2 Schematic diagram of the structure of an electronic device according to an embodiment of the present invention.

[0069] 1010, processor; 1020, memory; 1030, input / output interface; 1040, communication interface; 1050, bus. DETAILED DESCRIPTION

[0070] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0071] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0072] Embodiment 1

[0073] like Figure 1 As shown, the present invention discloses a method for high-quality 3D structure restoration of urban renewal landscape architecture, the method comprising: proposing a new end-to-end SfM framework based on deep learning, namely E2ESfM, extracting 2D trajectories from input images, reconstructing cameras using image and trajectory features, initializing point clouds based on these trajectories and camera parameters, and applying a bundle optimization layer for reconstruction refinement. The entire framework is fully differentiable, that is, each component is differentiable, can be trained in an end-to-end manner, can extract reliable pixel-accurate trajectories, can synchronously restore all camera poses, and can optimize cameras and triangulated 3D points through a differentiable bundle optimization layer.

[0074] E2ESfM eliminates the pairwise matching step in traditional SfM and uses deep learning technology to directly extract reliable pixel-accurate trajectories. The benefit of this is that it can reduce computational complexity and improve matching accuracy and robustness. By using a deep learning model, E2ESfM can directly learn feature representations from images, avoiding some limitations and problems in traditional matching methods. Therefore, E2ESfM can more effectively solve the matching problem in SfM.

[0075] E2ESfM improves upon the step-by-step camera registration in traditional SfM. First, E2ESfM avoids the pair-wise matching step in traditional methods by using deep learning techniques to extract reliable pixel-accurate trajectories directly from images. This reduces computational complexity, makes the entire process simpler and easier to differentiate, and improves the accuracy and robustness of matching. Second, E2ESfM uses a global optimization approach to simultaneously estimate the poses of all cameras and the positions of all points, rather than gradually registering cameras. This makes better use of information from all images and improves the accuracy and completeness of reconstruction. Finally, E2ESfM uses a differentiable bundle adjustment method to optimize the camera pose and point cloud, thus avoiding the limitations of non-differentiable optimization algorithms in traditional methods.

[0076] E2ESfM uses a differentiable approach to replace non-differentiable solvers. This allows the entire SfM process to be trained end-to-end through optimization algorithms such as gradient descent, thereby improving the efficiency and stability of training. Differentiable solvers can also directly handle complex loss functions without manually designing and implementing gradient calculations. In addition, differentiable solvers can automatically learn the hyperparameters of the optimization algorithm, reducing the workload of manual adjustment.

[0077] In this example, the input parameters are:

[0078] 1. Given a series of input images observing a scene E2ESfM can estimate camera parameters Using point cloud Represents the shape of the 3D scene.

[0079] (1) Input image in Where H×W represents the resolution of the input image, 3 represents the three channels of RGB, and N I represents the number of input images, i I Represents the image index.

[0080] (2) Camera projection matrix in Represents a 3×4 matrix consisting of a posture external parameter and a camera intrinsic Composition, among which represents the special Euclidean group, i.e., the rotation and translation of a rigid body in three-dimensional space; i P Represents the camera projection matrix index.

[0081] (3) Point cloud Where N x represents the number of 3D points in the point cloud, Represents the spatial position coordinates of a 3D point.

[0082] 2. A 3D point in a point cloud Can be projected to the i C cameras, resulting in a 2D screen coordinate in λ is a projection coefficient greater than zero, Represents world coordinates The corresponding i C The viewport coordinates of a camera (i.e., local coordinates),

[0083] 3. 3D Points Project the trajectories of all input cameras (i.e. one input image corresponds to one input camera) in is a binary indicator, indicating the i C The i-th camera X The visibility of a 3D point. We use Indicates the i C All tracks on a camera.

[0084] In this embodiment, E2ESfM can be reconstructed using a function f θ To represent SfM, Indicates input N I Image series And output camera projection parameters And scene point cloud

[0085] Reconstruction function f θ is fully differentiable and can be solved by minimizing the training loss To optimize the parameter θ. in They represent the real camera projection parameters, the real trajectory and the real point cloud respectively, S represents the total batch of training, i S Indicates the index of the total training batch.

[0086] E2ESfM will reconstruct the function f θ It can be broken down into four stages (also the four stages of SfM):

[0087] (1) Feature extraction and matching stage

[0088] The trajectory tracker T is based on the input image To estimate the 2D trajectory

[0089] In this stage, feature points are extracted from the input image sequence, and feature matching is performed to establish the correspondence between images.

[0090] (2) Camera pose estimation stage

[0091] Initialize the camera estimator According to the input image and 2D trajectory To estimate the initial camera projection parameters

[0092] In this stage, the camera pose of each image, including the position and orientation of the camera, is estimated by analyzing the geometric relationship of the feature points.

[0093] (3) 3D point cloud triangulation stage

[0094] Converter According to the 2D trajectory and the initial camera projection parameters To estimate the initial point cloud

[0095] In this stage, the correspondence between multiple images is used to calculate the position of points in three-dimensional space through triangulation methods.

[0096] (4) Bundle Optimization Phase

[0097] The bundle optimizer BA is based on the 2D trajectory Initial camera projection parameters and the initial point cloud Optimize the camera and 3D points together to improve accuracy.

[0098] In this stage, the camera posture and the position of the 3D points are optimized so that they can better satisfy the matching relationship of the feature points.

[0099] In this embodiment, 1. Trajectory tracker

[0100] (1) The first stage of SfM estimates the correspondence between paired images, which usually only involves point pair matching, but the linking of pair matching is still a manual process (the first stage of SfM is to estimate the correspondence between paired images and then link them into multi-image trajectories. In existing methods, only point pair matching benefits from the learning component, while the links between point pairs are still manually designed, so the pairwise correspondences need to be manually connected). E2ESfM greatly simplifies the correspondence of SfM trajectories through the deep feedforward trajectory function, making the entire trajectory process available for learning components.

[0101] (2) Reliable pixel-accurate trajectories: E2ESfM uses the latest deep learning methods to directly extract reliable pixel-accurate trajectories. This trajectory tracking method can process all frames in an unordered image sequence without relying on the temporal relationship between frames. Compared with the traditional pair-matching-based method in SfM, the trajectory module of E2ESfM is simpler and more accurate.

[0102] (3) Simplified correspondence estimation: In traditional SfM, correspondence is usually constructed through paired image matching. The trajectory module of E2ESfM directly extracts reliable trajectories, avoiding the pairwise matching process. This simplifies the correspondence estimation steps in traditional SfM and reduces errors in the matching process.

[0103] (4) Accurate sub-pixel trajectory: In order to obtain more accurate trajectory results, the trajectory module of E2ESfM adopts a coarse-to-fine trajectory tracking mechanism. By refining the coarse tracking, sub-pixel accuracy can be achieved, which improves the accuracy of trajectory tracking.

[0104] 2. Trajectory Tracker Architecture

[0105] (1) In the i C images In the example, N is given T Query points Here Represents the true value, that is, the manually marked query point.

[0106] (2) In the feature map of the image generated by the residual network, the descriptor First, a cross-attention mechanism is performed between the global image features and the track descriptors, and identifiers are generated for each image. The corresponding relationship is used as input, and the updated trajectory is obtained through bilinear sampling (8-point algorithm). And iteratively input into the cross attention mechanism.

[0107] (3) Compare each descriptor with all N I Input images (Note: N I The resolution of the input image can be different) and the feature map is associated with the identifier To represent, where C represents the i C The number of descriptors in an image.

[0108] (4) Identifier Input into the Transformer neural network to get the trajectory where i T Indicates the index of the track.

[0109] (5) Our trajectory tracker does not assume temporal continuity between input images (i.e., it does not assume that the input is a frame image of a video), which has greater flexibility. In addition, the predictions between each trajectory are relatively independent, which allows more points to be tracked during training, thereby increasing the density of the reconstructed point cloud.

[0110] (6) Our trajectory tracker is fully differentiable, which enables back-propagation of the gradients of the training loss to the trajectory tracker parameters through the entire framework, strengthening the synergy between the trajectory and subsequent stages.

[0111] 3. Trajectory confidence

[0112] (1) In SfM, it is crucial to filter out anomalous correspondences since these outliers can negatively impact the reconstruction stage. To this end, we enhance the trajectory tracker to estimate the confidence of each predicted trajectory point.

[0113] (2) We use an aleatoric uncertainty model to estimate the confidence of trajectory point predictions and assume that the covariance matrix is ​​a diagonal matrix, that is, the uncertainties in the horizontal and vertical directions are The arithmetic uncertainty model predicts each 2D trajectory point Variance With each 2D trajectory point together form a tightly clustered normal distribution in is the true value of each trajectory point. After training, the confidence metric Proportional to the inverse of the prediction variance.

[0114] In this embodiment, the trajectory tracking process is a step-by-step trajectory tracking strategy:

[0115] 1. The trajectory tracking process is divided into two stages: coarse tracking and fine tracking.

[0116] 2. Because SfM requires sub-pixel precise correspondence, through a step-by-step trajectory tracking strategy (i.e., a coarse-to-fine tracking strategy), E2ESfM can improve the density and reconstruction effect of the point cloud while maintaining high accuracy.

[0117] 3. In the coarse tracking stage, the approximate positions of the corresponding points are located. In the coarse tracking stage, a deep learning model is used for point tracking, which accepts a set of images as input and directly outputs reliable point trajectories in all images. This model uses the advanced technology of the nearest point tracking method to accurately estimate the trajectory of the point without the need for temporal continuity.

[0118] 4. Fine tracking stage, the initial prediction is further optimized. In the fine tracking stage, a shallow transformer is used to process the initial prediction. Specifically, the initial predicted position and visibility are used as input, and processed by several self-attention layers and multi-layer perceptrons to obtain more accurate point positions and confidences.

[0119] 5. In the coarse tracking process, the feature map generated by the entire residual network is bilinearly interpolated to generate a descriptor. In the fine tracking process, the feature map generated by the residual network is first cropped to a resolution of Then bilinear interpolation is performed to generate the descriptor.

[0120] In this example, the camera and point cloud are initialized:

[0121] 1. The classic SfM framework relies on incremental iterations, initializing with image pairs with rich correspondences, then gradually registering new input images, expanding the point cloud, and performing joint optimization. Although this process has become more robust and accurate after more than a decade of development, the complexity has greatly increased, and the incremental SfM process is not differentiable, which makes end-to-end learning from registered data impossible.

[0122] 2. In order to simplify the process, remove incremental iterations, and make the process differentiable, we use a pair of Transformer networks to initialize the camera and point cloud.

[0123] 3. Camera parameter initialization, The camera estimator is initialized According to the input image and 2D trajectory To estimate the initial camera projection parameters Camera Estimator is a neural network module, φ represents the ResNet-50 neural network, Indicates that the i C The input image is fed into ResNet-50. is a descriptor.

[0124] 4. Converter According to the 2D trajectory and the initial camera projection parameters To estimate the initial point cloud Descriptors Includes a trajectory tracker feature, as well as the initial point cloud midpoint The initial point cloud is formed by three-dimensional triangulation through closed multi-view direct linear transform (DLT). Specifically, given the initial camera parameters and the corresponding points in the image, the approximate position of the initial point cloud is calculated by the DLT algorithm. At the same time, the distance between the camera ray and the initial point cloud and the nearest point on the camera ray are also calculated. In this way, a shape of N T×3 The initial point cloud, where N T represents the number of points, and 3 represents the three-dimensional coordinates of each point. These vectors are connected to obtain a shape of N T ×N I×7 tensor and embed it into a 256-dimensional space through positional encoding to obtain a shape of N T ×N I × 256 tensor. These embedding vectors are further used in subsequent processing to predict the final point cloud.

[0125] In this embodiment, bundle optimization is an optimization algorithm that extracts the best 3D model and camera parameters from visual reconstruction, which can improve the accuracy and stability of the reconstruction results. It optimizes the estimated values ​​of the camera pose and the 3D point cloud by minimizing the reprojection error so that they better fit the observed image feature points.

[0126] In bundle optimization, we first need to provide the initial camera pose and the estimated value of the 3D point cloud. Then, we construct a loss function by calculating the error between the reprojected position of each feature point in the image and the actual observed position: Reprojection loss if Points with low visibility or low confidence are filtered out, and points with large projection errors are also filtered out. This loss function is the sum of the reprojection errors of all feature points. Next, an optimization algorithm (such as the Levenberg-Marquardt algorithm) is used to iteratively adjust the estimated values ​​of the camera pose and the 3D point cloud to minimize the loss function. In each iteration, the direction and size of the next parameter update are determined by calculating the partial derivative (i.e., gradient) of the loss function with respect to the camera pose and the 3D point cloud. Through multiple iterations, bundle optimization gradually optimizes the estimated values ​​of the camera pose and the 3D point cloud so that they more accurately reflect the real scene. The final optimization result can be used to generate a more accurate 3D reconstruction model.

[0127] In this embodiment, the training loss function

[0128] where λ1, λ2, and λ3 are the weight hyperparameters of the loss function, and the symbol |·| ∈ is the Pseudo-Huber loss function with threshold ∈. Symbol Indicates the true value or the marked value, the symbol represents an initialized and bundle-optimized representation, e.g. Represents a real 3D point. ‖@‖ ∈ represents the Huber loss function with threshold ∈. Represents a mean value of The variance is The normal distribution of Represents the true value The likelihood estimate of .

[0129] Embodiment 2

[0130] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a method for restoring the three-dimensional structure of a high-quality urban renewal landscape building as described in any of the above embodiments is implemented.

[0131] Figure 2 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.

[0132] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0133] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0134] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0135] The communication interface 1040 is used to connect a communication module (not shown in the figure) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB (Universal Serial Bus), network cable, etc.), or through a wireless mode (such as mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).

[0136] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0137] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.

[0138] The system of the above embodiment is used to implement a corresponding three-dimensional structure restoration method of a high-quality urban renewal landscape building in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0139] Embodiment 3

[0140] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute a three-dimensional structure restoration method for a high-quality urban renewal landscape building as described in any of the above embodiments.

[0141] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0142] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute a three-dimensional structure restoration method for high-quality urban renewal landscape buildings as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0143] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.

[0144] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the known power / ground connections to the integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it is apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with changes in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0145] Although the present disclosure has been described in conjunction with specific embodiments of the present disclosure, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0146] Therefore, the units of each example described in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present application.

[0147] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.

Claims

1. A high-quality three-dimensional structure restoration method for urban renewal landscape architecture, characterized in that: The method comprises: Build an end-to-end SfM framework E2ESfM based on deep learning; Using E2ESfM, 2D trajectories are extracted from the input image; The camera is reconstructed using image and trajectory features, a point cloud is initialized based on the 2D trajectory and camera parameters, and a bundle optimization layer is applied to optimize the camera and triangulated 3D points.

2. The method according to claim 1, characterized in that E2ESfM uses a reconstruction function to represent SfM; The reconstruction function is fully differentiable and its parameters are optimized by minimizing the training loss.

3. The method according to claim 2, characterized in that Input Image in Where H×W represents the resolution of the input image, 3 represents the three channels of RGB, and N I represents the number of input images, i I Represents the image index; Camera projection matrix in Represents a 3×4 matrix consisting of a posture external parameter and a camera intrinsic Composition, among which represents the special Euclidean group, i.e., the rotation and translation of a rigid body in three-dimensional space; i P Represents the camera projection matrix index; Scene point cloud Where N x represents the number of 3D points in the point cloud, Represents the spatial position coordinates of a 3D point.

4. The method according to claim 3, characterized in that E2ESfM decomposes the reconstruction function into four stages: feature extraction and matching stage, camera pose estimation stage, 3D point cloud triangulation stage, and bundle optimization stage; Feature extraction and matching stage: extract feature points from the input image sequence, perform feature matching, and establish the correspondence between images. The calculation formula is: The trajectory tracker T is based on the input image To estimate the 2D trajectory Camera pose estimation stage: By analyzing the geometric relationship of feature points, the camera pose of each image is estimated, including the position and direction of the camera. The calculation formula is: Initialize the camera estimator According to the input image and 2D trajectory To estimate the initial camera projection parameters 3D point cloud triangulation stage: Using the correspondence between multiple images, the position of points in 3D space is calculated by triangulation method. The calculation formula is: Converter According to the 2D trajectory and the initial camera projection parameters To estimate the initial point cloud Bundle optimization stage: By optimizing the camera posture and the position of the 3D points, the matching relationship of the feature points is satisfied. The calculation formula is: The bundle optimizer BA is based on the 2D trajectory Initial camera projection parameters and the initial point cloud Optimize the camera and 3D points together to improve accuracy.

5. The method according to claim 4, characterized in that The trajectory tracker T estimates the confidence of each predicted trajectory point including: The arithmetic uncertainty model is used to estimate the confidence of trajectory point prediction, and the covariance matrix is ​​assumed to be a diagonal matrix, that is, the uncertainty in the horizontal and vertical directions is The arithmetic uncertainty model predicts each 2D trajectory point Variance With each 2D trajectory point together form a tightly clustered normal distribution in is the true value of each trajectory point. After training, the confidence metric Proportional to the inverse of the prediction variance.

6. The method according to claim 5, characterized in that The trajectory tracking process is divided into two stages: coarse tracking stage and fine tracking stage; In the coarse tracking stage, the approximate positions of corresponding points are located. In the coarse tracking stage, a deep learning model is used for point tracking. The model accepts a set of images as input and directly outputs reliable point trajectories in all images. Among them, the model uses the advanced technology of the nearest point tracking method to accurately estimate the trajectory of the point without the need for temporal continuity. In the fine tracking stage, the initial prediction is further optimized; in the fine tracking stage, the initial prediction is processed using a shallow transformer; specifically, the initial predicted position and visibility are used as input and processed by several self-attention layers and multi-layer perceptrons to obtain more accurate point positions and confidences.

7. The method according to claim 6, characterized in that A pair of Transformer networks are used to initialize the camera and point cloud including: Camera parameter initialization, The camera estimator is initialized According to the input image and 2D trajectory To estimate the initial camera projection parameters Camera Estimator is a neural network module, φ represents the ResNet-50 neural network, Indicates that the i C The input image is fed into ResNet-50. is a descriptor; Point cloud initialization, Converter According to the 2D trajectory and the initial camera projection parameters To estimate the initial point cloud Descriptors Includes a trajectory tracker feature, as well as the initial point cloud midpoint Position harmonic embedding, the initial point cloud is formed by 3D triangulation via closed-form multi-view direct linear transformation.

8. The method according to claim 7, characterized in that Application bundle optimization layer optimization includes: Get the initial camera pose and 3D point cloud estimate; By calculating the error between the reprojected position of each feature point in the image and the actual observed position, a loss function is constructed: Reprojection loss if Points with low visibility or low confidence or projection error are filtered out, and the loss function is the sum of the reprojection errors of all feature points; An optimization algorithm is used to iteratively adjust the camera pose and 3D point cloud estimates to minimize the loss function. In each iteration, the direction and size of the next parameter update are determined by calculating the partial derivatives of the loss function with respect to the camera pose and the 3D point cloud. Through multiple iterations, bundle optimization gradually optimizes the estimated values ​​of the camera pose and the 3D point cloud, and the final optimization result is used to generate a more accurate 3D reconstruction model.

9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Urban live-action three-dimensional modeling method based on air-ground image consistency feature learning

    CN117830522A

  • Three-dimensional reconstruction method of arbitrary point tracking network based on dynamic perception

    CN118212364A