Three-dimensional scene reconstruction method and system based on improved SuperPoint feature points

By improving the three-dimensional scene reconstruction method of SuperPoint feature points, using SuperPoint-MSFF network for feature point detection and ICP algorithm for point cloud registration, the problem of three-dimensional reconstruction being susceptible to environmental impact and mismatch is solved, and the accuracy and robustness of three-dimensional reconstruction are improved.

CN120219610APending Publication Date: 2025-06-27HUZHOU ELECTRIC POWER SUPPLY CO OF STATE GRID ZHEJIANG ELECTRIC POWER CO LTD +2

Patent Information

Application Number
CN202510182879.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing three-dimensional reconstruction methods are susceptible to environmental influences and are prone to mismatches, resulting in large errors in matching results, affecting the accuracy and consistency of three-dimensional reconstruction.

Method used

The three-dimensional scene reconstruction method of improved SuperPoint feature points is adopted, feature point detection is used using the SuperPoint-MSFF network to generate SuperPoint-MSFF feature descriptors, and the point cloud is accurately registered through the improved ICP iterative nearest point algorithm.

Benefits of technology

Effectively capture the global context information and local detailed information of the image, improve the robustness of feature point matching and the overall performance of three-dimensional reconstruction, and improve the accuracy of three-dimensional reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219610A_ABST
    Figure CN120219610A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional scene reconstruction method and system based on improved SuperPoint feature points, and solves the problems of low precision and consistency of three-dimensional scene reconstruction in the prior art, and the method comprises the steps: collecting image information, and obtaining an RGB image and a depth image; a SuperPoint-MSFF network is utilized to carry out feature point detection on the RGB image of the current frame and the RGB image of the next frame, and corresponding SuperPoint-MSFF feature descriptors are generated; feature point matching is carried out, camera attitude estimation is carried out, and coarse registration is carried out on the source point cloud; based on the current frame depth map and the next frame depth map, generating a dense point set containing each pixel space position; and using an improved ICP iterative nearest point algorithm to carry out fine registration of the point cloud based on the dense point set to obtain a reconstructed three-dimensional point cloud. The global context information and the local detailed information of the image are effectively captured, and the feature point detection process is further optimized, so that the robustness of feature point matching and the overall performance of three-dimensional reconstruction are improved, and the precision of three-dimensional reconstruction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly relates to a three-dimensional scene reconstruction method and system based on improved SuperPoint feature points. Background Art

[0002] Currently, three-dimensional reconstruction algorithms are widely used in various fields of computer vision and are the basis for emerging technologies such as mobile robots, digital twins, and complex environment modeling. RGB-D cameras can directly obtain the position information of objects at video frame rates and have relatively low costs. At the same time, they also have significant advantages such as real-time performance, high precision, easy integration, and safety. However, in the three-dimensional reconstruction process based on RGB-D cameras, there is a problem of insufficient robustness in the matching of RGB image feature point pairs between adjacent frames.

[0003] On December 26, 2023, the Chinese Patent Office published patent CN118071918A, a three-dimensional reconstruction method and system for non-rigid dynamic targets based on RGB-D data, which acquires an RGB-D image sequence including RGB data and depth information; calculates the target mask of the reconstructed object according to the RGB data and the CNN model; predicts the three-dimensional coordinates of key points through the depth information and the Transfomer model; and upsamples the three-dimensional coordinates of key points according to the target mask and the three-dimensional coordinates of key points to achieve real-time dense three-dimensional reconstruction of the target. However, algorithms based on feature point matching are easily affected by illumination changes and noise problems in complex environments, thereby affecting the accuracy and consistency of three-dimensional reconstruction. In particular, sparse key point matching may produce local mis-matches, and these mis-matches may spread and cause error accumulation during the global optimization process. The accumulated errors will cause the positions and postures of subsequent frames to deviate further and further from the true values, thus affecting the reconstruction quality of the entire scene. Summary of the Invention

[0004] The purpose of the present invention is to solve the problems in the prior art that the three-dimensional reconstruction method is easily affected by the environment, prone to mis-matches, and resulting in large errors in the matching results. A three-dimensional scene reconstruction method and system based on improved SuperPoint feature points are provided, which can effectively capture the global context information and local detailed information of the image, further optimize the feature point detection process, thereby improving the robustness of feature point matching and the overall performance of three-dimensional reconstruction, and improving the accuracy of three-dimensional reconstruction.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A three-dimensional scene reconstruction method based on improved SuperPoint feature points, comprising the following steps: S1: Collect image information to obtain an RGB image and a depth image; S2: Use the SuperPoint-MSFF network to detect feature points in the current-frame RGB image and the next-frame RGB image respectively, and generate corresponding SuperPoint-MSFF feature descriptors; S3: Perform feature point matching and camera pose estimation, and conduct rough registration on the source point cloud; S4: Based on the current-frame depth image and the next-frame depth image, generate a dense point set containing the spatial position of each pixel; S5: Use the improved ICP iterative closest point algorithm to perform fine registration of the point cloud based on the dense point set, and obtain the reconstructed three-dimensional point cloud.

[0006] The present invention uses the SuperPoint-MSFF network for feature point matching, effectively captures the global context information and local detailed information of the image, further optimizes the feature point detection process, thereby improving the robustness of feature point matching and the overall performance of three-dimensional reconstruction.

[0007] Preferably, S2 includes: extracting the deep features of the RGB image to obtain a feature map; adjusting the size of the feature map to obtain a feature response map, and generating a feature probability map according to the feature response map; learning a semi-dense descriptor from the feature map to obtain a complete descriptor, and then performing bicubic polynomial calculation and L2 normalization on the descriptor to obtain a unit-length descriptor.

[0008] Preferably, when extracting the deep features of the RGB image, it includes three branches. The first branch retains the original hierarchical information; the second branch extracts high-level features; the third branch changes the size of the input feature map to half of the original, and then restores the feature map to the original size through a convolutional layer; fuse the output results of the three branches to obtain a feature map.

[0009] Preferably, S4 includes: defining a source point cloud and a target point cloud, for each point in the source point cloud, find its nearest neighbor point in the target point cloud, and establish the corresponding relationship between the two sets of point clouds; iterate the nearest point pairs between the two sets of point clouds, and minimize the distance error between the two sets of point clouds to complete the fine registration.

[0010] Preferably, S3 includes: estimating the camera pose by matching feature points and using the PnP algorithm, obtaining the position and orientation of the camera in the three-dimensional space during the shooting process, and obtaining the rotation matrix and translation vector of the camera in the three-dimensional space during the shooting process; transform the source point cloud through this rigid body transformation matrix to complete the rough registration of the point cloud.

[0011] Preferably, when performing feature point matching: obtain the Euclidean distance ratio between feature points, compare the Euclidean distance ratio with a preset matching threshold. If the Euclidean distance ratio is less than the matching threshold, it is a correct matching point; otherwise, it is an incorrect matching point.

[0012] Preferably, traverse all pairs of matching points, calculate the number of inliers, compare the number of inliers with a preset inlier number threshold. If the number of inliers is less than the inlier number threshold, continue the iteration; otherwise, end the iteration, output the current set of inliers, and complete the optimization of the matching point pairs.

[0013] A three-dimensional scene reconstruction system based on improved SuperPoint feature points, comprising: An image acquisition module that acquires image information and obtains the RGB image and depth map of the current frame and the next frame; A feature point detection module that automatically extracts feature points in the image, generates high-dimensional descriptors for each feature point, and uses the multi-branch structure of the SuperPoint-MSFF network to learn features at different scales; A pose estimation module that performs feature point matching and uses the PnP algorithm to estimate the pose of the image acquisition module, and obtains the position and orientation of the image acquisition module in the three-dimensional space during the shooting process; A three-dimensional point cloud reconstruction module that performs rough registration of the point cloud according to the output result of the pose estimation module, and at the same time uses an improved iterative closest point algorithm to perform fine registration of the point cloud.

[0014] Preferably, the feature point detection module includes: A shared encoder, including a convolutional layer and a multi-scale feature fusion module MSFF, extracts deep features, and inputs the obtained feature maps into a feature point extraction decoder and a descriptor calculation decoder respectively; A feature point extraction decoder that converts the feature map into a feature response map through convolution, and then obtains a feature probability map through a Softmax layer and a Reshape layer; A descriptor calculation decoder that learns semi-dense descriptors from the input feature map to obtain complete descriptors, and then obtains unit-length descriptors through bicubic polynomial calculation and L2 normalization.

[0015] Preferably, the objective function with the lowest cable connection cost includes: taking the sum of the cable laying cost and the sea occupation cost, the cable body cost and the cable loss cost as the cable connection cost, and establishing an optimal cable type matrix.

[0016] Preferably, the multi-scale feature fusion module MSFF includes three branches. The first branch includes an activation function layer, a normalization layer, and a convolutional layer. The second branch includes two 1×1 convolutional layers and one 3×3 convolutional layer, and a normalization layer and a Relu activation layer are provided before each convolutional layer. The third branch includes two 3×3 convolutional layers, and a normalization layer and an activation layer are provided before each convolutional layer. A pooling layer is also provided before the first convolutional layer, and a sampling layer is provided after the second convolutional layer.

[0017] Therefore, the present invention has the following beneficial effects: Using the SuperPoint-MSFF network for feature point detection, through the multi-branch structure, features at different scales can be learned, so as to effectively capture the global context information and local detailed information of the image, further optimize the feature point detection process, and thus improve the robustness of feature point matching and the overall performance of 3D reconstruction. Description of the Drawings

[0018] Figure 1 It is the overall step flow chart of the 3D scene reconstruction method based on the improved SuperPoint feature points in Embodiment 1.

[0019] Figure 2 It is the schematic diagram of the SuperPoint-MSFF network structure in Embodiment 2.

[0020] Figure 3 It is the schematic diagram of the architecture of the multi-scale feature fusion module in Embodiment 2. Detailed Embodiments

[0021] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments: Embodiment 1: This embodiment provides a 3D scene reconstruction method based on improved SuperPoint feature points. As Figure 1 shown, its operation process is as follows: Step 1, collect image information to obtain RGB images and depth maps; Step 2, use the SuperPoint-MSFF network to detect feature points on the current frame RGB image and the next frame RGB image respectively, and generate corresponding SuperPoint-MSFF feature descriptors; Step 3, perform feature point matching and camera pose estimation, and perform rough registration on the source point cloud; Step 4, based on the current frame depth map and the next frame depth map, generate a dense point set including the spatial position of each pixel; Step 5, use the improved ICP iterative closest point algorithm to perform fine registration of the point cloud based on the dense point set to obtain the reconstructed 3D point cloud.

[0022] The 3D scene reconstruction method based on improved SuperPoint feature points provided in this embodiment uses the SuperPoint-MSFF network for feature point detection. Through a multi-branch structure, it can learn features at different scales, effectively capture the global context information and local detailed information of the image, further optimize the feature point detection process, and thus improve the robustness of feature point matching and the overall performance of 3D reconstruction.

[0023] Next, through specific examples and specific application scenarios, the technical solutions and technical effects of the present invention will be further described. The following examples are explanations of the present invention, and the present invention is not limited to the following examples.

[0024] Specifically, as Figure 1 shown, a 3D scene reconstruction method based on improved SuperPoint feature points specifically includes the following steps: The first step: Collect image information to obtain RGB images and depth maps.

[0025] In this embodiment, an RGB-D camera is used to collect image information. The RGB-D camera can not only capture the color information (RGB) of the scene, but also obtain the depth information (D) corresponding to each pixel. At the same time, it can directly obtain the position information of the object at the video frame rate, and the cost is relatively low. At the same time, it also has significant advantages such as real-time, high-precision, easy integration, and safety.

[0026] The second step: Use the SuperPoint-MSFF network to perform feature point detection on the current frame RGB image and the next frame RGB image respectively, and generate corresponding SuperPoint-MSFF feature descriptors.

[0027] SuperPoint is a feature point detection and description method based on a convolutional neural network. It can automatically extract feature points in an image and generate a high-dimensional descriptor for each feature point for subsequent matching processes.

[0028] In order to better capture the global context information and local detailed information of the image and further optimize the feature point detection process, in this embodiment, a multi-scale feature fusion module (Multi scale featurefusion module) is integrated into the SuperPoint network, and the SuperPoint-MSFF network is used for feature point detection.

[0029] The detection process specifically includes: Step (1): Extract the deep features of the RGB image to obtain a feature map.

[0030] When extracting the deep features of the RGB image, the multi-branch structure within the SuperPoint-MSFF network is utilized to learn features at different scales, thereby better capturing the global context information and local detailed information of the image, and further optimizing the feature point detection process. At the same time, when extracting the deep features of the RGB image, features are extracted by reducing the image size, thereby reducing the computational amount.

[0031] Specifically, in this embodiment, three branches are adopted. In the first branch, when the number of input channels is equal to the number of output channels, it is an identity mapping, which can retain the original hierarchical information. The second branch is used to extract higher-level features; the third branch is used to reduce the input RGB image to half of its original size and then restore the RGB image to its original size. Finally, the output results of the three branches are subjected to feature fusion to obtain the processed feature map.

[0032] The processed feature map enters two branches for processing respectively. One branch is used to extract feature points, and the other branch is used to extract feature descriptors.

[0033] Step (2): Adjust the size of the feature map to obtain a feature response map, and generate a feature probability map based on the feature response map.

[0034] The input feature map becomes a feature response map with a size of H / 8×W / 8×65 after convolution, and then a feature probability map of H×W×1 is obtained after passing through the Softmax layer and Reshape layer within the SuperPoint-MSFF network, that is, the probability situation of whether each pixel point is a feature point.

[0035] Step (3): Learn semi-dense descriptors from the input feature map to obtain complete descriptors, and then perform bicubic polynomial calculation and L2 normalization on the descriptors to obtain unit-length descriptors, thus obtaining feature descriptors.

[0036] The third step: Perform feature point matching, perform camera pose estimation, and perform rough registration on the source point cloud.

[0037] In order to be able to perform three-dimensional reconstruction in a real and complex scene, this embodiment adopts a pose estimation method based on feature matching, and only uses a small number of but key matching feature points for pose calculation, successfully reducing the computational amount. It not only provides accurate pose information, but also significantly enhances the robustness of the system. Compared with other three-dimensional reconstruction methods, it has better anti-noise performance. In a complex environment, the enhancement of this anti-interference ability is crucial for coping with various noises and interferences that may exist in the scene.

[0038] Specifically, in this embodiment, the PROSAC (Progressive Sampling with Approximate Clustering) algorithm is used for feature point matching. PROSAC is an iterative method that estimates model parameters by randomly sampling in the subset of data points with the highest evaluation function values, calculates the number of inliers to evaluate the rationality of the model, and filters out feature point pairs with incorrect matches. Compared with the RANSAC (Random Sample Consensus) algorithm, the PROSAC algorithm reduces the number of model iterations, thereby saving computing resources and improving the operation speed. In other embodiments, other algorithms (such as the RANSAC algorithm) can also be used for feature point matching.

[0039] The PROSAC algorithm introduces a quality function and the probability that a data point is an inlier, and then calculates the Euclidean distance ratio between feature points. When calculating the Euclidean distance ratio between feature points, first calculate the Euclidean distances of the nearest neighbor and the second-nearest neighbor matching points through the feature descriptor, and then use the ratio of the two Euclidean distances as a quantification representing the quality function. Therefore, the smaller the Euclidean distance ratio, the greater the probability that this feature point is an inlier.

[0040] Compare the Euclidean distance ratio with the set matching threshold. If the Euclidean distance ratio is less than the set matching threshold, it is a correct matching point; otherwise, it is an incorrect matching point. Traverse all matching point pairs, calculate the number of inliers, and compare the number of inliers with the preset inlier number threshold. If the number of inliers is less than the inlier number threshold, continue the iteration; otherwise, output the current inlier set to complete the optimization of the matching point pairs.

[0041] After completing the feature point matching, use the PnP algorithm to estimate the camera pose, obtain the position and orientation of the camera in the three-dimensional space during the shooting process, and obtain the rotation matrix and translation vector of the camera in the three-dimensional space during the shooting process; transform the source point cloud through this rigid body transformation matrix to complete the rough registration of the point cloud.

[0042] Specifically, use the PnP algorithm to estimate the camera pose and obtain the position and orientation of the camera in the three-dimensional space during the shooting process. The core of the PnP algorithm is to estimate the camera pose, that is, the rotation matrix R and the translation vector t, through a known set of 3D space points and their projection points on the 2D image plane. EPnP is currently the most effective solution method. It transforms the 3D points in the world coordinate system into 3D points in the camera coordinate system and solves the rotation matrix R and the translation vector t by converting the 3D-2D problem into a 3D-3D problem.

[0043] Let the 3D matching point pairs be p and p', then: p = {p1,..., p n}, p' = {p'1,..., p'1}.

[0044] For an unknown rotation matrix R and translation vector t, we have:

[0045] Define the error term for the i-th pair of points: e i = p i - (Rp’ + t) e i = p i - (Rp’ i + t).

[0046] Calculate the sum of squared errors for the rotation matrix R and translation vector t:

[0047] Solve the above equation to obtain the optimal rotation matrix R and translation vector t: R = UV T , t = p - Rp’.

[0048] Where U and V are diagonal matrices.

[0049] A rigid body transformation only includes two functions: rotation and translation. That is, a rigid body can change its state in space through rotation and translation without deformation. The rigid body transformation matrix is usually a 4x4 matrix, where the rotation part is a 3x3 orthogonal matrix and the translation part is a 3D vector. In this embodiment, the rotation matrix is the orthogonal matrix of the rigid body transformation, and the translation vector is the 3D vector of the rigid body transformation.

[0050] Step 4: Based on the current frame depth map and the next frame depth map, generate a dense point set containing the spatial position of each pixel.

[0051] Using the depth image captured by an RGB-D camera, a dense point set containing the spatial position of each pixel can be generated.

[0052] Step 5: Use the improved ICP (Iterative Closest Point) algorithm to perform fine registration of the point cloud based on the dense point set, and obtain the reconstructed 3D point cloud.

[0053] Then, adopt the improved ICP (Iterative Closest Point) algorithm to register the point sets from multiple perspectives, iteratively find the closest point pairs between the two sets of point clouds, and minimize the distance error between them, so as to accurately align and fuse these point clouds.

[0054] The specific process is as follows: Define the source point cloud P and the target point cloud Q: P = {p i , i = 1, 2,..., N p}, Q = {q i , i = 1, 2,..., N q}.

[0055] For each point p i find its nearest neighbor point q in the target point cloud Q i By this step, a correspondence is established, which determines how each point in point cloud P aligns with points in point cloud Q under the current transformation.

[0056] Calculate the error function:

[0057] where p i and q i are corresponding points, representing the point cloud data in the source point cloud and the target point cloud.

[0058] Transform point cloud P by (R, t) to obtain P’, and repeat the above steps until the condition for stopping iteration is met.

[0059] Through the cooperation of coarse registration and fine registration, the registration of the source point cloud and the target point cloud is completed, so that the spatial positions of the two point clouds are best matched, and an accurate pose transformation matrix R and translation vector t are obtained. Furthermore, these point clouds are aligned and fused to obtain a three-dimensional reconstruction result.

[0060] Embodiment 2: This embodiment provides a three-dimensional scene reconstruction system based on improved SuperPoint feature points.

[0061] With the improvement of sensor technology and computing power, the KinectFusion algorithm first realized real-time rigid body reconstruction based on inexpensive consumer cameras, which greatly promoted the commercial process of real-time dense three-dimensional reconstruction. The BundleFusion algorithm is currently a relatively good method for dense three-dimensional reconstruction based on RGB-D cameras. The input color images and depth images first need to perform correspondence matching between frames. In terms of matching, a sparse-then-dense parallel global optimization method is used, that is, first use sparse SIFT feature points for relatively rough registration, and then use dense geometric and photometric continuity for more detailed registration, thus realizing a high-precision three-dimensional reconstruction system.

[0062] Although existing three-dimensional reconstruction methods have made significant progress in theory and application, algorithms based on feature point matching still have bottlenecks in terms of processing speed and hardware requirements. Therefore, a three-dimensional scene reconstruction system based on improved SuperPoint feature points provided by this embodiment effectively captures the global context information and local detailed information of the image, further optimizes the feature point detection process, thereby improving the robustness of feature point matching and the overall performance of three-dimensional reconstruction.

[0063] Specifically, a three-dimensional scene reconstruction system based on improved SuperPoint feature points provided in this embodiment includes an image acquisition module that acquires image information and obtains the RGB image and depth image of the current frame and the next frame. In this embodiment, the image acquisition module is an RGB-D camera. The RGB-D camera can not only capture the color information (RGB) of the scene, but also obtain the depth information (D) corresponding to each pixel, greatly enhancing the accuracy and effect of three-dimensional reconstruction.

[0064] The feature point detection module uses a multi-branch structure to learn features at different scales, automatically extracts feature points in the image, and generates high-dimensional descriptors for each feature point.

[0065] The pose estimation module uses the PnP algorithm to estimate the pose of the RGB-D camera, including the rotation matrix R and translation vector t in the three-dimensional space during the camera shooting process, and obtains the position and orientation of the image acquisition module in the three-dimensional space during the shooting process.

[0066] The three-dimensional point cloud reconstruction module, according to the output result of the pose estimation module, performs rough registration of the point cloud by transforming the source point cloud through this rigid body transformation matrix. At the same time, it uses an improved iterative closest point algorithm to iteratively find the closest point pairs between two groups of point clouds and minimize the distance error between them, so as to accurately align and fuse these point clouds and perform fine registration of the point cloud.

[0067] Specifically, the feature point detection module includes an improved SuperPoint-MSFF network. The structure diagram of the SuperPoint-MSFF network is as Figure 2 shown, including a shared encoder, a feature point extraction decoder, and a descriptor calculation decoder. The shared encoder includes a convolutional layer and a multi-scale feature fusion module; the feature point extraction decoder includes a convolutional layer, a Softmax layer, and a Reshape layer. The convolutional layer is connected to the shared encoder, the Softmax layer is connected to the convolutional layer, and the Reshape layer is connected to the Softmax layer; the descriptor calculation decoder includes a convolutional layer, a bicubic polynomial calculation module, and an L2 normalization module. The convolutional layer is connected to the shared decoder, the bicubic polynomial calculation module is connected to the convolutional layer, and the L2 normalization module is connected to the bicubic polynomial calculation module.

[0068] The working process of the improved SuperPoint-MSFF network structure is as follows: The SuperPoint-MSFF network inputs an RGB image and extracts deep features through the shared encoder. The shared encoder includes a convolutional layer, a pooling layer, a non-linear activation function, and a multi-scale feature fusion (MSFF) module to reduce the computational amount through dimensionality reduction. The feature maps extracted by the shared encoder are respectively input into the feature point extraction decoder and the descriptor calculation decoder.

[0069] As Figure 3 shown, the multi-scale feature fusion (MSFF) module mainly includes three branches. The first branch includes an activation function layer, a normalization layer, and a 1×1 convolutional layer. The normalization layer is connected to the activation function layer, and the activation function layer is connected to the convolutional layer. When the resolution of the feature map decreases or the dimension of the feature channels increases, it can effectively avoid large variance responses.

[0070] In the second branch, there are two convolutional layers with a kernel size of 1×1 and a 3×3 convolutional layer. In addition, there is also a normalization layer and an activation function layer respectively. That is, the first normalization layer is connected to the first activation function layer, the first activation function layer is connected to the first 1×1 convolutional layer; the 1×1 convolutional layer is connected to the second normalization layer, the second normalization layer is connected to the second activation function layer, the second activation function layer is connected to the 3×3 convolutional layer, the 3×3 convolutional layer is connected to the third normalization layer, the third normalization layer is connected to the third activation function layer, and the third activation function layer is connected to the second 1×1 convolutional layer.

[0071] The third branch includes two convolutional layers with a kernel size of 3×3. There are a normalization layer and an activation layer before both of these convolutional layers, and a pooling layer is added before the first convolutional layer, and a Spatial Dropout layer is added before the second convolutional layer. An upsampling layer is added after the second convolutional layer. That is, the first normalization layer is connected to the first activation function layer, the first activation function layer is connected to the pooling layer, and the pooling layer is connected to the first 3×3 convolutional layer; the first 3×3 convolutional layer is connected to the second normalization layer, the second normalization layer is connected to the second activation function layer, the second activation function layer is connected to the Spatial Dropout layer, and the Spatial Dropout layer is connected to the second 3×3 convolutional layer, and the second 3×3 convolutional layer is connected to the upsampling layer. Spatial Dropout randomly sets some channels of the feature map to zero with a certain probability during the training process, which can force the network to learn independent features between different channels. Using the Spatial Dropout technique can effectively prevent overfitting, improve the generalization ability of the network, and at the same time can accelerate the training and inference processes of the model.

[0072] In this embodiment, the activation function layers are all Relu activation layers.

[0073] Specifically, when using the multi-scale feature fusion module provided in this embodiment to process the feature map, the multi-scale feature fusion module is divided into three branches, the number of input channels is M, the number of output channels is N, and the kernel size of the convolutional kernel is k.

[0074] Among them, the first branch is a 1×1 convolutional layer, which is an identity mapping when the number of input channels is equal to the number of output channels.

[0075] In the second branch, there are two convolutional layers with a kernel size of 1×1 and a 3×3 convolutional layer. In addition, there is a normalization layer and a Relu activation layer respectively. The multi-scale feature fusion module operates on images of any size by controlling the input depth and output depth. It can not only extract higher-level features through the second branch, but also retain the original hierarchical information through the first branch.

[0076] In the third branch, the size of the input feature map is changed to 1 / 2 of the original. After passing through two 3×3 convolutional layers, the feature map is restored to the original size. Finally, the features of the three branches are fused and input into the next process. Through this multi-branch structure, the multi-scale feature fusion module can learn features at different scales and better capture the global context information and local detailed information of the image.

[0077] The overall expression of the multi-scale feature fusion module is as follows: In the formula, h(p i ) represents the first branch, which is a 1×1 convolutional layer; represents the second branch, which is composed of two 1×1 convolutions and a 3×3 convolution; represents the third branch, which is composed of 1 pooling layer, 2 3×3 convolutional layers and 1 upsampling layer.

[0078] In the feature point extraction decoder, the input feature map becomes a feature response map with a size of H / 8×W / 8×65 after convolution. After passing through the Softmax layer and the Reshape layer, a feature probability map of H×W×1 is obtained, that is, the probability of whether each pixel point is a feature point.

[0079] Among them, the calculation method of the Softmax layer is as follows: In the formula, means the value of the r-th channel of the feature probability map at the (i,j) position, means the value of the r-th channel of the input feature response map at the (i,j) position.

[0080] The descriptor calculation decoder first learns semi-dense descriptors from the input feature map to obtain complete descriptors, and then obtains unit-length descriptors through bicubic polynomial calculation and L2 normalization.

[0081] The calculation process of the bicubic polynomial calculation module is as follows: p(x, y) represents the pixel value at (x, y), and W represents the weight function. In the formula,

[0082] Among them, the calculation method of the weight function is: Among them, a is a hyperparameter, and its value in this embodiment is -0.5.

[0083] The pose estimation module uses the PnP algorithm to estimate the pose of the RGB-D camera. By a known set of 3D spatial points and their projection points on the 2D image plane, the pose of the RGB-D camera is estimated, that is, the rotation matrix R and the translation vector t. The 3D points in the world coordinate system are converted into 3D points in the camera coordinate system, and the problem of 3D-2D is converted into 3D-3D to solve the rotation matrix R and the translation vector t.

[0084] The three-dimensional point cloud reconstruction module, after the pose estimation module obtains the rotation matrix R and the translation vector t in the three-dimensional space during the shooting process of the RGB-D camera, transforms the source point cloud through this rigid body transformation matrix to complete the rough registration of the point cloud.

[0085] Then, the improved ICP iterative closest point algorithm is adopted to register the point sets from multiple viewpoints, iteratively find the closest point pairs between two sets of point clouds, and minimize the distance error between them, so as to accurately align and fuse these point clouds.

[0086] Specifically, first define the source point cloud P and the target point cloud Q: P = {p i , i = 1, 2,..., N p}, Q = {q i , i = 1, 2,..., N q}.

[0087] For each point p i , find its closest point q i in the target point cloud Q, and establish a corresponding relationship, which determines how each point in the point cloud P aligns with the points in the point cloud Q under the current transformation.

[0088] Among them, the error function is: p i and q i are corresponding points, representing the point cloud data in the source point cloud and the target point cloud. The point cloud P is transformed through (R, t) to obtain P', and repeat the above steps until the condition for stopping iteration is met.

[0089] Through the cooperation of rough registration and fine registration, the registration of the source point cloud and the target point cloud is completed, so that the spatial positions of the two point clouds are optimally matched, and the accurate pose transformation matrices R and t are obtained, and then these point clouds are aligned and fused.

[0090] In this embodiment, the SuperPoint-MSFF deep neural network model that improves the SuperPoint feature points is used to accurately obtain the feature points of the images between adjacent frames, which can effectively capture the global context information and local detailed information of the images, further optimize the feature point detection process, and thus improve the robustness of feature point matching and the overall performance of 3D reconstruction.

[0091] The above-described embodiments are only a preferred solution of the present invention, and do not impose any form of limitation on the present invention. There are other variations and modifications without exceeding the technical solutions recorded in the claims.

Claims

1. A three-dimensional scene reconstruction method based on improved SuperPoint feature points, characterized in that: include: S1: Collect image information and obtain RGB image and depth image; S2: Use the SuperPoint-MSFF network to detect feature points on the current frame RGB image and the next frame RGB image, and generate the corresponding SuperPoint-MSFF feature descriptor; S3: Perform feature point matching, camera pose estimation, and rough registration of the source point cloud; S4: Based on the current frame depth map and the next frame depth map, generate a dense point set containing the spatial position of each pixel; S5: Using the improved ICP iterative closest point algorithm, based on the dense point set, the point cloud is precisely aligned to obtain the reconstructed three-dimensional point cloud.

2. The three-dimensional scene reconstruction method based on improved SuperPoint feature points according to claim 1, characterized in that: The S2 includes: extracting deep features of the RGB image to obtain a feature map; resizing the feature map to obtain a feature response map, and generating a feature probability map based on the feature response map; learning a semi-dense descriptor based on the feature map to obtain a complete descriptor, and then performing bicubic multinomial calculation and L2 normalization on the descriptor to obtain a descriptor of unit length.

3. The three-dimensional scene reconstruction method based on improved SuperPoint feature points according to claim 2, characterized in that: When extracting deep features of RGB images, it includes three branches. The first branch retains the original hierarchical information; the second branch extracts high-level features; The third branch changes the size of the input feature map to half of the original size, and then restores the feature map to its original size through the convolution layer; the output results of the three branches are feature fused to obtain the feature map.

4. A three-dimensional scene reconstruction method based on improved SuperPoint feature points according to claim 1, 2 or 3, characterized in that: The S4 includes: defining a source point cloud and a target point cloud, finding the nearest point in the target point cloud for each point in the source point cloud, and establishing a corresponding relationship between the two groups of point clouds; iterating the nearest point pair between the two groups of point clouds, and minimizing the distance error between the two groups of point clouds to complete precise alignment.

5. The three-dimensional scene reconstruction method based on improved SuperPoint feature points according to claim 1, characterized in that: The S3 includes: estimating the camera posture by matching feature points and using the PnP algorithm to obtain the position and direction of the camera in the three-dimensional space during the shooting process, and obtaining the rotation matrix and translation vector in the three-dimensional space during the camera shooting process; transforming the source point cloud through the rigid body transformation matrix to complete the rough alignment of the point cloud.

6. The three-dimensional scene reconstruction method based on improved SuperPoint feature points according to claim 5, characterized in that: When performing feature point matching: obtain the Euclidean distance ratio between the feature points, and compare the Euclidean distance ratio with a preset matching threshold. If the Euclidean distance ratio is less than the matching threshold, it is a correct matching point, otherwise it is an incorrect matching point.

7. A three-dimensional scene reconstruction method based on improved SuperPoint feature points according to claim 5 or 6, characterized in that: Traverse all matching point pairs, calculate the number of inliers, compare the number of inliers with the preset inlier threshold, if the number of inliers is less than the inlier threshold, continue iteration, otherwise end iteration, output the current inlier set, and complete the optimization of matching point pairs.

8. A three-dimensional scene reconstruction system based on improved SuperPoint feature points, characterized in that: include: Image acquisition module, collects image information and obtains the RGB image and depth image of the current frame and the next frame; The feature point detection module uses a multi-branch structure to learn features at different scales, automatically extract feature points in the image, and generate a high-dimensional descriptor for each feature point; The pose estimation module performs feature point matching and uses the PnP algorithm to estimate the pose of the image acquisition module, and obtains the position and direction of the image acquisition module in the three-dimensional space during the shooting process; The 3D point cloud reconstruction module performs coarse registration of the point cloud based on the output results of the pose estimation module, and uses the improved iterative closest point algorithm to perform fine registration of the point cloud.

9. The three-dimensional scene reconstruction system based on improved SuperPoint feature points according to claim 8, characterized in that: The feature point detection module comprises: Shared encoder, including convolutional layers and multi-scale feature fusion module MSFF, extracts deep features; The feature point extraction decoder converts the feature map output by the shared encoder into a feature response map through convolution, and then passes it through the Softmax layer and the Reshape layer to obtain the feature probability map; The descriptor calculation decoder learns a semi-dense descriptor based on the feature map output by the shared encoder to obtain a complete descriptor, and then obtains a descriptor of unit length through bicubic multinomial calculation and L2 normalization.

10. The three-dimensional scene reconstruction system based on improved SuperPoint feature points according to claim 9, characterized in that: The multi-scale feature fusion module MSFF includes three branches. The first branch includes an activation function layer, a normalization layer and a convolution layer; the second branch includes two 1×1 convolution layers and one 3×3 convolution layer, and each convolution layer is preceded by a normalization layer and a Relu activation layer; the third branch includes two 3×3 convolution layers, and each convolution layer is preceded by a normalization layer and an activation layer, a pooling layer is also provided before the first convolution layer, and a sampling layer is provided after the second convolution layer.

Citation Information

Patent Citations

  • Non-rigid dynamic target three-dimensional reconstruction method and system based on RGB-D data

    CN118071918A

Cited By

  • Quality detection system based on three-dimensional point cloud scanning and RGB image fusion

    CN121095712A