Scene three-dimensional reconstruction method based on SFM technology

By combining the deep feature matching and SFM technology of OmniGlue and LightGlue, semantic segmentation and global optimization are performed, and the problems of large computing resource consumption, poor adaptability and insufficient global optimization in the existing technology are solved, and a three-dimensional reconstruction of scenarios with high precision and low computing resource consumption is realized.

CN120014162APending Publication Date: 2025-05-16BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510074285.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing SFM technology consumes a lot of computing resources during three-dimensional reconstruction of high-precision scenarios, making it difficult to adapt to extreme perspective differences and scale changes, and lacks effective global optimization methods, resulting in insufficient reconstruction accuracy and real-time performance.

Method used

The deep feature matching methods of OmniGlue and LightGlue are used, and image matching and three-dimensional reconstruction are combined with SFM technology, and reconstruction accuracy and robustness are improved through semantic segmentation and global optimization algorithms.

Benefits of technology

It improves the accuracy and robustness of image matching, reduces computing resource consumption, enhances the adaptability to viewing angle and scale changes, and improves the accuracy and stability of three-dimensional reconstruction through global optimization algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014162A_ABST
    Figure CN120014162A_ABST
Patent Text Reader

Abstract

The invention relates to scene three-dimensional reconstruction, in particular to a scene three-dimensional reconstruction method based on the SFM technology, and the method comprises the steps: carrying out the depth feature extraction and depth feature matching of an input image through the cooperation of OmniGlue and LightGlue; based on a depth feature matching result, introducing confidence coefficient weighting to carry out image matching and image registration on the input image; performing semantic segmentation on the registered input image, and dividing a semantic region so as to separate different objects in the input image from a background; performing multi-view data fusion on the input image subjected to semantic segmentation by using an SFM technology to realize scene three-dimensional reconstruction to obtain an initial three-dimensional point cloud; performing scene three-dimensional reconstruction optimization by adopting a global optimization algorithm to obtain an optimized three-dimensional point cloud; according to the technical scheme provided by the invention, the defect that high-precision scene three-dimensional reconstruction is difficult to carry out in the prior art can be effectively overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to three-dimensional reconstruction of a scene, and in particular to a three-dimensional reconstruction method of a scene based on SFM technology. Background Art

[0002] SFM (Structure from Motion) is a widely used technology for 3D reconstruction. SFM extracts feature points from multiple 2D images and infers the 3D structure based on the spatial distribution of these feature points. Similarly, deep learning methods such as LightGlue, SuperGlue, and LINOv2 are also used for feature matching and image matching. These deep learning methods can find similar feature points between images, thereby improving the accuracy of image matching and having a positive impact on the efficiency of 3D reconstruction.

[0003] However, although the above methods have good performance in feature matching, they still have some disadvantages:

[0004] 1) High computing resource requirements

[0005] SFM usually requires a lot of computing resources for feature extraction, feature matching, and spatial structure calculation. Especially in the processing of large-scale data sets or high-resolution images, the required computing time and memory consumption often increase dramatically, seriously affecting real-time performance and system scalability.

[0006] Although deep learning methods such as LightGlue and SuperGlue have improved efficiency by optimizing the computing process, there are still computing bottlenecks, especially when processing large-scale data, which are still limited by computing resources and difficult to meet the needs of real-time or large-scale application scenarios;

[0007] 2) Poor adaptability to changes in viewing angle and scale

[0008] Although the above deep learning methods can cope with perspective changes in most cases, they still have the problem of low matching accuracy when facing extreme perspective differences and scale changes. Especially when dealing with multi-scale scenes, matching error is still a problem that cannot be ignored.

[0009] 3) Lack of effective global optimization methods

[0010] Existing image matching and 3D reconstruction technologies usually focus on local matching and often find it difficult to effectively integrate global information. Even when the local feature matching accuracy is high, the final global reconstruction result may still have large deviations due to the accumulation of local errors. Summary of the invention

[0011] 1. Technical issues to be solved

[0012] In view of the above-mentioned shortcomings of the prior art, the present invention provides a scene 3D reconstruction method based on SFM technology, which can effectively overcome the defect of the prior art that it is difficult to perform high-precision scene 3D reconstruction.

[0013] (II) Technical solution

[0014] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0015] A scene 3D reconstruction method based on SFM technology comprises the following steps:

[0016] S1. Use OmniGlue and LightGlue together to extract and match deep features of input images.

[0017] S2. Based on the deep feature matching results, confidence weighting is introduced to perform image matching and image registration on the input image;

[0018] S3, performing semantic segmentation on the registered input image, dividing the semantic regions to separate different objects in the input image from the background;

[0019] S4, using SFM technology to perform multi-view data fusion on the input image after semantic segmentation to achieve three-dimensional reconstruction of the scene and obtain the initial three-dimensional point cloud;

[0020] S5. Use a global optimization algorithm to optimize the three-dimensional reconstruction of the scene to obtain an optimized three-dimensional point cloud.

[0021] Preferably, S1 uses OmniGlue and LightGlue in synergistic manner to extract and match deep features of the input image, including:

[0022] S11. Use OmniGlue to extract global features of the input images to determine the overall correspondence between the input images, maintain robustness in the face of illumination changes, perspective differences and occlusions, and perform feature matching on the global features of the input images to obtain global feature matching results;

[0023] S12, using LightGlue to extract local features of the input image to supplement the detail information of the input image, and performing feature matching on the local features of the input image to obtain local feature matching results;

[0024] Among them, the global feature matching results of OmniGlue can guide the search of LightGlue and accelerate the feature matching of local features.

[0025] Preferably, in S2, based on the deep feature matching result, confidence weighting is introduced to perform image matching and image registration on the input image, including:

[0026] S21, using the RANSAC algorithm to remove outliers in the global feature matching results and the local feature matching results;

[0027] S22. The global feature matching results based on OmniGlue are more reliable, while the local feature matching results of LightGlue are affected by illumination changes and perspective differences. Confidence weighting is introduced to give higher weight to the global feature matching results of OmniGlue.

[0028] S23, calculating the matching degree between the input images according to the weighted global feature matching results and the local feature matching results, and performing image matching on the input images using a multi-scale matching strategy to determine whether the input images belong to the same scene and the relative positions between the input images;

[0029] S24: Perform image registration on the input images to align input images from different perspectives into the same three-dimensional space.

[0030] Preferably, in S3, semantic segmentation is performed on the registered input image to divide semantic regions so as to separate different objects in the input image from the background, including:

[0031] The DeepLabv3+ model is used to perform semantic segmentation on the registered input image and divide the semantic regions to separate different objects in the input image from the background.

[0032] Preferably, in S4, the SFM technology is used to perform multi-view data fusion on the input image after semantic segmentation to achieve three-dimensional reconstruction of the scene and obtain an initial three-dimensional point cloud, including:

[0033] S41, extracting matching feature points from the input image after semantic segmentation, and calculating the position coordinates of the feature points in the three-dimensional space;

[0034] S42, calibrating cameras of different viewing angles and calculating camera calibration parameters;

[0035] S43. According to the position coordinates of the feature points in the three-dimensional space and the camera calibration parameters, the feature points are projected from the two-dimensional image into the three-dimensional space, and multi-view data fusion is performed to achieve three-dimensional reconstruction of the scene and obtain an initial three-dimensional point cloud.

[0036] Preferably, in S5, a global optimization algorithm is used to optimize the three-dimensional reconstruction of the scene to obtain an optimized three-dimensional point cloud, including:

[0037] S51, calculating the global error based on the minimization error function;

[0038] S52, by adjusting the position and viewing angle of the camera, so that all the matched feature points in the three-dimensional space are consistent with the actual situation to the greatest extent, so as to avoid local error accumulation and improve the accuracy of three-dimensional reconstruction of the scene;

[0039] S53, repeat S1-S4 to obtain an optimized three-dimensional point cloud.

[0040] Preferably, in S5, a global optimization algorithm is used to optimize the scene 3D reconstruction, and after obtaining the optimized 3D point cloud, the following steps are included:

[0041] Use 3D visualization tools to display the optimized 3D point cloud.

[0042] (III) Beneficial effects

[0043] Compared with the prior art, the scene 3D reconstruction method based on SFM technology provided by the present invention has the following beneficial effects:

[0044] 1) Improve the accuracy and robustness of image matching

[0045] The deep learning image matching algorithm OmniGlue can perform more accurate and robust feature matching in more complex environments. Compared with traditional image matching technologies (such as SIFT and SURF), OmniGlue can better adapt to lighting changes, perspective differences and occlusion in images, effectively improving image matching accuracy.

[0046] Existing image matching methods are often affected by environmental factors (such as lighting, viewing angle, etc.), resulting in reduced image matching accuracy. Omniglue can effectively improve the accuracy and stability of image matching under these challenges, thereby improving the quality of the final 3D reconstruction.

[0047] 2) Reduce computing resource consumption

[0048] The combination of OmniGlue and SFM can process image data more efficiently by optimizing the calculation process and feature matching strategy, reducing the calculation burden of traditional SFM algorithms in image matching and spatial reconstruction;

[0049] The present invention reduces computational complexity and improves the speed and efficiency of image processing through more accurate feature matching. This optimization not only improves computational efficiency, but also makes large-scale or real-time 3D reconstruction possible, meeting the use requirements of a variety of large-scale application scenarios (such as autonomous driving, augmented reality, etc.);

[0050] 3) Enhance adaptability to scale and perspective changes

[0051] The present invention combines the deep feature matching of OmniGlue with the global optimization of SFM, which can better adapt to the perspective differences and scale changes between images;

[0052] In traditional image matching and SFM technologies, perspective differences and scale inconsistencies are common challenges that easily lead to errors in image matching. However, the present invention uses deep feature extraction, and OmniGlue can handle these problems more accurately. SFM further ensures the accuracy and consistency of 3D reconstruction through global optimization. Therefore, the present invention can handle a wider range of image matching in complex scenes and improve the 3D reconstruction effect under different perspectives and scales.

[0053] 4) Global optimization algorithm improves reconstruction accuracy

[0054] The present invention solves the problem of insufficient global optimization in the prior art by combining the fine matching of OmniGlue and the global optimization of SFM;

[0055] Traditional methods mostly focus on local matching, which may lead to local error accumulation during global reconstruction. However, the present invention uses a global optimization algorithm to optimize the scene 3D reconstruction, so that the 3D model can better maintain consistency and reduce the impact of local errors on the overall result, thereby improving the accuracy and stability of the final 3D model.

[0056] In addition, the improved global optimization scheme can better handle data of different perspectives and scales in large data sets and complex scenes, greatly improving the reconstruction accuracy in practical application scenarios;

[0057] 5) Efficient multi-view data fusion

[0058] In the existing technology, although SFM can reconstruct images from multiple perspectives, it usually faces the problem of data fusion difficulties for a large number of images with different and inconsistent perspectives;

[0059] OmniGlue provides an efficient feature matching solution that can help better perform multi-view data fusion, enhance the ability to integrate image data collected from different viewpoints, and improve the integrity and consistency of 3D reconstruction. This advantage makes the present invention particularly suitable for application scenarios with dynamic changes, sparse data, or extreme changes in viewpoints (such as urban modeling, virtual reality, augmented reality, etc.). BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0061] Figure 1 It is a schematic diagram of the process of the present invention;

[0062] Figure 2 A schematic diagram of performing deep feature matching on an input image in the present invention;

[0063] Figure 3 Schematic diagram of scene 3D reconstruction using SFM technology;

[0064] Figure 4 The first input image captured by the camera at a certain viewing angle (the image contains matching feature points);

[0065] Figure 5 A second input image captured by a camera from another viewing angle (the image contains matching feature points);

[0066] Figure 6 In order to adopt the technical solution of the present invention Figure 4 and Figure 5 The 3D model obtained by 3D reconstruction of the scene (the figure contains the position and pose of each camera). DETAILED DESCRIPTION

[0067] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0068] A scene 3D reconstruction method based on SFM technology, such as Figure 1 As shown in S1, OmniGlue and LightGlue are used together to perform deep feature extraction and deep feature matching on the input image, such as Figure 2 As shown, specifically including:

[0069] S11. Use OmniGlue to extract global features of the input images to determine the overall correspondence between the input images, maintain robustness in the face of illumination changes, perspective differences and occlusions, and perform feature matching on the global features of the input images to obtain global feature matching results;

[0070] S12, using LightGlue to extract local features of the input image to supplement the detail information of the input image, and performing feature matching on the local features of the input image to obtain local feature matching results;

[0071] Among them, the global feature matching results of OmniGlue can guide the search of LightGlue and accelerate the feature matching of local features.

[0072] The above technical solution has the following advantages:

[0073] 1) Stronger robustness: The coordinated use of OmniGlue and LightGlue can better handle complex situations such as lighting changes, perspective differences, and occlusion;

[0074] 2) Higher accuracy: Global features and local features complement each other, making the matching more accurate;

[0075] 3) Higher efficiency: OmniGlue's global feature matching results can guide LightGlue's search and optimize the feature matching efficiency of local features;

[0076] 4) Greater adaptability: It can be applied to scenarios of different scales and complexities without making assumptions about specific scenarios.

[0077] S2. Based on the deep feature matching results, confidence weighting is introduced to perform image matching and image registration on the input image, specifically including:

[0078] S21, using the RANSAC algorithm to remove outliers in the global feature matching results and the local feature matching results;

[0079] S22. The global feature matching results based on OmniGlue are more reliable, while the local feature matching results of LightGlue are affected by illumination changes and perspective differences. Confidence weighting is introduced to give higher weight to the global feature matching results of OmniGlue.

[0080] S23, calculating the matching degree between the input images according to the weighted global feature matching results and the local feature matching results, and performing image matching on the input images using a multi-scale matching strategy to determine whether the input images belong to the same scene and the relative positions between the input images;

[0081] S24: Perform image registration on the input images to align input images from different perspectives into the same three-dimensional space.

[0082] The above technical solution has the following advantages:

[0083] 1) Fewer mismatches: RANSAC outlier removal and the introduction of confidence weighting reduce mismatches;

[0084] 2) Higher matching accuracy: Using multi-scale matching strategy for image matching can better adapt to scale changes;

[0085] 3) More stable registration: The image registration results are more stable.

[0086] S3. Perform semantic segmentation on the registered input image and divide the semantic regions to separate different objects in the input image from the background, including:

[0087] The DeepLabv3+ model is used to perform semantic segmentation on the registered input image and divide the semantic regions to separate different objects in the input image from the background.

[0088] The above technical solution has the following advantages:

[0089] In complex scenes, feature points of the same object can be identified and matched more accurately.

[0090] S4, using SFM technology to perform multi-view data fusion on the input image after semantic segmentation to achieve 3D reconstruction of the scene and obtain the initial 3D point cloud, such as Figure 3 As shown, specifically including:

[0091] S41, extracting matching feature points from the input image after semantic segmentation, and calculating the position coordinates of the feature points in the three-dimensional space;

[0092] S42, calibrating cameras of different viewing angles and calculating camera calibration parameters;

[0093] S43. According to the position coordinates of the feature points in the three-dimensional space and the camera calibration parameters, the feature points are projected from the two-dimensional image into the three-dimensional space, and multi-view data fusion is performed to achieve three-dimensional reconstruction of the scene and obtain an initial three-dimensional point cloud.

[0094] S5. Use a global optimization algorithm to optimize the scene 3D reconstruction to obtain an optimized 3D point cloud, specifically including:

[0095] S51, calculating the global error based on the minimization error function;

[0096] S52, by adjusting the position and viewing angle of the camera, so that all the matched feature points in the three-dimensional space are consistent with the actual situation to the greatest extent, so as to avoid local error accumulation and improve the accuracy of three-dimensional reconstruction of the scene;

[0097] S53, repeat S1-S4 to obtain an optimized three-dimensional point cloud.

[0098] The above technical solution has the following advantages:

[0099] Improved reconstruction accuracy, especially in large-scale or complex scenes;

[0100] It effectively reduces mismatching and local errors and ensures global consistency.

[0101] Specifically, S5 uses a global optimization algorithm to optimize the scene 3D reconstruction, and after obtaining the optimized 3D point cloud, it includes:

[0102] Use 3D visualization tools to display the optimized 3D point cloud.

[0103] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A scene 3D reconstruction method based on SFM technology, characterized by: The following steps are involved: S1. Use OmniGlue and LightGlue together to extract and match deep features of input images. S2. Based on the deep feature matching results, confidence weighting is introduced to perform image matching and image registration on the input image; S3, performing semantic segmentation on the registered input image to divide the semantic regions so as to separate different objects in the input image from the background; S4, using SFM technology to perform multi-view data fusion on the input image after semantic segmentation to achieve three-dimensional reconstruction of the scene and obtain the initial three-dimensional point cloud; S5. Use a global optimization algorithm to optimize the three-dimensional reconstruction of the scene to obtain an optimized three-dimensional point cloud.

2. The scene 3D reconstruction method based on SFM technology according to claim 1, characterized in that: S1 uses OmniGlue and LightGlue to extract and match deep features of input images, including: S11. Use OmniGlue to extract global features of the input images to determine the overall correspondence between the input images, maintain robustness in the face of illumination changes, perspective differences and occlusions, and perform feature matching on the global features of the input images to obtain global feature matching results; S12, using LightGlue to extract local features of the input image to supplement the detail information of the input image, and performing feature matching on the local features of the input image to obtain local feature matching results; Among them, the global feature matching results of OmniGlue can guide the search of LightGlue and accelerate the feature matching of local features.

3. The scene 3D reconstruction method based on SFM technology according to claim 2, characterized in that: In S2, based on the deep feature matching results, confidence weighting is introduced to perform image matching and image registration on the input image, including: S21, using the RANSAC algorithm to remove outliers in the global feature matching results and the local feature matching results; S22. The global feature matching results based on OmniGlue are more reliable, while the local feature matching results of LightGlue are affected by illumination changes and perspective differences. Confidence weighting is introduced to give higher weight to the global feature matching results of OmniGlue. S23, calculating the matching degree between the input images according to the weighted global feature matching results and the local feature matching results, and performing image matching on the input images using a multi-scale matching strategy to determine whether the input images belong to the same scene and the relative positions between the input images; S24: Perform image registration on the input images to align the input images from different perspectives into the same three-dimensional space.

4. The scene 3D reconstruction method based on SFM technology according to claim 3, characterized in that: S3 performs semantic segmentation on the registered input image and divides the semantic regions to separate different objects in the input image from the background, including: The DeepLabv3+ model is used to perform semantic segmentation on the registered input image and divide the semantic regions to separate different objects in the input image from the background.

5. The scene 3D reconstruction method based on SFM technology according to claim 4, characterized in that: In S4, SFM technology is used to perform multi-view data fusion on the input image after semantic segmentation to achieve 3D reconstruction of the scene and obtain the initial 3D point cloud, including: S41, extracting matching feature points from the input image after semantic segmentation, and calculating the position coordinates of the feature points in the three-dimensional space; S42, calibrating cameras of different viewing angles and calculating camera calibration parameters; S43. According to the position coordinates of the feature points in the three-dimensional space and the camera calibration parameters, the feature points are projected from the two-dimensional image into the three-dimensional space, and multi-view data fusion is performed to achieve three-dimensional reconstruction of the scene and obtain an initial three-dimensional point cloud.

6. The scene 3D reconstruction method based on SFM technology according to claim 5, characterized in that: S5 uses a global optimization algorithm to optimize the scene 3D reconstruction and obtain an optimized 3D point cloud, including: S51, calculating the global error based on the minimization error function; S52, by adjusting the position and viewing angle of the camera, so that all the matched feature points in the three-dimensional space are consistent with the actual situation to the greatest extent, so as to avoid local error accumulation and improve the accuracy of three-dimensional reconstruction of the scene; S53, repeat S1-S4 to obtain an optimized three-dimensional point cloud.

7. The scene 3D reconstruction method based on SFM technology according to claim 6, characterized in that: S5 uses a global optimization algorithm to optimize the scene 3D reconstruction. After obtaining the optimized 3D point cloud, it includes: Use 3D visualization tools to display the optimized 3D point cloud.