Hierarchical feature detection method based on semantic-geometric coupling

Through the semantic-geometric coupling hierarchical feature detection method, Fast-SCNN and SuperPoint networks are used for feature detection and screening, which solves the problem of insufficient image matching accuracy in three-dimensional reconstruction, and improves the accuracy of camera pose estimation and the geometric stability of the three-dimensional model.

CN120526263APending Publication Date: 2025-08-22UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510552706.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The prior art lacks image matching accuracy in three-dimensional reconstruction, especially in complex scenarios, poor robustness, dynamic interference and lighting changes, and other problems lead to large feature matching errors, affecting the camera pose estimation and the accuracy of the three-dimensional model.

Method used

The hierarchical feature detection method based on semantic-geometric coupling is adopted, and semantic segmentation is used to obtain semantic information, and geometric information is extracted in combination with the SuperPoint network. Dynamic feature masks are generated through local reliability and global stability evaluation, dynamic regional features are eliminated, and stable regional feature points are filtered out.

Benefits of technology

It improves the accuracy and stability of feature point detection, reduces the impact of dynamic interference and illumination changes, enhances the robustness of image matching, and improves the accuracy of camera pose estimation and the quality of three-dimensional reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526263A_ABST
    Figure CN120526263A_ABST
Patent Text Reader

Abstract

The invention discloses a hierarchical feature detection method based on semantic-geometric coupling. In order to overcome the problem that feature distribution is easily interfered by dynamic objects (such as the sky and vehicles) due to the fact that a local feature method in a feature detection link depends on local texture response, the invention provides a hierarchical feature detection method based on semantic-geometric coupling. Based on a SuperPoint network framework, a multi-view image set is used as input data, a two-channel feature fusion mechanism is adopted, the scene understanding ability of a semantic segmentation network is fused into a feature detector through transfer learning, and fusion of semantic information and geometric features is achieved. According to the mechanism, by means of understanding of semantic information on scene elements (such as a scene structure and object attributes), feature response of a low-stability area is restrained, meanwhile, a dynamic area and a static area are effectively distinguished, dynamic interference is overcome, and reprojection errors are reduced from the source.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image matching in computer three-dimensional reconstruction, and is based on a deep learning method to achieve image matching of multi-view images. Background Art

[0002] In the field of 3D reconstruction, image matching accuracy is crucial to the accuracy of camera pose estimation and the quality of the final reconstruction. Image matching, as a precursor to camera pose estimation, provides a geometric foundation for pose estimation through feature detection, description, and matching. In the feature detection stage, traditional algorithms (such as SIFT and ORB) face multiple challenges in complex scenes: They lack temporal consistency in the face of dynamic interference; they independently process point and line features, failing to fully exploit geometric complementarity; and under extreme conditions (such as day-night illumination transitions and seasonal changes), the performance of manual feature descriptors based on gradients and grayscale degrades significantly. In the feature matching stage, traditional methods rely on preset metrics (such as squared error and correlation), which can easily generate pseudo feature points in low-texture areas and generate redundant features on the surfaces of dynamic objects, leading to matching ambiguity. In scenes with changing lighting, repetitive textures, and motion blur, these methods are prone to numerous mismatches and matching sparsity defects due to problems such as decreased discriminability of feature descriptors, mismatching of similar local features, and loss of local structural information of feature points. When these problems are input into pose estimation algorithms (such as RANSAC), even if the algorithm parameters are optimized, the pose estimation accuracy is still limited by the quality of the input data.

[0003] Although deep learning methods provide a new path for image matching, they have some limitations, such as insufficient robustness in complex scenarios and limited adaptability to problems such as illumination changes, perspective differences, occlusion or background interference; models usually require a large amount of labeled data for training, and their generalization ability is limited in small samples or special domain scenarios; they consume a lot of computing resources, and the inference speed of deep models is slow, making it difficult to meet real-time requirements; some methods rely on manually designed feature extraction modules and fail to fully achieve end-to-end learning optimization.

[0004] Further analysis of the error chain in the flowchart shows that image matching errors are amplified during the pose estimation stage. For example, feature mismatches can lead to errors in the intrinsic matrix estimation, causing the rotation and translation parameters of the camera pose to deviate from their true values. Errors accumulate during multi-view fusion, ultimately leading to geometric distortion in the 3D model. Therefore, the image matching algorithm directly determines the feasibility and accuracy of camera pose estimation and is an indispensable core link in the 3D reconstruction technology chain. Summary of the Invention

[0005] In order to overcome the problem that the local feature method in the feature detection process relies on local texture response, which makes the feature distribution susceptible to interference from dynamic objects (such as the sky and vehicles), this paper proposes a hierarchical feature detection method based on semantic-geometric coupling. This method is based on the SuperPoint network. The network structure of SuperPoint is as follows: Figure 1 As shown in the figure, it uses a multi-view image set as input data. SuperPoint, a feature point detection network based on self-supervised training, extracts feature points and descriptors through two branches, both of which share the same encoder. However, because the network does not incorporate semantic information into the system and lacks a mechanism to enhance the discriminative power of feature extraction through semantic capabilities, it has certain limitations in terms of accuracy and efficiency in the 3D reconstruction feature detection subtask. The Fast-SCNN network is used to embed the scene understanding capabilities of semantic segmentation into the feature detector, achieving feature fusion of semantic and geometric information, making the feature points both apparent recognizability and physically interpretable. The concept of global stability is also proposed to generate dynamic feature masks for dynamic feature screening. Figure 2 The framework diagram of the proposed algorithm is shown.

[0006] The technical solution adopted by the present invention is to perform feature detection based on semantic guidance, and the method includes:

[0007] Step 1: Use Fast-SCNN as the semantic segmentation network to generate semantic feature maps to obtain semantic information, and use SuperPoint to obtain geometric information;

[0008] Step 2: For the two extracted features, the concepts of local reliability and global stability are proposed to achieve multi-level feature fusion and obtain the final feature response map;

[0009] Step 3: To achieve dynamic feature screening, a dynamic feature mask is generated to mark and eliminate dynamic area features, retain stable area features, and filter out interference features generated by dynamic changes;

[0010] Step 4: Use non-maximum suppression to screen candidate points and form the candidate point geometry, and use SoftNMS to adjust the scores to extract the final set of feature points.

[0011] Compared with the prior art, the present invention has the following beneficial effects:

[0012] (1) Previous feature point detection algorithms only used local detection methods and relied on local texture features. They did not incorporate semantic information into the system and lacked a mechanism to enhance the distinguishing power of feature extraction with the help of semantic capabilities. The present invention implicitly integrates semantic information into the feature detection process and enhances the system's ability to embed semantic information through the fusion and layered processing of semantic information and geometric features.

[0013] (2) Previous feature point detection algorithms lack the ability to distinguish between dynamic objects and static objects. The present invention introduces the generation of dynamic feature masks in the feature screening part to mark and eliminate dynamic area features, retain stable area features, and filter out interference features generated by dynamic changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Attachment Figure 1 : SuperPoint's network framework diagram.

[0015] Attachment Figure 2 : Flowchart of hierarchical feature detection algorithm based on semantic-geometric coupling. DETAILED DESCRIPTION

[0016] The present invention will be further described below with reference to the accompanying drawings.

[0017] Step 1: Take the original image As input, Fast-SCNN is used to obtain the semantic segmentation result F seg That is, semantic information; at the same time, the SuperPoint network is used to obtain the geometric information S in the image rel ;

[0018] Step 2: Propose the concept of local reliability to map the geometric information obtained in step 1 and obtain the local reliability R local (p);

[0019] Step 3: Propose a global stability evaluation model based on semantic segmentation prior and construct a pixel-level stability weight matrix

[0020] Step 4: Semantic features F extracted in step 1 seg , use SoftMax to calculate the probability that each pixel belongs to each semantic category. For each pixel position (i, j) in the semantic feature map, its feature vector is recorded as The calculation formula is:

[0021]

[0022] Among them, M c (p) represents the probability that pixel p belongs to category C;

[0023] Step 5: Based on steps 3-4, generate a global stability score S by fusing the semantic category weights with the linear combination of the probability distribution global (p);

[0024] Step 6: Based on steps 2 and 5, a dynamic weighting mechanism is used to integrate the local reliability R local and global stability S global Generate the final feature response map Sfinal (p);

[0025] Step 7: Generate the final feature response map In addition, in order to realize dynamic feature screening, a dynamic feature mask M is generated dynamic (p), in order to mark and remove dynamic area features.

[0026] Step 8: Apply local maximum suppression on the feature response graph to screen out candidate points and form a set B = {(xi,yi,si)}. Finally, use SoftNMS to adjust the score and screen the final set of detection points.

Claims

1. A hierarchical feature detection method based on semantic-geometric coupling, characterized by the following steps: Step 1: Take the original image As input, Fast-SCNN is used to obtain the semantic segmentation result F seg That is, semantic information; at the same time, the SuperPoint network is used to obtain the geometric information S in the image rel ; Step 2: Propose the concept of local reliability to map the geometric information obtained in step 1. The local reliability of the mapping is expressed as Where N(p) represents the neighborhood range of point p, ω q is the spatial weighting coefficient based on the Gaussian kernel function; Step 3: Propose a global stability evaluation model based on semantic segmentation prior, the core of which is to construct a pixel-level stability weight matrix The generation of this matrix relies on the semantic annotation system of the ADE20K dataset. By analyzing the spatiotemporal characteristics of 150 types of scene objects in the dataset, a four-element classification model is established: volatile category (time stability weight 0.1), dynamic category (weight 0.5), short-term stable category (weight 0.5), and long-term stable category (weight 1.0). Step 4: Semantic features F extracted in step 1 seg , use SoftMax to calculate the probability that each pixel belongs to each semantic category. For each pixel position (i, j) in the semantic feature map, its feature vector is recorded as The calculation formula is: Among them, M c (p) represents the probability that pixel p belongs to category C. Step 5: Based on steps 3-4, generate a global stability score S by fusing the semantic category weights with the linear combination of the probability distribution global (p). The specific formula is as follows: Step 6: Based on steps 2 and 5, a dynamic weighting mechanism is used to integrate the local reliability R local and global stability S global Generate the final feature response map S final (p). S final (p)=R local ·S global (p)·(1+γV) Where V is a parameter related to the dynamic changes of the image, and γ is a weight coefficient used to adjust the degree of influence of dynamic factors on the final stability score. The dynamic modulation factor V is obtained by calculating the difference in spatiotemporal gradients. Assuming a frame in an image video sequence, the dynamic change can be measured by the gradient difference between adjacent frames. The weight coefficient γ is used to adjust the degree of influence of dynamic factors on the final stability score. The present invention uses a cross-validation method to determine the optimal value of γ. Step 7: Generate the final feature response map In addition, in order to realize dynamic feature screening, it is necessary to generate a dynamic feature mask M dynamic (p), in order to mark and remove the dynamic region features, retain the stable region features, and filter the interference features generated by dynamic changes. final , if S final (p)<τ), then the dynamic feature mask M dynamic (p) = 1 (marked as dynamic area); if S final (p)≥τ, then M dynamic (p) = 0 (marked as a stable area), and the final output dynamic feature mask M dynamic ∈{0,1} H×W , complete dynamic feature screening. Step 8: Apply local maximum suppression on the feature response map to screen out candidate points, collect the coordinates (x, y) of all candidate points and their corresponding response values ​​(scores) to form a set B = {(xi, yi, si)}, and finally use SoftNMS to adjust the scores and screen the final set of detection points.

2. The method according to claim 1, wherein: The local reliability obtained in step 2 is the result of the combined effect of geometric information and spatial weighting.

3. The method according to claim 1, wherein: The global stability score obtained in step 5 is the result of fusing the generated semantic category weights with the linear combination of the probability distribution.

4. The method according to claim 1, wherein: The final characteristic response map obtained in step 6 is the integration of local reliability R using a dynamic weighting mechanism local and global stability S global results.