Lightweight image feature point detection and matching method and related equipment thereof

Through the lightweight feature extraction, segmentation and matching optimization of SuperPoint and SuperGlue models, the computing resource consumption problem of image feature point detection and matching on resource-constrained devices is solved, efficient and accurate feature point detection and matching is achieved, and the positioning and mapping capabilities of the visual inertial system are improved.

CN120339646APending Publication Date: 2025-07-18QINGDAO INST OF COMPUTING TECH XIDIAN UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510302882.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art cannot effectively reduce computing resource consumption while ensuring accuracy in image feature point detection and matching, especially on resource-constrained devices.

Method used

The lightweight SuperPoint neural network model is used for multi-scale feature extraction, combined with RTSS-DBSCAN superpixel segmentation and lightweight SuperGlue neural network model for initial matching, geometric consistency optimization is performed through the adaptive exponential moving average algorithm, and the SuperPoint and SuperGlue models are pruned to reduce the computational amount and memory usage.

Benefits of technology

While maintaining high accuracy, the consumption of computing resources is significantly reduced, the system's operating efficiency and robustness are improved, and it is suitable for practical application scenarios such as visual inertial systems, robot navigation, and autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339646A_ABST
    Figure CN120339646A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision, and relates to a lightweight image feature point detection and matching method and related equipment thereof. The method comprises the following steps: performing feature extraction on a target image through a lightweight SuperPoint neural network model to generate feature information; performing super-pixel segmentation processing on the image to generate a super-pixel segmentation result; performing initial matching on the feature information through a lightweight SuperGlue neural network model to generate an initial matching pair; performing outlier filtering on the initial matching pair based on a superpixel segmentation result, and determining an optimized target matching pair; performing geometric consistency optimization on the optimized target matching pair by adopting an adaptive exponential moving average algorithm to generate a final matching result; the lightweight SuperPoint neural network model is generated by carrying out hierarchical progressive pruning processing on an original SuperPoint neural network model; the lightweight SuperGlue neural network model is generated by carrying out attention head pruning processing and graph network layer pruning processing on an original SuperGlue neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and particularly to a lightweight image feature point detection and matching method and related devices thereof. Background Art

[0002] Computer vision technology has promoted the development of visual inertial systems (VIS) in fields such as autonomous driving and robot navigation. VIS combines a camera and an inertial measurement unit (IMU) to achieve environmental perception and motion estimation through image feature point detection technology. Traditional feature point detection methods such as SIFT, SURF, and ORB have large computational amounts, are sensitive to environmental changes, and have poor real-time performance. Although deep learning methods can extract higher-quality feature points, the model has a huge computational amount and unstable accuracy, and performs poorly in resource-constrained systems.

[0003] In a deep learning-driven visual SLAM system, the SuperPoint model is used for feature extraction, and the SuperGlue algorithm is used for feature matching, but these methods have deficiencies in practical applications. First, they have insufficient adaptability to dynamic environments. When facing dynamic objects and rapid changes in light, scale, and rotation, they are prone to false matching and loss of tracking. Second, traditional feature matching methods have low computational efficiency and high memory requirements, resulting in long single-frame processing time and high system complexity, making it difficult to meet real-time requirements. In summary, the existing technology lacks a systematic method to comprehensively optimize feature detection, matching, and model lightweight processing, and cannot effectively reduce computational resource consumption while ensuring accuracy, which limits its application scope on resource-constrained devices. Summary of the Invention

[0004] A lightweight image feature point detection and matching method and related devices provided by an embodiment of the present invention at least solve the problem in the related technology that computational resource consumption cannot be effectively reduced while ensuring accuracy.

[0005] According to a first aspect of an embodiment of the present invention, a lightweight image feature point detection and matching method is provided, including:

[0006] Performing multi-scale feature extraction on a target image through a lightweight SuperPoint neural network model to generate feature information including key point coordinates and descriptors;

[0007] Based on the feature information, performing superpixel segmentation processing on the target image through an RTSS-DBSCAN superpixel segmentation algorithm to generate a superpixel segmentation result;

[0008] Performing initial matching on the feature information through a lightweight SuperGlue neural network model to generate initial matching pairs;

[0009] Based on the superpixel segmentation results, outlier filtering is performed on the initial matching pairs to remove inconsistent matching pairs in the initial matching pairs and determine the optimized target matching pairs;

[0010] The adaptive exponential moving average algorithm is used to optimize the geometric consistency of the optimized target matching pairs to generate the final matching results;

[0011] Among them, the lightweight SuperPoint neural network model is generated by performing hierarchical progressive pruning on the original SuperPoint neural network model. The hierarchical progressive pruning includes coarse-grained pruning, fine-grained pruning, and global fine-tuning and recovery processing; the lightweight SuperGlue neural network model is generated by performing attention head pruning and graph network layer pruning on the original SuperGlue neural network model.

[0012] According to the second aspect of the embodiments of the present invention, a lightweight image feature point detection and matching device is provided, including:

[0013] An extraction module for performing multi-scale feature extraction on a target image through a lightweight SuperPoint neural network model to generate feature information including key point coordinates and descriptors;

[0014] A superpixel segmentation module for performing superpixel segmentation on the target image based on the feature information through the RTSS-DBSCAN superpixel segmentation algorithm to generate superpixel segmentation results;

[0015] A matching module for initially matching the feature information through a lightweight SuperGlue neural network model to generate initial matching pairs;

[0016] A filtering module for performing outlier filtering on the initial matching pairs based on the superpixel segmentation results to remove inconsistent matching pairs in the initial matching pairs and determine the optimized target matching pairs;

[0017] An optimization module for optimizing the geometric consistency of the optimized target matching pairs by using the adaptive exponential moving average algorithm to generate the final matching results;

[0018] Among them, the lightweight SuperPoint neural network model is generated by performing hierarchical progressive pruning on the original SuperPoint neural network model. The hierarchical progressive pruning includes coarse-grained pruning, fine-grained pruning, and global fine-tuning and recovery processing; the lightweight SuperGlue neural network model is generated by performing attention head pruning and graph network layer pruning on the original SuperGlue neural network model.

[0019] According to a third aspect of an embodiment of the present invention, there is provided an electronic device, including: a processor, and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to execute the method according to the first aspect.

[0020] According to a fourth aspect of an embodiment of the present invention, there is provided a non-transitory machine-readable medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to the first aspect.

[0021] Advantageous effects of the embodiments of the present invention:

[0022] The lightweight image feature point detection and matching method provided by the embodiments of the present invention can efficiently generate feature information including key point coordinates and descriptors by performing multi-scale feature extraction on a target image through a lightweight SuperPoint neural network model, providing accurate basic data for subsequent image matching. The lightweight SuperPoint neural network model is generated by performing hierarchical progressive pruning on the original SuperPoint neural network model. The hierarchical progressive pruning includes coarse-grained pruning, fine-grained pruning, and global fine-tuning recovery processing. This pruning method not only effectively reduces the computational amount and memory occupation of the model, but also ensures the recovery and improvement of the model performance through step-by-step optimization, enabling it to adapt to resource-constrained devices while maintaining high-precision feature extraction and meeting the low-power and high-efficiency operation requirements in practical applications.

[0023] Then, based on the extracted feature information, the RTSS-DBSCAN superpixel segmentation algorithm is used to perform superpixel segmentation on the image, which can quickly and accurately divide the image into multiple small regions with similar attributes, simplify the image representation and retain important edge information, providing a clearer and more structured image basis for subsequent feature matching. Then, the lightweight SuperGlue neural network model is used to perform initial matching on the feature information to generate initial matching pairs. The lightweight SuperGlue neural network model is generated by performing attention head pruning and graph network layer pruning on the original SuperGlue neural network model. This pruning strategy reduces the model complexity while retaining key matching information and improving the matching efficiency.

[0024] Furthermore, outlier filtering is performed on the initial matching pairs based on the superpixel segmentation results, which can effectively remove inconsistent matching pairs in the initial matching pairs and determine the optimized target matching pairs. This process utilizes the regional information of superpixels to finely screen the matching results, greatly improving the accuracy and robustness of the matching and reducing the impact of incorrect matching pairs on subsequent processing. Finally, the adaptive exponential moving average algorithm is used to optimize the geometric consistency of the optimized target matching pairs to generate the final matching results. This algorithm can dynamically adjust the decay coefficient according to different stages of the training process, making the update of model parameters more reasonable, further enhancing the accuracy and consistency of the matching results, and ensuring the high-efficiency localization and mapping capabilities in complex environments.

[0025] In summary, the method of the embodiment of the present invention significantly reduces the consumption of computing resources through lightweight processing such as model pruning while ensuring the accuracy of image feature point detection and matching, improves the operating efficiency and robustness of the system, and makes it more suitable for practical application scenarios such as visual inertial systems, robot navigation, autonomous driving, and augmented reality. It can provide more efficient and accurate image feature point detection and matching support for the technological development in these fields.

[0026] The details of one or more embodiments of the present invention are set forth in the following drawings and description to make the other features, objects, and advantages of the present invention more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other embodiments based on these drawings without creative efforts.

[0028] Figure 1 It is a flowchart of a lightweight image feature point detection and matching method provided by an embodiment of the present invention.

[0029] Figure 2 It is a flowchart of an image feature matching system for a visual inertial system provided by an embodiment of the present invention.

[0030] Figure 3 It is a schematic diagram of pruning lightweight processing provided by an embodiment of the present invention.

[0031] Figure 4 It is a schematic structural diagram of a lightweight SuperPoint neural network model provided by an embodiment of the present invention.

[0032] Figure 5A schematic diagram of a network pruning process provided by an embodiment of the present invention.

[0033] Figure 6 A schematic diagram of an improved SuperGlue feature matching framework provided by an embodiment of the present invention.

[0034] Figure 7 A schematic structural diagram of a lightweight image feature point detection and matching device provided by an embodiment of the present invention.

[0035] Figure 8 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0036] The embodiments of the present embodiment will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present embodiment are shown in the drawings, it should be understood that the present embodiment can be implemented in various forms and should not be construed as limited to the embodiments described herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present embodiment. It should be understood that the drawings and embodiments of the present embodiment are only for exemplary purposes and are not used to limit the protection scope of the present embodiment.

[0037] The visual inertial system combines the advantages of cameras and inertial measurement units. Even in complex or GPS-denied environments, it can achieve three-dimensional perception of the environment and device motion estimation through image feature point detection technology, and is widely used in fields such as autonomous driving, robot navigation, and augmented reality, providing stable and accurate position perception and motion tracking capabilities, greatly enhancing the robustness and reliability of the system.

[0038] In the field of computer vision, image feature point detection technology is crucial. Traditional feature point detection methods such as SIFT, SURF, and ORB, although they can provide relatively stable feature point detection effects, generally face problems such as large computational complexity, sensitivity to environmental changes, and poor real-time performance. With the introduction of deep learning technology, methods based on deep convolutional neural networks (CNNs) have gradually become mainstream and can extract higher-quality feature points, but there are still defects such as large model computational complexity and unstable accuracy, especially prominent in resource-constrained visual inertial systems. In summary, the prior art lacks a systematic method to comprehensively optimize feature detection, matching, and model lightweight processing, and cannot effectively reduce computational resource consumption while ensuring accuracy, restricting its application scope on resource-constrained devices.

[0039] To solve the above problems, an embodiment of the present invention provides a lightweight image feature point detection and matching method and related devices.

[0040] Figure 1The flowchart of a lightweight image feature point detection and matching method provided by an embodiment of the present invention. As Figure 1 shown, the method includes the following steps.

[0041] Step S101, perform multi-scale feature extraction on the target image through a lightweight SuperPoint neural network model to generate feature information including key point coordinates and descriptors.

[0042] Step S102, based on the feature information, perform superpixel segmentation processing on the image through the RTSS-DBSCAN superpixel segmentation algorithm to generate a superpixel segmentation result.

[0043] Step S103, perform initial matching on the feature information through a lightweight SuperGlue neural network model to generate initial matching pairs.

[0044] Step S104, based on the superpixel segmentation result, filter outlier from the initial matching pairs to remove inconsistent matching pairs in the initial matching pairs and determine the optimized target matching pairs.

[0045] Step S105, adopt an adaptive exponential moving average algorithm to optimize the geometric consistency of the optimized target matching pairs to generate the final matching result.

[0046] Among them, the lightweight SuperPoint neural network model is generated by performing hierarchical progressive pruning processing on the original SuperPoint neural network model. The hierarchical progressive pruning processing includes coarse-grained pruning processing, fine-grained pruning processing, and global fine-tuning recovery processing; the lightweight SuperGlue neural network model is generated by performing attention head pruning processing and graph network layer pruning processing on the original SuperGlue neural network model.

[0047] In this embodiment, before performing multi-scale feature extraction on the target image, it is necessary to first construct a lightweight SuperPoint neural network model. The lightweight SuperPoint neural network model in this embodiment is generated by performing hierarchical progressive pruning processing on the original SuperPoint neural network model.

[0048] After generating the lightweight SuperPoint neural network model, the multi-scale feature extraction can be performed on the target image through the lightweight SuperPoint neural network model to generate feature information including key point coordinates and descriptors.

[0049] In this embodiment, the purpose of extracting multi-scale features from the target image through the lightweight SuperPoint neural network model is to capture the features of the image at different scales to adapt to the detection of feature points of different sizes and distances. To improve the effect of feature extraction, the target image can be preprocessed first, that is, the target image is adjusted at multiple scales to generate an image pyramid with different resolutions. For example, the original image is scaled to 100%, 80%, 60% of the original image respectively to generate multi-scale images.

[0050] In practical applications, a scale adjustment layer can be added to the output of the lightweight SuperPoint neural network model to convert the feature point coordinates of the target image into predicted coordinates. Specifically, as shown in the following formula (1):

[0051] x pred = x orig ·2 s , s ∈ {0, 1, 2, 3} (1)

[0052] where x orig is the feature point coordinate of the target image, x pred is the predicted coordinate, and s is the scale parameter with a value range of {0, 1, 2, 3}.

[0053] Based on the above formula (1), the position mapping of feature points at different scales is achieved by multiplying the original coordinates by 2 s . For example, when s = 0, the feature point coordinates remain unchanged; when s = 1, the feature point coordinates are doubled; when s = 2, the feature point coordinates are quadrupled; when s = 3, the feature point coordinates are octupled.

[0054] When performing multi-scale feature extraction on the target image through the lightweight SuperPoint neural network model, multi-scale convolution and pooling operations can be performed on the target image to generate a feature pyramid.

[0055] Specifically, convolution and pooling operations are respectively performed on images of different scales to extract local and global features of the image. In this embodiment, the convolution operation slides the convolution kernel on the image and calculates the dot product of the convolution kernel and the local area of the image to obtain a feature map; the pooling operation downsamples the feature map by methods such as taking the maximum value and the average value to reduce the data volume while retaining important features.

[0056] The feature maps obtained through convolution and pooling operations at different scales are combined into a feature pyramid. Each layer of the feature pyramid corresponds to a feature map of a different scale. The bottom layer feature map has a high resolution and rich details, and the top layer feature map has a low resolution and rich semantic information.

[0057] Then, the local optimal feature points at each scale in the feature pyramid are screened by the non-maximum suppression algorithm; the coordinates of the screened local optimal feature points are mapped to the original image space through the scale adjustment layer, and multi-scale feature points and their descriptors are output.

[0058] Specifically, non-maximum suppression processing is performed on the feature maps at each scale in the feature pyramid to screen out the local optimal feature points. Specifically, the response value of each pixel point is calculated, and then the response values of the pixel points in its neighborhood are checked. If there are pixel points with higher response values in the neighborhood, the current pixel point is suppressed, that is, its response value is set to 0. This can remove redundant feature points and retain the feature points with the highest response value, improving the efficiency of subsequent matching.

[0059] Optionally, in order to ensure that the feature points extracted at each scale have high quality and representativeness, the number of feature points retained at each scale can be set. For example, the top 100 feature points with the highest response values can be retained at each scale. This can avoid the matching difficulties caused by overly dense feature points and ensure the stability and usability of feature points in different scenarios.

[0060] The coordinates of the screened local optimal feature points are mapped to the original image space through the scale adjustment layer, and multi-scale feature points and their descriptors are output. This process needs to consider the proportional relationship between images of different scales and convert the coordinates of the feature points from the feature maps of different scales back to the coordinate system of the original image.

[0061] After generating the feature information including the key point coordinates and descriptors, based on the feature information, the image can be subjected to superpixel segmentation processing by the RTSS-DBSCAN superpixel segmentation algorithm to generate the superpixel segmentation result. In this embodiment, the RTSS-DBSCAN superpixel segmentation algorithm is a real-time superpixel segmentation method (Real-time Superpixel Segmentation by DBSCAN Clustering Algorithm, RTSS-DBSCAN) based on the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm.

[0062] In this embodiment, the RTSS-DBSCAN superpixel segmentation algorithm is a density-based clustering method aimed at segmenting the image into superpixels with similar features. A superpixel refers to a group of pixels in the image that are similar in color, texture, and position. This algorithm combines the advantages of region growing and density clustering and can effectively handle the target segmentation in complex images.

[0063] When performing superpixel segmentation on an image, the similarity between pixels can be calculated based on the key point coordinates and response values in the feature information. Adjacent pixels with a similarity higher than a preset threshold are merged into superpixels through a density clustering algorithm to generate superpixel identifiers. Finally, based on the superpixel identifiers, the superpixel boundaries are dynamically adjusted to retain the image edge information, and the superpixel segmentation result is output.

[0064] In this embodiment, the similarity between pixels can be calculated based on the key point coordinates and response values in the feature information. The calculation formula for the similarity can be expressed as:

[0065]

[0066] where S ij represents the similarity between pixel p i and p j ; C is a constant used to adjust the scale of the similarity, which can take an empirical value such as 0.5 or be determined through cross-validation; ||p i -p j || is the distance between pixels, and the distance can be the Euclidean distance or other distance metrics suitable for image features.

[0067] Adjacent pixels with a similarity higher than a preset threshold are merged into superpixels through a density clustering algorithm. Specifically, for each pixel point, the pixel points within its neighborhood are determined. If the number of pixels within the neighborhood exceeds the preset minimum number of points λ (such as λ = 10), these pixels are grouped into the same superpixel, and the superpixel is continuously expanded until all adjacent pixels are clustered. This process is similar to region growing, starting from a seed pixel and continuously expanding into the surrounding pixel regions with high similarity.

[0068] For superpixel initialization, key points in the image are first selected, and a unique superpixel identifier (superpixel_id) is assigned to each key point. The selection of key points can be based on image saliency, corner detection, or other feature extraction methods. For example, the Harris corner detection algorithm can be used to select key points.

[0069] Next, by clustering these key points into the corresponding superpixels, a feature point map (sp_feat_map) is formed, where each superpixel contains its corresponding feature points. This process not only ensures that each superpixel can accurately represent a specific region in the image but also provides a basis for further optimization.

[0070] The generated superpixels are then adjusted to ensure that their shape and size meet the requirements of specific applications. This adjustment process helps to improve the stability and consistency of the superpixels, enabling the finally generated superpixels to not only accurately reflect the local features in the image but also better adapt to the requirements of subsequent processing steps. For example, the segmentation results can be optimized by smoothing the superpixel boundaries or adjusting the superpixel size.

[0071] In practical applications, based on the superpixel identifiers, the superpixel boundaries can be dynamically adjusted to retain the image edge information. For example, by checking the pixel gradients near the superpixel boundaries, if an edge with a large gradient change is found, the superpixel boundary is adjusted to fit the edge, preventing pixels from different objects or regions from being divided into the same superpixel. After adjustment, the superpixel segmentation result is output, which divides the image into multiple small regions (superpixels) with similar properties. Each superpixel contains a group of pixels that are similar in color, texture, and position.

[0072] After generating the superpixel segmentation result, the feature information can be initially matched through a lightweight SuperGlue neural network model to generate initial matching pairs.

[0073] Before that, a lightweight SuperGlue neural network model needs to be constructed first. In this embodiment, the lightweight SuperGlue neural network model is generated by performing attention head pruning and graph network layer pruning on the original SuperGlue neural network model.

[0074] After obtaining the lightweight SuperGlue neural network model, the feature information can be initially matched to generate initial matching pairs. The extracted feature information is input into the lightweight SuperGlue neural network model, which can calculate the similarity and perform matching based on the descriptors of the feature points. Specifically, the lightweight SuperGlue neural network model calculates the similarity between the descriptors of feature points in different images and determines the matching relationship according to the magnitude of the similarity. For example, if the similarity between two feature point descriptors is higher than a certain threshold (such as 0.7), they are considered a pair of matching points, generating initial matching pairs. This process may produce some incorrect matching pairs, so subsequent outlier filtering steps are needed for optimization.

[0075] After that, based on the superpixel segmentation results, outlier filtering can be performed on the initial matching pairs to remove inconsistent matching pairs in the initial matching pairs and determine the optimized target matching pairs. In this embodiment, the purpose of outlier filtering is to remove abnormal matching points that do not conform to the expected pattern or significantly deviate from the normal distribution in the initial matching pairs, thereby improving the stability and reliability of the matching results. Through this process, false matches can be effectively reduced, providing a higher-quality data basis for subsequent pose estimation and map construction.

[0076] Specifically, based on the superpixel segmentation results, the geometric consistency measure of each initial matching pair within the corresponding superpixel neighborhood can be calculated; a double-threshold decision mechanism is adopted to filter out the matching pairs with a geometric consistency measure lower than 0.6 or a confidence level lower than 0.4; then, the confidence levels of the remaining matching pairs are adjusted through dynamic confidence calibration technology, and the optimized target matching pairs are output.

[0077] Based on the superpixel segmentation results, the geometric consistency measure of each initial matching pair within the corresponding superpixel neighborhood is calculated. The geometric consistency measure reflects the rationality of the matching pair in terms of spatial position. For example, if the relative position relationship between two matching points in two images is consistent with the relative position relationships of other matching points, then its geometric consistency measure is relatively high.

[0078] In this embodiment, to improve the accuracy of feature point matching, a local consistency constraint model can be defined. This local consistency constraint model evaluates the quality of the match by calculating the geometric consistency measure of the matching pairs within the superpixel neighborhood. Specifically, for each matching point, the geometric distance differences between all matching pairs within its superpixel and its adjacent regions are calculated, and these differences are used to define a consistency measure. The consistency measure formula is as follows:

[0079]

[0080] where Consistency(i,j) represents the consistency measure between matching points i and j, N(i) is the set of matching points within the superpixel where matching point i is located and its adjacent regions, Δp ik and Δp jk are the geometric distance differences between matching points i and j and matching point k respectively, and σ is a constant used to adjust the sensitivity of the consistency measure.

[0081] To screen high-quality matching points, a dual-threshold decision mechanism is adopted to filter out matching pairs with a geometric consistency measure lower than 0.6 or a confidence level lower than 0.4. The geometric consistency measure threshold of 0.6 means that only when the geometric consistency measure of a matching pair is higher than 0.6 is it considered reasonable in terms of spatial position; the confidence level threshold of 0.4 indicates that only when the confidence level of a matching pair is higher than 0.4 is its matching considered credible. These two thresholds can be adjusted according to the actual application scenario and the characteristics of the dataset.

[0082] To further flexibly adjust the confidence of the matching points, the confidence of the remaining matching pairs is adjusted through dynamic confidence calibration technology. Specifically, the confidence value of a matching pair is dynamically adjusted according to its geometric consistency measure, and the formula is as follows:

[0083] Confidence adj = Confidence·(1 + 0.2·Consistency) (4)

[0084] where Confidence adj is the adjusted confidence level, Confidence is the original confidence level, and Consistency is the consistency measure.

[0085] This can enable those matches with higher consistency to obtain higher confidence scores. Even if their consistency measures are close to but do not reach the threshold standard, they can still be retained by increasing the confidence level, thereby improving the overall quality of the final matching result.

[0086] After the above outlier filtering steps, the retained matching pairs are the optimized target matching pairs. These matching pairs are highly reliable in terms of geometric position and confidence level, and can provide a higher-quality data basis for subsequent pose estimation and map construction.

[0087] Finally, the adaptive exponential moving average (EMA) algorithm is used to optimize the geometric consistency of the optimized target matching pairs to generate the final matching result. In this embodiment, the following steps can be included: Based on the target matching pairs, the decay coefficient is dynamically adjusted according to the increase in the number of training iterations to update the shadow weights, where the decay coefficient is smaller in the initial stage of training for rapid convergence and gradually increases with the increase in the number of iterations to stabilize the parameters; the stability of the matching result is constrained by the consistency loss function, and the consistency loss function combines the geometric consistency measure of the matching pair and the dynamic confidence calibration result to quantify the geometric error and confidence deviation of the target matching pair; according to the updated shadow weights and the consistency loss function, the spatial distribution of the matching pairs is optimized, and the abnormal matching pairs that do not meet the geometric constraints are removed to generate the final matching result.

[0088] In this embodiment, the adaptive exponential moving average algorithm updates the model parameters by introducing a dynamically decaying coefficient that changes over time, enabling it to quickly follow parameter changes in the early stage of training and stabilize the parameters in the later stage, reducing fluctuations. Specifically, the dynamically decaying coefficient is small in the initial stage of training, allowing the model to quickly adjust the parameters to adapt to the data; as training progresses, the decay coefficient gradually increases and approaches 1, making the parameter updates smoother and avoiding drastic fluctuations.

[0089] Based on the above embodiment, the geometric consistency optimization specifically includes the following steps: dynamically adjusting the decay coefficient, updating the shadow weights, constraining the consistency loss function, and optimizing the spatial distribution of the matching pairs. Specifically:

[0090] Dynamically adjusting the decay coefficient means dynamically adjusting the decay coefficient based on the increase in the number of training iterations. For example, in the initial stage of training (such as the first 10 iteration cycles), the decay coefficient can be set to 0.1, and as the number of iterations increases, it gradually increases to 0.9. In this way, in the early stage of training, the model parameters can be quickly updated to adapt to the data; in the later stage of training, the amplitude of parameter updates decreases and tends to be stable.

[0091] In this embodiment, the calculation method of the dynamically decaying coefficient is shown in the following formula (5):

[0092]

[0093] where decay n is the dynamically decaying coefficient at the nth iteration.

[0094] In the initial stage of training, n is small and the decay coefficient is also small, allowing the shadow weights to quickly follow the changes in the model weights. As training progresses, n increases and the decay coefficient gradually approaches 1, thus stabilizing the shadow weights and reducing the fluctuations of the model parameters.

[0095] Updating the shadow weights means updating the shadow weights according to the dynamically decaying coefficient. The update formula for the shadow weights is:

[0096] SV n = decay n ·SV n-1 +(1 - decay n )·V n (6)

[0097] where SV n represents the shadow weights at the nth iteration, decay n is the dynamically decaying coefficient at the nth iteration, and V n represents the current model weights at the nth iteration.

[0098] In this way, the shadow weights can smoothly track the changes in the model weights and reduce parameter fluctuations.

[0099] The consistency loss function constraint is to constrain the stability of the matching results through the consistency loss function. The consistency loss function combines the geometric consistency measure of the matching pairs and the dynamic confidence calibration results, and quantifies the geometric error and confidence deviation of the target matching pairs. The consistency loss function can be expressed by the following formula (7):

[0100]

[0101] where L Consistent represents the consistency loss, which measures the stability of the model in feature matching; λ is a hyperparameter used to balance the weights of the consistency loss and other losses, and can take a value of 0.5; N is the number of matching results; Consistency i represents the consistency measure of the i-th matching result.

[0102] By minimizing the consistency loss function, the model pays more attention to the geometric consistency and confidence of the matching pairs during the optimization process, improving the stability of the matching results.

[0103] Optimizing the spatial distribution of the matching pairs means optimizing the spatial distribution of the matching pairs according to the updated shadow weights and the consistency loss function. Specifically, abnormal matching pairs that do not meet the geometric constraints are removed, and the positions and parameters of the remaining matching pairs are adjusted to make them more reasonable and consistent in space. For example, if some matching pairs have a large deviation in geometric position from other matching pairs under the optimized shadow weights and the consistency loss is large, they will be removed; at the same time, the remaining matching pairs are fine-tuned to make their spatial distribution more in line with the geometric structure of the actual scene, and finally the final matching results are generated.

[0104] To clarify how to obtain the lightweight SuperPoint neural network model and the lightweight SuperGlue neural network model, the following further explains the training processes of the lightweight SuperPoint neural network model and the lightweight SuperGlue neural network model.

[0105] Training of the lightweight SuperPoint neural network model:

[0106] The training of the lightweight SuperPoint neural network model aims to improve the computational efficiency and memory utilization of the SuperPoint algorithm in the feature extraction process through a systematic pruning strategy, while maintaining its high-performance feature matching ability. The entire training process mainly includes two core parts: channel importance evaluation and hierarchical progressive pruning.

[0107] Channel importance evaluation is the basis of the pruning strategy. By using an L1-norm-based method, the importance of each channel in the output feature map of each convolutional layer is evaluated to identify and retain key channels and remove redundant channels. It specifically includes L1-norm calculation, normalization processing, and channel selection.

[0108] During the training process, calculate the L1-norm value of the output feature map of each convolutional layer. For the feature map F generated by each convolutional layer, the L1-norm value of its i-th channel is defined as follows:

[0109] L1-norm(F i ) = ∑ x,y |F i (x,y)| (8)

[0110] where x and y respectively represent the spatial position coordinates of the feature map.

[0111] Normalize these L1-norm values to ensure the comparability of channel importance between different layers. The normalization formula is as follows:

[0112]

[0113] In this way, a normalized value between 0 and 1 is obtained, representing the relative importance of each channel.

[0114] Based on the normalized L1-norm values, sort the channels and select the most important channels according to a preset threshold or ratio. For example, the top 50% of the channels can be selected as important channels, and the remaining channels are considered redundant.

[0115] The hierarchical progressive pruning method includes three main steps: coarse-grained pruning, fine-grained pruning, and global fine-tuning recovery.

[0116] In the coarse-grained pruning stage, sort each layer according to the comprehensive score (importance index based on L1-norm), and remove the 30% of the channels with the lowest scores in each layer, but ensure that at least 50% of the channels in each layer are retained to avoid performance degradation caused by excessive pruning. The specific operations are as follows:

[0117] Sort the channels of each layer from low to high according to the L1-norm score.

[0118] Remove the 30% of the channels with the lowest scores. For example, if a layer has 100 channels, then remove the 30 channels with the lowest scores and retain 70 channels; if a layer has only 50 channels, then remove the 15 channels with the lowest scores and retain 35 channels.

[0119] Fine-tune the pruned model for 10 epochs, and set the learning rate to 1×10-4 , to recover the performance loss caused by pruning.

[0120] In the fine-grained pruning stage, the model is further optimized by pruning the 20% channels with the lowest scores in each layer again. For layers containing skip connections, additional constraints are imposed to ensure that the number of output channels remains unchanged, thus maintaining the integrity of the model structure. The specific operations are as follows:

[0121] Sort the channels in each layer according to the L1-norm scores from low to high.

[0122] Remove the 20% channels with the lowest scores. For example, if a layer has 70 channels after coarse-grained pruning, then another 14 channels with the lowest scores are removed, leaving 56 channels.

[0123] For layers containing skip connections, ensure that the number of output channels remains unchanged. For example, if a layer contains a skip connection, during the pruning process, adjust the number of channels in other layers or use interpolation and other methods to compensate to ensure that the number of output channels is still 70.

[0124] Fine-tune the pruned model for 20 epochs with the learning rate reduced to 1.5×10 -5 , to further recover and optimize the model performance.

[0125] In the global fine-tuning recovery stage, the model is jointly trained on the TUM RGB-D and COCO mixed datasets to enhance its generalization ability. The specific operations are as follows:

[0126] Merge the images and annotation information of the TUM RGB-D dataset and the COCO dataset to form a larger training set.

[0127] Use the joint training set to train the pruned model and adjust the weight parameters of the model so that the model can adapt to the feature distributions of different datasets.

[0128] Enhance the generalization ability of the model through more data samples, and finally output a lightweight SuperPoint neural network model.

[0129] Training of the lightweight SuperPoint neural network model:

[0130] To improve the computational efficiency and memory utilization of the SuperGlue model while ensuring its high performance in feature matching tasks, the embodiments of the present invention provide a systematic pruning strategy, including two core components: attention head pruning and graph network layer pruning.

[0131] By removing one attention head from the SuperGlue model, the computational load can be significantly reduced while maintaining the model's performance as much as possible. The original model contains four attention heads, and each head undertakes a certain proportion of the computational work. After removing one head, the overall computational load is reduced by approximately 25%. The selection of which attention heads to retain is based on their matching contribution on the validation set. Specifically, the contribution of each attention head is measured by its average attention weight across all samples. The contribution Contributionh of each attention head is defined by the formula:

[0132]

[0133] where N is the number of samples, is the weight of the h-th attention head in the i-th sample.

[0134] The importance of each attention head is calculated according to formula (10), and the three attention heads with the highest contribution are retained. This ensures that the model maintains its original performance as much as possible while reducing the computational load.

[0135] To optimize the graph neural network structure in the original model, it can be reduced from 6 layers to 4 layers by deleting the intermediate redundant layers to achieve this goal. This process involves evaluating the importance of each layer to ensure that removing certain layers does not significantly affect the overall performance of the model. The importance of a layer is measured based on the gradient propagation amplitude value. Based on this metric, those layers that have less impact on the final loss are identified and removed. The importance of each layer is defined as follows:

[0136]

[0137] where N is the number of samples, is the gradient of the loss function L with respect to the weight W (k) of the k-th layer, and ||·|| represents the Frobenius norm.

[0138] To further optimize the pruned SuperGlue model, an embodiment of the present invention provides a composite loss function, aiming to simultaneously consider the accuracy, consistency of feature matching, and the sparsity of the model.

[0139] Specifically, the composite loss function Ltotal is defined as:

[0140] L total = L match + 0.5·L consist + 0.1·L sparsity (12)

[0141] where L matchis the basic matching loss, which is used to measure the performance of the model in the feature matching task. It directly reflects the prediction accuracy of the model for the corresponding point pairs in the input image. L consist is the consistency loss, which comes from the outlier filtering module. This module improves the robustness of the model by detecting and filtering out inconsistent matching point pairs. The consistency loss ensures the stability of the model when processing noisy data. L sparsity is the model sparsity regularization term, which is used to control the complexity of the model. By introducing the sparsity regularization term, the number of model parameters can be reduced, thereby reducing the computational cost and memory occupancy.

[0142] To effectively utilize the composite loss function and further improve the performance of the pruned model, the present invention also provides a training strategy. First, in the fine-tuning stage, the structure of the pruned model is frozen, and only the weight parameters are adjusted to maintain the simplicity of the model structure while restoring or improving its performance. Then, the AdamW optimizer is selected. This improved version of the Adam optimizer better controls the generalization ability of the model by adding weight decay to ensure stable convergence during the training process. In addition, the cosine annealing scheduling strategy is adopted to dynamically adjust the learning rate. This method uses a higher learning rate at the beginning of training to accelerate the learning process, and then gradually reduces the learning rate, enabling the model to finely adjust the parameters in the later stage and finally achieving a better optimization effect.

[0143] Through the above pruning and fine-tuning strategies, while maintaining the high performance of the SuperGlue model, the computational efficiency and memory utilization rate can be significantly improved, making it more suitable for running on resource-constrained devices.

[0144] According to a lightweight image feature point detection and matching method provided by an embodiment of the present invention, based on the ORB-SLAM3 framework, by introducing lightweight SuperPoint and SuperGlue models, an efficient and accurate image feature point detection and matching system is realized, significantly improving the positioning and mapping capabilities of ORB-SLAM3 in complex dynamic environments.

[0145] Specifically, as an advanced visual SLAM system, ORB-SLAM3 has the ability to process multiple sensor inputs and can realize functions such as camera positioning, map construction, and loop detection. However, there is room for improvement in its feature detection and matching module in dynamic environments. To address this issue, in practical applications, lightweight SuperPoint and SuperGlue models can be integrated into ORB-SLAM3 to replace the original ORB feature detection and matching methods, thereby improving the adaptability of the system to dynamic environments and reducing the cases of false matches and tracking loss.

[0146] In terms of implementing the positioning and mapping functions of the system, first, the lightweight SuperPoint model can be used to extract multi-scale features from the input image, generating feature information including key point coordinates and descriptors. Then, the image is subjected to superpixel segmentation processing through the RTSS-DBSCAN superpixel segmentation algorithm to generate the superpixel segmentation result. Next, the lightweight SuperGlue model is used to perform initial matching on the extracted feature information to generate initial matching pairs. After that, based on the superpixel segmentation result, outlier filtering is performed on the initial matching pairs to remove inconsistent matching pairs and determine the optimized target matching pairs. Finally, the adaptive exponential moving average algorithm is used to optimize the geometric consistency of the optimized target matching pairs to generate the final matching result.

[0147] Through the above steps, not only the accuracy and robustness of feature matching are improved, but also a higher-quality data basis is provided for the pose estimation and map construction modules of ORB-SLAM3. The experimental results show that this method reduces the absolute trajectory error from 0.041 meters to 0.025 meters, a decrease of 28.1%, significantly improving the matching accuracy of the system. At the same time, model pruning and optimization reduce the single-frame processing time to 31.4 milliseconds and the memory occupancy to 720MB, meeting the real-time and low-power requirements in practical applications.

[0148] In summary, the embodiments of the present invention provide an efficient image feature point detection and matching solution by combining lightweight SuperPoint and SuperGlue models with the ORB-SLAM3 framework, significantly improving the positioning and mapping performance of ORB-SLAM3 in complex environments and expanding its application scope on resource-constrained devices.

[0149] Based on the above lightweight image feature point detection and matching method provided by the embodiments of the present invention, the embodiments of the present invention also provide an image feature matching system for a visual inertial system based on the above method.

[0150] The image feature matching system for a visual inertial system provided in this embodiment is particularly suitable for a deep learning-driven visual SLAM (simultaneous localization and mapping) environment. The system aims to achieve efficient and accurate image feature point detection and matching through a series of optimization steps, thereby improving the positioning and mapping capabilities of the visual inertial system in complex environments.

[0151] Figure 2 It is a flowchart of an image feature matching system for a visual inertial system provided by the embodiments of the present invention. As Figure 2 shown, the method includes the following steps.

[0152] Step S201, input image acquisition.

[0153] The system process starts with obtaining the original input images through a camera or sensors. These images serve as the basis for subsequent processing, and their quality and stability directly affect the final matching effect. In practical applications, the input images may come from various sensors, such as monocular cameras, binocular cameras, or RGB-D cameras, etc., to adapt to different scene requirements and device limitations.

[0154] Step S202, SuperPoint feature extraction.

[0155] After obtaining the input images, use the SuperPoint model to extract features from them. This model can detect key points in the images and generate corresponding descriptors, providing high-quality and discriminative feature information for the subsequent matching steps. In this process, the SuperPoint model uses deep learning techniques to automatically learn the essential representation of image features, thus maintaining the stability of features and the accuracy of matching under different perspectives and lighting conditions.

[0156] Step S203, RTSS-DBSCAN superpixel segmentation.

[0157] After completing feature extraction, apply the RTSS-DBSCAN algorithm to perform superpixel segmentation on the images. This algorithm divides the images into multiple small regions (superpixels) with similar attributes, and each superpixel contains a group of pixels that are similar in color, texture, and position. In this way, not only the representation of the images is simplified, but also important edge information can be retained, providing more favorable structured data support for subsequent feature matching. The result of superpixel segmentation helps to reduce the data volume, while maintaining the key features of the images, improving the processing efficiency and the robustness of matching.

[0158] Step S204, SuperGlue initial matching.

[0159] Based on the extracted feature points, use the SuperGlue algorithm to perform initial matching and generate preliminary matching results.

[0160] The SuperGlue algorithm automatically establishes the matching correspondence between different images by learning the geometric relationships between feature points and the similarity of descriptors. However, the initial matching results may contain a certain number of incorrect matching pairs, and these outliers need to be further filtered and optimized in subsequent steps to improve the accuracy and reliability of matching.

[0161] Step S205, outlier filtering based on superpixels.

[0162] To improve the accuracy and robustness of matching, outlier filtering is performed on the initial matching pairs based on the information of superpixels. By analyzing the geometric consistency measure and confidence of the matching pairs within the superpixel neighborhood, a dual-threshold decision mechanism is adopted to filter out abnormal matching points that do not conform to the expected pattern or deviate significantly from the normal distribution. This process effectively reduces false matches and provides a higher-quality data basis for subsequent pose estimation and map construction. The filtered matching pairs are more reliable in terms of geometric position and confidence, and can better reflect the feature correspondence in the actual scene.

[0163] Step S206, adaptively optimize the matching with EMA.

[0164] Adopt an adaptive exponential moving average (EMA) optimization algorithm to optimize the geometric consistency of the optimized target matching pairs. This algorithm updates the model parameters by introducing a dynamically decaying coefficient that changes over time, enabling it to quickly follow parameter changes in the early stage of training and stabilize the parameters in the later stage to reduce fluctuations. Specifically, the decay coefficient is dynamically adjusted based on the increase in the number of training iterations to update the shadow weights; the stability of the matching results is constrained by the consistency loss function, and the geometric error and confidence deviation of the target matching pairs are quantified by combining the geometric consistency measure and the dynamic confidence calibration results of the matching pairs; according to the updated shadow weights and the consistency loss function, the spatial distribution of the matching pairs is optimized, and abnormal matching pairs that do not conform to the geometric constraints are eliminated, finally generating accurate and consistent matching results.

[0165] Step S207, pose estimation and map construction.

[0166] Based on the finally optimized matching results, pose estimation and map construction are carried out to realize the positioning and mapping functions of the system. Accurate feature matching provides reliable data support for the visual SLAM system, enabling the system to estimate its own pose in real time in an unknown environment and gradually construct a map of the environment. This process is crucial for application fields such as autonomous driving, robot navigation, and augmented reality, and can provide accurate spatial perception capabilities for devices to ensure their stable operation and task execution in complex environments.

[0167] To further improve the efficiency and applicability of the system, the above image feature matching system for the visual inertial system can be subjected to pruning and lightweight processing. Figure 3 This is a schematic diagram of the pruning and lightweight processing provided by the embodiment of the present invention. As Figure 3 shown, the pruning and lightweight processing can include SuperPoint progressive pruning and SuperGlue attention head pruning and graph network layer pruning.

[0168] Progressive pruning is performed on the SuperPoint model. Specifically, it includes three stages: coarse-grained pruning, fine-grained pruning, and global fine-tuning recovery. In the coarse-grained pruning stage, each layer is sorted according to the comprehensive score (importance index based on L1-norm), and the 30% channels with the lowest scores in each layer are removed, but at least 50% of the channels in each layer are retained to avoid performance degradation caused by excessive pruning; in the fine-grained pruning stage, the model is further optimized, and the 20% channels with the lowest scores in each layer are pruned again, and additional constraints are imposed on the layers containing skip connections to ensure that the number of output channels remains unchanged; in the global fine-tuning recovery stage, the model is jointly trained on the TUM RGB-D and COCO mixed datasets to enhance its generalization ability, and finally a lightweight SuperPoint neural network model is output. Through this pruning strategy, while reducing the number of model parameters and computational volume, the feature extraction performance of the model is maintained as much as possible, making it more suitable for running on resource-constrained devices.

[0169] For the SuperGlue model, optimization is carried out by means of attention head pruning and graph network layer pruning. By removing one attention head in the model, the computational volume can be significantly reduced while maintaining the performance of the model as much as possible. Which attention heads to retain is selected based on their matching contribution on the validation set; at the same time, the graph neural network structure in the original model is reduced from 6 layers to 4 layers, and this goal is achieved by deleting the intermediate redundant layers, and the importance of each layer is evaluated based on the gradient propagation amplitude value to ensure that removing some layers will not significantly affect the overall performance of the model.

[0170] In addition, a composite loss function is introduced, which aims to simultaneously consider the accuracy, consistency of feature matching, and the sparsity of the model, and the pruned model structure is frozen in the fine-tuning stage, and only the weight parameters are adjusted. The AdamW optimizer is selected and the cosine annealing scheduling strategy is used to dynamically adjust the learning rate to further improve the performance and generalization ability of the pruned model.

[0171] Through the above pruning and lightweight processing, not only the computational complexity and memory occupancy of the model are reduced, but also it can better adapt to resource-constrained visual inertial systems on the premise of maintaining a high matching accuracy, expanding its scope and scenarios in practical applications.

[0172] Figure 4 This is a schematic structural diagram of a lightweight SuperPoint neural network model provided by an embodiment of the present invention. As Figure 4 shown, the neural network model mainly includes an encoder, a feature point decoder, and a descriptor decoder.

[0173] Input image: The size of the input image is H×W×1, where H and W respectively represent the height and width of the image, and 1 represents a single-channel grayscale image.

[0174] The encoder performs convolution and pooling operations: The encoder consists of multiple convolutional layers and pooling layers. The convolutional layer slides a convolutional kernel over the image, calculates the dot product between the convolutional kernel and the local region of the image, and extracts the local features of the image. The pooling layer downsamples the feature map by methods such as taking the maximum value or average value, reducing the amount of data while retaining important features. After a series of convolution and pooling operations, the spatial resolution of the image gradually decreases, and finally a feature map with a size of (H / 8)×(W / 8)×64 is obtained. This process not only extracts the high-level features of the image, but also greatly reduces the amount of computation and the number of parameters, improving the efficiency of the model.

[0175] The keypoint decoder performs a convolution operation: The keypoint decoder receives the feature map output by the encoder and performs a convolution operation on it. The purpose of this convolution operation is to further extract the keypoint information in the feature map and generate a feature map with a size of (H / 8)×

[0176] (W / 8)×65. Among them, 65 represents the number of channels of the feature map, and each channel corresponds to different types of features.

[0177] Softmax and reshaping: After the convolution operation, the softmax function is used to process the feature map, converting the value of each pixel point in the feature map into a probability value between 0 and 1, indicating the possibility that the pixel point is a keypoint. Then, through the reshaping operation, the feature map is converted into a keypoint heatmap with a size of H×W×1. The value of each pixel point in the keypoint heatmap represents the confidence that the position is a keypoint, and the larger the value, the more likely it is to be a keypoint. This process realizes the mapping from the high-level features extracted by the encoder to the specific keypoint positions, providing a basis for subsequent keypoint matching.

[0178] The descriptor decoder performs a convolution operation: The descriptor decoder also receives the feature map output by the encoder and performs a convolution operation on it. The purpose of this convolution operation is to generate a descriptor feature map with a size of (H / 8)×(W / 8)×D, where D represents the dimension of the descriptor, usually a relatively large value, such as 256. Each pixel point in the descriptor feature map corresponds to a D-dimensional descriptor vector, which is used to describe the features of the image region around that position.

[0179] Bicubic interpolation and L2 norm: To restore the resolution of the descriptor feature map to the same size as the input image, the bicubic interpolation method is used to upsample the feature map. Bicubic interpolation generates new pixel values by calculating the weighted average of surrounding pixel points, thereby increasing the resolution of the feature map. The size of the upsampled feature map is H×W×D. Then, the L2 norm operation is applied to the upsampled feature map to normalize the norm of each descriptor vector to 1, obtaining the final descriptor map. The normalized descriptor map has better robustness and discriminability and can be more effectively used for feature point matching.

[0180] SuperPoint is trained in a self-supervised learning manner without manually labeled feature point data. During training, images from different perspectives of the same scene are used as input, allowing the model to automatically learn feature point detection and description. Specifically, for a pair of images from different perspectives, the model extracts their feature points and descriptors respectively, and then finds the corresponding feature point pairs between the two images through a matching algorithm. Based on these matching pairs, a matching loss function is defined to measure the difference between the feature points and descriptors predicted by the model and the true matches. By minimizing the matching loss function, the model can continuously adjust its own parameters, optimize the performance of feature point detection and description, so that it can accurately extract feature points with high discriminability and stability in new images and generate corresponding descriptors. This self-supervised learning method not only reduces the cost of manual annotation but also enables the model to be trained on large-scale data, improving its generalization ability and adaptability.

[0181] Figure 5 This is a schematic diagram of a network pruning process provided by an embodiment of the present invention. As Figure 5 shown, it includes a comparison of the network structures before and after pruning, and the process of removing convolutional kernels that have less impact on the model performance by evaluating the importance of each layer of convolutional kernels.

[0182] Before pruning: The network structure contains redundant convolutional kernels.

[0183] Initial network: The left side shows the network structure before pruning, which contains multiple convolutional layers. There are multiple convolutional kernels in each convolutional layer. For example, in the i-th convolutional layer, there are multiple convolutional kernels such as C_{i1}, C_{i2}, C_{i3}, C_{i4},..., C_{in}.

[0184] Channel scaling factors: Between convolutional layers, there are channel scaling factors used to evaluate the importance of each convolutional kernel. These factors determine their importance by quantifying the impact of each convolutional kernel on the model performance. For example, the figure shows the channel scaling factors from the i-th layer to the i+1-th layer, including values such as 1.170, 0.001, 0.290, 0.003, 0.820. The larger these values, the greater the contribution of the corresponding convolutional kernel to the model performance.

[0185] Pruning process:

[0186] Evaluating the importance of convolutional kernels: By calculating the channel scaling factors, the importance of each convolutional kernel is evaluated. For example, in the i-th convolutional layer, the channel scaling factor of C_{i1} is 1.170, indicating its relatively large contribution to the model performance; while the channel scaling factor of C_{i2} is 0.001, indicating its relatively small contribution and it may be a redundant convolutional kernel.

[0187] Removing redundant convolutional kernels: According to the magnitudes of the channel scaling factors, the convolutional kernels with less impact on the model performance are removed. For example, in the i-th convolutional layer, convolutional kernels such as C_{i2} and C_{i4} with relatively small channel scaling factors may be removed, while important convolutional kernels such as C_{i1}, C_{i3}, C_{in} are retained.

[0188] Automatically deleting unnecessary parts: The pruning process automatically deletes unnecessary parts by quantifying the impact degree of each convolutional kernel. This method not only reduces the computational amount of the network but also maintains the model performance and even improves the generalization ability of the model in some cases.

[0189] After pruning: A streamlined network structure is obtained.

[0190] Compact network: The right side shows the network structure after pruning. By removing redundant convolutional kernels, the network structure becomes more compact. For example, in the i-th convolutional layer, important convolutional kernels such as C_{i1}, C_{i3}, C_{in} are retained, while redundant convolutional kernels such as C_{i2}, C_{i4} are removed.

[0191] Adjustment of channel scaling factors: In the pruned network, the channel scaling factors are also adjusted accordingly to reflect the importance of the remaining convolutional kernels. For example, the channel scaling factors from the i-th layer to the i+1-th layer are adjusted to values such as 1.170, 0.290, 0.820 to ensure that the network maintains efficient feature extraction ability while being streamlined.

[0192] Through the above pruning process, the computational complexity of the network is significantly reduced, and the efficiency and performance of the model are optimized, making it more suitable for running on resource-constrained devices. This method is of great significance in the optimization of deep learning models, especially in application scenarios that require efficient computing and low memory occupancy, such as mobile devices, embedded systems, etc.

[0193] Figure 6 It is a schematic diagram of an improved SuperGlue feature matching framework provided by an embodiment of the present invention. As Figure 6 shown, this framework combines a local consistency constraint model, a dual-threshold decision mechanism, and a dynamic confidence calibration technique to effectively identify and remove unreasonable matching points, thereby improving the performance of the entire system.

[0194] Input images: The two images to be matched come from different perspectives or different time points, and feature matching is required to determine the corresponding relationship between them.

[0195] Feature encoder: The feature encoder is responsible for extracting feature points and descriptors in the image. It processes the input image through a multi-layer convolutional neural network to generate feature information containing the coordinates and descriptors of key points. The output of the feature encoder is the basis for the subsequent matching process, and its quality directly affects the final matching effect.

[0196] In the feature matching module, self-attention or cross-attention mechanisms are used to analyze the relationships between feature points. The self-attention mechanism allows the model to analyze the interactions between feature points within a single image, while the cross-attention mechanism is used to analyze the corresponding relationships between feature points in two images. Through these mechanisms, the model can better understand the context information of feature points and improve the accuracy of matching.

[0197] The descriptors of feature points are mapped to an assignment matrix through a multi-layer perceptron (MLP). Each element of this matrix represents the matching probability or similarity between two feature points. The assignment matrix is the core of the matching process, which determines which feature point pairs are potential matching pairs.

[0198] The outlier filtering module identifies and removes abnormal matching points that do not conform to the expected pattern or deviate significantly from the normal distribution by analyzing the consistency of each matching point with its superpixel region. Specifically, a local consistency constraint model is used to calculate the geometric consistency measure of the matching point within the superpixel neighborhood to evaluate the rationality of the spatial position of the matching point.

[0199] A dual-threshold decision mechanism is adopted to filter out matching pairs with a geometric consistency measure lower than 0.6 or a confidence level lower than 0.4. By setting these two thresholds, false matches can be effectively removed, and high-quality matching point pairs can be retained.

[0200] Real-time supervised segmentation algorithm (such as DBSCAN): Use a real-time supervised segmentation algorithm (such as DBSCAN) to perform clustering analysis on feature points to further optimize the matching results. The DBSCAN algorithm can automatically identify and remove isolated matching points that do not conform to the overall distribution according to the distribution of feature points, improving the stability and reliability of the matching results.

[0201] Finally, the adaptive exponential moving average (EMA) algorithm is used to optimize the geometric consistency of the optimized target matching pairs. The EMA algorithm updates the model parameters by introducing a dynamically decaying coefficient that changes over time, enabling it to quickly follow parameter changes in the early stage of training and stabilize the parameters in the later stage to reduce fluctuations. Specifically, the decay coefficient is dynamically adjusted based on the increase in the number of training iterations to update the shadow weights; the stability of the matching results is constrained by the consistency loss function, combining the geometric consistency metric of the matching pairs and the dynamically calibrated confidence results to quantify the geometric error and confidence deviation of the target matching pairs; according to the updated shadow weights and the consistency loss function, the spatial distribution of the matching pairs is optimized, and abnormal matching pairs that do not conform to the geometric constraints are removed to generate the final matching results.

[0202] Through the accurate matching results after outlier filtering and EMA optimization, it can be seen that the matching point pairs are more concentrated and reasonable, and the situation of false matches is significantly reduced. These accurate matching results will provide a high-quality data basis for subsequent pose estimation and map construction, ensuring the efficient positioning and mapping capabilities of the system in complex environments.

[0203] Through the improved SuperGlue feature matching framework above, it can not only effectively reduce false matches, but also enhance the stability and reliability of the matching results, significantly improving the overall performance of the system.

[0204] Based on the above lightweight image feature point detection and matching method provided by the embodiments of the present invention, the embodiments of the present invention also provide a lightweight image feature point detection and matching device, such as Figure 7 shown, the lightweight image feature point detection and matching device includes an extraction module 701, a superpixel segmentation module 702, a matching module 703, a filtering module 704, and an optimization module 705.

[0205] The extraction module 701 is used to perform multi-scale feature extraction on the target image through a lightweight SuperPoint neural network model to generate feature information including key point coordinates and descriptors.

[0206] The superpixel segmentation module 702 is used to perform superpixel segmentation processing on the image based on the feature information through the RTSS-DBSCAN superpixel segmentation algorithm to generate a superpixel segmentation result.

[0207] A matching module 703, configured to perform initial matching on the feature information through a lightweight SuperGlue neural network model to generate initial matching pairs.

[0208] A filtering module 704, configured to perform outlier filtering on the initial matching pairs based on the superpixel segmentation result to remove inconsistent matching pairs in the initial matching pairs and determine optimized target matching pairs.

[0209] An optimization module 705, configured to perform geometric consistency optimization on the optimized target matching pairs by using an adaptive exponential moving average algorithm to generate a final matching result.

[0210] Wherein, the lightweight SuperPoint neural network model is generated by performing hierarchical progressive pruning on the original SuperPoint neural network model, and the hierarchical progressive pruning includes coarse-grained pruning, fine-grained pruning, and global fine-tuning and recovery processing; the lightweight SuperGlue neural network model is generated by performing attention head pruning and graph network layer pruning on the original SuperGlue neural network model.

[0211] According to an embodiment of the present invention, the extraction module 701 is further configured to perform grayscale processing on the initial image to generate a single-channel grayscale image corresponding to the initial image; perform normalization processing on the single-channel grayscale image to map the pixel values of the single-channel grayscale image to a preset range to generate a normalized image; input the normalized image into a preprocessing layer of a convolutional neural network, and extract local features through a convolutional kernel and perform spatial smoothing processing to generate a target image.

[0212] According to an embodiment of the present invention, the extraction module 701 is specifically configured to perform multi-scale convolution and pooling operations on the target image to generate a feature pyramid; screen local optimal feature points at each scale in the feature pyramid through a non-maximum suppression algorithm; map the coordinates of the screened local optimal feature points to the original image space through a scale adjustment layer, and output multi-scale feature points and their descriptors.

[0213] According to an embodiment of the present invention, the superpixel segmentation module 702 is specifically configured to calculate the similarity between pixels based on the key point coordinates and response values in the feature information; merge adjacent pixels with similarity higher than a preset threshold into superpixels through a density clustering algorithm to generate superpixel identifiers; dynamically adjust the superpixel boundaries based on the superpixel identifiers to retain image edge information, and output a superpixel segmentation result.

[0214] According to an embodiment of the present invention, the filtering module 704 is specifically configured to calculate the geometric consistency metric of each initial matching pair within the corresponding superpixel neighborhood based on the superpixel segmentation result; adopt a dual-threshold decision mechanism to filter out the matching pairs with a geometric consistency metric lower than 0.6 or a confidence level lower than 0.4; adjust the confidence levels of the remaining matching pairs through a dynamic confidence calibration technique, and output the optimized target matching pairs.

[0215] According to an embodiment of the present invention, the optimization module 705 is specifically configured to, based on the target matching pairs, dynamically adjust the attenuation coefficient according to the increase in the number of training iterations to update the shadow weights, where the attenuation coefficient is smaller in the initial stage of training for rapid convergence and gradually increases with the increase in the number of iterations to stabilize the parameters; constrain the stability of the matching result through a consistency loss function, and the consistency loss function combines the geometric consistency metric and the dynamic confidence calibration result of the matching pairs to quantify the geometric error and confidence deviation of the target matching pairs; optimize the spatial distribution of the matching pairs according to the updated shadow weights and the consistency loss function, eliminate the abnormal matching pairs that do not conform to the geometric constraints, and generate the final matching result.

[0216] According to an embodiment of the present invention, the lightweight SuperPoint neural network model is trained based on the following method: remove 30% of the channels with the lowest L1 norm scores in each layer of the original SuperPoint neural network model, retain at least 50% of the channels and perform the first fine-tuning; remove 20% of the channels with the lowest scores in each layer, keep the number of output channels unchanged for the layers containing skip connections, and perform the second fine-tuning; jointly train the pruned model on the TUM RGB-D and COCO mixed datasets, and output the lightweight SuperPoint neural network model.

[0217] According to an embodiment of the present invention, the lightweight SuperGlue neural network model is trained based on the following method: evaluate the contribution degree of each attention head in the original SuperGlue neural network model to the matching result through the validation set, and remove the attention heads with a contribution degree lower than the preset threshold; identify redundant graph neural network layers based on the gradient propagation amplitude value, and remove the layers with an impact on the matching accuracy less than 5%; perform fine-tuning on the pruned model, and output the lightweight SuperGlue neural network model.

[0218] An embodiment of the present invention further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program that can be executed by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method of the embodiment of the present invention.

[0219] An embodiment of the present invention also provides a non-transitory machine-readable medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiment of the present invention.

[0220] An embodiment of the present invention also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiment of the present invention.

[0221] Referring Figure 8 , a block diagram of an electronic device that can be a server or a client according to an embodiment of the present invention will now be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0222] As Figure 8 shown, the electronic device includes a computing unit 801, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0223] Multiple components in the electronic device are connected to the I / O interface 805, including: an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 can be any type of device capable of inputting information into the electronic device. The input unit 806 can receive input digital or character information and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 807 can be any type of device capable of presenting information and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 can include, but is not limited to, a magnetic disk, an optical disk. The communication unit 809 allows the electronic device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0224] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a CPU, a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above. For example, in some embodiments, the method embodiments of the present invention can be implemented as a computer program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device via the ROM 802 and / or the communication unit 809. In some embodiments, the computing unit 801 can be configured to execute the above-described method in any other suitable manner (e.g., by means of firmware).

[0225] The computer program for implementing the method of the embodiments of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0226] In the context of embodiments of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable signal medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0227] It should be noted that the term "including" and its variations used in embodiments of the present invention are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The modifiers "a" and "multiple" mentioned in embodiments of the present invention are illustrative rather than restrictive, and those skilled in the art should understand that, unless clearly specified otherwise in the context, it should be understood as "one or more".

[0228] In the method embodiments provided by the present invention, the various steps recorded can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of protection of the present invention is not limited in this regard.

[0229] The term "embodiment" in this specification means that the specific features, structures, or characteristics described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various positions in the specification and does not necessarily mean the same embodiment, nor does it mean being independent or alternative to other embodiments and mutually exclusive. The various embodiments in this specification are described in a related manner, and the same or similar parts among the various embodiments are cross-referred to. In particular, for device, apparatus, and system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts refer to the partial description of the method embodiments.

[0230] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the appended claims.

Claims

1. A lightweight image feature point detection and matching method, characterized in that Including: Performing multi-scale feature extraction on a target image through a lightweight SuperPoint neural network model to generate feature information including key point coordinates and descriptors; Based on the feature information, performing superpixel segmentation processing on the target image through an RTSS-DBSCAN superpixel segmentation algorithm to generate a superpixel segmentation result; Performing initial matching on the feature information through a lightweight SuperGlue neural network model to generate initial matching pairs; Based on the superpixel segmentation result, filtering outliers from the initial matching pairs to remove inconsistent matching pairs in the initial matching pairs and determining optimized target matching pairs; Using an adaptive exponential moving average algorithm to perform geometric consistency optimization on the optimized target matching pairs to generate a final matching result; Among them, the lightweight SuperPoint neural network model is generated by performing hierarchical progressive pruning processing on the original SuperPoint neural network model, and the hierarchical progressive pruning processing includes coarse-grained pruning processing, fine-grained pruning processing, and global fine-tuning and recovery processing; the lightweight SuperGlue neural network model is generated by performing attention head pruning processing and graph network layer pruning processing on the original SuperGlue neural network model.

2. The method according to claim 1, wherein The method further includes: Performing grayscale processing on an initial image to generate a single-channel grayscale image corresponding to the initial image; Performing normalization processing on the single-channel grayscale image to map the pixel values of the single-channel grayscale image to a preset range to generate a normalized image; Inputting the normalized image into a preprocessing layer of a convolutional neural network, extracting local features through a convolution kernel and performing spatial smoothing processing to generate the target image.

3. The method according to claim 1, wherein The performing multi-scale feature extraction on a target image through a lightweight SuperPoint neural network model to generate feature information including key point coordinates and descriptors includes: Performing multi-scale convolution and pooling operations on the target image to generate a feature pyramid; Screening local optimal feature points at each scale in the feature pyramid through a non-maximum suppression algorithm; Mapping the coordinates of the screened local optimal feature points to the original image space through a scale adjustment layer and outputting the multi-scale feature points and their descriptors.

4. The method according to claim 1, characterized in that The performing superpixel segmentation processing on an image through an RTSS-DBSCAN superpixel segmentation algorithm based on the feature information to generate a superpixel segmentation result includes: Calculating the similarity between pixels based on the key point coordinates and response values in the feature information; Merging adjacent pixels with a similarity higher than a preset threshold into superpixels through a density clustering algorithm to generate superpixel identifiers; Based on the superpixel identifiers, dynamically adjusting the superpixel boundaries to retain image edge information and outputting the superpixel segmentation result.

5. The method according to claim 1, wherein The filtering outliers from the initial matching pairs based on the superpixel segmentation result to remove inconsistent matching pairs in the initial matching pairs and determining optimized target matching pairs includes: Based on the superpixel segmentation results, calculate the geometric consistency measure of each initial matching pair within the corresponding superpixel neighborhood; Adopt a double-threshold decision mechanism to filter out the matching pairs with a geometric consistency measure lower than 0.6 or a confidence level lower than 0.4; Adjust the confidence levels of the remaining matching pairs through a dynamic confidence calibration technique and output the optimized target matching pairs.

6. The method according to claim 1, wherein The geometric consistency optimization of the optimized target matching pairs by using the adaptive exponential moving average algorithm to generate the final matching result includes: Based on the target matching pairs, dynamically adjust the decay coefficient according to the increase in the number of training iterations to update the shadow weights, where the decay coefficient is smaller in the initial stage of training for rapid convergence and gradually increases with the increase in the number of iterations to stabilize the parameters; Constrain the stability of the matching result through a consistency loss function, where the consistency loss function combines the geometric consistency measure and the dynamic confidence calibration result of the matching pairs to quantify the geometric error and confidence deviation of the target matching pairs; According to the updated shadow weights and the consistency loss function, optimize the spatial distribution of the matching pairs, eliminate the abnormal matching pairs that do not meet the geometric constraints, and generate the final matching result.

7. The method according to claim 1, wherein The lightweight SuperPoint neural network model is trained based on the following method: Remove 30% of the channels with the lowest L1 norm scores in each layer of the original SuperPoint neural network model, retain at least 50% of the channels and perform the first fine-tuning; Remove 20% of the channels with the lowest scores in each layer, keep the number of output channels unchanged for the layers containing skip connections, and perform the second fine-tuning; Jointly train the pruned model on the TUM RGB-D and COCO mixed datasets and output the lightweight SuperPoint neural network model.

8. The method according to claim 1, characterized in that, The lightweight SuperGlue neural network model is trained based on the following method: Evaluate the contribution degree of each attention head in the original SuperGlue neural network model to the matching result through the validation set, and remove the attention heads with a contribution degree lower than the preset threshold; Identify redundant graph neural network layers based on the gradient propagation amplitude value, and remove the layers with an impact on the matching accuracy less than 5%; Fine-tune the pruned model and output the lightweight SuperGlue neural network model.

9. A lightweight image feature point detection and matching device, characterized in that, Including: An extraction module for multi-scale feature extraction of the target image through a lightweight SuperPoint neural network model to generate feature information including key point coordinates and descriptors; A superpixel segmentation module for performing superpixel segmentation processing on the target image through the RTSS-DBSCAN superpixel segmentation algorithm based on the feature information to generate superpixel segmentation results; A matching module for initial matching of the feature information through a lightweight SuperGlue neural network model to generate initial matching pairs; A filtering module for outlier filtering of the initial matching pairs based on the superpixel segmentation results to remove the inconsistent matching pairs in the initial matching pairs and determine the optimized target matching pairs; An optimization module for geometric consistency optimization of the optimized target matching pairs by using an adaptive exponential moving average algorithm to generate a final matching result; Among them, the lightweight SuperPoint neural network model is generated by performing hierarchical progressive pruning on the original SuperPoint neural network model. The hierarchical progressive pruning includes coarse-grained pruning, fine-grained pruning, and global fine-tuning and recovery processing; the lightweight SuperGlue neural network model is generated by performing attention head pruning and graph network layer pruning on the original SuperGlue neural network model.

10. An electronic device, comprising: A processor and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to execute the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Transform model lightweight method and power transmission line image analysis method

    CN120910703A

  • Stereo matching method based on hierarchical graph convolution

    CN121330024A

  • Real-time feature point extraction and matching method based on star operation and GateMLP

    CN121438013A

  • Real-time feature point extraction and matching method based on star operation and gateMLP

    CN121438013B

  • Bimodal image registration method based on deep learning and GPU parallel optimization

    CN122289336A