Key point detection method based on deep learning

By designing a lightweight key point detection method based on deep learning, the problems of small number of key point matching, low accuracy and high computational complexity in the existing technology are solved, and efficient and accurate key point detection and descriptor extraction are achieved, meeting the needs of real-time operation.

CN120088510APending Publication Date: 2025-06-03HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510053565.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing key point detection methods have small number of key point matches, low accuracy and high computational complexity, which cannot meet the needs of real-time operation.

Method used

A lightweight key point detection method based on deep learning is designed, including feature encoder, multi-level feature fusion module, reliability detector and repeatability detector and descriptor extractor. The reliability and repeatability of key points were jointly optimized by using a fusion non-maximum suppression NMS sampling strategy and four loss functions (bidirectional negative log likelihood loss, bidirectional L1 distance loss, local similarity loss and local peak loss).

Benefits of technology

It realizes efficient key point detection and descriptor extraction, improves key point matching performance and accuracy, reduces computational complexity, and maintains efficient performance in real-time operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088510A_ABST
    Figure CN120088510A_ABST
Patent Text Reader

Abstract

The invention provides a key point detection method based on deep learning, and the method comprises the steps: designing a lightweight network structure which comprises a feature encoder, a multistage feature fusion module, a reliability detector, a repeatability detector and a descriptor extractor, and outputting a reliability score graph, a repeatability score graph and dense descriptors; a non-maximum suppression sampling strategy is designed and fused, coordinates of a limited number of key points are sampled, bidirectional negative logarithm likelihood loss and bidirectional L1 distance loss are proposed to carry out coupling training on descriptors and reliability of the key points, and local similarity loss and local peak loss are proposed to optimize the positions of the key points. Compared with the prior art, the method not only achieves the performance equivalent to that of the most advanced SOTA method in feature matching, but also greatly shortens the reasoning time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image detection and recognition, and particularly relates to a key point detection method based on deep learning. Background Art

[0002] In autonomous driving, the Simultaneous Localization and Mapping (SLAM) technology has become a crucial link. The goal of SLAM is to obtain information about the surrounding environment through sensors (such as lidar, cameras, IMUs, etc.) in an unknown environment, construct a map of the environment, and simultaneously update the vehicle's self-position in real time. The application of SLAM technology is not limited to autonomous driving, but also widely used in fields such as robot navigation, autonomous flight of drones, and augmented reality. Especially in the autonomous driving scenario, SLAM allows the vehicle to perceive and understand the environment without a pre-constructed map, thereby achieving autonomous navigation. SLAM can be divided into sparse SLAM and dense SLAM. The former is used for feature point mapping in the scene, while the latter focuses more on the overall dense point cloud mapping of the scene. The core technologies of SLAM include sensor data fusion, motion estimation, feature extraction, data association, and map construction, etc.

[0003] In the SLAM system, feature point detection, as one of its key modules, plays a crucial role. The purpose of feature point detection is to identify a set of stable, unique, and easily recognizable points (such as corner points, edges, etc.) in the environment to ensure that the system can perform matching and tracking in consecutive frames. Through these feature points, the SLAM system can calculate motion estimation, thereby realizing the self-positioning of the device. At the same time, these feature points also provide marker points for map construction, helping the system to associate the same environmental structure from different perspectives. The detection effect of feature points directly affects the accuracy and stability of the SLAM system. Therefore, a variety of feature detection algorithms are widely used in SLAM, such as traditional algorithms SIFT and ORB, and deep learning models such as SuperPoint.

[0004] Current mainstream methods include SuperPoint, DISK, and R2D2, etc. They have made innovations in key point detection and descriptor matching respectively, but there are also some limitations. SuperPoint first trains a MagicPoint model on a synthetic dataset, then adopts a homography adaptive strategy to bootstrap the score map on real images, and trains its descriptor using triplet loss. Although its method has a certain degree of robustness, the triplet loss cannot capture global features, resulting in a small number of generated key points. DISK relaxes key point detection and descriptor matching into a probabilistic process and optimizes the network through reinforcement learning, improving the overall matching performance. However, this method has more model parameters and higher computational complexity, making it difficult to meet the requirements of real-time operation. R2D2 defines key points as reliable and repeatable positions in the image and uses the AP loss function to train the network to improve the reliability of detection. Nevertheless, due to the lack of a constraint mechanism for the reliability of key points, this method may still generate key points in unreliable regions, affecting the detection accuracy. Summary of the Invention

[0005] Object of the Invention: Aiming at the problems of few key point matches and low accuracy, as well as high computational complexity and inability to complete real-time operation tasks in existing key point detection methods, the present invention proposes a deep learning-based key point detection method.

[0006] Technical Solution: The present invention discloses a deep learning-based key point detection method, including the following steps:

[0007] Step 1: Obtain a publicly available image dataset, where the image dataset includes two images of the same scene from different perspectives, and preprocess the image dataset;

[0008] Step 2: Design a lightweight network structure, where the lightweight network structure includes a feature encoder, a multi-level feature fusion module, a reliability detector, a repeatability detector, and a descriptor extractor; the lightweight network structure inputs two images of the same scene from different perspectives respectively, and outputs a reliability score map R, a repeatability score map S, and a dense descriptor D correspondingly;

[0009] Step 3: Propose a fused non-maximum suppression (NMS) sampling strategy, and use the fused non-maximum suppression (NMS) sampling strategy to sample a finite number of key points with the highest local scores on the reliability score map, and constrain the reliability score map to help the model focus only on locally reliable key points;

[0010] Step 4: Propose four loss functions to jointly optimize the key point reliability and repeatability, namely bidirectional negative log-likelihood loss, bidirectional L1 distance loss, local similarity loss, and local peak loss; for key point reliability, the bidirectional L1 distance loss is used to train the reliability score map; for key point repeatability, a local peak loss is designed to avoid the phenomenon of dense crowding on the repeatability score map; for key point descriptors, the bidirectional negative log-likelihood loss is adopted to transform the descriptor similarity problem into a probability problem;

[0011] Step 5: Use the preprocessed image dataset to train the lightweight network structure constructed in Steps 2 to 4 to obtain the trained lightweight network structure;

[0012] Step 6: Input the reference image I 1 and the image I 2 to be matched into the trained lightweight network structure respectively, and output the reliability score maps R 1 and R 2 , the repeatability score maps S 1 and S 2 , as well as the dense descriptors D 1 and D 2 , and repeat the content of Steps 4 and 5;

[0013] Step 7: Obtain the key point coordinates in the two images according to the reliability score maps R 1 and R 2 , the repeatability score maps S 1 and S 2 , obtain the pre-matched key points in the two images through the K-nearest neighbor matching algorithm, and then screen out the correct matching pairs through the RANSAC algorithm.

[0014] Furthermore, the lightweight network structure is specifically as follows:

[0015] The feature encoder encodes the input image into a feature map. H is the height of the input image, and W is the width of the input image. This encoder consists of 4 modules. The input image of the first module is a grayscale image, i.e., C = 1, which contains 2 basic layers. The latter 3 modules all contain 1 MaxPooling layer and 2 basic layers. The basic layer includes a two-dimensional convolution Conv2d, Relu + BatchNorm; the downsampling rates are 1 / 2, 1 / 4, 1 / 4 respectively. The channel depths of the 4 modules are {32, 64, 128, 128}, and the final spatial resolution is

[0016] The multi-level feature fusion module fuses the multi-level features of the feature encoder, selects the output features of 4 modules, uses 1×1 convolution and bilinear interpolation upsampling to adjust the channel depth, and then concatenates them together;

[0017] The two detectors of the reliability detector and the repeatability detector respectively output H×W×1 feature maps. The repeatability map is normalized by softplus to the repeatability score map S, and the reliability map is activated by softmax to the reliability score map R;

[0018] The descriptor extractor outputs an H×W×(dim + 1) feature map, where the first dim channels are the dense descriptor D.

[0019] Furthermore, the implementation process of the fusion non-maximum suppression NMS sampling strategy is as follows:

[0020] Step 3.1: Search for the pixel with the maximum score within the local window. This operation is equivalent to argmax in the local window N*N:

[0021]

[0022] Among them, R(i,j) represents the reliability score of the local coordinates (i,j), and W(x,y) represents the local window centered on each pixel position (x,y) in the graph, including all neighboring pixels (i,j);

[0023] Step 3.2: NMS map The score map first suppresses the non-maximum score r = R(u,v) in the local N×N window:

[0024]

[0025] In the formula, r max is the local highest score. To avoid falling into the local optimal solution, a threshold th is applied to the NMS map score map to filter out low response scores, and the top top scores are selected. Finally, the index values indice of these scores are obtained, and the key point coordinates (x 1 kpt ,y kpt ) in the graph I are extracted as:

[0026]

[0027] Step 3.3: If the graph I 2 is synthesized after the change of I 1 , then the optical flow map F∈R H×W×2 is obtained, and the corresponding key point coordinates in I 2 are obtained through F ​Each position (x, y) stores a two-dimensional vector F(x, y) = (u(x, y), v(x, y)), where u and v respectively represent the 1 horizontal displacement and vertical displacement of the pixel (x, y) in 2 Figure I:

[0028]

[0029] Furthermore, the implementation process of the bidirectional negative log-likelihood loss is as follows:

[0030] Step 4.1: Set the dense descriptors D 1 and D 2 to be obtained from the dense mappings D(·,·) of Figure I 1 and Figure I 2 respectively. And the dense descriptors D 1 and the i-th row D 1 (i,·) and D 2 (i,·) respectively correspond to the descriptors of the corresponding points in Figure I 1 and Figure I 2 . Calculate the corresponding descriptors according to the key-point coordinates obtained by the fusion non-maximum suppression NMS sampling strategy in Step 3:

[0031]

[0032] Step 4.2: The similarity matrix is obtained from . Considering the symmetry of the matching, take two matching directions simultaneously to obtain the bidirectional negative log-likelihood loss L des , where the similarity of the corresponding features is located on the main diagonal M ii of M, and softmax is calculated along the "columns":

[0033]

[0034] Furthermore, the implementation process of the bidirectional L1 distance loss is as follows:

[0035] Step 5.1: During the training process, supervise the reliability score map through bidirectional softmax, and use and as the matching confidence, denoted as

[0036]

[0037] Step 5.2: Given the matching confidence and In the case of, directly use the L1 distance loss to supervise the reliability mapping to obtain the bidirectional L1 distance loss L rel :

[0038]

[0039] where λ is the sigmoid activation function, ⊙ is the Hadamard product, and R 1 = R 1 (x kpts , y kpts ), R 1 and R 2 are the reliability score maps obtained from Figure I 1 and I 2 respectively. For the bidirectional L1 distance loss L rel , only perform backpropagation gradients on R 1 and R 2 .

[0040] Furthermore, the implementation process of the local similarity loss is as follows:

[0041] Assume that Figure I 1 and I 2 are two images of the same scene. In the fusion non-maximum suppression (NMS) sampling strategy, the corresponding point relationship between the two images has been obtained through the optical flow map . Let H be the height of the input image and W be the width of the input image. According to the optical flow map, deform S 2 to generate

[0042] To make all local maxima in the repeatability score map S 1 correspond to the maxima in , the key idea is to maximize the local similarity similar between S 1 and . When is maximized, the two repeatability score maps are the same; define the local block set P = {p}, which contains local blocks of size N×N in the image size {1,..., W}×{1,..., H}, and define the local similarity loss L sim (I 1 , I 2 , F):

[0043]

[0044] where represents the flattened N×N local block extracted from the repeatability score map S 1 , represents the one extracted from The flattened N×N local blocks extracted from

[0045] Further, the local peak loss L peak is defined as follows:

[0046]

[0047] where P = {p} is the set of local blocks, S ij is the local region on S, is the maximum value within the local region, is the average value within the local region.

[0048] Further, during the training process of the lightweight network structure, the overall loss is:

[0049] L = ω des L des + ω rel L rel + ω sim L sim + ω peak L peak

[0050] where ω des = 1, ω rel = 3, ω sim = 1, ω peak = 1, L des is the bidirectional negative log-likelihood loss, L rel is the bidirectional L1 distance loss, L sim is the local similarity loss, L peak is the local peak loss; in the bidirectional negative log-likelihood loss and the bidirectional L1 distance loss, σ = 0.5, and in the local similarity loss and the local peak loss, the local block size N = 16.

[0051] Beneficial effects:

[0052] (1) The present invention designs a lightweight network structure, which improves the speed of key point detection and descriptor extraction by minimizing the number of convolutional layers and rapid downsampling, and at the same time adopts multi-level feature fusion to improve the feature expression ability. This model runs at a speed of 36 frames per second on the GPU and achieves performance comparable to that of the state-of-the-art (SOTA) method.

[0053] (2) The proposed fusion non-maximum suppression (NMS) sampling strategy in the present invention uses this strategy to constrain the reliability score map, thereby helping the model to only focus on locally reliable key points, thus avoiding the model from detecting key points in unreliable regions and improving the matching performance of key points.

[0054] (3) The present invention designs a bidirectional L1 distance loss to train the reliability score map, which helps the network avoid detecting key points in unreliable regions, especially in unreliable regions such as the sky and rivers. Moreover, with fewer computing resources, its performance in unreliable regions such as the sky and rivers is superior to existing methods.

[0055] (4) The present invention designs a negative log-likelihood loss, which transforms the descriptor similarity problem into a probability problem, fully learning the rich feature information of key points, so as to maximize the descriptor similarity of matching points and provide more stable convergence.

[0056] (5) The present invention designs a local similarity loss, which makes the positions of local maxima in two images correspond to each other, improving the repeatability score of key points.

[0057] (6) When the local similarity between two images is too high, there will be a phenomenon of dense distribution of key points in the local area. This phenomenon will cause the key points to not meet the global distribution, significantly reducing the matching performance. To solve the problem of local dense distribution, the present invention designs a local peak loss function, whose purpose is to maximize the local maximum while suppressing non-key points, improving the network's attention to locally sparse key points, thereby improving the matching performance of the model. Description of the Drawings

[0058] Figure 1 It is the system flow chart of the present invention.

[0059] Figure 2 It is the lightweight network structure diagram.

[0060] Figure 3 It is the fused non-maximum suppression sampling strategy diagram.

[0061] Figure 4 It is the training effect diagram of the bidirectional L1 distance loss function.

[0062] Figure 5 It is the training effect diagram of the local similarity loss function and the local peak loss function.

[0063] Figure 6 It is the interface diagram of the key point detection system. Detailed Embodiment

[0064] The present invention will be further described in detail below with reference to the drawings.

[0065] As Figure 1 shown, the present invention proposes a lightweight and reliable key point detection method based on deep learning, and the specific steps are as follows:

[0066] Step 1: Input the reference image I 1 and the image I to be matched 2, through preprocessing methods such as image rotation, image blurring, and random cropping.

[0067] Step 2: Construct a lightweight network structure. The lightweight network structure has three parts, namely a feature encoder, a multi-level feature fusion module, and a generated mapping graph. The lightweight network structure learns the points and descriptors at the same positions in two images, as Figure 2 shown below:

[0068] Feature encoder: The present invention designs a basic layer, including a two-dimensional convolution Conv2d, Relu + BatchNorm. The encoder encodes the input image into a feature map, where H is the height of the input image and W is the width of the input image. This encoder consists of 4 main modules (Blocks). The input image of the first module is a grayscale image, i.e., C = 1, and it contains 2 basic layers. The latter 3 modules all contain 1 MaxPooling layer and 2 basic layers, and the downsampling rates are 1 / 2, 1 / 4, and 1 / 4 respectively. The channel depths of the 4 modules are {32, 64, 128, 128}, and the final spatial resolution is

[0069] Multi-level feature fusion module: Fuse the multi-level features of the encoder. Select the output features of the 4 modules, use 1×1 convolution and bilinear interpolation upsampling to adjust their channel depths, and then simply concatenate them together.

[0070] Reliability detector and repeatability detector: The two detectors respectively output H×W×1 feature mapping graphs. The repeatability mapping graph is normalized by softplus to a repeatability score graph S, whose purpose is to provide sparse but repeatable key point positions. To achieve sparsity, only the key points of local maxima are extracted in S. The reliability mapping graph is activated by softmax to a reliability score graph R, which represents the reliability of estimating the descriptors of each pixel, that is, the possibility of being suitable as a matching point.

[0071] Descriptor extractor: The descriptor extractor outputs an H×W×(dim + 1) feature map, where the first dim channels are dense descriptors D.

[0072] Step 3: Input the processed reference image I 1 and the image I to be matched 2 into the trained lightweight network structure respectively to output the reliability score graph R 1 and R 2 , the repeatability score graph S 1 and S 2 , as well as the dense descriptors D 1 and D 2 , and then, R 1After integrating the non-maximum suppression sampling strategy, for the integrated non-maximum suppression (NMS) sampling strategy, as Figure 3 shown, the specific steps are as follows:

[0073] Step 3.1: NMS map The score map first suppresses the non-maximum score r = R(u, v) in a local N×N window, as Figure 3 (c) shown

[0074]

[0075] where r max is the local highest score.

[0076] Step 3.2: Apply a threshold th on the NMS map score map to filter out low response scores, and select the top top scores, as Figure 3 (d) shown. Finally, obtain the index values indice of these scores, so that the key point coordinates (x 1 , y kpt , kpt ) in Figure I are extracted in this step:

[0077]

[0078] If Figure I 2 is synthesized through a transformation of I 1 , such as a homography transformation, then the optical flow map F ∈ R H ×W×2 can be accurately obtained. By F, the corresponding key point coordinates in I 2 are obtained where each position (x, y) stores a two-dimensional vector F(x, y) = (u(x, y), v(x, y)), where u and v respectively represent the horizontal displacement and vertical displacement of the pixel (x, y) in Figure I 1 in Figure I 2 :

[0079]

[0080] Four loss functions are proposed to jointly optimize the key-point reliability and repeatability, namely the bidirectional negative log-likelihood loss, the bidirectional L1 distance loss, the local similarity loss, and the local peak loss. For key-point reliability, the bidirectional L1 distance loss is proposed to train the reliability score map, thus directly improving the key-point reliability; for key-point repeatability, the local peak loss is designed to avoid the dense and crowded phenomenon on the repeatability score map, thus forcing the key points in the local area of the score map to show a sparse distribution and making the key-point repeatability score reach the peak; in addition, for key-point descriptors, the bidirectional negative log-likelihood loss is adopted, which transforms the descriptor similarity problem into a probability problem, thus maximizing the descriptor similarity of the matching points to provide more stable convergence. The four loss functions are as follows:

[0081] Step 4: Dense descriptor D 1 and D 2 are obtained from the dense mappings D(·,·) of Figure I 1 and Figure I 2 respectively, and the dense descriptor D 1 and the i-th row D 1 (i,·) and D 2 (i,·) correspond to the descriptors of the corresponding points in Figure I 1 and Figure I 2 respectively. The corresponding descriptors are calculated according to the key-point coordinates obtained by the sampling strategy in Step 3:

[0082]

[0083] Similarity matrix is obtained from . Considering the symmetry of the matching, two matching directions are taken simultaneously to obtain the double negative log-likelihood loss L des , where the similarity of the corresponding features is located on the main diagonal M ii of M, and softmax is calculated along the "columns":

[0084]

[0085] Step 5: The reliability score map is supervised by bidirectional softmax,, using and as the matching confidence, denoted as

[0086]

[0087] Given the matching confidence and , the bidirectional L1 distance loss L is directly used to supervise the reliability map with the L1 distance lossrel :

[0088]

[0089] where λ is the sigmoid activation function, ⊙ is the Hadamard product, and R 1 = R 1 (x kpts , y kpts ). R 1 and R 2 are the reliability score maps obtained from I 1 and I 2 respectively. Additionally, for the bidirectional L1 distance loss L rel , only the backpropagation gradients are calculated for R 1 and R 2 . The effect of the bidirectional L1 distance loss is as Figure 4 shown.

[0090] Step 6: Assume that images I 1 and I 2 are two images of the same scene. In the fusion non-maximum suppression (NMS) sampling strategy, the corresponding point relationship between the two images has been obtained through the optical flow map . Let H be the height of the input image and W be the width of the input image. According to the optical flow map, S 2 is warped to generate

[0091] To ensure that all local maxima in the repeatability score map S 1 correspond to the maxima in , the key idea is to maximize the local similarity similar between S 1 and . When is maximized, the two repeatability score maps are the same, and their maxima can fully correspond. The present invention defines a local block set P = {p}, which contains local blocks of size N×N in the image size {1,..., W}×{1,..., H}, and defines the loss:

[0092]

[0093] where represents the flattened N×N local block extracted from S 1 , and represents the flattened N×N local block extracted from the repeatability score map . The local similarity loss and the local peak effect are as Figure 5 shown.

[0094] When S1 and When the similarity with is too high, the phenomenon of dense distribution of key points in a local area will occur. As shown in Figure 5 (a) and (b), at the same time, if a limited number of key points are extracted, this phenomenon will cause the key points to not meet the global distribution, as shown in Figure 5 (c) and (d).

[0095] Step 7: To solve the phenomenon of local dense distribution, maximize the maximum value within the local area while suppressing non-key points, improve the network's attention to local sparse key points, and avoid the situation where L sim approaches 0. Design the following local peak loss:

[0096]

[0097] where \(P = \{p\}\) is the set of local blocks, \(S\) ij is the local area on \(S\), is the maximum value within this local area, is the average value within this local area.

[0098] Step 8: Input the image pairs \(I\) 1 and \(I\) 2 into the trained lightweight network structure respectively. Obtain the key point coordinates in the two images according to the reliability score map and the repeatability score map. After the K-nearest neighbor matching algorithm, obtain the pre-matched key points in the two images, and then screen out the correct matching pairs through the RANSAC algorithm. Finally, conduct an accuracy evaluation experiment on the publicly available dataset HPatches dataset.

[0099] When training the lightweight network structure, use the publicly available dataset Aachen-Day-Night to train the lightweight network structure. The details of loss calculation are to use NMS with a window size of \(N = 5\) to detect 1000 key points. In the bidirectional negative log-likelihood loss and the bidirectional L1 distance loss, \(\sigma = 0.5\). In the local similarity loss and the local peak loss, the local block size \(N = 16\). The overall loss is:

[0100] \(L=\omega\) des \(L\) des +\(\omega\) rel \(L\) rel +\(\omega\) sim \(L\) sim +\(\omega\) peak \(L\) peak (10)

[0101] where \(\omega\) des =1, \(\omega\) rel =3, \(\omega\) sim =1, \(\omega\) peak =1, \(L\)des is the bidirectional negative log-likelihood loss, L rel is the bidirectional L1 distance loss, L sim is the local similarity loss, L peak is the local peak loss.

[0102] During training, the images are randomly cropped to 480×480. The present invention uses the Adam optimizer to optimize the network, with the learning rate fixed at 0.0001, the batch size set to 4, and it converges after about 12 hours of training on the device RTX2080TI.

[0103] The interface of the key point detection system is as Figure 6 shown.

[0104] The performance comparison experiment with the existing method is shown in Table 1:

[0105] Table 1 Performance comparison experiment table with the existing method

[0106]

[0107] The present invention compares the proposed method with the prior art in homography estimation and average matching accuracy tasks. The method proposed by the present invention is denoted as "LRFeat".

[0108] 1) Network complexity: Table 1 reports the network complexity of different methods, including the number of network structure parameters, GFLOPs (billions of floating-point operations), and real-time running FPS. Intuitively, it is found that R2D2 has the fewest network parameters (484K), followed by LRFeat (611K), and finally ALIKE-L (653K). However, a model with fewer network parameters does not necessarily mean it has less computational cost. For example, although R2D2 has only 484K parameters, its GFLOPs is 464.55, while LRFeat's GFLOPs is only 14.22. For a more intuitive comparison, Table 4 also shows that LRFeat has the highest frame rate (FPS) during real-time operation, which is 36.8FPS. Therefore, the method proposed by the present invention has high real-time operation efficiency and also provides accurate matching accuracy.

[0109] 2) Homography estimation:

[0110] To evaluate the performance of the proposed method in the homography estimation task, the present invention conducts experiments based on the HPatches dataset. The present invention introduces the Mean Homography Accuracy (MHA) as a key metric and evaluates MHA at thresholds of 1, 2, and 3. The results show that the proposed LRFeat method exhibits significant performance advantages over existing methods at low thresholds (especially MHA@3). Compared with the best-performing ASLFeat (MS), LRFeat (SS) achieves a 1.28% accuracy improvement in MHA@3 and reduces the computational complexity by approximately 3.1 times, indicating its excellent performance and efficiency. In addition, LRFeat (MS) reaches accuracies of 45.37%, 70.37%, and 80.37% in the evaluations of MHA@1, MHA@2, and MHA@3, respectively, significantly outperforming other existing methods, where MS (Multi-scale) represents multi-scale keypoint detection and SS (Single-scale) represents single-scale keypoint detection. Although LRFeat uses fewer computational resources (14.22 GFLOPs), it still has high competitiveness in the homography estimation task, further verifying its advantage in low-resource scenarios.

[0111] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it should not be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.

Claims

1. A key point detection method based on deep learning, characterized in that: The following steps are involved: Step 1: Obtain a public image dataset, which includes two images of the same scene from different perspectives, and preprocess the image dataset; Step 2: Design a lightweight network structure, which includes a feature encoder, a multi-level feature fusion module, a reliability detector and a repeatability detector, and a descriptor extractor; The lightweight network structure inputs two images of the same scene from different perspectives, and outputs a reliability score map R, a repeatability score map S, and a dense descriptor D accordingly; Step 3: Propose a fusion non-maximum suppression NMS sampling strategy, which uses the fusion non-maximum suppression NMS sampling strategy to sample a limited number of key points with the highest local scores on the reliability score map. The constrained reliability score map helps the model focus only on locally reliable key points; Step 4: Four loss functions are proposed to jointly optimize the reliability and repeatability of key points, namely, bidirectional negative log-likelihood loss, bidirectional L1 distance loss, local similarity loss, and local peak loss. For key point reliability, bidirectional L1 distance loss is used to train the reliability score map. For key point repeatability, local peak loss is designed to avoid dense crowding on the repeatability score map. For key point descriptors, bidirectional negative log-likelihood loss is used to convert the descriptor similarity problem into a probability problem. Step 5: Use the preprocessed image data set to train the lightweight network structure constructed in steps 2 to 4 to obtain a trained lightweight network structure; Step 6: Input the reference image I1 and the image to be matched I2 into the trained lightweight network structure respectively, and output the reliability score map R 1 and R 2 , repeatability score graph S 1 With S 2 and dense descriptors D1 and D2, and repeat steps 4 and 5; Step 7: Based on the reliability score graph R 1 and R 2 , repeatability score graph S 1 With S 2 The coordinates of the key points in the two images are obtained, the pre-matched key points in the two images are obtained through the K nearest neighbor matching algorithm, and then the correct matching pairs are screened out through the RANSAC algorithm.

2. A key point detection method based on deep learning according to claim 1, characterized in that: The lightweight network structure is as follows: The feature encoder takes as input an image The code is a feature map, H is the height of the input image, W is the width of the input image, the encoder consists of 4 modules, the input image of the first module is a grayscale image, that is, C = 1, containing 2 basic layers, the next three modules contain 1 MaxPooling layer and 2 basic layers, the basic layer includes a two-dimensional convolution Conv2d, Relu + BatchNorm; the downsampling rates are 1 / 2, 1 / 4, 1 / 4, respectively, the channel depths of the four modules are {32, 64, 128, 128}, respectively, and the final spatial resolution is The multi-level feature fusion module fuses the multi-level features of the feature encoder, selects the output features of 4 modules, adjusts the channel depth using 1×1 convolution and bilinear interpolation upsampling, and then splices them together; The two detectors of the reliability detector and the repeatability detector respectively output H×W×1 feature maps, the repeatability map is normalized by softplus to a repeatability score map S, and the reliability map is activated by softmax to a reliability score map R; The descriptor extractor outputs H×W×(dim+1) feature maps, where the first dim channels are dense descriptors D.

3. The key point detection method based on deep learning according to claim 1, characterized in that: The implementation process of the fusion non-maximum suppression NMS sampling strategy is as follows: Step 3.1: Search for the pixel with the maximum score in the local window. This operation is equivalent to argmax in the local window N*N: Where R(i,j) represents the reliability score of the local coordinate (i,j), and W(x,y) represents the local window centered at each pixel position (x,y) in the image, including all neighboring pixels (i,j); Step 3.2: NMS map The score map is first constructed by suppressing the non-maximum scores r = R(u,v) in a local N × N window: In the formula, r max is the local maximum score. In order to avoid falling into the local optimal solution, in NMS map The threshold th is applied to the score map to filter out low response scores, and the top scores are selected. Finally, the index values ​​of these scores are obtained. The coordinates of the key points (x kpt ,y kpt ) is extracted as: Step 3.3: If image I2 is synthesized by changing I1, then the optical flow map F∈R is obtained H×W×2 , use F to obtain the corresponding key point coordinates in I2 Each position (x, y) stores a two-dimensional vector F(x, y) = (u(x, y), v(x, y)), where u and v represent the horizontal and vertical displacement of the pixel (x, y) in Figure I1 in Figure I2, respectively:

4. The key point detection method based on deep learning according to claim 1, characterized in that: The two-way negative log-likelihood loss is implemented as follows: Step 4.1: Assume that dense descriptors D1 and D2 are obtained from dense mappings D(·,·) of Figure I1 and Figure I2, respectively, and dense descriptors D1 and D2 are obtained from dense mappings D(·,·) of Figure I1 and Figure I2, respectively. The i-th row D1(i,·) and D2(i,·) correspond to the descriptors of the corresponding points in Figures I1 and I2 respectively. The corresponding descriptors are calculated based on the key point coordinates obtained by the fusion non-maximum suppression NMS sampling strategy in step 3: Step 4.2: Similarity Matrix Depend on Considering the symmetry of the matching, we take two matching directions at the same time and get the bidirectional negative log-likelihood loss L des , where the similarity of the corresponding features is located on the main diagonal M of M ii , and the softmax is calculated along the "columns":

5. The key point detection method based on deep learning according to claim 1, characterized in that: The bidirectional L1 distance loss implementation process is as follows: Step 5.1: During training, supervise the reliability score map through bidirectional softmax, using and As the matching confidence, it is recorded as Step 5.2: Given a matching confidence and In the case of L1 distance loss, the reliability mapping is directly supervised to obtain the bidirectional L1 distance loss L rel : Among them, λ is the sigmoid activation function, ⊙ is the Hadamard product, R1=R 1 (x kpts ,y kpts ), R 1 and R 2 Reliability score graphs obtained from Figures I1 and I2, respectively, for the bidirectional L1 distance loss L rel , only for R 1 and R 2 Back propagate gradients.

6. The key point detection method based on deep learning according to claim 1, characterized in that: The local similarity loss implementation process is as follows: Assume that Figures I1 and I2 are two pictures of the same scene, and the optical flow map has been used in the fusion non-maximum suppression NMS sampling strategy. The corresponding point relationship between the two images is obtained, H is the height of the input image, W is the width of the input image, and S is converted according to the optical flow map. 2 Deformation Generation In order to make the reproducibility score graph S 1 All local maxima in The key idea is to maximize S 1 and The local similarity between When the two repeatability score maps are maximized, they are the same; define a local block set P = {p}, which contains N × N local blocks in the image size {1, ... W} × {1, ... H}, and define a local similarity loss L sim (I1,I2,F): in, Representation from the reproducibility score map S 1 The flattened N×N local block extracted from Indicates from The flattened N×N local block extracted from .

7. The key point detection method based on deep learning according to claim 1, characterized in that: The local peak loss L peak The definition is as follows: Among them, P = {p} is the local block set, S ij is a local area on S, is the maximum value in the local area, is the average value in the local area.

8. The key point detection method based on deep learning according to claim 1, characterized in that: During the training of the lightweight network structure, the overall loss is: L=ω des L des +oh rel L rel +oh sim L sim +oh peak L peak Among them, ω des =1,ω rel =3,ω sim =1,ω peak =1,L des is the two-way negative log-likelihood loss, L rel is the bidirectional L1 distance loss, L sim is the local similarity loss, L peak is the local peak loss; in the bidirectional negative log-likelihood loss and the bidirectional L1 distance loss, σ = 0.5, and in the local similarity loss and the local peak loss, the local block size N = 16.

Citation Information

Cited By

  • Real-time image splicing method and device suitable for robot embedded platform

    CN121392449A