Image matching method and device based on complementary descriptors
By constructing a twin orthogonal descriptor network, using different loss functions and orthogonal loss functions to train the descriptor network to make its output features orthogonal, thereby learning two complementary descriptor features simultaneously in a single CNN, solving the problem that sparse feature matching methods in the prior art is difficult to balance in accuracy and computational overhead, and achieving efficient and robust image matching.
Patent Information
- Application Number
- CN202510273006.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-10
AI Technical Summary
In the prior art, sparse feature matching methods are difficult to balance the matching accuracy and calculation overhead, and the performance of a single low-resolution descriptor is limited. The method of combining multiple models is due to the lack of complementarity of the descriptors and the accumulation of wrong matching pairs, resulting in unstable matching performance.
A twin orthogonal descriptor network is constructed, including two identical descriptor network branches, using different loss functions and orthogonal loss functions during training, making the output features of the two descriptor networks orthogonal, thereby learning two complementary descriptor features simultaneously in a single CNN and using them in combination.
It significantly improves the performance of the descriptor, can effectively filter out wrong matches while retaining more correct matches, generate more robust matching relationships, improve matching accuracy and maintain efficient matching efficiency.
Smart Images

Figure CN119762813B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image matching, and in particular to an image matching method and device based on complementary descriptors. Background Art
[0002] In various 3D computer vision tasks, it is necessary to establish a local feature matching relationship between images, such as motion structure recovery (SFM), simultaneous localization and mapping (SLAM), and visual positioning. In the existing technology, the local feature matching methods based on learning can be divided into two methods: sparse matching and dense matching, depending on whether there is a detector or not. Among them, most sparse local feature matching methods include three stages: feature detection, feature description, and feature matching. In the detection stage, pixels with high scores are selected as feature points by using non-maximum suppression (NMS) on the output score map, and then a descriptor is coupled to the feature point by interpolation. Finally, the correspondence between points is found by calculating the similarity between descriptors. The dense local feature matching method enhances the dense local features extracted from the convolutional backbone through the self-attention layer and cross-attention layer with Transformer. Although this type of method can improve the accuracy of feature matching to a certain extent, it also brings extremely high computational overhead and is not suitable for real-time tasks with high computational overhead requirements. In contrast, the sparse local feature matching method can achieve a better balance between matching accuracy and computational overhead.
[0003] In the prior art, sparse feature matching methods usually only predict a single descriptor feature. Although the learning method for descriptor training is constantly optimized and improved, the performance improvement of a single descriptor is limited. In addition, in order to meet the needs of real-time tasks, the descriptor features of most methods in the prior art are usually low-resolution, and the performance of a single low-resolution descriptor feature is difficult to improve. If the descriptors of multiple models are directly combined for matching, although the matching performance will usually be improved to a certain extent due to the increase in the number of generated matching pairs, since the descriptors between different models are not complementary, the wrong matching pairs will accumulate when used in combination, so the actual matching performance is not stable, thus affecting the overall performance. Summary of the invention
[0004] The technical problem to be solved by the present invention is: in view of the technical problems existing in the prior art, the present invention provides an image matching method and device based on complementary descriptors, which has a simple implementation method, low cost, high matching efficiency and accuracy.
[0005] In order to solve the above technical problems, the technical solution proposed by the present invention is:
[0006] An image matching method based on complementary descriptors, comprising the following steps:
[0007] Constructing an image matching model, the image matching model includes a twin orthogonal descriptor network for generating feature descriptors and a detector network for extracting feature points, the twin orthogonal descriptor network includes two identical descriptor network branches, each of the descriptor network branches is composed of an encoder and a decoder, the encoder in each of the descriptor network branches includes a multi-layer residual network module, and the decoder includes more than two upsampling networks;
[0008] The constructed image matching model is trained using a training data set, and different loss functions are used to supervise the training of the two descriptor network branches during the training process, and an orthogonal loss function is used to train the two descriptor network branches so that the output descriptor features of the two orthogonal descriptor networks are decoupled. The final similarity matrix is generated by combining the similarity matrices of the two descriptor network branches, and finally, the nearest neighbor algorithm is used to obtain the matching relationship.
[0009] The image pairs to be matched are obtained and input into the trained image matching model to obtain the matching relationship results between the image pairs to be matched.
[0010] Furthermore, during the training of the twin orthogonal descriptor network, one descriptor network is supervised and trained using negative log-likelihood loss, and the other descriptor network is supervised and trained using triplet loss, and the two descriptor features are decoupled by adding orthogonal loss.
[0011] Furthermore, the expression of the negative log-likelihood loss is:
[0012]
[0013] in, represents the negative log-likelihood loss, N represents the number of feature points, Represents the descriptor similarity matrix of the two images to be matched, ,in , According to the two images to be matched The corresponding feature point information is obtained from the dense feature map generated by the first descriptor network The descriptor set sampled from Respectively represent the two images to be matched The corresponding feature point numbers extracted from the hour Represents the descriptor similarity matrix of the two images to be matched S The main diagonal elements in , descriptor similarity matrix SThe main diagonal elements of are the similarity measures between corresponding feature points.
[0014] Further,
[0015]
[0016]
[0017]
[0018]
[0019] in, represents the triplet loss, N represents the number of feature points, Represents the Euclidean distance between the corresponding descriptors of the two images to be matched, represents the most difficult negative sample distance, , According to the two images to be matched The corresponding feature point information is generated from the dense feature map generated by the second descriptor description The descriptor set sampled from is the margin parameter, N represents the number of feature points, Respectively represent the two images to be matched The extracted corresponding feature point numbers that match each other, k Indicates removal = The serial number of the point with the greatest similarity outside the corresponding points.
[0020] Furthermore, the expression of the orthogonal loss is:
[0021]
[0022] in, represents the orthogonal loss, Represents the absolute value after calculating the dot product between two descriptors, N represents the number of feature points, , According to the two images to be matched The corresponding feature point information is obtained from the dense feature map generated by the first descriptor network The descriptor set sampled from , According to the two images to be matched The corresponding feature point information is generated from the dense feature map generated by the second descriptor description The descriptor set sampled from They respectively represent the serial numbers of the corresponding feature points that match each other extracted from the two images to be matched.
[0023] Furthermore, the similarity matrices of the two descriptor network branches are combined to generate the final similarity matrix according to the following formula: :
[0024]
[0025]
[0026]
[0027] in, represents the weight, , Represent the image pairs to be matched The characteristic points of They represent the corresponding feature point numbers extracted from the two images to be matched, , They represent the descriptor sets obtained by sampling the dense descriptor features output by the two descriptor networks using the corresponding feature points. , Respectively represent the similarity matrices calculated using the two descriptor network branches.
[0028] Furthermore, it also includes using the specified model as the teacher network to perform distillation learning on the detector network, and learning to detect repeated feature points by learning the distribution of the score map of the teacher model. The loss function is defined as:
[0029]
[0030] in, represents the binary cross entropy loss, Output score map for the teacher model, Output score map of the student model.
[0031] An image matching device based on complementary descriptors, comprising:
[0032] A model construction module, used to construct an image matching model, wherein the image matching model includes a twin orthogonal descriptor network for generating feature descriptors and a detector network for extracting feature points, wherein the twin orthogonal descriptor network includes two identical descriptor network branches, each of which is composed of an encoder and a decoder, wherein the encoder in each of the descriptor network branches includes a multi-layer residual network module, and the decoder includes more than two upsampling networks;
[0033] A model training module, used to train the constructed image matching model using a training data set, and supervise the training of the two descriptor network branches using different loss functions during the training process, and train the two descriptor network branches using an orthogonal loss function to decouple the output descriptor features of the two orthogonal descriptor networks, generate a final similarity matrix by combining the similarity matrices of the two descriptor network branches, and finally, obtain a matching relationship using a nearest neighbor algorithm;
[0034] The matching detection module is used to input into the trained image matching model to obtain the matching relationship results between the image pairs to be matched.
[0035] A computer device comprises a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to perform the above method.
[0036] A computer-readable storage medium storing a computer program, wherein the computer program implements the above method when executed by a processor.
[0037] Compared with the prior art, the advantages of the present invention are: the present invention constructs a twin orthogonal network model, which includes two identical descriptor network branches. During the training process, different loss functions are used to supervise the training of the two descriptor network branches, and an orthogonal loss function is used to train the two descriptor network branches so that the output descriptor features of the two orthogonal descriptor networks are orthogonal. The similarity matrices of the two descriptor network branches are combined to generate the final similarity matrix, so that two complementary descriptor features can be learned and used in combination in a single CNN at the same time, which greatly improves the performance of the descriptor, can effectively filter out incorrect matches while retaining more correct matches, generate a more robust matching relationship, and can significantly improve the matching accuracy while ensuring the matching efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic diagram of the implementation flow of the image matching method based on complementary descriptors in this embodiment.
[0039] Figure 2 It is a schematic diagram of the structure of the twin orthogonal descriptor network and the structural principle of model reasoning in this embodiment.
[0040] Figure 3 It is a schematic diagram of the average matching accuracy evaluation result obtained on the HPatches data set in a specific application embodiment of the present invention. DETAILED DESCRIPTION
[0041] The present invention is further described below in conjunction with the accompanying drawings and specific preferred embodiments, but the protection scope of the present invention is not limited thereby.
[0042] For ease of understanding, the relevant technical background of the present invention is first introduced by way of example.
[0043] In the prior art, sparse local feature matching methods usually learn a single discriminative descriptor feature by encouraging positive sample points in the description space to be close to each other and negative sample points to be far away from each other. For example, the SuperPoin method is a method that uses a guided training strategy to train a model to detect key points and jointly trains its descriptors using triplet loss; the XFeat method is a method that introduces maximum likelihood loss based on a lightweight and accurate architecture and provides sparse and dense matching options at the same time; the Alike method is a method that uses a partially differentiable key point detection module and uses NRE loss to learn dense descriptor features. The AUC (Area Under the Curve) results obtained by performing posture estimation using the above three methods and the methods formed by their combination in a specific application embodiment are shown in Table 1.
[0044] Table 1: Test results of each method
[0045]
[0046] As can be seen from Table 1, the performance improvement of a single model is limited, and the performance of the SuperPoint and XFeat group methods is lower than that of a single SuperPoint at a threshold of 5°, indicating that directly combining multiple model descriptor features cannot stably improve the performance.
[0047] We further distinguish the existing SuperPoint detector, CAPS, XFeat, Alike and our method descriptors in the HPatches dataset and conduct evaluation tests in the homography estimation task. CAPS extracts descriptors from weak supervision of camera poses. The test results of the percentage of corner point errors under 1 / 3 / 5 pixel thresholds, the number of matches and the average matching time of homography estimation by different methods are shown in Table 2, where "SP" represents the SuperPoint method. As can be seen from Table 2, the performance of descriptors of different models is not much different, that is, the optimization loss function has limited improvement on the performance of descriptors. However, our method combined with the SuperPoint method greatly improves the accuracy of homography estimation. This proves the effectiveness of our complementary descriptor design method.
[0048] Table 2: Evaluation results of homography estimation on the HPatches dataset, where “NN” stands for the nearest neighbor algorithm
[0049]
[0050] In addition, the resolution of the descriptor feature map is also an important factor affecting the discriminability of the descriptor. A high-resolution descriptor feature map means that all pixels have their corresponding descriptor features, while a low-resolution descriptor feature map needs to obtain the descriptor of the corresponding pixel through interpolation. Therefore, it is easily affected by the descriptor features of other pixels in the field, thereby affecting the discriminability of the descriptor. However, since sparse local feature matching is usually applied to real-time applications, the descriptor features generated by most methods in the prior art are low-resolution, and high-resolution descriptor features will bring higher computational overhead.
[0051] In summary, the performance of a single low-resolution descriptor in the existing technology is limited, and it is difficult to improve the performance of the descriptor by directly optimizing the loss function or combining multiple models.
[0052] In order to efficiently combine multiple descriptor features and further improve the descriptor performance, the present invention constructs a twin orthogonal network model, which includes two identical descriptor network branches. During the training process, different loss functions are used to supervise the training of the two descriptor network branches, and an orthogonal loss function is used to train the two descriptor network branches so that the output descriptor features of the two orthogonal descriptor networks are orthogonal. The similarity matrices of the two descriptor network branches are combined to generate the final similarity matrix, so that two complementary descriptor features can be simultaneously learned and used in a single CNN (convolutional neural network), which greatly improves the performance of the descriptor, can effectively filter out incorrect matches while retaining more correct matches, and generate a more robust matching relationship, which can significantly improve the matching accuracy while ensuring the matching efficiency.
[0053] The present invention will be further described below in conjunction with specific embodiments.
[0054] like Figure 1 As shown, the steps of the image matching method based on complementary descriptors in this embodiment include:
[0055] Step S01. Construct an image matching model, which includes a twin orthogonal descriptor network for generating feature descriptors and a detector network for extracting feature points. The twin orthogonal descriptor network includes two identical descriptor network branches, each of which is composed of an encoder and a decoder. In each descriptor network branch, the encoder includes a multi-layer residual network module, and the decoder includes more than two upsampling networks.
[0056] like Figure 2As shown, in this embodiment, the twin orthogonal descriptor network is designed by designing two identical descriptor network branches in a single CNN. Each descriptor branch consists of an encoder and a decoder. The encoder network consists of a multi-layer (for example, 5 layers can be used) residual network module (ResNet). The feature resolution of each layer is , number of channels ,in Represents the number of network layers. The decoder includes more than two upsampling networks to upsample the output features of the encoder. For example, when the decoder uses two upsampling networks, the output features of the encoder can be upsampled to 1 / 4 resolution. Furthermore, in order to retain more image information, this embodiment inputs the 1 / 4 and 1 / 8 resolution features of the encoder into the decoder for feature fusion.
[0057] This embodiment is based on the sparse local feature matching method, and designs a detector network in the network. At the same time, in order to retain more original image information, feature point detection is performed at the original image resolution. Considering that the convolution operation on the high-resolution image is very time-consuming, this embodiment designs the detector network as a simple 2-layer network, and the activation function of the last layer is set to the Sigmoid activation function to limit the score to the interval (0,1).
[0058] Step S02. Use the training data set to build an image matching model for training, and use different loss functions to supervise the training of the two descriptor network branches during the training process, and use an orthogonal loss function to train the two descriptor network branches so that the output descriptor features of the two orthogonal descriptor networks are orthogonal, and the final similarity matrix is generated by combining the similarity matrices of the two descriptor network branches.
[0059] In the training process of this embodiment, by using different loss functions for the two descriptor networks and adding orthogonal loss to make the output descriptor features of the two networks orthogonal, different descriptor network branches can establish matching relationships in different directions, thereby achieving efficient performance improvement by combining the two descriptor branches. Among them, in the process of training the twin orthogonal descriptor network, one descriptor network is supervised by negative log-likelihood loss. The purpose of using negative log-likelihood loss is to optimize the descriptor similarity score between corresponding points to make the two corresponding points close to each other; the other descriptor network is supervised by triplet loss. The purpose of using triplet loss is to make positive samples closer and push the most difficult negative samples farther. Finally, the two descriptor features are decoupled by adding orthogonal loss.
[0060] Specifically, the descriptor network is trained with pixel-level point correspondence supervision. , , assuming that the image pair have Matching pixels ,in The first two columns correspond to Medium pixel Coordinates , the last two columns correspond to Medium pixel Coordinates .
[0061] For the first description subnetwork , using negative log-likelihood loss To supervise the training. Among them, the descriptor set , is based on and The corresponding point information of the graph is obtained from the first description subnetwork Generated dense descriptor feature map The sampling is obtained, that is, , According to the two images to be matched The corresponding feature point information is obtained from the dense feature map generated by the first descriptor network The descriptor set sampled from Indicates that including 128-dimensional descriptors. Since the purpose of the negative log-likelihood loss is to optimize the descriptor similarity score between corresponding feature points to make the two corresponding points close to each other, the descriptor similarity matrix of the two images is calculated ,in , The main diagonal elements of are the similarity measures of the corresponding features. Considering the symmetry of the matching, this embodiment specifically samples two matching directions and performs row-by-row matching. Generate double Loss, specifically the negative log-likelihood loss The calculation expression is:
[0062] (1)
[0063] in, represents the negative log-likelihood loss, N represents the number of feature points, Respectively represent the two images to be matched The corresponding feature point numbers extracted from the hour Represents the two images to be matched The descriptor similarity matrix S The main diagonal elements in , Represents the descriptor similarity matrix of the two images to be matched S The main diagonal elements in the descriptor similarity matrix S The main diagonal elements of are the similarity measures between corresponding feature points.
[0064] For the second descriptor branch network , this embodiment uses triple loss to supervise the training. Similarly, from the second descriptor branch Generated dense descriptor feature map Medium Sampling and The corresponding feature point information of the graph obtains the descriptor set , ,Right now , According to the two images to be matched The corresponding feature point information is generated from the dense feature map generated by the second descriptor description The descriptor set sampled from is expressed as The purpose of the triplet loss is to make the positive samples closer and push the hardest negative samples farther away. First, define the positive sample distance is the Euclidean distance between the corresponding descriptors of the two images to be matched, that is:
[0065] (2)
[0066] Then define the most difficult negative sample as the point that is most easily confused except the corresponding point, that is, the most difficult negative sample distance Defined as:
[0067] (3)
[0068] in:
[0069] (4)
[0070] Right now k Indicates removal = The number of the point with the greatest similarity outside the corresponding points, which represents the most confusing point.
[0071] according to and Defining triplet loss for:
[0072] (5)
[0073] in, represents the triplet loss, N represents the number of feature points, Represents the Euclidean distance between the corresponding descriptors of the two images to be matched, represents the most difficult negative sample distance, is a margin parameter. Considering that the descriptors are L2 normalized and the Euclidean distance between descriptors is less than 4, M can preferably be set to 1.
[0074] In order to effectively combine multiple descriptor features to further improve the performance of the descriptor, this embodiment simultaneously learns two complementary descriptor features in a single network, and by combining the two complementary descriptor features, a more robust matching relationship can be established. For two descriptor networks, it is hoped that their output descriptor features are decoupled so that the two descriptor features can generate matching relationships in different directional dimensions, thereby achieving the complementarity of the two descriptor features. Therefore, this embodiment uses orthogonal loss during the training process to force the output descriptor features of the two networks to be decoupled.
[0075] Specifically, use the point set , For dense descriptor features , Interpolate and sample to obtain the descriptor features corresponding to each point , , , calculate the similarity score of the descriptor features between the same points of the two branches, and force its score to decrease to achieve orthogonality of the two branches, thereby achieving the decoupling of the two descriptor branches. The orthogonal loss function can be defined as:
[0076] (6)
[0077] in, represents the orthogonal loss, Represents the absolute value after calculating the dot product between two descriptors.
[0078] The purpose of learning a detector is to detect repeated and reliable feature points in two images. However, due to the ambiguity of the definition of feature points, few data sets have feature point labels, so the training of feature points is relatively difficult. The traditional method of training detectors is to use epipolar geometry or homography transformation to find the correspondence between feature points, and train the detector by optimizing the scores of corresponding feature points, but this method usually requires relying on large data sets and is difficult to train. In this embodiment, the detector network is distilled and learned by using the existing model Alike as a teacher network. By learning the distribution of the score map of the teacher model, it learns to detect repeated and reliable feature points, which can simplify the training stage of feature points as much as possible and obtain better generalization. The loss function can be defined as:
[0079] (7)
[0080] where represents the binary cross entropy loss, Output score map for the teacher model, The output score graph of the student model is shown in Figure 2. The score graphs of the teacher model and the model of the present invention are both restricted to the interval (0, 1) by the Sigmoid function.
[0081] Step S03. Obtain the image pairs to be matched, input them into the trained twin orthogonal descriptor network, and the twin orthogonal descriptor network outputs the similarity matrix between the image pairs to be matched to determine the matching relationship between the image pairs to be matched.
[0082] Since orthogonal loss is added to the two descriptor networks, the two complementary descriptor branches establish matching relationships in different subspaces. In order to retain more correct matches and filter out incorrect matches, and to achieve an effective combination of two complementary descriptors, this embodiment jointly considers the similarity matrices of the two descriptor features, and combines the similarity matrices of the two descriptor features by setting different weights, so that a more robust matching relationship can be established.
[0083] Specifically, suppose there is a pair of images , feature points are obtained through the trained image matching model , And dense descriptor features , ,in , , represents the negative log-likelihood loss training branch, Represents the ternary loss training branch, using feature points to sample the descriptor features , ,Right now , They represent the descriptor sets obtained by sampling the dense descriptor features output by the two descriptor networks using the corresponding feature points, normalizing the descriptors and calculating the similarity matrix , , then according to and The final similarity matrix is obtained according to the following formula :
[0084] (8)
[0085] in, represents the weight, , Represent the image pairs to be matched The characteristic points of They represent the corresponding feature point numbers extracted from the two images to be matched, , They represent the descriptor sets obtained by sampling the dense descriptor features output by the two descriptor networks using the corresponding feature points. , Respectively represent the similarity matrices calculated using the two descriptor network branches.
[0086] This embodiment obtains the final similarity matrix by combining the similarity matrices of the two descriptor branches, and uses the nearest neighbor algorithm to perform joint matching in multiple directions, thereby establishing a more robust matching relationship and achieving efficient model combination.
[0087] The above method of the present invention simultaneously learns two complementary descriptor features in a single CNN and adds orthogonal loss to make multiple descriptor features mutually orthogonal. At the same time, by jointly considering the similarity matrix of the two descriptor branches, a local feature matching relationship between images is established in different subspaces. It can effectively combine multiple descriptor features to further improve the descriptor performance and obtain a more robust matching result. The present invention can be applied to various task scenarios such as homography estimation, pose estimation, and visual positioning.
[0088] To verify the effectiveness of the present invention, the present invention and the sparse feature matching method in the prior art were used in a number of task scenarios for evaluation and comparison in specific application examples, including image matching, homography estimation, relative pose estimation, and visual positioning. Furthermore, the inference time of the present invention method and the sparse feature matching method in the prior art were tested under the same conditions.
[0089] Figure 3 is the mean matching accuracy (MMA) evaluation result obtained by different methods at 1-10 pixel threshold on the HPatches dataset, where Figure 3 (a) shows the MMA results under overall, illumination changes, and viewpoint changes. Figure 3 (b) also shows the key points and matching numbers of each method. The results show that the proposed model performs best overall.
[0090] Table 3 shows the percentage results of corner point errors of different methods for estimating homography on the HPatches dataset at 1 / 3 / 5 pixel thresholds. The table also shows the number of matches and average matching time of different methods. The D2Net method is a method that first extracts dense descriptors and then detects key points from dense descriptors through special rules; the R2D2 method is a method that considers the repeatability and reliability of key point detection by deploying an effective loss function; and the Patch2Pix method is a dense matching algorithm that establishes robust pixel correspondences through multiple matching refinements.
[0091] Table 3: Evaluation results of homography estimation on the HPatches dataset.
[0092]
[0093] Table 4 shows the AUC of the pose error within the 5° / 10° / 20° thresholds in the indoor pose estimation of ScanNet. Among them, Silk is a method that re-evaluates the elements of learning feature extraction and adopts an effective and simple key point and descriptor learning strategy, and the D2Net method is a method that introduces triplet loss. It can be seen from the results that the present invention is superior to other sparse feature matching methods in all thresholds.
[0094] Table 4: ScanNet indoor pose estimation evaluation results
[0095]
[0096] Table 5 shows the accuracy results obtained in Aachen Day-Night visual positioning at thresholds of (0.5m, 2), (1m, 5), and (5m, 10).
[0097] Table 5 Aachen Day-Night visual positioning evaluation results
[0098]
[0099] Table 6 shows the test results of computational complexity (Gflops) and average inference time (MIT) when the batch-size is 1 and 16 respectively, using input images of the same size for testing. Among them, PosFeat is a method that follows the line-to-window search strategy proposed by epipolar supervision and decouples the training of descriptors from detectors, Silk is a method that re-evaluates the elements of learning feature extraction and adopts an effective and simple key point and descriptor learning strategy, and DISK is a method that relaxes key point detection and descriptor matching into a probabilistic process and trains the network through reinforcement learning. The results show that the method of the present invention can well meet the real-time task requirements.
[0100] Table 6 Computational complexity and inference time evaluation results
[0101]
[0102] The image matching device based on complementary descriptors of the present invention comprises:
[0103] A model building module, used to build an image matching model, the image matching model includes a twin orthogonal descriptor network for generating feature descriptors and a detector network for extracting feature points, the twin orthogonal descriptor network includes two identical descriptor network branches, each descriptor network branch consists of an encoder and a decoder, the encoder in each descriptor network branch includes a multi-layer residual network module, and the decoder includes more than two upsampling networks;
[0104] A model training module, used to train the constructed image matching model using the training data set, and supervise the training of the two descriptor network branches using different loss functions during the training process, and train the two descriptor network branches using an orthogonal loss function so that the output descriptor features of the two orthogonal descriptor networks are orthogonal, and generate a final similarity matrix by combining the similarity matrices of the two descriptor network branches;
[0105] The matching detection module is used to input into the trained image matching model to obtain the matching relationship results between the image pairs to be matched.
[0106] The image matching device based on complementary descriptors in this embodiment corresponds one to one with the above-mentioned image matching method based on complementary descriptors, and will not be described in detail here.
[0107] This embodiment further provides a computer device, including a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to perform the above method.
[0108] It is understandable that the above method of this embodiment can be executed by a single device, such as a computer or server, etc., and can also be applied to a distributed scenario and completed by multiple devices in cooperation with each other. In the case of a distributed scenario, one of the multiple devices can only execute one or more steps in the above method of this embodiment, and multiple devices interact to complete the above method. The processor can be implemented in the form of a general-purpose CPU, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, etc., for executing related programs to implement the above method of this embodiment. The memory can be implemented in the form of a read-only memory ROM, a random access memory RAM, a static storage device, and a dynamic storage device. The memory can store an operating system and other applications. When the above method of this embodiment is implemented by software or firmware, the relevant program code is stored in the memory and called and executed by the processor.
[0109] This embodiment further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above method is implemented.
[0110] Those skilled in the art should understand that the above-mentioned embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the process Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0111] The above is only a preferred embodiment of the present invention, and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment, it is not intended to limit the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the scope of protection of the technical solution of the present invention.
Claims
1. An image matching method based on complementary descriptors, characterized in that the steps include: Constructing an image matching model, the image matching model includes a twin orthogonal descriptor network for generating feature descriptors and a detector network for extracting feature points, the twin orthogonal descriptor network includes two identical but complementary descriptor network branches, each of the descriptor network branches is composed of an encoder and a decoder, the encoder in each of the descriptor network branches includes a multi-layer residual network module, and the decoder includes more than two upsampling networks; The constructed image matching model is trained using a training data set, and during the training process, different loss functions are used to supervise the training of the two descriptor network branches, and an orthogonal loss function is used to train the two descriptor network branches so that the output descriptor features of the two orthogonal descriptor networks are orthogonal, and a final similarity matrix is generated by combining the similarity matrices of the two descriptor network branches, and finally a nearest neighbor algorithm is used to obtain a matching relationship; The image pairs to be matched are obtained and input into the trained image matching model to obtain the matching relationship results between the image pairs to be matched.
2. The image matching method based on complementary descriptors according to claim 1, characterized in that: During the training of the twin orthogonal descriptor network, one descriptor network is supervised and trained using negative log-likelihood loss, and the other descriptor network is supervised and trained using triplet loss, and the two descriptor features are decoupled by adding orthogonal loss.
3. The image matching method based on complementary descriptors according to claim 2, characterized in that: The expression of the negative log-likelihood loss is: in, represents the negative log-likelihood loss, N represents the number of feature points, Represents the descriptor similarity matrix of the two images to be matched, ,in , According to the two images to be matched The corresponding feature point information is obtained from the dense feature map generated by the first descriptor network The descriptor set sampled from Respectively represent the two images to be matched The corresponding feature point numbers extracted from the hour Represents the two images to be matched The descriptor similarity matrix S The main diagonal elements in , descriptor similarity matrix S The main diagonal elements of are the similarity measures between corresponding feature points.
4. The image matching method based on complementary descriptors according to claim 2, characterized in that: The expression of the triplet loss is: in, represents the triplet loss, N represents the number of feature points, Represents the Euclidean distance between the corresponding descriptors of the two images to be matched, represents the most difficult negative sample distance, , According to the two images to be matched The corresponding feature point information is generated from the dense feature map generated by the second descriptor description The descriptor set sampled from is the margin parameter, N represents the number of feature points, Respectively represent the two images to be matched The extracted corresponding feature point numbers that match each other, k Indicates removal = The serial number of the point with the greatest similarity outside the corresponding points.
5. The image matching method based on complementary descriptors according to claim 1, characterized in that: The expression of the orthogonal loss is: in, represents the orthogonal loss, Represents the absolute value after calculating the dot product between two descriptors, N represents the number of feature points, , According to the two images to be matched The corresponding feature point information is obtained from the dense feature map generated by the first descriptor network The descriptor set sampled from , According to the two images to be matched The corresponding feature point information is generated from the dense feature map generated by the second descriptor description The descriptor set sampled from They respectively represent the serial numbers of the corresponding feature points that match each other extracted from the two images to be matched.
6. The image matching method based on complementary descriptors according to any one of claims 1 to 5, characterized in that: The similarity matrices of the two descriptor network branches are combined to generate the final similarity matrix according to the following formula: : in, represents the weight, , Represent the image pairs to be matched The characteristic points of They represent the corresponding feature point numbers extracted from the two images to be matched, , They represent the descriptor sets obtained by sampling the dense descriptor features output by the two descriptor networks using the corresponding feature points. , Respectively represent the similarity matrices calculated using the two descriptor network branches.
7. The image matching method based on complementary descriptors according to any one of claims 1 to 5, characterized in that: It also includes using the specified model as the teacher network to perform distillation learning on the detector network, learning to detect repeated feature points by learning the distribution of the score map of the teacher model, and the loss function is defined as: in, represents the binary cross entropy loss, Output score map for the teacher model, Output score map of the student model.
8. An image matching device based on complementary descriptors, characterized in that: include: A model construction module, used to construct an image matching model, wherein the image matching model includes a twin orthogonal descriptor network for generating feature descriptors and a detector network for extracting feature points, wherein the twin orthogonal descriptor network includes two identical descriptor network branches, each of which is composed of an encoder and a decoder, wherein the encoder in each of the descriptor network branches includes a multi-layer residual network module, and the decoder includes more than two upsampling networks; A model training module, used to train the constructed image matching model using a training data set, and supervise the training of the two descriptor network branches using different loss functions during the training process, and train the two descriptor network branches using an orthogonal loss function to decouple the output descriptor features of the two orthogonal descriptor networks, and generate a final similarity matrix by combining the similarity matrices of the two descriptor network branches; finally, a nearest neighbor algorithm is used to obtain a matching relationship; The matching detection module is used to input into the trained image matching model to obtain the matching relationship results between the image pairs to be matched.
9. A computer device comprising a processor and a memory, wherein the memory is used to store a computer program, wherein: The processor is configured to execute the computer program to perform the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Pavement crack identification method based on point cloud-RGB (Red, Green, Blue) heterogenous image multistage registration mapping
CN117036300A
Semi-supervised medical image segmentation method based on data enhancement strategy
CN117710681A