Feature matching method, system and equipment based on multi-geometry collaborative learning and storage medium

Through the feature matching method based on multi-geometric collaborative learning, the problem of difficult balance in the speed and accuracy of feature matching in the prior art is solved, efficient matching in different texture scenarios is achieved, and the robustness and accuracy of matching are improved through geometric constraints.

CN120182375AActive Publication Date: 2025-06-20WUHAN ZHIXINFEI TECHNOLOGY PARTNERSHIP (LLP)
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510255910.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-20
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

The prior art is difficult to achieve a better balance of feature matching in speed and accuracy, especially in different texture scenarios. Traditional methods are inefficient in strong texture scenarios and insufficient accuracy in weak texture scenarios.

Method used

A feature matching method based on multi-geometric collaborative learning is adopted to build a sparse network by acquiring image pairs, extracting sparse local features and inputting to the multi-geometric collaborative network model. This method enhances local features through attention mechanism, embeds affine geometric information, calculates pose and homographic geometric information, combines manual annotation to perform loss calculations, and optimizes the multi-geometric collaborative network model.

Benefits of technology

The better balance of feature matching in speed and accuracy is achieved, the matching performance in different texture scenarios is improved, the feature robustness is improved through geometric constraints, the geometric solution ambiguity is reduced, and the forward cycle of end-to-end joint optimization is formed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182375A_ABST
    Figure CN120182375A_ABST
Patent Text Reader

Abstract

The invention provides a feature matching method, system and device based on multi-geometry collaborative learning and a storage medium, and relates to the technical field of robot positioning and navigation.The method comprises the steps that an image pair is acquired, and two sets of sparse local features in the image pair are extracted and input into a multi-geometry collaborative network model; local features are enhanced through an attention mechanism, and iterative updating is performed on the sparse network through a multi-layer perceptron and / or linear projection; generating and screening reliable matching points through affine transformation estimation, position coding and key point neighborhood extension; constraining a matching process through pose information, and adaptively adjusting a matching strategy to carry out matching prediction; carrying out loss calculation in combination with real matching information and homography geometric information so as to train the multi-geometric network model; repeating the steps until a preset number of iterations is reached; through collaborative optimization of affine geometry, epipolar geometry and homographic geometry, mutual promotion of local feature discrimination and geometric consistency is realized, and better balance of feature matching in speed and precision can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot positioning and navigation, and particularly to a feature matching method, system, device and storage medium based on multi-geometry collaborative learning. Background Art

[0002] Feature matching is a basic task in intelligent robot positioning and navigation, aiming to find corresponding points from image pairs of overlapping scenes and eliminate non-matchable points.

[0003] However, feature matching faces challenges such as scene differences like illumination, perspective, texture, and appearance changes, as well as real-time requirements; traditional methods extract sparse features based on key point detectors and use Transformer for matching, performing excellently in strong texture scenes but having limited effects in weak texture scenes; while dense matching methods solve the weak texture matching problem but have redundant calculations and low efficiency in strong texture scenes; therefore, it is difficult to achieve a better balance between the speed and accuracy of feature matching. Summary of the Invention

[0004] The purpose of the present invention is to provide a feature matching method, system, device and storage medium based on multi-geometry collaborative learning to solve the problem that the prior art in the above-mentioned background art is difficult to achieve a better balance between the speed and accuracy of feature matching.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A feature matching method based on multi-geometry collaborative learning, applied to a robot positioning and navigation system, the method steps include: S1: Obtain an image pair, extract two sets of sparse local features in the image pair and input them into a multi-geometry collaborative network model to construct a sparse network; S2: Enhance the local features through an attention mechanism, and iteratively update the sparse network through a multi-layer perceptron and / or linear projection; S3: Embed affine geometry information into the local feature representation, and expand and screen reliable matching points in the key point neighborhood in the sparse network; S4: Calculate the pose information of the mutually corresponding reliable matching points, use epipolar geometry information to constrain the matching process, and adaptively adjust the matching strategy for matching prediction; S5: Calculate the homography geometry information of the mutually corresponding reliable matching points, combine the true matching information obtained by manual annotation and the homography geometry information for loss calculation to train the multi-geometry network model; S6: Repeat steps S2 - S5 until a preset number of iterations is reached to optimize the multi-geometry collaborative network model.

[0006] Optionally, the S1 step specifically includes: obtaining two images of the same scene, respectively extracting sparse local features in the two images, where each local feature in the sparse local features includes a key point and a descriptor; pre-predicting a partial assignment matrix between two sets of sparse local features through a multi-geometric collaborative network model, and outputting a set of matching matrices, and constructing a sparse network through the assignment matrix and the matching matrix.

[0007] Optionally, the S2 step specifically includes: updating the intermediate state vectors of the two images, aggregating information to enhance local features by combining self-attention and cross-attention mechanisms, and iteratively updating the intermediate state vectors of different layers through non-linear transformations of a multi-layer perceptron and / or linear projection.

[0008] Optionally, the S3 step specifically includes: estimating an affine transformation through the corresponding relationship of key point positions in two sets of sparse local features to obtain affine geometric information; where the affine transformation estimation equation is: In the formula, the is the affine geometric information, the is the rotation component to be solved, and the is the translation component to be solved; generating dense matching points by expanding in the key point neighborhood, and performing multiple rounds of iteration through embedding affine geometric information in the attention scores to screen the dense matching points to obtain reliable matching points; where the cross-attention score embeds the affine geometric information between the image pair I 1 and I 2 , and its calculation formula is: The self-attention score encodes the self-relative position in the image I 1 , and its calculation formula is: In the formula, the is the cross-attention score, the is the self-attention score, the is the transposed attention query, the R(p) is the rotation encoding, the kp j is the key point position of the image I 2 , the kp i is the key point position of the image I 1 , the is the attention key of the image I 1 , and the is the attention key of the image I 2 .

[0009] Optionally, the S4 step specifically includes: converting the state into an enhanced descriptor through learned linear projection, predicting a partial assignment matrix by combining the matching likelihood score and the similarity score, then outputting a matching matrix according to a set threshold, and removing untrusted key points based on a confidence classifier and a decay threshold.

[0010] Optionally, the S4 step specifically further includes: estimating an essential matrix through the corresponding reliable matching points, calculating the epipolar distance, pruning the sparse matching and the dense matching respectively, and deciding whether to exit the network layer iteration according to the confidence of the matching points.

[0011] Optionally, the S5 step specifically includes: calculating the homography geometric information of the corresponding reliable matching points, obtaining the true matching information through manual annotation, determining the unmatched points according to the relationship between the predicted matching information obtained by the matching prediction and the true matching information, calculating the matching loss, estimating the homography using the predicted matching information and calculating the homography loss, and obtaining the total loss by integrating the matching losses and the homography losses of each layer.

[0012] On the other hand, the present invention also provides a feature matching system based on multi - geometry collaborative learning, including: a feature extraction module, configured to obtain an image pair and extract two sets of sparse local features in the image pair and input them into a multi - geometry collaborative network model to construct a sparse network; a feature enhancement module, configured to enhance the local features through an attention mechanism and iteratively update the sparse network through a multi - layer perceptron and / or linear projection; a feature screening module, configured to generate and screen reliable matching points through affine transformation estimation, position encoding, and key - point neighborhood expansion; a matching prediction module, configured to calculate the pose information of the corresponding reliable matching points, constrain the matching process through the pose information, and adaptively adjust the matching strategy for matching prediction; a loss calculation module, configured to calculate the homography geometric information of the corresponding reliable matching points, obtain the true matching information through manual annotation, and calculate the loss by combining the true matching information and the homography geometric information; a model optimization module, configured to repeat the steps until a preset number of iterations is reached to optimize the multi - geometry collaborative network model.

[0013] On the other hand, the present invention also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above - mentioned feature matching method based on multi - geometry collaborative learning are implemented.

[0014] On the other hand, the present invention also provides a computer - readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above - mentioned feature matching method based on multi - geometry collaborative learning are implemented.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0016] 1. In order to achieve a better balance between speed and accuracy in feature matching and improve the matching performance in different texture scenes, this application proposes a scheme to enhance the sparse matching ability of the multi-geometric collaborative network by obtaining image pairs, extracting two sets of sparse local features in the image pairs and inputting them into the multi-geometric collaborative network model to construct a sparse network; through the collaborative optimization of affine geometry, epipolar geometry and homography geometry, the mutual promotion of local feature discriminability and geometric consistency is realized, the feature robustness is improved through geometric constraints, and the high-quality features reduce the geometric solution ambiguity, forming a positive cycle of end-to-end joint optimization; in the inference stage, the fusion of multi-geometric constraints can be quickly completed, and a better balance between speed and accuracy of feature matching can be achieved.

[0017] 2. At the front end of the multi-geometric collaborative network model, affine geometric information is embedded in the local feature representation, and reliable matching points are generated and screened in the neighborhood of key points in the sparse network. By using affine enhancement to assist in extracting enhanced local features, the matching effect in challenging scenes such as weak texture can be improved.

[0018] 3. At the middle end of the multi-geometric collaborative network model, the pose information of the corresponding reliable matching points is calculated, the matching process is constrained by epipolar geometric information, and the matching strategy is adaptively adjusted for matching prediction.

[0019] 4. At the back end of the multi-geometric collaborative network model, the homography geometric information of the corresponding reliable matching points is calculated, and the loss is calculated by combining the true matching information obtained by manual annotation and the homography geometric information to train the multi-geometric network model. A homography supervision mechanism is introduced, and pixel-level correspondence is established through homography geometric constraints, driving the network to learn more accurate matching capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram of the overall architecture of the method of the present invention.

[0021] Figure 2 It is a schematic diagram of the method step flow of the present invention.

[0022] Figure 3 It is a schematic diagram of the system function modules of the present invention. Figure 4 It is a schematic diagram of the device structure of the present invention.

[0023] In the figure, 10 - feature extraction module, 20 - feature enhancement module, 30 - feature screening module, 40 - matching prediction module, 50 - loss calculation module, 60 - model optimization module, 70 - processor, 80 - memory. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] The following will clearly and completely describe the solution of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments.

[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of the present application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0026] Those skilled in the art of this technology can understand that, unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of this application means the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0027] Those skilled in the art of this technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which this application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0028] It should be understood that the sequence numbers and magnitudes of the steps in this embodiment do not mean the order of execution. The order of execution of each process is determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0029] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0030] Please refer to Figure 1 - Figure 2 A feature matching method based on multi-geometry collaborative learning of the present invention is applied to a robot positioning and navigation system, and the method steps include:

[0031] S1. Obtain an image pair, extract two groups of sparse local features from the image pair, and input them into a multi-geometric collaborative network model to construct a sparse network.

[0032] Specifically, the image pair is two images with a specific correlation relationship. In the present application, it can be understood as two pictures in the same scene, such as images of the same scene from different perspectives, images of the same object at different time points, images with similar content but large differences in style, or an original image and an image enhanced by data, all of which have a certain degree of correlation; by extracting sparse local features of the two images and inputting them into a multi-geometric collaborative network model to construct a sparse network, a large amount of calculations can be avoided.

[0033] S2. Enhance local features through an attention mechanism, and iteratively update the sparse network through a multi-layer perceptron and / or linear projection.

[0034] Specifically, the attention mechanism includes a self-attention mechanism and a cross-attention mechanism, which can enhance the representation capability of local features by dynamically focusing on key areas in the image, thereby improving the representation accuracy and robustness of local features; it can mine deep local feature correspondences through a multi-layer perceptron, and reduce the amount of calculation and control the computational complexity by mapping high-dimensional sparse features to low-dimensional space.

[0035] S3, embed affine geometric information into local feature representation, and expand the neighborhood of key points in the sparse network to generate and screen reliable matching points.

[0036] Specifically, the affine transformation is estimated through the correspondence between the key point positions in two sets of sparse local features to obtain affine geometric information, dense matching points are generated by expansion in the key point neighborhood, and multiple rounds of iterations are performed by embedding the affine geometric information in the attention score to screen the dense matching points to obtain reliable matching points, which can improve the quantity and quality of image feature matching. At the same time, the rotation encoding in the affine geometric information is used to help the model capture global and relative position information.

[0037] S4. Calculate the position and posture information of the corresponding reliable matching points, use the epipolar geometry information to constrain the matching process, and adaptively adjust the matching strategy to perform matching prediction.

[0038] Specifically, based on adaptive prediction embedding pose guidance for guided matching prediction to complete the matching prediction task, the state is transformed into an enhanced descriptor by learning linear projection, the partial assignment matrix is predicted by combining the matching possibility score and the similarity score, and then the matching matrix is output according to the set threshold. The fundamental matrix is estimated from the corresponding reliable matching points, the epipolar distance is calculated, outliers in the initial corresponding relationship are detected and removed based on the epipolar geometry constraint, the matching process is optimized using epipolar geometry, pruning is performed on sparse and dense matches respectively, and whether to exit the network layer iteration is determined according to the confidence of the matching points. The matching points are screened and optimized through epipolar constraint, and the guided matching prediction is carried out quickly, reducing redundancy and making the network layer iteration from coarse to fine.

[0039] S5. Calculate the homography geometric information of the corresponding reliable matching points, and calculate the loss by combining the true matching information obtained from manual annotation and the homography geometric information to train the multi-geometric network model.

[0040] Specifically, calculate the homography geometric information of the corresponding reliable matching points, obtain the true matching information through manual annotation, determine the unmatched points according to the relationship between the predicted matching information obtained from the matching prediction and the true matching information, calculate the matching loss, estimate the homography using the predicted matching information and calculate the homography loss, and obtain the total loss by integrating the matching loss and the homography loss of each layer. During the multi-supervised training of the multi-geometric collaborative network model, the loss is calculated by combining the true matching information and the homography geometric information, greatly improving the performance of the multi-geometric collaborative network model.

[0041] S6. Repeat steps S2 - S5 until the preset number of iterations is reached to optimize the multi-geometric collaborative network model.

[0042] Specifically, first pre-train the network on a synthetic homography large dataset sampled from the Oxford-Paris retrieval-1M images; then continue to fine-tune on the MegaDepth dataset; the model is implemented using PyTorch, with the Adam optimizer. The batch size is set to 128 during pre-training and 32 during fine-tuning. The initial learning rate is set to 0.0001 on a 2x3090 GPU with 24GB video memory; the number of iterative layers L = 9 is adaptive in depth and width and is also affected by the method of affine augmentation and pose guidance; it should be noted that the multi-geometric matrix estimated at the L-th layer is approximated to the estimated value of some estimators; after training, a well-performing MGC-Net model can be obtained for subsequent image matching and other related tasks.

[0043] It can be understood that, in order to achieve a better balance between speed and accuracy in feature matching and improve the matching performance in different texture scenarios, the present application proposes a scheme for enhancing the sparse matching ability of a multi-geometric collaborative network by obtaining an image pair, extracting two sets of sparse local features in the image pair and inputting them into a multi-geometric collaborative network model to construct a sparse network; enhancing local features through an attention mechanism, and iteratively updating the sparse network through a multi-layer perceptron and / or linear projection, which greatly improves the computational efficiency; at the front end of the multi-geometric collaborative network model, embedding affine geometric information into the local feature representation, generating and screening reliable matching points in the neighborhood of key points in the sparse network, and extracting enhanced local features through affine enhancement assistance, which can improve the matching effect in challenging scenarios such as weak texture; at the middle end of the multi-geometric collaborative network model, calculating the pose information of the corresponding reliable matching points, constraining the matching process through epipolar geometric information, and adaptively adjusting the matching strategy for matching prediction; at the back end of the multi-geometric collaborative network model, calculating the homography geometric information of the corresponding reliable matching points, obtaining real matching information through manual annotation, and calculating the loss by combining the real matching information and the homography geometric information to train the multi-geometric network model, introducing a homography supervision mechanism, establishing pixel-level correspondence through homography geometric constraints, and driving the network to learn more accurate matching capabilities; the multi-geometric collaborative network model realizes the mutual promotion of local feature discriminability and geometric consistency through the collaborative optimization of affine geometry, epipolar geometry and homography geometry, improves the robustness of features through geometric constraints, and high-quality features reduce the ambiguity of geometric solution, forming a positive cycle of end-to-end joint optimization; in the inference stage, multi-geometric constraint fusion can be quickly completed, and a better balance between speed and accuracy of feature matching can be achieved.

[0044] In some embodiments, the S1 step specifically includes: obtaining two images of the same scene, respectively extracting sparse local features in the two images, wherein each local feature in the sparse local features includes a key point and a descriptor; predicting a partial assignment matrix between two sets of sparse local features in advance through a multi-geometric collaborative network model, and outputting a set of matching matrices, and constructing a sparse network through the assignment matrix and the matching matrix.

[0045] Specifically, for two images I 1 and I 2 of the same given scene, processing them with a feature detector and a descriptor to obtain two sets of sparse local features F 1 and F 2 ; each local feature is composed of a normalized key point position and a descriptor; wherein, the key point position calculation formula is: kp i =(x, y) i ∈[0, 1] 2, where \(K_p\) i is the key point position; the descriptor calculation formula is: where \(d\) i is the descriptor and \(d\) is the vector dimension.

[0046] Furthermore, the multi - geometric collaborative network model pre - predicts a partial assignment matrix \(P\) between two sets of sparse local features, and then outputs a set of matching matrices \(M\). The \(l\) - th layer of the stacked sparse network is iteratively updated with each other through an attention mechanism, and a classifier is set to adaptively determine the depth and width to stop the update.

[0047] Furthermore, at each layer, the similarity and matching scores are combined to predict the \(l\) - th layer assignment matrix \((l)P\). Based on this prediction result, the top \(N\) t matches with higher scores are further selected to estimate the multi - geometric relationship, because excellent local matches can approximate the global correspondence.

[0048] Furthermore, the subsequent multi - geometric collaborative network model estimates the multi - geometric relationship of each layer, including the affine transformation fundamental matrix and the homography matrix where the fundamental matrix can be decomposed into pose, and these geometric relationships are collaboratively embedded into the network to learn the matching.

[0049] Through the collaborative optimization of affine geometry, epipolar geometry, and homography geometry, the mutual promotion between local feature discriminability and geometric consistency is realized. The geometric constraints are used to enhance the feature robustness, while the high - quality features reduce the geometric solution ambiguity, forming a positive cycle of end - to - end joint optimization; in the inference stage, the multi - geometric constraint fusion can be quickly completed, and a better balance between the speed and accuracy of feature matching can be achieved. The multi - geometric collaborative learning strategy proposed in this application can be effectively applied to the positioning and navigation of intelligent robots.

[0050] In some embodiments, the step S2 specifically includes: updating the intermediate state vectors of two images, aggregating information by combining self - attention and cross - attention mechanisms to enhance local features, and iteratively updating the intermediate state vectors of different layers through non - linear transformations of a multi - layer perceptron and / or linear projection.

[0051] Specifically, from all the intermediate states 1 and 2 corresponding to two images \(I\) and an intermediate state (l) \(x\) i at layer \(l\lt10\) is randomly selected. When \(l = 0\), it is initialized as the descriptor \(d\) i, a multi - layer perceptron is used to combine the messages obtained from the aggregation of the entire local features. The intermediate state of the $l$-th layer is updated according to the calculation formula. Among them, the calculation formula for the intermediate state of the $l$-th layer is: (l) x i (l-1) x i +mlp( (l -1) x i |m i ), where, the (l) x i is the intermediate state, the [·|·] is the concatenation of two vectors, and the m i is the message.

[0052] Furthermore, the message m i consists of two parts: the self - attention aggregated message and the cross - attention aggregated message . Assuming that the state x i is in the image I 1 , the self - attention aggregated message is aggregated by each local feature in the image I 1 combined with the self - attention score . The calculation formula is where, the is the self - attention aggregated message, and the is the self - state attention value; among them, the self - state attention value is calculated by a learnable linear projection of the self - state 1 in the image I . The calculation formula is where, the is the self - state attention value, the is the weight term, the is the self - state, and the is the bias term; while the cross - attention aggregated message is aggregated by each local feature in the image I 2 combined with the cross - attention score . Its calculation formula is where, the is the cross - attention aggregated message, is the cross - state attention value; among them, the cross - state attention value is calculated by a learnable linear projection of the cross - state 2 in the image I . The calculation formula is where, the is the cross - state attention value, the is a weight term, and the is a cross state, and the is a bias term; meanwhile, use linear projection to learn the parameter to obtain the corresponding attention key and calculate the attention query includes the weight term W q and the bias term b q ; the attention score α ij is the similarity between the attention query q i and the attention key k j and different position encoders for self-attention and cross-attention units are embedded inside; by dynamically focusing on the key regions in the image, the representation ability of local features is enhanced, and the representation accuracy and robustness of local features can be improved; through the multi-layer perceptron, the corresponding relationship of deep local features is mined, and by mapping the high-dimensional sparse features to the low-dimensional space, the amount of calculation is reduced and the computational complexity is controlled.

[0053] In some embodiments, the S3 step specifically includes: estimating the affine transformation through the corresponding relationship of key point positions in two sets of sparse local features to obtain the affine geometric information, expanding and generating dense matching points in the key point neighborhood, and performing multiple rounds of iteration through the affine geometric information embedded in the attention score to screen the dense matching points to obtain reliable matching points.

[0054] Specifically, select the first N t > 3 corresponding relationships of matching key point positions {(x, y) i , (x, y) j} from the image, estimate the 6-DOF affine transformation Establish the affine transformation estimation equation as In the formula, is the rotation component to be solved, is the translation component to be solved, and the two together form the affine transformation between images

[0055] Furthermore, embed the affine geometric information into the attention score. For the cross-attention score use the position encoding to embed the affine geometric information between the image I 1 and the image I 2 . Its calculation formula is: For the self-attention score use the self-relative position in the image I 1 for encoding, and its calculation formula is: Among them, the is the transpose of the attention query, and the is the image I 1The attention key, the is the image I 2 The attention key, the kp i is the image I 1 The key point position of the image I, the kp j is the image I 2 The key point position of the image I, the is the affine transformation estimated by the affine transformation estimation equation; the rotation encoding calculation formula is: In the formula, the R(p) is the rotation encoding, following the Fourier feature, and the k-th rotation angle θ k is determined by the learnable angular frequency ω k and the parameter p, enabling the model to capture the global and relative positions.

[0056] Furthermore, in each sparse key point kp i A set of dense points is extended in the neighborhood to generate the dense matching M dense , with the key point position kp i as the center, and a number N e of dense points np i,k =(x,y) i,k ∈[0,1] 2 , k = 1,2,...N e , and its calculation formula is: Taking the image I 1 as the source image, through the affine transformation project these dense points np i,k onto the target image I 2 to obtain another set of dense points np j,k , thereby generating the dense matching M dense ={(np i,k , np j,k )}, that is:

[0057] Furthermore, filter the reliable matching points; remove the extended points projected outside the image I 2 at each layer, and use the threshold λ 0→L to retain the reliable matches from the 0th layer to the Lth layer; by calculating the error error between the extended points (l) np j,k in the l-th iteration layer and the extended points (L) np j,k in the L-th iteration layer, and comparing it with the threshold λ 0→L to screen out the final reliable matching points, where the error calculation formula is:

[0058] In some embodiments, the step S4 specifically includes: converting the state into an enhanced descriptor through learned linear projection, predicting a partial assignment matrix by combining the matching likelihood score and the similarity score, and then outputting a matching matrix according to a set threshold, and removing untrusted key points based on a confidence classifier and a decay threshold.

[0059] Specifically, at each layer l, the learned linear projection is used to convert each state x i into an enhanced descriptor y i , where the weight term W and the bias term b are learned parameters; calculating the matching likelihood score based on these enhanced descriptors This score represents the probability that a key point matches its corresponding point, and its calculation formula is

[0060] Furthermore, calculate the similarity score which reflects the similarity degree between the key points in the image I 1 and the image I 2 , and its calculation formula is In the formula, is the enhanced descriptor on the image I 1 , is to traverse the enhanced descriptor on the image I 2 .

[0061] Furthermore, jointly calculate the assignment matrix P ij by combining the matching likelihood score and the similarity score, and its calculation formula is When the assignment matrix P ij is greater than the threshold τ and is the maximum value in its row and column, output a set of matches, denoted as the sparse match M sparse .

[0062] Furthermore, set a confidence classifier c f = sigmoid(mlp(x f ))), where f is the index of the local feature, and x f is the intermediate state corresponding to the local feature, and infer the confidence of the predicted assignment of each key point; set the threshold λ l at the l-th layer. When the confidence classifier c f is less than the threshold λ l , prune and remove the untrusted key points, and this threshold will be attenuated and adjusted in each layer according to the verification accuracy of each classifier.

[0063] In some embodiments, the step S4 specifically further includes: estimating the fundamental matrix through the corresponding reliable matching points, calculating the epipolar distance, pruning the sparse match and the dense match respectively, and deciding whether to exit the network layer iteration according to the confidence of the matching points.

[0064] Specifically, estimate the fundamental matrix which indirectly reflects the pose information between images; starting from the correspondence relationships of the top N t > 7 matching points {(x i , y i ), (x j , y j )}, use the eight-point algorithm for estimation, and through iterative optimization, establish the equation as: Calculate the epipolar line e from the source image I 1 to the target image I 2 , and use the function dist(·) to calculate the distance d from the corresponding point p 2 =(x j , y j ) in the target image I j to the epipolar line e.

[0065] Furthermore, on the basis of key-point pruning, further use to prune the matching set M containing sparse matching M sparse and dense matching M dense to quickly guide matching prediction; at each layer, use the sparse threshold λ sparse and the dense threshold λ dense to prune the matching, further reducing redundancy; for sparse matching pruning, since key-point pruning has been performed, the sparse threshold λ sparse is set relatively loosely; after prediction and key-point pruning, the pose guidance module further removes the outliers in the sparse matching M sparse , and calculates the confidence c f of each inlier, and compares it with the threshold λ l at the l-th layer to quickly and adaptively determine the network depth; n1 + n2 calculates the total number of key points of the two images, and the local feature f belongs to the local feature sets {F 1 , F 2} of the two images. The condition for judging whether to exit the network layer iteration is: In the formula, is the indicator function, and when the exit parameter exit satisfies being greater than the threshold parameter ρ, the network layer iteration can be exited.

[0066] Furthermore, for dense matching pruning, since rough matching expansion is performed in the affine enhancement module, the dense threshold λ dense is set relatively strictly; after removing the extended points projected outside the target image, generate the rough dense matching M dense ; from the extended rough dense matching M denseAmong them, at each layer, the matches inconsistent with the global epipolar geometry are eliminated, and then a threshold λ is used for screening from layer 0 to layer L; this is because the estimated fundamental matrix 0→L has a higher degree of freedom than the affine transformation and can guide the matches from coarse to fine, reducing the iterative burden of the network layer.

[0067] In some embodiments, the step S5 specifically includes: calculating the homography geometric information of the corresponding reliable matching points, obtaining the true matching information through manual annotation, determining the unmatched points according to the relationship between the predicted matching information obtained by the matching prediction and the true matching information, calculating the matching loss, estimating the homography using the predicted matching information and calculating the homography loss, and obtaining the total loss by integrating the matching loss and the homography loss of each layer.

[0068] Specifically, in the predicted sparse matching, if some key points ( or ) are not close to the true matching M gt , they are marked as unmatched points; their unmatched scores are calculated separately including the unmatched score 1 in image I and the unmatched score 2 in image I After prediction and fine pruning, P ij is assigned by minimizing the negative log-likelihood, and the matching loss (l) L m of the l-th layer is calculated, and its calculation formula is:

[0069] Furthermore, similar to the affine transformation and the fundamental matrix , a homography of 8 degrees of freedom is estimated from the top N t > 4 matches for adding geometric supervision; the final prediction M is obtained by the affine enhancement module and the pose guidance module, and includes the sparse matching M pred from the negative log-likelihood minimization assignment P ij and the dense matching M sparse ; the homography loss dense L (l) of the l-th layer is calculated, and its calculation formula is: h wherein, the PE(·) is the projection error between the homography and the matching.

[0070] Furthermore, different weights α are applied to the multi-supervision, and the total loss Loss of the L-th layer is calculated, and its calculation formula is: ​​

[0071] In some embodiments, steps S2 - S5 are repeated until a preset number of iterations is reached to optimize the multi - geometry collaborative network model.

[0072] Specifically, first, the network is pre - trained on a synthetic homography large - scale dataset sampled from the Oxford - Paris retrieval - 1M images; then, it is further fine - tuned on the MegaDepth dataset. The model is implemented using PyTorch, with an Adam optimizer. The batch size is set to 128 during pre - training and 32 during fine - tuning. The initial learning rate is set to 0.0001 on a 2x3090 GPU with 24GB video memory. The number of iterative layers L = 9 is adaptive in depth and width and is also affected by the method of affine augmentation and pose guidance. It should be noted that the multi - geometry matrix estimated at the L - th layer is approximated to the estimated value of some estimators. After training, a well - performing MGC - Net model can be obtained for subsequent related tasks such as image matching.

[0073] Please refer to Figure 3 , on the other hand, the present invention also provides a feature matching system based on multi - geometry collaborative learning, including: a feature extraction module 10 for obtaining an image pair and extracting two sets of sparse local features in the image pair and inputting them into the multi - geometry collaborative network model to construct a sparse network; a feature enhancement module 20 for enhancing local features through an attention mechanism and iteratively updating the sparse network through a multi - layer perceptron and / or linear projection; a feature screening module 30 for generating and screening reliable matching points through affine transformation estimation, position encoding, and key - point neighborhood expansion; a matching prediction module 40 for calculating the pose information of the mutually corresponding reliable matching points, constraining the matching process through the pose information, and adaptively adjusting the matching strategy for matching prediction; a loss calculation module 50 for calculating the homography geometric information of the mutually corresponding reliable matching points, obtaining real matching information through manual annotation, and calculating the loss by combining the real matching information and the homography geometric information; a model optimization module 60 for repeating the steps until a preset number of iterations is reached to optimize the multi - geometry collaborative network model.

[0074] Please refer to Figure 4 , on the other hand, the present invention also provides a computer device, including a memory 80 and a processor 70. The memory stores a computer program, and when the processor executes the computer program, the steps of the above - mentioned feature matching method based on multi - geometry collaborative learning are implemented.

[0075] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned feature matching method based on multi-geometry collaborative learning are implemented.

[0076] Specifically, Figure 4 FIG. is a schematic structural diagram of a feature matching device based on multi-geometry collaborative learning provided by an embodiment of the present invention. The feature matching device based on multi-geometry collaborative learning may vary greatly due to configuration or performance differences, and may include one or more processors 70 (central processing units, CPUs) (for example, one or more processors 70) and a memory 80, and one or more storage media for storing application programs or data (for example, one or more mass storage devices); wherein, the memory 80 and the storage media may be transient storage or persistent storage; further, the processor 70 may be configured to communicate with the storage media and execute a series of instruction operations in the storage media on the feature matching device based on multi-geometry collaborative learning to implement the steps of the feature matching method based on multi-geometry collaborative learning provided by the above-mentioned method embodiments.

[0077] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The foregoing storage media include: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs and other media that can store program codes.

[0078] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to the memory 80, storage, database, or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memories 80. The non-volatile memory 80 can include read-only memory 80 (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. The volatile memory 80 can include random access memory 80 (RAM) or an external cache memory 80. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0079] The above are only embodiments of the present invention and do not limit the patent scope of the present invention. All equivalent transformations made using the specification and drawings of the present invention, directly or indirectly applied in related technical fields, are equally included in the patent protection scope of the present invention.

Claims

1. A feature matching method based on multi-geometry collaborative learning, applied to a robot positioning and navigation system, characterized in that: The method steps include: S1: Obtain an image pair, extract two groups of sparse local features from the image pair and input them into a multi-geometry collaborative network model to construct a sparse network; S2: Enhance local features through an attention mechanism and iteratively update the sparse network through a multi-layer perceptron and / or linear projection; S3: Embed affine geometric information into local feature representation, and expand the neighborhood of key points in the sparse network to generate and filter reliable matching points; S4: calculating the position and posture information of the corresponding reliable matching points, using the epipolar geometry information to constrain the matching process, and adaptively adjusting the matching strategy to perform matching prediction; S5: Calculating homography geometric information of the mutually corresponding reliable matching points, and performing loss calculation based on the real matching information obtained by manual annotation and the homography geometric information, so as to train the multi-geometry network model; S6: Repeat steps S2-S5 until a preset number of iterations is reached to optimize the multi-geometry collaborative network model.

2. The feature matching method based on multi-geometric collaborative learning according to claim 1 is characterized in that: The S1 step specifically includes: Acquire two images of the same scene, and extract sparse local features from the two images respectively, wherein each local feature in the sparse local features includes a key point and a descriptor; A partial allocation matrix between two groups of sparse local features is predicted in advance through a multi-geometry collaborative network model, and a group of matching matrices is output, and a sparse network is constructed through the allocation matrix and the matching matrix.

3. The feature matching method based on multi-geometric collaborative learning according to claim 1 is characterized in that: The S2 step specifically includes: The intermediate state vectors of the two images are updated, and the self-attention and cross-attention mechanisms are combined to aggregate information and enhance local features. The intermediate state vectors of different layers are iteratively updated through nonlinear transformations of multi-layer perceptrons and / or linear projections.

4. The feature matching method based on multi-geometry collaborative learning according to claim 1, characterized in that: The S3 step specifically includes: Affine geometric information is obtained by estimating the affine transformation through the position correspondence of key points in two sets of sparse local features; Among them, the affine transformation estimation equation is: In the formula, is the affine geometric information, is the rotation component to be solved, is the translation component to be solved; Dense matching points are generated by expanding the neighborhood of key points, and multiple rounds of iterations are performed by embedding affine geometric information in the attention scores to screen the dense matching points to obtain reliable matching points; Among them, the cross attention score uses the position encoding to embed the image pair I 1 and I 2 The affine geometric information between is calculated as follows: Self-attention score using image I 1 The relative position in is encoded, and the calculation formula is: In the formula, is the cross attention score, is the self-attention score, is the attention query transpose, R(p) is the rotation code, and kp j For image I 2 The key point position, the kp i For image I 1 The key point position of For image I 1 The attention key, For image I 2 Attention key.

5. The feature matching method based on multi-geometry collaborative learning according to claim 1, characterized in that: The S4 step specifically includes: The state is converted into an enhanced descriptor by learning linear projection, and the matching possibility score and similarity score are combined to predict the partial allocation matrix. Then, the matching matrix is ​​output according to the set threshold, and untrustworthy key points are removed based on the confidence classifier and the attenuation threshold.

6. The feature matching method based on multi-geometry collaborative learning according to claim 4 is characterized in that: The S4 step specifically includes: The basic matrix is ​​estimated by the corresponding reliable matching points, the epipolar distance is calculated, the sparse matching and the dense matching are pruned respectively, and whether to exit the network layer iteration is determined according to the confidence of the matching points.

7. The feature matching method based on multi-geometry collaborative learning according to claim 1, characterized in that: The S5 step specifically includes: Calculate the homography geometric information of the corresponding reliable matching points, obtain the real matching information through manual annotation, determine the unmatched points according to the relationship between the predicted matching information obtained by matching prediction and the real matching information, calculate the matching loss, estimate the homography using the predicted matching information and calculate the homography loss, and obtain the total loss by combining the matching loss and homography loss of each layer.

8. A feature matching system based on multi-geometry collaborative learning, characterized in that: include: A feature extraction module is used to obtain an image pair, extract two groups of sparse local features from the image pair and input them into a multi-geometry collaborative network model to construct a sparse network; A feature enhancement module, used to enhance local features through an attention mechanism and iteratively update the sparse network through a multi-layer perceptron and / or linear projection; Feature screening module, used to generate and screen reliable matching points through affine transformation estimation, position encoding, and key point neighborhood expansion; A matching prediction module, used to calculate the position and posture information of the corresponding reliable matching points, constrain the matching process through the position and posture information, and adaptively adjust the matching strategy to perform matching prediction; A loss calculation module, used to calculate homography geometric information of the mutually corresponding reliable matching points, obtain real matching information through manual annotation, and perform loss calculation based on the real matching information and the homography geometric information; The model optimization module is used to repeat the steps until a preset number of iterations is reached to optimize the multi-geometry collaborative network model.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the feature matching method based on multi-geometric collaborative learning described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the feature matching method based on multi-geometric collaborative learning described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • System and method of hybrid scene representation for visual simultaneous localization and mapping

    CA3202821A1

  • 6D pose estimation method and device based on geometric constraint collaborative attention network

    CN113269830A

  • Image matching method combining sparse and dense neighborhood consistency

    CN115631354A

  • Composite visual angle infrared target space reconstruction method based on deep learning

    CN118397064A

  • Systems and methods for image processing based on optimal transport and epipolar geometry

    US20230052582A1