Double-view matching pair optimization method and device based on context-aware network

By adopting a context-aware network-based method in dual-view matching pairing optimization, multiple rounds of feature updates are used to use the mapping layer, local feature reconstruction layer and global feature update layer to perform multiple rounds of feature updates, which solves the problem of poor optimization of dual-view matching pairing in the existing technology, and achieves more efficient matching pair set optimization.

CN120032148APending Publication Date: 2025-05-23ANHUI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510117697.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art has not yet met the requirements when using neural networks to optimize the original two-view matching pairing set, and there is room for improvement.

Method used

The dual-view matching pair optimization method based on context-aware network is adopted. The original dual-view matching pair set is mapped to a high-dimensional space through the mapping layer, and the local feature reconstruction layer and the global feature update layer are combined to perform multiple rounds of update global feature processing, and the probability set is finally output to optimize the matching pair set.

Benefits of technology

It significantly improves the optimization effect of the dual-view matching pair set, can accurately filter out the correct point matching pairs, and improves the accuracy and robustness of image matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032148A_ABST
    Figure CN120032148A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision processing, in particular to a double-view matching pair optimization method and device based on a context-aware network. The invention designs a context sensing network based on global and local combination, and the context sensing network is used for optimizing an original double-view matching pair set; according to the method, an original double-view matching pair set is mapped to a high-dimensional space through a mapping layer to obtain a high-dimensional feature set, and then the high-dimensional feature set is preliminarily enhanced and screened through a local feature reconstruction layer to obtain local features; performing multi-round updating global feature processing through N global feature updating layers to output a final probability set; the final probability set is high in accuracy, correct point matching pairs can be accurately screened out, and optimization of the original double-view matching pair set is achieved. Through simulation comparison, compared with an existing method, the method has the advantage that the optimization effect is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision processing technology, and more specifically to: 1. A dual-view matching pair optimization method based on a context-aware network; 2. A dual-view matching pair optimization device based on a context-aware network. Background Art

[0002] In computer vision, dual-view image matching technology aims to find reliable feature correspondences between two images of the same point taken from different perspectives. This technology plays a key role in multiple tasks such as image stitching, map construction, motion structure recovery, and visual positioning. The traditional process usually starts with local key point detection and feature extraction in each image, using methods such as SIFT and SuperPoint to extract the attributes of the key points (mainly coordinate information) and descriptors, and then establishes the initial correspondence through cross-image nearest neighbor search in the feature space, that is, to form the original dual-view matching pair set.

[0003] However, due to the limitations of local feature descriptors in terms of view invariance, the original two-view matching pair set often contains a large number of false matches, which will affect the image matching effect.

[0004] Therefore, it is very important to optimize the original dual-view matching pair set with a large number of false matches to identify the correct matches. The existing optimization method introduces neural network for processing, which has achieved certain results, but still fails to meet the requirements and has room for improvement. Summary of the invention

[0005] Based on this, it is necessary to provide a dual-view matching pair optimization method and device based on a context-aware network to address the problem that the existing method of optimizing the original dual-view matching pair set using a neural network fails to meet the requirements.

[0006] The present invention is implemented by the following technical solutions:

[0007] In a first aspect, the present invention discloses a dual-view matching pair optimization method based on a context-aware network, comprising:

[0008] Step 1: Obtain the original dual-view matching pair set COR; COR includes: M pairs of point matching pairs;

[0009] Step 2: Input COR into the trained context-aware network for processing to obtain the Nth probability set P N ;

[0010] The context-aware network includes: 1 mapping layer Emdedding, 1 local feature reconstruction layer LFR, N global feature update layers GFU 1 ~GFUN ;

[0011] Emdedding is used to map COR to a high-dimensional space to obtain a high-dimensional feature set F COR ;

[0012] LFR is used for F COR Perform preliminary enhancement and screening to obtain local features F l ;

[0013] GFU n For GFU-based n-1 The resulting output feature F n-1 To update the global features, we get the nth probability set P n ; n∈[2,N]; where F l As GFU 1 Input; P n Including: M probability values ​​corresponding to M pairs of point matching pairs;

[0014] Step 3: Based on P N Optimizing COR to remove erroneous point matching pairs in COR;

[0015] Among them, if P N If the probability of the mth point matching pair in is greater than 0, the mth point matching pair is the correct point matching pair and is retained; otherwise, the mth point matching pair is removed from COR; m∈[1,M].

[0016] The dual-view matching pair optimization method based on a context-aware network implements a method or process according to an embodiment of the present disclosure.

[0017] In a second aspect, the present invention discloses a dual-view matching pair optimization device based on a context-aware network, which uses the dual-view matching pair optimization method based on a context-aware network disclosed in the first aspect.

[0018] The dual-view matching pair optimization device based on context-aware network includes: a data acquisition module, a network processing module, and a matching pair optimization module.

[0019] The data acquisition module is used to obtain the original dual-view matching pair COR; the network processing module is used to input the COR into the trained context-aware network for processing to obtain the Nth probability set P N ; The matching pair optimization module is used to N The COR is optimized to remove erroneous point matching pairs in the COR.

[0020] The dual-view matching pair optimization device based on a context-aware network implements the method or process according to the embodiment of the present disclosure.

[0021] Compared with the prior art, the present invention has the following beneficial effects:

[0022] 1. The present invention designs a context-aware network based on the combination of global and local, and uses it to optimize the original dual-view matching pair set; wherein, the present invention first uses a mapping layer to map the original dual-view matching pair set to a high-dimensional space to obtain a high-dimensional feature set, and then uses a local feature reconstruction layer to preliminarily enhance and screen the high-dimensional feature set to obtain local features, and then passes through N global feature update layers to perform multiple rounds of global feature processing to output the final probability set; the final probability set has a high accuracy rate, and can accurately screen out the correct point matching pairs, thereby optimizing the original dual-view matching pair set. After simulation comparison, the present invention significantly improves the optimization effect compared with the existing method.

[0023] 2. The local feature reconstruction layer in the context-aware network of the present invention adopts a local attention mechanism to maintain the consistency of local features, and reconstructs the central features through neighboring features, thereby achieving preliminary enhancement of matching features and laying a more stable correspondence foundation for the subsequent matching process.

[0024] 3. The global feature update layer in the context-aware network of the present invention utilizes the core point features selected according to the confidence level to finely reconstruct the correct point features through the global cross-attention mechanism, effectively distinguishing the correct point matches from the outliers, and significantly improving the accuracy and robustness of the matching. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0026] Figure 1 A structural diagram of a context-aware network provided in Example 1 of the present invention;

[0027] Figure 2 for Figure 1 The structure diagram of the local feature reconstruction layer LFR;

[0028] Figure 3 for Figure 1 The nth global feature update layer GFU n Structure diagram of

[0029] Figure 4 for Figure 3 The structural diagram of the cross-attention mechanism Aggregation1;

[0030] Figure 5 for Figure 3 Structural diagram of the cross-attention mechanism Aggregation2. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0032] It should be noted that when a component is referred to as being "mounted on" another component, it may be directly on the other component or there may be a central component. When a component is considered to be "set on" another component, it may be directly set on the other component or there may be a central component at the same time. When a component is considered to be "fixed to" another component, it may be directly fixed on the other component or there may be a central component at the same time.

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "or / and" used herein includes any and all combinations of one or more of the related listed items.

[0034] Example 1

[0035] See also Figure 1 , Figure 1 The data flow diagram of the dual-view matching pair optimization method based on the context-aware network is shown, which actually also shows the process of the dual-view matching pair optimization method based on the context-aware network and the overall structure of the context-aware network.

[0036] like Figure 1 As shown, the dual-view matching pair optimization method based on context-aware network includes the following steps:

[0037] Step 1: Get the original dual-view matching pair set COR.

[0038] COR includes: M pairs of point matching pairs. That is, the specification of COR is M×4, where 4 represents the coordinates of the point matching pairs.

[0039] COR can be obtained through the traditional process described in the background art. As described in the background art, COR contains a large number of wrong matching pairs, and the subsequent steps are to remove the wrong matching pairs and improve the matching pair accuracy of COR.

[0040] Step 2: Input COR into the trained context-aware network for processing to obtain the Nth probability set P N .

[0041] The context-aware network used in step 2 is the inventive point of the present invention, and is constructed based on the idea of ​​combining global and local.

[0042] For the context-aware network, it includes: 1 mapping layer Emdedding, 1 local feature reconstruction layer LFR, N global feature update layers GFU 1 ~GFU N .

[0043] like Figure 1 As shown:

[0044] Emdedding is used to map COR to a high-dimensional space to obtain a high-dimensional feature set F COR ;

[0045] LFR is used for F COR Perform preliminary enhancement and screening to obtain local features F l ;

[0046] GFU n For GFU-based n-1 The resulting output feature F n-1 To update the global features, we get the nth probability set P n ; n∈[2,N]; where F l As GFU 1 Input.

[0047] The following is a detailed introduction to each part:

[0048] 1. Emdedding aims to map COR to a high-dimensional space.

[0049] In this embodiment 1, the Emdedding can be designed to include: a multi-layer perceptron MLP0. MLP0 is used to map COR to obtain F COR .

[0050] For ease of understanding, the emdedding process can be formulated as follows:

[0051] F COR =MLP 0 (COR);

[0052] In the formula, MLP 0 (.) indicates the processing process of MLP0.

[0053] Therefore, F COR The specifications are M×C, where C represents the dimension.

[0054] Of course, emdedding can also adopt other designs, but it should meet the above functional requirements.

[0055] 2. The core task of LFR is to identify, build and utilize the proximity relationship between features.

[0056] In this embodiment 1, see Figure 2 , LFR can be designed to include: 1 local graph creation layer KNN, 2 multi-layer perceptrons MLP1~MLP2, 1 transposer Trans1, 2 product layers, 1 normalization layer Scale1, 2 activation function layers softmax1~softmax2, 2 convolution layers Conv1~Conv2, 1 maximum average pooling layer maxpool, 1 linear transformation layer linear, 1 weighted processing layer SUM, and 1 superposition layer.

[0057] In LFR:

[0058] 201. KNN is used to calculate the K-nearest neighbor algorithm based on F COR Construct a local graph G COR ; G COR The specification is M×C×K; K represents the number of nearest neighbor features;

[0059] 202, MLP1 is used for G COR Perform feature enhancement to obtain the enhanced local graph G' COR ;

[0060] 203, Trans1 is used to convert G' COR Perform transposition processing;

[0061] 204. The first product layer is used to combine the output of Trans with G' COR Perform product processing;

[0062] 205. Scale1 is used to normalize the output of the first product layer;

[0063] 206. softmax1 is used to process the output of the Scale layer through the softmax activation function to obtain a similarity score;

[0064] 207. The second product layer is used to combine the output of softmax1 with G' CORPerform product processing to obtain the enhanced local image The specifications are M×C×K;

[0065] Regarding the processing of 201 to 207, most correct point matches are more likely to maintain a closer distance in the feature space than outliers, while incorrect matches do not have this advantage due to feature differences. Therefore, it is necessary to construct and utilize the proximity relationship between features and use self-attention to ensure the consistency of local neighboring features.

[0066] 208. Conv1 is used for Perform convolution processing;

[0067] 209. Conv2 is used to perform convolution on the output of Conv1 to obtain the branch feature f 1 ;f 1 The specifications are M×C×1.

[0068] It should be noted that Conv1 can use a 1×K / g convolution kernel and Conv2 can use a 1×g convolution kernel to achieve K-dimensional aggregation into 1; g represents the number of groups, which means that the K neighboring features are divided into g groups, and the neighboring features in each group are first aggregated by Conv1, and then the neighboring features of all groups are aggregated by Conv2.

[0069] 210, MLP2 is used for Perform feature enhancement;

[0070] 211. maxpool is used to perform maximum average pooling on the output of MLP2 to obtain the branch feature f 2 ;f 2 The specifications are M×C×1;

[0071] 212. linear is used for Perform linear transformation;

[0072] 213. Softmax2 is used to process the output of linear through the softmax activation function to obtain the weight score K weight ;

[0073] 214. SUM is used for K weight and Perform weighted summation to obtain the branch feature f 3 ;f 3 The specifications are M×C×1;

[0074] 215. The overlay is used to 1 、f 2 、f 3 Superposition is performed to obtain Fl ; F l The specifications are M×C×1.

[0075] in other words:

[0076] Conv1 and Conv2 can be regarded as processing branch 1, which aggregates adjacent features through convolution operations;

[0077] MLP2 and maxpool can be regarded as processing branch 2, which extracts the neighboring features that are most similar to the central feature, retains the significant features, and suppresses noise;

[0078] Linear, softmax2, and SUM can be regarded as processing branch three, which calculates the contribution of adjacent features to the central feature and reduces the impact of irrelevant features.

[0079] For ease of understanding, the above process from 201 to 215 can be formulated as follows:

[0080] G COR =KNN(F COR );

[0081] G' COR =MLP 1 (G COR );

[0082]

[0083]

[0084]

[0085]

[0086]

[0087] F l =f 1 +f 2 +f 3 ;

[0088] In the formula, KNN(.) represents the processing of KNN; MLP 1 (.) indicates the processing of MLP1; d 1 Indicates the processing parameters of Scale1 (that is, the dimension of the feature to be processed); softmax 1 (.) indicates the processing of softmax1; conv1(.) indicates the processing of Conv1; conv2(.) indicates the processing of Conv2; MLP 2(.) indicates the processing of MLP2; maxpool(.) indicates the processing of maxpool; softmax 2 (.) represents the processing process of softmax2; linear(.) represents the processing process of linear; SUM(.) represents the processing process of SUM.

[0089] In other words, LFR first uses the K-nearest neighbor algorithm to determine the set of neighboring features for each feature point, and then borrows the concept of the self-attention mechanism to locally enhance the central feature and its neighboring features, thereby strengthening the connection between the central feature and the neighboring features; then, in view of the similarity of the neighboring features to the central feature in high-dimensional space, three different methods are used to aggregate the neighboring features to the central feature, and the aggregation results of these three methods are combined to form the central feature to reconstruct the local feature (that is, the core point feature).

[0090] Through the above processing process, LFR ensures that each feature point can be effectively enhanced by its neighboring features, which not only improves the representation ability of the feature, but also enhances the robustness of the feature to noise. Through this local feature reconstruction strategy, the quality and reliability of local features are significantly improved, which is also conducive to subsequent processing.

[0091] 3. GFU 1 ~GFU N The same strategy is adopted: a certain proportion of high-confidence matching points (identified as relatively reliable matching points) are screened out and used to reconstruct the correct matching pairs. This strategy not only reduces the load in the self-attention calculation, but also reduces the influence of outliers, thereby improving the accuracy of the matching process.

[0092] GFU 1 ~GFU N In this embodiment 1, N is set to 6 in consideration of both network performance and computing load.

[0093] Below is GFU n Take this as an example to illustrate:

[0094] In this embodiment 1, see Figure 3 , GFU n It can be designed to include: 2 multi-layer perceptrons MLP3~MLP4, 1 activation function layer Sigmoid, 1 selector Topk, 1 filter Filter, 2 cross attention mechanisms Aggregation1~Aggregation2, G self-attention mechanisms SA 1 ~SA G , 1 activation function layer Relu.

[0095] at GFUn middle:

[0096] 301, MLP3 is used for F n-1 Perform channel adjustment; where F 0 F l ;

[0097] 302. Sigmoid is used to process the output of MLP3 through the Sigmoid activation function to obtain the confidence P score,n-1 ;

[0098] 303. Topk is used to score,n-1 Select M' higher confidence values ​​and construct the same n-1 Corresponding index n-1 ; M'=M*r, r represents the selection ratio;

[0099] 304. Filter is used based on Index n-1 From F n-1 Filter out the corresponding matching features in 1,n ;in 1,n The specifications are M'×C×1;

[0100] It should be noted that considering that the number of matching pairs and points between different image pairs may be different, using only Topk and Filter processing may result in the loss of a large amount of original matching context information. Therefore, the cross attention mechanism is used to n-1 The original context information is aggregated into 1,n middle.

[0101] 305. Aggregation1 is used for in 1,n 、F n-1 The cross attention mechanism is used to process and obtain the matching features in 2,n ;in 2,n The specifications are M'×C×1;

[0102] For Aggregation1, see Figure 4 , which includes: 3 matrix converters, 1 transposer Trans2, 2 product layers, 1 normalization layer Scale2, 1 activation function layer softmax3, 1 merging layer concat1, 1 multi-layer perceptron MLP5, and 1 superposition layer.

[0103] In Aggregation1:

[0104] 3500, the first matrix converter is used to convert in 1,n Converted into query matrix Q 1,n ;

[0105] 3501, the second matrix converter is used to convert F n-1 Convert to key matrix K 1,n ;

[0106] 3502, the third matrix converter is used to convert F n-1 Convert to value matrix V 1,n ;

[0107] 3503, Trans2 is used to convert Q 1,n Perform transposition processing;

[0108] 3504, the first product layer is used to transform the output of Trans2, K 1,n Perform product processing;

[0109] 3505, Scale2 is used to normalize the output of the first product layer;

[0110] 3506, softmax3 is used to process the output of Scale2 through the softmax activation function;

[0111] 3507, the second product layer is used to transform the output of softmax3, V 1,n Perform product processing;

[0112] 3508, concat1 is used to convert in 1,n , the output of the second product layer is merged in the channel dimension;

[0113] 3509, MLP5 is used to adjust the channel of the output of concat1;

[0114] 3510, the overlay is used to 1,n , the output of MLP5 is superimposed to obtain in 2,n .

[0115] It should be noted that in 2,n While retaining the original context information, it also contains redundant information of outliers, and the impact of erroneous information needs to be further reduced.

[0116] 306、SA g For SA g-1 The output of is processed using the self-attention mechanism; g∈[2,G];

[0117] Among them, 2,n As SA 1 Input; SA G The output is the matching feature in 3,n ;in 3,nThe specifications are M'×C×1.

[0118] SA 1 ~SA G Both use a Transformer-based self-attention mechanism, which can further eliminate the influence of erroneous information. In this embodiment 1, G is set to 3 in consideration of both network performance and computing load.

[0119] 307. Aggregation2 is used for F n-1 、in 3,n The cross attention mechanism is used to process and obtain the matching features in 4,n ;in 4,n The specifications are M×C×1.

[0120] For Aggregation2, see Figure 5 , which includes: 3 matrix converters, 1 transposer Trans3, 2 product layers, 1 normalization layer Scale3, 1 activation function layer softmax4, 1 merging layer concat2, 1 multi-layer perceptron MLP6, and 1 superposition layer.

[0121] In Aggregation2:

[0122] 3700, the first matrix converter is used to convert F n-1 Converted into query matrix Q 3,n ;

[0123] 3701, the second matrix converter is used to convert in 3,n Convert to key matrix K 3,n ;

[0124] 3702, the third matrix converter is used to convert in 3,n Convert to value matrix V 3,n ;

[0125] 3703, Trans3 is used to convert Q 3,n Perform transposition processing;

[0126] 3704, the first product layer is used to transform the output of Trans3, K 3,n Perform product processing;

[0127] 3705, Scale3 is used to normalize the output of the first product layer;

[0128] 3706, softmax4 is used to process the output of Scale3 through the softmax activation function;

[0129] 3707, the second product layer is used to transform the output of softmax4, V 3,n Perform product processing;

[0130] 3708, concat2 is used to convert F n-1 , the output of the second product layer is merged in the channel dimension;

[0131] 3709, MLP5 is used to adjust the channels of the output of concat2;

[0132] 3710, the overlay is used to n-1 , the output of MLP5 is superimposed to obtain in 4,n .

[0133] It should be emphasized that the output specifications of Aggregation1 and Aggregation2 are different. 2,n The specifications are M'×C×1, and the latter outputs in 4,n The specifications are M×C×1.

[0134] 308, MLP4 is used for in 4,n Make channel adjustments.

[0135] MLP4 will be in 4,n The number of channels is adjusted to 1 to facilitate subsequent predictions.

[0136] 309. Relu is used to process the output of MLP4 through the Relu activation function to obtain P n ;P n The specifications are M×1.

[0137] For ease of understanding, the above processing steps 301 to 309 can be formulated as follows:

[0138] P score,n-1 =Sigmoid(MLP 3 (F n-1 ));

[0139] Index n-1 =Topk(P score,n-1 );

[0140] in 1,n =Filter(Index n-1 ,F n-1 );

[0141] in 2,n =Aggregation 1 (in 1,n ,F n-1 )

[0142] in 3,n =SA G (...(SA 1 (in 2,n )));

[0143] in 4,n =Aggregation 2 (F n-1 ,in 3,n );

[0144] P n =Relu(MLP 4 (in 4,n ));

[0145] In the formula, Sigmoid(.) represents the processing of Sigmoid; MLP 3 (.) indicates the processing of MLP3; Topk(.) indicates the processing of Topk; Filter(.) indicates the processing of Filter; Aggregation 1 (.) indicates the processing of Aggregation1; SA G (.) indicates SA G The processing of SA 1 (.) indicates SA 1 Aggregation 2 (.) indicates the processing process of Aggregation2.

[0146] Among them, the above process from 3500 to 3510 can be formulated as:

[0147] Q 1,n =in 1,n *W q ;

[0148] K 1,n =F n-1 *W k ;

[0149] V 1,n =F n-1 *W v ;

[0150]

[0151] Where W q represents the query transformation matrix; W k represents the key transformation matrix; W v represents the value conversion matrix; d 2Represents the processing parameter of Scale2 (i.e., the dimension of the feature to be processed); softmax 3 (.) represents the processing process of softmax3; MLP 5 (.) represents the processing process of MLP5.

[0152] Among them, the process from 3700 to 3710 above can be formulated as:

[0153] Q 3,n = F n-1 * W q ;

[0154] K 3,n = in 3,n * W k ;

[0155] V 3,n = in 3,n * W v ;

[0156]

[0157] In the formula, W q represents the query transformation matrix; W k represents the key transformation matrix; W v represents the value transformation matrix; d 3 represents the processing parameter of Scale3 (i.e., the dimension of the feature to be processed); softmax 4 (.) represents the processing process of softmax4; MLP 6 (.) represents the processing process of MLP6.

[0158] In summary, step two actually obtains N probability sets P 1 ~ P N ; among them, P n includes: M probability values corresponding to M point matching pairs. But P N is the most accurate, so only the last P N is used to optimize COR.

[0159] Step three, optimize COR based on P N to remove the incorrect point matching pairs in COR;

[0160] Among them, if the probability of the m-th point matching pair in P N is greater than 0, the m-th point matching pair is a correct point matching pair and is retained; otherwise, the m-th point matching pair is removed from COR; m ∈ [1, M].

[0161] In this way, the probabilities are compared to perform screening, thereby removing the wrong point matching pairs in COR.

[0162] In addition, considering the network training and subsequent matching requirements, the P n , COR is processed to obtain the nth essential matrix between the two views

[0163] It should be noted that this method requires the use of a trained context-aware network with optimal network parameters.

[0164] Generally, the context-aware network is trained based on an existing known matching pair data set to ensure the network training effect.

[0165] When training the context-aware network, a combined loss function Loss is used to simultaneously supervise the optimization of probability and matrix. Among them, Loss is expressed as:

[0166]

[0167] Where, L cls (.) represents the binary cross entropy loss function, which is used to supervise the probability set calculated by the network to approach the true probability; P true represents the probability true label used in the training phase; P training represents the probability set generated by the network during the training phase;

[0168] L reg (.) represents the geometric loss function, which is used to measure the difference between the essential matrix obtained by transforming the probability set calculated based on the network and the true essential matrix; Represents the true label of the essential matrix used in the training phase; represents the essential matrix generated by the network during the training phase; λ represents the weight coefficient.

[0169] That is to say, during the model training phase, the N probability sets and N essential matrices obtained by the network need to be compared with the corresponding true labels to calculate the loss value to update the network parameters.

[0170] Simulation comparison

[0171] In order to verify and illustrate the effectiveness and superiority of this method, this embodiment 1 introduces other existing networks for comparative experiments.

[0172] Existing models include: Point-Net++, DFE, ACNe, T-Net, CNe, OANet, MS 2 DG-Net, MSA-Net, PGFNet, UMatch, DeMatch.

[0173] The context-aware network provided in Example 1 (abbreviated as ours, N is 6, G is 3) is selected for comparison with the above-mentioned existing model.

[0174] The data sets used are the outdoor data set YFCC100M and the indoor data set SUN3D, and the data were divided according to the rules of OANet to ensure the fairness of the comparative experiment.

[0175] All comparative experiments were conducted on a device equipped with an NVIDIA RTX 4090 GPU to ensure sufficient computing resources and the accuracy of the experimental results.

[0176] When evaluating the performance of the corresponding predictions, three key indicators are used: precision (P), recall (R), and F-score. Precision measures the proportion of true positive examples among the samples predicted by the model as positive examples, that is, the accuracy of the model's predictions. Recall measures the proportion of all actual positive examples that are correctly predicted as positive by the model, reflecting the model's ability to find true positive examples. The F-score is the harmonic mean of precision and recall. By taking these two indicators into consideration, it can provide a more comprehensive evaluation of model performance.

[0177] For the performance evaluation of cross-view pose estimation, the mean average precision mAP5° is used. The calculation method of mAP5° involves considering the angular difference between the predicted vector and the true value vector on the rotation and translation vectors. Only when the angular difference is less than or equal to 5°, these data will be included in the calculation of mAP5°. The setting of this threshold is based on the consideration of the acceptable error range in actual application scenarios to ensure the accuracy of the evaluation. mAP5° can evaluate the overall performance of the model in the cross-view pose estimation task, especially the prediction accuracy when the angle difference is small.

[0178] See Table 1, Table 2 and Table 3 for comparative experimental results.

[0179] Table 1 Comparative experimental results

[0180]

[0181] Table 2 Comparative experimental results II

[0182]

[0183]

[0184] Table 3 Comparative experimental results

[0185]

[0186] As can be seen, Tables 1 and 2 present the accuracy evaluation of the network prediction results. The results show that Ours has achieved the optimal level in all three indicators in four environments (known indoor environment, unknown indoor environment, known outdoor environment, and unknown outdoor environment).

[0187] Table 3 presents the evaluation of the camera pose estimation performance. The results show that Ours has the best mAP5° in four environments (known indoor environment, unknown indoor environment, known outdoor environment, and unknown outdoor environment).

[0188] Based on the above simulation results, the effectiveness and superiority of the method in Example 1 are proved.

[0189] Example 3

[0190] This embodiment 3 discloses a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps of the dual-view matching pair optimization method based on a context-aware network disclosed in embodiment 1 are implemented.

[0191] This embodiment 3 also discloses a readable storage medium, in which computer program instructions are stored. When the computer program instructions are read and executed by a processor, the steps of the dual-view matching pair optimization method based on context-aware network disclosed in embodiment 1 are executed.

[0192] This embodiment 3 also discloses a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the dual-view matching pair optimization method based on a context-aware network disclosed in embodiment 1 are implemented.

[0193] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A dual-view matching pair optimization method based on context-aware network, characterized in that: include: Step 1: Obtain the original dual-view matching pair set COR; COR includes: M pairs of point matching; Step 2: Input COR into the trained context-aware network for processing to obtain the Nth probability set P N ; The context-aware network includes: 1 mapping layer Emdedding, 1 local feature reconstruction layer LFR, N global feature update layers GFU1~GFU N ; Emdedding is used to map COR to a high-dimensional space to obtain a high-dimensional feature set F COR ; LFR is used for F COR Perform preliminary enhancement and screening to obtain local features F l ; GFU n For GFU-based n-1 The resulting output feature F n-1 To update the global features, we get the nth probability set P n ; n∈[2,N]; where F l As the input of GFU1; P n Including: M probability values ​​corresponding to M pairs of point matching pairs; Step 3: Based on P N Optimizing COR to remove erroneous point matching pairs in COR; Among them, if P N If the probability of the mth point matching pair in is greater than 0, the mth point matching pair is the correct point matching pair and is retained; otherwise, the mth point matching pair is removed from COR; m∈[1,M].

2. The dual-view matching pair optimization method based on context-aware network according to claim 1, characterized in that: Emdedding includes: 1 multi-layer perceptron MLP0; The specification of COR is M×4, where 4 represents the coordinates of the point matching pair; F COR The specifications are M×C, where C represents the dimension.

3. The dual-view matching pair optimization method based on context-aware network according to claim 2, characterized in that: LFR includes: 1 local graph creation layer KNN, 2 multi-layer perceptrons MLP1~MLP2, 1 transposer Trans1, 2 product layers, 1 normalization layer Scale1, 2 activation function layers softmax1~softmax2, 2 convolution layers Conv1~Conv2, 1 maximum average pooling layer maxpool, 1 linear transformation layer linear, 1 weighted processing layer SUM, and 1 superposition layer; KNN is used to calculate the K-nearest neighbor algorithm based on F COR Construct a local graph G COR ; G COR The specification is M×C×K; K represents the number of nearest neighbor features; MLP1 is used to COR Perform feature enhancement to obtain the enhanced local graph G' COR ; Trans1 is used to convert G' COR Perform transposition processing; The first product layer is used to combine the output of Trans with G' COR Perform product processing; Scale1 is used to normalize the output of the first product layer; Softmax1 is used to process the output of the Scale layer through the softmax activation function to obtain the similarity score; The second product layer is used to combine the output of softmax1 with G' COR Perform product processing to obtain the enhanced local graph G COR ; G COR The specifications are M×C×K; Conv1 is used to COR Perform convolution processing; Conv2 is used to perform convolution processing on the output of Conv1 to obtain branch feature f1; the specification of f1 is M×C×1; MLP2 is used to COR Perform feature enhancement; maxpool is used to perform maximum average pooling on the output of MLP2 to obtain branch feature f2; the specification of f2 is M×C×1; Linear is used for G COR Perform linear transformation; Softmax2 is used to process the output of linear through the softmax activation function to obtain the weight score K weight ; SUM is used to calculate K weight and G COR Perform weighted summation to obtain branch feature f3; the specification of f3 is M×C×1; The superposition layer is used to superimpose f1, f2, and f3 to obtain F l ; F l The specifications are M×C×1.

4. The dual-view matching pair optimization method based on context-aware network according to claim 1, characterized in that: GFU1~GFU N The structures of are the same, including: 2 multi-layer perceptrons MLP3~MLP4, 1 activation function layer Sigmoid, 1 selector Topk, 1 filter Filter, 2 cross attention mechanisms Aggregation1~Aggregation2, G self-attention mechanisms SA1~SA G , 1 activation function layer Relu; at GFU n middle: MLP3 is used for F n-1 Perform channel adjustment; where F0 is F l ; Sigmoid is used to process the output of MLP3 through the Sigmoid activation function to obtain the confidence P score,n-1 ; Topk is used to score,n-1 Select M' higher confidence values ​​and construct the same n-1 Corresponding index n-1 ; M'=M*r, r represents the selection ratio; Filter is used based on Index n-1 From F n-1 Filter out the corresponding matching features in 1,n ;in 1,n The specifications are M'×C×1; Aggregation1 is used for in 1,n 、F n-1 The cross attention mechanism is used to process and obtain the matching features in 2,n ;in 2,n The specifications are M'×C×1; S g For SA g-1 The output of is processed by the self-attention mechanism; g∈[2,G]; where in 2,n As the input of SA1; SA G The output is the matching feature in 3,n ;in 3,n The specifications are M'×C×1; Aggregation2 is used to n-1 、in 3,n The cross attention mechanism is used to process and obtain the matching features in 4,n ;in 4,n The specifications are M×C×1; MLP4 is used for in 4,n Make channel adjustments; Relu is used to process the output of MLP4 through the Relu activation function to obtain P n ;P n The specifications are M×1.

5. The dual-view matching pair optimization method based on context-aware network according to claim 4, characterized in that: N=6; G=3.

6. The dual-view matching pair optimization method based on context-aware network according to claim 4 or 5, characterized in that: Aggregation1 includes: 3 matrix converters, 1 transposer Trans2, 2 product layers, 1 normalization layer Scale2, 1 activation function layer softmax3, 1 merging layer concat1, 1 multi-layer perceptron MLP5, and 1 superposition layer; at GFU n Aggregation1: The first matrix converter is used to convert the in 1,n Converted into query matrix Q 1,n ; The second matrix converter is used to convert F n-1 Convert to key matrix K 1,n ; The third matrix converter is used to convert F n-1 Convert to value matrix V 1,n ; Trans2 is used to convert Q 1,n Perform transposition processing; The first product layer is used to transform the output of Trans2, K 1,n Perform product processing; Scale2 is used to normalize the output of the first product layer; Softmax3 is used to process the output of Scale2 through the softmax activation function; The second product layer is used to transform the output of softmax3, V 1,n Perform product processing; concat1 is used to convert in 1,n , the output of the second product layer is merged in the channel dimension; MLP5 is used to adjust the channels of the output of concat1; Overlays are used to place in 1,n , the output of MLP5 is superimposed to obtain in 2,n .

7. The dual-view matching pair optimization method based on context-aware network according to claim 4 or 5, characterized in that: Aggregation2 includes: 3 matrix converters, 1 transposer Trans3, 2 product layers, 1 normalization layer Scale3, 1 activation function layer softmax4, 1 merging layer concat2, 1 multi-layer perceptron MLP6, and 1 superposition layer; at GFU n Aggregation2: The first matrix converter is used to convert F n-1 Converted into query matrix Q 3,n ; The second matrix converter is used to convert the in 3,n Convert to key matrix K 3,n ; The third matrix converter is used to convert the in 3,n Convert to value matrix V 3,n ; Trans3 is used to convert Q 3,n Perform transposition processing; The first product layer is used to transform the output of Trans3, K 3,n Perform product processing; Scale3 is used to normalize the output of the first product layer; Softmax4 is used to process the output of Scale3 through the softmax activation function; The second product layer is used to transform the output of softmax4, V 3,n Perform product processing; concat2 is used to convert F n-1 , the output of the second product layer is merged in the channel dimension; MLP6 is used to adjust the channels of the output of concat2; The overlay is used to convert F n-1 , the output of MLP6 is superimposed to obtain in 4n .

8. The dual-view matching pair optimization method based on context-aware network according to claim 1, characterized in that: Step three also includes: The eight-point weighted algorithm is used to n , COR is processed to obtain the nth essential matrix between the two views 9. The dual-view matching pair optimization method based on context-aware network according to claim 8, characterized in that: The loss function Loss used by the context-aware network during the training phase is: Where, L cls (.) represents the binary cross entropy loss function; P true represents the probabilistic true label used in the training phase; P training represents the probability set generated by the network during the training phase; L reg (.) represents the geometric loss function; Represents the true label of the essential matrix used in the training phase; represents the essential matrix generated by the network during the training phase; λ represents the weight coefficient.

10. A dual-view matching pair optimization device based on a context-aware network, characterized in that: It uses the dual-view matching pair optimization method based on a context-aware network as described in any one of claims 1 to 9; The dual-view matching pair optimization device based on context-aware network includes: A data acquisition module, which is used to acquire an original dual-view matching pair COR; The network processing module is used to input COR into the trained context-aware network for processing to obtain the Nth probability set P N ; as well as Matching pair optimization module, which is used to N The COR is optimized to remove erroneous point matching pairs in the COR.