A method for adaptive dense matching calculation between two frames of images
By constructing a confidence-driven adaptive matching network model, iteratively evaluates confidence and match distrust points, the ambiguity problem in dense matching is solved, and efficient and accurate intensive matching calculations are realized at low time and parameter costs, which is especially suitable for geometric multi-view image processing.
Patent Information
- Application Number
- CN202210427447.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-04-21
AI Technical Summary
There is a problem of ambiguity in dense matching, which makes the flow prediction results difficult to be accurate, and the existing methods lack performance when dealing with geometric multi-view images.
An adaptive intensive matching calculation method is proposed. By constructing a confidence-driven adaptive matching network model, iteratively evaluates the confidence and matching distrust points, and realizes adaptive intensive matching calculation. Specific steps include using a twin network to extract image features, calculate cost and initial optical flow, iteratively predict confidence and optical flow, selecting unconfidence points for feature matching, and updating the optical flow.
Achieve excellent intensive matching performance at low time and parameter costs, especially in geometric multi-view image processing, improving matching accuracy and stability.
Smart Images

Figure CN114743069B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of video and scene understanding, and in particular relates to a method for performing dense matching calculation on two frames of images. Background Art
[0002] Matching calculation methods can be divided into three categories: sparse matching, optical flow estimation, and dense matching. Sparse matching generally includes three stages: key point detection, feature description, and feature matching. Detection methods based on artificial features are widely used, such as SIFT[1], SURF[2], ORB[3], etc. Recently, researchers have tried to use neural networks to detect and describe key points, and have also achieved good results. Feature matching methods have been proposed as a separate stage to obtain corresponding points by matching descriptors. Methods include neighborhood consensus and RANSAC[4]. There are also some recent works that are learned through neural networks. LoFTR[5] proposes to use Transformer to learn end-to-end feature detection and matching. Although the point matching steps of the present invention are similar to feature matching, the method of the present invention can predict dense flow and select matching points based on confidence prediction rather than key points.
[0003] Optical flow estimation is the problem of finding pixel-level correspondences between consecutive frames in a video. Early optical flow estimation used handcrafted features to match keypoints detected in video sequences, such as the Lukas-Kanade method [6] and many subsequent works. The present invention notes that some early methods also focused on combining sparse matching and optical flow. However, these methods used handcrafted features and their performance was not comparable to that of recent methods. FlowNet [7] was the first method to use convolutional neural networks for optical flow estimation, achieving good results and inspiring a large number of subsequent studies. PWC-Net [8] was an influential work that used a multi-layer feature pyramid to predict optical flow at multiple levels, and there were also many subsequent studies. Recently, RAFT [9] proposed an iterative update strategy that does not require multi-level computation, but instead iteratively updates neighboring points. These methods are specialized for optical flow data and are limited to small displacements. The present method aims to solve general matching problems, especially in geometric multi-view images.
[0004] The goal of dense matching estimation is to find the pixel-to-pixel correspondence between a pair of images, which includes the above-mentioned optical flow, geometric matching, disparity and many other tasks. This is an important and common problem in video and scene understanding. There are many downstream tasks in video interpretation and processing based on dense matching between frames, such as video interpolation, video motion capture, expression recognition, etc. In addition, VR / VR tasks such as 3D reconstruction, perspective synthesis, and 3D positioning also require dense matching between multi-view inputs. GLU-Net
[10] and a series of subsequent articles are one of the representatives in this field. Through extensive experiments, the matching ability of feature pyramid-based matching methods on various tasks has been demonstrated. Recently, some articles have been devoted to combining sparse and dense matching. For example, RANSAC-Flow
[11] uses the sparse method RANSAC to predict flow, and COTR
[12] uses implicit representation to unify dense and sparse matching.
[0005] Ambiguity is a common problem inherent in dense matching. This ambiguity comes from textureless areas, motion boundaries, occlusions, and so on. It is difficult for the flow prediction result model to find an accurate match in one iteration, and the ambiguity will accumulate with iterations, thus affecting the final performance. Therefore, many methods begin to estimate confidence or uncertainty, but they only consider unconfident points as outlines and ignore the subsequent matching information. The method of the present invention mines matching information from unconfident points and realizes adaptive matching by iteratively evaluating confidence and matching unconfident points. Summary of the invention
[0006] The purpose of the present invention is to propose a method for adaptive dense matching calculation of two frames of images, which can obtain advanced and excellent performance with lower time and parameter costs.
[0007] The method for performing adaptive dense matching calculation on two frames of images provided by the present invention includes constructing a confidence-driven adaptive matching network model, and implementing adaptive dense matching calculation by iteratively evaluating confidence and matching unconfidence points; the specific steps are:
[0008] (1) Use a deep learning network to extract features of the input image;
[0009] Specifically adopt the twin network: matching feature network N m and content feature network N c To extract two frames of input image I 1 and I 2 The matching feature M 1 and M 2 and content feature C 1 and C 2; Then use the matching features to calculate the cost volume V, which represents the similarity between the previous and next frames; then use the initial optical flow F to extract the motion feature T from the cost volume V; then send it together with the content feature C into a prediction network N p , predict matching confidence CF and optical flow F';
[0010] (2) Selecting untrusted points based on different confidence evaluations in each iteration of each image matching is an adaptive process for matching under various conditions, especially for geometric matching. To this end, three strategies for evaluating confidence are adopted, each of which is accompanied by different confidence representations, point selections, and loss functions, so as to adaptively propose untrusted regions (i.e., untrusted point sets) P. 1 and P 2 ;
[0011] (3) For the untrusted point set P 1 and P 2 , extract feature vectors and adaptively match these points; specifically, the most advanced feature matching strategy is used for matching, and then the mutual flow MF with high matching probability is used to update the optical flow to provide guidance for uncertain points;
[0012] (4) Iteratively execute steps (1)-(3), starting from the initial optical flow, iteratively predict the confidence CF and optical flow F', select untrusted points and match them, and then update the original optical flow, and finally obtain the iteratively optimized optical flow and confidence map;
[0013] (5) Training the network. At this time, the flow prediction results and confidence map results of each iteration are output to calculate the loss function; the back propagation algorithm is used to iteratively update the neural network to achieve a result that converges on a given data set and generalizes to other data sets;
[0014] (6) When testing with a trained network, only two frames of images need to be input, and after the neural network is inferred (i.e., multi-scale inference), the corresponding dense matching is output; in actual use, the multi-scale inference method is used to improve the accuracy of the predicted matching.
[0015] Further:
[0016] In step (1), the twin network is used to extract the matching features M and content features C of the two input images; then the matching features M are used to calculate the cost V. The specific process is as follows: Figure 2 As shown, the two input RGB images I 1 and I 2 , respectively, through the matching feature network N m and content feature network N c Perform feature extraction to obtain matching feature M 1 , M 2and content feature C 1 and C 2 ; The twin network means that the network processing the two frames of images is the same and shares the same parameters. The cost V is calculated based on the extracted matching features M, and feature matching is performed in the feature matching stage. The calculation of the cost V can be expressed as the following formula:
[0017] V ijkl =∑ h (M 1 ) ijh ·(M 2 ) klh , (1)
[0018] Among them, M 1 and M 2 is the extracted matching feature, whose size is H×W×C, which are height, width and number of channels respectively; (M 1 ) ijh Represents the matching feature M 1 As the value of the matrix at position (i, j, h); the operator · represents a multiplication operation, where each product of the third dimension h is summed. V has four dimensions and its size is H×W×H×W, V ijkl Represents the matching feature M 1 The (i, j) position and matching feature M 2 The similarity of the (k, l) position.
[0019] In step (1), the process of extracting motion features T from the cost volume V using the initial optical flow F can be expressed as follows:
[0020] x=(u,v), x′=(u+f 1 (u), v+f 2 (v)), (2)
[0021]
[0022] Among them, x represents the feature map M 1 A point in the, (u, v) is its coordinates; f 1 (u), f 2 (2) represents the flow value of the initial flow F at (u, v), x′ represents the coordinates of the point x after the initial flow transformation; the motion feature T extracts the cost volume V at That is, the characteristics of the neighborhood of x′ with radius r; dx is the coordinate offset, represents integer coordinates, ||dx|| 1 ≤r means the offset is less than the radius r. Then the motion feature T and content feature C are input into the prediction network N p , generating matching confidence CF and optical flow F'. Here for two frames of image I1 and I 2 Both calculate the forward optical flow (I 1 To I 2 )’s matching confidence CF 1 and optical flow F′ 1 , and backward optical flow (I 2 To I 1 )’s matching confidence CF 2 and optical flow F′ 2 .
[0023] In step (2), the three strategies for evaluating confidence are each accompanied by different confidence representations, point selections, and loss functions, so as to adaptively propose untrustworthy areas. This is a key step in achieving adaptation. The specific process is as follows:
[0024] (1) Introduce a pair of values (f i , c i ) is used to represent the flow prediction result and its confidence of a point, and the confidence evaluation problem is expressed in a unified way:
[0025]
[0026] Among them, D is the point set on the confidence map, there are n points in total, and each point i has a flow prediction result value f i and confidence c i , the c value of the confidence point is between 0 and 1.
[0027] (2) Probability-based confidence assessment: The goal is to classify the points on the feature map into two categories, confident and unconfident. In this way, the selection of confidence predictions and unconfident points becomes a binary classification problem; the classification standard for these two categories is a deviation threshold. The deviation is the error of the flow prediction result. The confidence is taken as the probability expectation that the deviation is less than the threshold:
[0028] c i =P(δ i ≤δ s ), (5)
[0029] Among them, δ i is the deviation value, δ s is the threshold and P is the probability. The threshold is set to the neighbor radius of the cost volume in the realization process. The prediction network is used as a binary classifier. Set the minimum confidence c s To distinguish untrustworthy points. Confidence c i Below c s The points are selected as untrusted points. Here c s The value of is 0.5.
[0030] (3) Value-based confidence assessment: The size of optical flow varies greatly in different scenes, especially when the view changes. Therefore, using a fixed threshold on a single pixel cannot adapt to each scene. The present invention proposes a strategy based on the confidence information of the entire image. The equation is extended to a regression problem to avoid the classification threshold problem:
[0031]
[0032] Among them, δ i is the deviation. The ratio of the maximum deviation of a prediction is used as a variable. Through such a scheme, the confidence can be evaluated and unconfident points can be selected in the case of large deviations. Here, the confidence threshold c is also used s To select untrusted points, c s The value of is 0.5.
[0033] (4) There are still certain risks in evaluating by probability and value. They rely on the selection of thresholds and proportions, which may lead to the selection of too many points, thereby reducing the matching efficiency. Therefore, sorting is further used to evaluate confidence, which can be expressed as the following formula.
[0034]
[0035] Among them, rank is the calculation of the i-th deviation δ in descending order i The confidence threshold c s The number of matching points is also controlled. Therefore, this method can control the amount of calculation, making the reasoning more stable and faster. However, it may also bring a disadvantage that the untrusted area is incomplete, making it impossible to find paired points. As a result, the matching confidence CF 1 CF 2 The untrusted region P of the two frames for the predicted optical flow is obtained 1 and P 2 , to facilitate subsequent matching optimization.
[0036] In step (3), the feature matching strategy is used for matching, and then the mutual flow MF with high matching probability is used to update the optical flow to provide guidance for uncertain points. This is a key step in achieving matching. The specific process is as follows:
[0037] (1) Figure 3 As shown, the selected points are adaptively matched to produce a more accurate flow. Assuming that for an image pair, semantically similar regions have the same confidence, then the unconfident point P in the image pair can be used 1 and P 2 After completing the selection of untrusted points, take out the matching feature M of each point extracted by the matching network 1 and M2 .
[0038] (2) The present invention also applies self-attention and cross-attention
[13] to the extracted features. In this way, the network can learn the relationship with other features before calculating the correlation. Self-attention and cross-attention are calculated separately and connected together for later matching. The attention used here is linear attention to save time.
[0039] (3) An efficient method is used to perform matching. First, the features are normalized and the correlation matrix is calculated.
[0040] R(i, j)= <M 1 (i), M 2 (j)>, (8)
[0041] Where R(i, j) is the image I 1 Point i in image I 1 The correlation between points j in M. 1 and M 2 are the matching features, and 〈·,·〉 is the inner product.
[0042] (4) Then, use the following conditions to filter the matching point pairs. Specifically, use softmax to perform matching filtering. Calculate the score of each pair of matching items. Image I 1 The untrusted point set P 1 Point i and image I 2 The untrusted point set P 2 The point j satisfies the following conditions: 2 In the point set P, the most likely matching point of i is j. 1 In the example, the most likely matching point of j is i; the final score P for each point pair i and j matching c (i, j) is calculated using the following formula:
[0043] P c (i, j) = softmax(R(i, ·)) j softmax(R(., j)) i , (9)
[0044] Where R(i, ) represents the i-th column of the correlation matrix, and softmax(R(i, )) j represents the value of the i-th column and j-th row after softmax calculation. The operator · represents the product. Softmax can be expressed as the following formula:
[0045]
[0046] Among them, z jis the jth number in vector z, e represents the natural exponent. Σ represents the sum of the natural exponents of all numbers in vector z.
[0047] (5) The features of untrusted points come from the matching features M 1 and M 2 . Set two thresholds for matching. First, the mutual probability P of matching pair i, j c (i, j) must be above a threshold to reduce matching error points. This criterion is called mutual nearest neighbor. Secondly, the correlation R(i, j) of matching pair i, j must also be above a threshold to reduce untrusted matches. The matching selection strategy can be expressed by the following formula.
[0048]
[0049] Where MF is the resulting match or mutual flow, θ p is the mutual probability threshold, θ r is the correlation threshold. Set and θ r =0.7. (i, j)∈(P1, P 2 ) indicates that i, j are from the point set P 1 , P 2 The point set pairs generated in .
[0050] (6) Update the original prediction F' of the optical flow using the mutual flow MF before the next iteration. In this way, incorrect matches can be corrected in advance, which is a guide for the dense flow estimation and is more accurate before subsequent iterations. The present invention only performs feature matching in the first α percentage of iterations, where α is taken as 0.5. After multiple iterations, the confidence increases and the dense flow estimation can handle the matching and refine the correspondence. This also saves the calculation of redundant matches.
[0051] In step (5), the loss function is divided into flow prediction result loss and confidence loss, which is the basis of adaptive matching network optimization. The specific calculation is as follows:
[0052] (1) The flow prediction result loss uses the first norm of the flow prediction result error and the weight that increases with the number of iterations:
[0053]
[0054] Among them, f i is the predicted flow value of the i-th iteration, γ is the weight coefficient, which can be 0.8. M is the total number of iterations. ||f i -f gt || 1 Represents the prediction flow f i and the real flow f gtThe one-dimensional norm of the difference between , that is, the absolute deviation ∈.
[0055] (2) Depending on the evaluation strategy, the confidence loss can be expressed in two different forms. For probability-based strategies, a binary classification problem needs to be solved, so the cross entropy loss is used. As the loss function. Apply the sigmoid curve to scale the predicted confidence to the interval (0, 1), and use whether the absolute deviation ∈ is greater than the threshold δ s as labels. The loss can be expressed as follows.
[0056]
[0057] Among them, P s (·|δ) means δ is greater than or less than the threshold δ s The probability of i,j and δ i,j is the absolute deviation and prediction deviation of point j in the i-th iteration, and N is the total number of points. In order to be consistent with the flow prediction results, the present invention also adds a weighting factor to adjust the learning scale at each iteration. It can avoid learning from too large a deviation, especially at the beginning of training.
[0058] For both value-based and ranking-based strategies, the first norm of δ with respect to ∈ is used to constrain the estimation; the loss function can be expressed as:
[0059]
[0060] Among them, ||δ i,j -∈ i,j || 1 Denotes the prediction deviation δ i,j and the true deviation ∈ i,j The one-dimensional norm of the difference.
[0061] (3) The final loss function L is the linear sum of these two losses, that is, L = L f +βL c , where β is a weight factor, usually taken as 1. The loss function L is used to perform backpropagation and update the parameters of the neural network.
[0062] In step (6), the reasoning through the neural network is the key process for actual use. The specific process is as follows: the flow estimation process is divided into two processes: the first process estimates a smaller flow at a lower resolution, which is then used as the initial for the second process, and the second process uses the original resolution for prediction; it can be described as follows:
[0063] φ l (x) = φ(x, up 2 (φ l-1 (down2 (x)))), (15)
[0064] Among them, φ l is the matching network with the lth layer, up 2 Indicates 2× bilinear upsampling, down 2 And it is 2× downsampling.
[0065] The purpose of step (6) is that it is particularly difficult to infer the correct optical flow in one iteration, especially for high-resolution images and videos and extreme viewpoint changes. The present invention uses multi-scale reasoning to improve high-resolution performance. Multi-scale reasoning can not only improve reasoning performance, but also save time. The present invention can reduce time costs by using low-resolution images in the first few iterations.
[0066] The characteristics and advantages of the present invention are:
[0067] (1) This paper proposes a novel confidence-driven adaptive matching architecture. It is the first to solve the matching ambiguity in dense matching by adaptively performing feature matching;
[0068] (2) In order to adapt to ambiguity, three confidence-driven region extraction strategies are proposed, each of which is accompanied by different confidence evaluation, untrustworthy point selection and loss function; untrustworthy regions can be adaptively proposed;
[0069] (3) This paper proposes a new adaptive matching network that combines dense matching with feature matching to provide guidance for the next iteration. The feature matching method improves the matching of paired points in uncertain areas;
[0070] (4) The present invention conducts extensive experiments on optical flow and geometric matching datasets. The final model can achieve state-of-the-art and excellent performance at a lower time and parameter cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 It is a flow chart of the present invention.
[0072] Figure 2 This is a structural diagram of the adaptive matching network used in the present invention.
[0073] Figure 3 This is a structural diagram of the feature matching module used in the present invention.
[0074] Figure 4 To calculate the matching results of real image pairs using this method.
[0075] Figure 5 Performance comparison results of other methods on the KITTI dataset. DETAILED DESCRIPTION
[0076] (1) The structure of the present invention is as follows Figure 1 As shown, the present invention simultaneously evaluates the confidence of the flow and adaptively selects untrusted points for feature matching. Due to large motion, occlusion, and textureless areas, the selected points can change with each prediction. The adaptive matching of the present invention can correct the match and provide guidance for the next iteration. Figure 2 As shown, the present invention uses a twin network to extract matching features and context features of the image. Then, the present invention iteratively performs adaptive matching to predict optical flow and confidence maps. The present invention selects untrusted points in the confidence map and performs feature matching to provide guidance for the next iteration. Finally, the present invention outputs optical flow and confidence maps.
[0077] (2) The present invention uses Pytorch to implement the model of the present invention. The present invention uses a residual convolutional network in feature extraction and a GRU in iterative matching. During training, the present invention uses an AdamW optimizer and a 1cycle learning rate strategy. The confidence threshold c of the point is selected by probability, value, and ranking s , are 0.5, 0.5 and 0.1 respectively. The loss factor β is 1, and the feature matching rate α = 0.5. The present invention adopts two layers of self and cross attention in the attention mechanism, with 4 attention heads.
[0078] (3) The present invention uses different training sets to optimize for different application scenarios. For optical flow estimation, the present invention trains the model on the FlyingChairs dataset with a learning rate of 4*10 -4 , and iterates 100,000 times with a resolution of 368*496. Then the network of the present invention is iteratively trained on the FlyingThings3D dataset for 100,000 times, with a learning rate of 10 -4 , the resolution is 400*720. DPED-CityScape-ADE and MegaDepth datasets are designed for geometric matching of natural images. The present invention trains 100,000 iterations on DPED-CityScape-ADE and MegaDepth respectively, with a batch size of 6, in order to perform better on scene datasets. The present invention trains image pairs on a resolution of 520*520, with a learning rate of 10 -4 .
[0079] (4) Figure 5The time cost, parameters and accuracy of different methods on KITTI are compared. The shapes represent different training sets, the triangles represent the FlyingChairs and FlyingThings3D datasets, and the circles represent the MegaDepth dataset. The size of the graph represents the average running time, different colors represent different models, and Ours represents the model of the present invention. The final model of the present invention can achieve state-of-the-art performance with lower time and parameter costs.
[0080] References
[0081] [1]DGLowe.1999.Object recognition from local scale-invariantfeatures.ICCV(1999),1150–1157vol.2.
[0082] [2]Herbert Bay, Tinne Tuytelaars, and Luc Van Gool. 2006.SURF: SpeededUpRobust Features. In Computer Vision–ECCV 2006, Leonardis, Horst Bischof, and Axel Pinz (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 404–417
[0083] [3]Ethan Rublee, Vincent Rabaud, Kurt Konolige, and Gary Bradski. 2011. ORB: An efficient alternative to SIFT or SURF. (2011), 2564–2571
[0084] [4]Martin A.Fischler and Robert C.Bolles.1981.Random SampleConsensus:AParadigm for Model Fitting with Applications to Image Analysis andAutomatedCartography.Commun.ACM 24,6(jun 1981),381–395.
[0085] [5]Jiaming Sun,Zehong Shen,Yuang Wang,Hujun Bao,and XiaoweiZhou.2021.LoFTR:Detector-Free Local Feature Matching with Transformers.2021IEEE / CVFConference on Computer Vision and Pattern Recognition(CVPR)(2021),8918–8927.
[0086] [6]Bruce D.Lucas and Takeo Kanade.1981.An Iterative ImageRegistration Technique with an Application to Stereo Vision.In Proceedings ofthe 7th InternationalJoint Conference on Artificial Intelligence,IJCAI’81,Vancouver,BC,Canada,August24-28,1981,Patrick J.Hayes(Ed.).William Kaufmann,674–679
[0087] [7]Alexey Dosovitskiy,Philipp Fischer,Eddy Ilg,Philip Hausser,CanerHazirbas,VladimirGolkov,Patrick Van Der Smagt,Daniel Cremers,and ThomasBrox.2015.Flownet:Learning optica lflow with convolutional networks.InProceedingsof the IEEE international conference on computer vision.2758–2766
[0088] [8]Deqing Sun,Xiaodong Yang,Ming-Yu Liu,and Jan Kautz.2018.PWC-Net:CNNsfor Optical Flow Using Pyramid,Warping,and Cost Volume.2018 IEEE / CVFConference on Computer Vision and Pattern Recognition(2018),8934–8943
[0089] [9]Zachary Teed and Jia Deng.2020.RAFT:Recurrent All-Pairs FieldTransformsfor Optical Flow.ArXiv abs / 2003.12039(2020).
[0090]
[10] Prune Truong,Martin Danelljan,and Radu Timofte.2020.GLU-Net:GlobalLocal Universal Network for Dense Flow and Correspondences.2020 IEEE / CVFConference on Computer Vision and Pattern Recognition(CVPR)(2020),6257–6267
[0091]
[11] Xi Shen, Darmon,Alexei A.Efros,and MathieuAubry.2020.RANSAC-Flow:Generic Two-Stage Image Alignment.In Computer Vision-ECCV 2020-16th European Conference,Glasgow,UK,August 23-28,2020,Proceedings,Part IV(Lecture Notes in Computer Science,Vol.12349),Andrea Vedaldi,HorstBischof,Thomas Brox,and Jan-Michael Frahm(Eds.).Springer,618–637
[0092]
[12] Wei Jiang,Eduard Trulls,Jan Hosang,Andrea Tagliasacchi,and KwangMoo Yi.2021.COTR:Correspondence Transformer for Matching Across Images.(2021),6187–6197
[0093]
[13] Ashish Vaswani,Noam Shazeer,Niki Parmar,Jakob Uszkoreit,LlionJones,Aidan N.Gomez,Lukasz Kaiser,and Illia Polosukhin.2017.Attention isAllyou Need.IEEE / CVFConference on Computer Vision and Pattern Recognition(2017),5998。
Claims
1. A method for performing adaptive dense matching calculation on two frames of images, characterized in that: Construct a confidence-driven adaptive matching network model to achieve adaptive dense matching calculation by iteratively evaluating confidence and matching unconfident points; the specific steps are: (1) Use a deep learning network to extract features of the input image; Specifically adopt the twin network: matching feature network N m and content feature network N c To extract the matching features M1 and M2 and content features C1 and C2 of the two input images I1 and I2; then use the matching features to calculate the cost volume V, which represents the similarity between the two frames before and after; then use the initial optical flow F to extract the motion feature T from the cost volume V; then send it together with the content feature C to a prediction network N p , predict matching confidence CF and optical flow F'; (2) Selecting untrusted points based on different confidence evaluations in each iteration of each image match is an adaptive process for matching under various conditions; To this end, three strategies for evaluating confidence are adopted, each of which is accompanied by different confidence representations, point selections, and loss functions, so as to adaptively propose unconfident point sets P1 and P2; (3) For the untrusted point sets P1 and P2, extract feature vectors and adaptively match these points; specifically, use a feature matching strategy for matching, and then use the mutual flow MF with high matching probability to update the optical flow to provide guidance for the uncertain points: (4) Iteratively execute steps (1)-(3), starting from the initial optical flow, iteratively predict the confidence CF and optical flow F', select untrusted points and match them, and then update the original optical flow, and finally obtain the iteratively optimized optical flow and confidence map; (5) Train the network and output the flow prediction results and confidence map results of each iteration to calculate the loss function; use the back propagation algorithm to iteratively update the neural network to achieve convergence results on a given data set and generalize to other data sets; (6) To test the trained network, only two frames of images need to be input, and after the neural network is inferred, the corresponding dense matches are output; In step (1): The twin network is used to extract the matching features M and content features C of the two input images; then the matching features M are used to calculate the cost V. The specific process is as follows: The two input RGB images I1 and I2 are matched through the feature matching network N m and content feature network N c Feature extraction is performed to obtain matching features M1, M2 and content features C1, C2; the twin network represents that the networks processing the two frames of images are the same and share the same parameters; the cost V is calculated based on the extracted matching feature M, and feature matching is performed in the feature matching stage; the calculation of the cost V is expressed as the following formula: V ijkl =∑ h (M1) ijh ·(M2) klh ,(1) Among them, M1 and M2 are the extracted matching features, and their size is H×W×C, which are height, width, and number of channels respectively; (M1) ijh represents the matching feature M1 as the value of the matrix at position (i, j, h); the operator · represents the multiplication operation, where each product of the third dimension h is summed; V has four dimensions, and its size is H×W×H×W, V ijkl Indicates the similarity between the (i, j) position of the matching feature M1 and the (k, l) position of the matching feature M2; The process of extracting motion features T from the cost volume V using the initial optical flow F is expressed as follows: x=(u, v), x′=(u+f1(u), v+f2(v)), (2) Among them, x represents a point in the feature map M1, (u, v) is its coordinates; f1(u), f2(v) represent the flow values of the initial flow F at (u, v), and x′ represents the coordinates of the point x after the initial flow transformation; the motion feature T extracts the cost volume V at That is, the characteristics of the neighborhood of x′ with radius r; dx is the coordinate offset, Represents integer coordinates, ||dx||1≤r means the offset is less than the radius r; Input motion features T and content features C into the prediction network N p , generate matching confidence CF and optical flow F', here for the two frames I1 and I2, the forward optical flow is calculated respectively: the matching confidence CF1 and optical flow F'1 from I1 to I2, and the backward optical flow is calculated: the matching confidence CF2 and optical flow F'2 from I2 to I1; In step (2), the three strategies for evaluating confidence are each accompanied by different confidence representations, point selections, and loss functions, so as to adaptively propose untrusted point sets P1 and P2. The specific process is as follows: (1) Introduce a pair of values (f i , c i ) is used to represent the flow prediction result and its confidence of a point, and the confidence evaluation problem is expressed in a unified way: Among them, D is the point set on the confidence map, there are n points in total, and each point i has a flow prediction result value f i and confidence c i , the c value of the confidence point is between 0 and 1; (2) Probability-based confidence assessment: The goal is to classify the points on the feature map into two categories, confident and unconfident. In this way, the confidence prediction and the selection of unconfident points become a binary classification problem. The classification standard for these two categories is a deviation threshold. The deviation is the error predicted by the optical flow network. The confidence is taken as the probability expectation that the deviation is less than the threshold: c i =P(δ i ≤δ s ), (5) Among them, δ i is the deviation value, δ s is the threshold, P is the probability; the threshold is set to the neighbor radius of the cost amount in the implementation process; the prediction network is used as a binary classifier, and the minimum confidence c is set s To distinguish untrustworthy points, the confidence level c i Below c s The points are selected as untrusted points; (3) Value-based confidence assessment: A strategy based on the confidence information of the entire image is adopted to expand the equation to a regression problem to avoid the classification threshold problem: Among them, δ i is the deviation, and the ratio of the maximum deviation of a prediction is used as a variable to evaluate the confidence and select unconfident points in the case of large deviations. The confidence threshold c is also used here. s To select untrusted points; (4) To improve matching efficiency, we further use ranking to evaluate confidence, which can be expressed as follows: Among them, rank is the calculation of the i-th deviation δ in descending order i rank, confidence threshold c s Controlled the number of matching points; Obtain the untrusted point sets P1 and P2 of the two frames for the predicted optical flow from the matching confidences CF1 and CF2 to facilitate subsequent matching optimization; In step (3), the feature matching strategy is used for matching, and then the mutual flow MF with high matching probability is used to update the optical flow to provide guidance for the uncertain points. The specific process is as follows: (1) Adaptively match the selected points to produce a more accurate flow; assume that for an image pair, semantically similar regions have the same confidence, and then use the untrusted points P1 and P2 in the image pair for accurate matching; after completing the selection of untrusted points, take out the matching features M1 and M2 of each point extracted by the matching network; (2) Apply self-attention and cross-attention to the extracted features; make the network learn the relationship with other features before calculating the correlation; calculate self-attention and cross-attention separately and connect them together for later matching; (3) Use an efficient method to perform matching; first normalize these features and calculate the correlation matrix: R(i, j)=<M1(i),M2(j)> , (8) Where R(i, j) is the correlation between point i in image I1 and point j in image I2; M1 and M2 are matching features, and <·, ·> is the inner product; (4) Then, the matching point pairs are filtered using the following conditions: Use softmax for matching filtering; Calculate the score of each pair of matching items; Point o in the untrusted point set P1 of image I1 and point j in the untrusted point set P2 of image I2 meet the following conditions: In point set P2, the most likely matching point of i is j, and in point set P1, the most likely matching point of j is i; The final score P for each point pair i and j matching c (i,j) is calculated using the following formula: P c (i,j)=softmax(R(i,·)) j ·softmax(R(·,j)) i , (9) Where R(i, ) represents the i-th column of the correlation matrix, and softmax(R(i, )) j It represents the value of the i-th column and j-th row after softmax calculation. The operator · represents the product. Softmax is expressed as the following formula: Among them, z j is the jth number of vector z, e represents the natural exponent, and Σ represents the sum of the natural exponents of all numbers in vector z; (5) The features of the untrusted points come from the matching features M1 and M2; two thresholds are set for matching: First, the mutual probability P of the matching pair i and j c (i, j) is higher than a threshold to reduce matching error points. This criterion is called mutual nearest neighbor. Secondly, the correlation R(i, j) of matching pair i and j is also higher than a threshold to reduce untrusted matches. The matching selection strategy is expressed by the following formula: Where MF is the resulting match or mutual flow, θ p is the mutual probability threshold, θ r is the threshold of correlation; (i, j)∈(P1, P2) means that i, j are point pairs generated from point sets P1, P2; (6) The original prediction F' of the optical flow is updated using the mutual flow MF before the next iteration to correct incorrect matches in advance, which is a guide for the dense flow estimation and is more accurate before subsequent iterations; feature matching is only performed in the first α percentage of iterations, with α being 0.5; after multiple iterations, the confidence increases, and the dense flow estimation can handle the matching and refine the correspondence, saving the calculation of redundant matching.
2. The method for performing adaptive dense matching calculation on two frames of images according to claim 1, characterized in that: In step (5), the loss function is divided into flow prediction result loss and confidence loss, which is the basis of adaptive matching network optimization. The specific calculation is as follows: (1) The flow prediction result loss uses the first norm of the flow prediction result error and the weight that increases with the number of iterations: Among them, f i is the predicted flow value of the oth iteration, γ is the weight coefficient, M is the total number of iterations, ||f i -f gt ||1 indicates the prediction flow f i and the real flow f gt The one-dimensional norm of the difference between , that is, the absolute deviation ∈; (2) Depending on the evaluation strategy, the confidence loss is expressed in two different forms; for probability-based strategies, a binary classification problem needs to be solved, so cross entropy is used. As the loss function; apply the sigmoid curve to scale the predicted confidence to the interval (0, 1), and use whether the absolute deviation ∈ is greater than the threshold δ s as labels; the loss is expressed as follows; Among them, P s (·|δ) means δ is greater than or less than the threshold δ s The probability of i,j and δ i,j is the absolute deviation and prediction deviation of point j in the i-th iteration, and N is the total number of points. In order to be consistent with the flow prediction results, a weighting factor is also added to adjust the learning scale at each iteration to avoid learning from too large a deviation. For both value-based and ranking-based strategies, the first norm of δ with respect to ∈ is used to constrain the estimate; the loss function is expressed as: Among them, ||δ i,j -∈ i,j ||1 indicates the prediction deviation δ i,j and the true deviation ∈ i,j The one-dimensional norm of the difference between ; (3) The final loss function L is the linear sum of these two losses, that is, L = L f +βL c , where β is the weight factor and the loss function L is used to perform back propagation and update the parameters of the neural network.
3. The method for performing adaptive dense matching calculation on two frames of images according to claim 2, characterized in that: In step (6), the reasoning through the neural network is specifically described as follows: the flow estimation process is divided into two processes: the first process estimates a smaller flow at a lower resolution, which is then used as the initialization of the second process, and the second process uses the original resolution for prediction; the specific description is as follows: f l (x)=φ(x,up2(φ l-1 (down2(x)))), (15) Among them, φ l It is the matching network with the lth layer, up2 means 2× bilinear upsampling, and down2 means 2× downsampling.
Citation Information
Patent Citations
Medical image segmentation method based on 3D dynamic edge insensitivity loss function
CN111968138A
Dense matching method and system based on deep learning view self-selection network
CN113807417A