An optical flow calculation method combining image pyramid subnet guidance and recurrent cross attention
Through the combined image pyramid subnet and cyclic cross-attention method, the problem of insufficient utilization of shallow information in optical flow calculation is solved, and the accuracy and robustness of optical flow calculation is improved, especially in complex motion scenarios and large displacement areas.
Patent Information
- Application Number
- CN202210480358.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-05
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2042-05-05
AI Technical Summary
The existing optical flow calculation model is insufficiently utilized in complex motion scenarios, resulting in insufficient accuracy and robustness of the estimation of motion edges and large displacement optical flows.
The combined image pyramid subnet and cyclic cross attention method are used to extract features through the image pyramid subnet and fuse with the feature pyramid, and context information is extracted in combination with the cyclic cross attention module to perform optical flow calculation.
The accuracy and robustness of optical flow estimation are significantly improved, especially the calculation accuracy of moving edges and large displacement regions, and are suitable for complex edge image sequences and large displacement image sequences.
Smart Images

Figure CN114821105B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an optical flow calculation method combining image pyramid subnet guidance and cyclic cross attention. Background Art
[0002] Optical flow is the instantaneous velocity of a moving object in the pixel observation plane. It is a method for calculating the motion of an object between adjacent frames. It is generated by the relative velocity of the object and the camera, and reflects the direction and speed of the object's corresponding image pixels in a very small timeframe. Recovering an object's three-dimensional structure and motion from optical flow is one of the most meaningful and challenging tasks facing current computer vision research. Optical flow plays a crucial role in computer vision, with applications in target segmentation, recognition, tracking, robotic navigation, and shape information recovery.
[0003] Currently, most feature extraction methods for optical flow calculation models use feature pyramids. However, simply using convolution for feature extraction prevents the effective utilization of spatial information in shallow layers, resulting in insufficient context extraction in complex motion scenes, which in turn reduces the accuracy of optical flow estimation for moving edges and large displacements. Introducing an image pyramid as a guide and adding recurrent cross-attention as auxiliary context extraction effectively balances information between deep and shallow layers, potentially improving the accuracy and robustness of optical flow calculation for moving edges and large displacements. Summary of the Invention
[0004] The purpose of the present invention is to provide an optical flow calculation method that combines image pyramid subnet guidance and cyclic cross attention to solve the problems involved in the above background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] The present invention provides an optical flow calculation method combining image pyramid subnet guidance and cyclic cross attention, the method comprising the following steps:
[0007] (1) Input two consecutive frames of the image sequence into the image pyramid subnet and feature pyramid respectively;
[0008] (2) Use the image pyramid subnet to process the image:
[0009] (3) The features extracted by the image pyramid subnet are added and fused with the features extracted by the feature pyramid at the same level as the input of the next level of feature pyramid;
[0010] (4) The feature map obtained by adding and fusing the features extracted by the fourth-layer image pyramid subnet and the features extracted by the same-layer feature pyramid, the feature map obtained by adding and fusing the features extracted by the fifth-layer image pyramid subnet and the features extracted by the same-layer feature pyramid, and the feature map extracted by the sixth-layer feature pyramid are used as inputs of the cyclic cross attention module to obtain the context information of the image;
[0011] (5) After deformation and correlation calculation of the feature map, the feature map is input into the shared optical flow decoder to calculate the optical flow to obtain the initial optical flow;
[0012] (6) The initial optical flow output in step (5) is refined by the context network and then optimized by the bilateral filter to obtain the final refined optical flow calculation result.
[0013] Furthermore, the input of the image pyramid subnet in step (2) is a set of downsampled image pyramid pictures; after the image pyramid is downsampled, the features of the image pyramid are extracted through a shallow network, namely the image pyramid subnet.
[0014] Furthermore, in step (4), two feature maps Q and K are obtained by dimensionality reduction through two 1×1 convolutions respectively. After obtaining Q and K, an attention map A is obtained by an association operation, and then a softmax operation is performed to obtain an attention map A'.
[0015] The optical flow calculation method of the present invention, which is guided by a joint image pyramid subnet and cyclic cross-attention, first inputs two consecutive frames of images into a feature extraction network guided by a joint image pyramid subnet and cyclic cross-attention for feature extraction; secondly, deforms and calculates the correlation of the feature map; then, the feature map after the correlation calculation is sent to a shared optical flow decoder for initial optical flow estimation; finally, the initial optical flow is refined by a context network and then bilaterally refined to obtain the final optical flow calculation result. The optical flow calculation method of the present invention, which is guided by a joint image pyramid subnet and cyclic cross-attention, extracts feature information of moving edges and large displacement areas of the image sequence by supplementing shallow information and accurately extracting context information, significantly improving the accuracy and robustness of optical flow estimation.
[0016] The optical flow calculation method of the present invention combines image pyramid subnet guidance and cyclic cross attention, which improves the accuracy and robustness of optical flow estimation of moving edges and large displacement areas by supplementing shallow information and accurately extracting context information.
[0017] The optical flow calculation method of the present invention combines image pyramid subnet guidance and cyclic cross attention, introduces shallow spatial information in deep convolution, and performs lightweight extraction of global context information, which significantly improves the accuracy of optical flow calculation and overcomes the problems of imbalance between deep and shallow layer information and large computational complexity. It has higher computational accuracy and better practicality for complex edge image sequences and large displacement image sequences, and has very important applications in target object segmentation, recognition, tracking, robot navigation, and shape information recovery. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 The 16th frame image in the cave_3 image sequence of the present invention;
[0019] Figure 2 The 17th frame image in the cave_3 image sequence of the present invention;
[0020] Figure 3 This is a diagram of the feature extraction network structure of the present invention's example combined with image pyramid subnet guidance and cyclic cross attention;
[0021] Figure 4 This is a correlation calculation graph of the feature graph of the present invention;
[0022] Figure 5 This is the shared decoder graph for optical flow and occlusion of the present invention;
[0023] Figure 6 Bilateral refinement graph of optical flow and occlusion for the present invention example;
[0024] Figure 7 The optical flow map of the cave_3 image sequence calculated by the present invention;
[0025] Figure 8 Flowchart of the calculation method of the present invention. Specific implementation methods
[0026] The following will provide a clear and complete description of the technical solutions in the examples of the present invention, in conjunction with the accompanying drawings. The examples described are only some embodiments of the present invention, not all embodiments. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.
[0027] See also Figures 1-8 , this paper provides an optical flow calculation method that combines image pyramid subnet guidance and cyclic cross attention, and uses cave_3 sequence images for experimental explanation:
[0028] (1) Input Figure 1 and Figure 2are two consecutive frames of the cave_3 image sequence; where: Figure 1 is the first frame image, Figure 2 is the second frame image;
[0029] (2) Figure 1 and Figure 2 Input to the image pyramid subnet and feature pyramid respectively;
[0030] (3) Figure 3 As shown, the image is first processed using the image pyramid subnet:
[0031] The input of the image pyramid subnet is a set of simple downsampled image pyramid images, represented as:
[0032]
[0033] Where H and W are the resolutions of the image, i represents the number of image pyramid layers, represents the resolution of the i-th layer image pyramid. For example, in one embodiment, i=5. After downsampling through the image pyramid, features of the image pyramid are extracted through a shallow network, namely, the image pyramid subnet:
[0034]
[0035] In formula (2), f(·) refers to the features extracted by the image pyramid subnet at layer i;
[0036] (4) In the first to fifth layers of the feature pyramid, the features extracted by the image pyramid subnet are added and fused with the features extracted by the feature pyramid at the same layer and directly used as the input of the next layer of the feature pyramid;
[0037] (5) The feature map obtained by adding and fusing the features extracted by the fourth-layer image pyramid subnet and the features extracted by the same-layer feature pyramid, the feature map obtained by adding and fusing the features extracted by the fifth-layer image pyramid subnet and the features extracted by the same-layer feature pyramid, and the feature map extracted by the sixth-layer feature pyramid are used as the input of the cyclic cross attention module to obtain richer global context information of the image;
[0038] Two 1×1 convolutions are used to reduce the dimension to obtain the two feature maps Q and K. After obtaining Q and K, the attention map A is obtained through the association operation, and then the softmax operation is performed to obtain the attention map A'. The association operation is as follows:
[0039] d i,u =Q u Ω i,u T (3)
[0040] Where d i,u Measured Q u and Ω i,u For each position u of the feature map Q, a vector Q with the same number of dimensions and channels as Q can be obtained. u , and then get an Ω on the feature map K i,u The set of vectors in the same row or column corresponding to position u, where i refers to Ω u The i-th element in;
[0041] Then, a 1×1 convolution is performed to obtain V, and the features of each position u in the horizontal and vertical directions of V are multiplied with the features of each position u in the horizontal and vertical directions of A, and the residual aggregation features of the position are added together, and the original features H are added. u Get a more powerful feature H' u , the aggregation operations used to collect context information are as follows:
[0042]
[0043] In formula (4), Φ i,u is the feature vector of the i-th layer in V in the same row or column as position u, A i,u is a scalar value in A at channel i and position u, H' u is the output feature map H u The eigenvector at position u, H u The feature map output by the sixth layer of feature pyramid;
[0044] (6) Figure 4 and Figure 5 As shown, to obtain the initial optical flow, the feature map is deformed and correlated, and then input into the shared optical flow decoder to calculate the optical flow. The specific operations are as follows:
[0045] x l 1_warp =warp l (x2,up2(flow l+1 )) (5)
[0046] Formula (5) represents the deformation operation of the image, where l represents the number of layers of the pyramid, x2 represents the second image, and warp l Represents the deformation operation of the image at the lth level of the pyramid, x l 1_warp It is the feature map of the second frame image pixel after the deformation operation by the previous layer upsampling optical flow, up2 is the optical flow upsampling using bilinear interpolation, flow l+1 Represents the upsampled optical flow output by the l+1th layer pyramid; then, the correlation between the deformed feature map and the original feature map is calculated:
[0047]
[0048] Formula (6) is the calculation process of optical flow, where Represents the first and second feature maps at the i-th level of the pyramid, respectively, x 1_warp represents the deformation operation in equation (5), Indicates that the optical flow of the previous layer is doubled bilinear upsampling, corr indicates the correlation calculation between the second feature map and the deformed feature map, cat represents the stacking operation, and finally the three stacks are input into the shared optical flow decoder D for optical flow estimation;
[0049] (7) Figure 6 As shown, the initial optical flow output in step (6) is refined by the context network and then optimized by the bilateral filter to obtain the final refined optical flow calculation result:
[0050]
[0051] Formula (7) is the bilateral optimization process of optical flow. Respectively represent the horizontal and vertical optical flows after bilateral filtering, g(x,y) represents the bilateral filter kernel that can be learned at the pixel point (x,y). Represents the horizontal and vertical optical flow image patches centered at (x,y);
[0052] like Figure 7 As shown, the method of the present invention has higher calculation accuracy and better applicability for moving edges and large displacement motion image sequences, and has very important applications in target object segmentation, recognition, tracking, robot navigation, and shape information recovery.
[0053] The optical flow calculation method of the present invention, which is guided by a joint image pyramid subnet and cyclic cross attention, first inputs two consecutive frames of images into a feature extraction network guided by a joint image pyramid subnet and cyclic cross attention for feature extraction; secondly, the correlation calculation is performed with the deformed feature map by changing the input order of the feature map; then the original feature map, the deformed feature map and the upsampled optical flow are stacked and sent to a shared occlusion and optical flow decoder for initial optical flow estimation; finally, the initial optical flow is refined by a context network and then bilaterally refined to obtain the final optical flow calculation result. The optical flow calculation method of the present invention, which is guided by a joint image pyramid subnet and cyclic cross attention, extracts feature information of moving edges and large displacement areas of an image sequence by supplementing shallow information and accurately extracting context information, thereby significantly improving the accuracy and robustness of optical flow estimation.
[0054] The optical flow calculation method of the present invention combines image pyramid subnet guidance and cyclic cross attention, which improves the accuracy and robustness of optical flow estimation of moving edges and large displacement areas by supplementing shallow information and accurately extracting context information.
[0055] The optical flow calculation method of the present invention combines image pyramid subnet guidance and cyclic cross attention, introduces shallow spatial information in deep convolution, and performs lightweight extraction of global context information, which significantly improves the accuracy of optical flow calculation and overcomes the problems of imbalance between deep and shallow layer information and large computational complexity. It has higher computational accuracy and better practicality for complex edge image sequences and large displacement image sequences, and has very important applications in target object segmentation, recognition, tracking, robot navigation, and shape information recovery.
[0056] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not limited to the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent replacements for some of the technical features therein. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An optical flow calculation method combining image pyramid subnet guidance and cyclic cross attention, characterized in that: The method comprises the following steps: (1) Input two consecutive frames of the image sequence into the image pyramid subnet and feature pyramid respectively; (2) Use the image pyramid subnet to process the image; The input of the image pyramid subnet in step (2) is a set of downsampled image pyramid pictures. After the image pyramid is downsampled, the image pyramid features are extracted by the image pyramid subnet; wherein the image pyramid subnet is a shallow network; (3) The features extracted by the image pyramid subnet are added and fused with the features extracted by the feature pyramid at the same level as the input of the next level of feature pyramid; (4) The feature map obtained by adding and fusing the features extracted by the fourth-layer image pyramid subnet and the features extracted by the same-layer feature pyramid, the feature map obtained by adding and fusing the features extracted by the fifth-layer image pyramid subnet and the features extracted by the same-layer feature pyramid, and the feature map extracted by the sixth-layer feature pyramid subnet are used as inputs of the cyclic cross attention module to obtain the context information of the image; Two 1×1 convolutions are used to reduce the dimension to obtain the two feature maps Q and K. After obtaining Q and K, the attention map A is obtained through the association operation, and then the softmax operation is performed to obtain the attention map A'. The association operation is as follows: d i,u =Q u Oh i,u T Where d i,u Measured Q u and Ω i,u For each position u of the feature map Q, a vector Q with the same number of dimensions and channels as Q is obtained. u , and then get an Ω on the feature map K i,u The set of vectors in the same row or column corresponding to position u, where i refers to Ω u The i-th element in; Then, a 1×1 convolution is performed to obtain V, and the features of each position u in the horizontal and vertical directions of V are multiplied with the features of each position u in the horizontal and vertical directions of A, and the residual aggregation features of the position are added together, and the original features H are added. u Get a more powerful feature H' u , the aggregation operations used to collect context information are as follows: Where, Φ i,u is the feature vector of the i-th layer in V in the same row or column as position u, A i,u is a scalar value in A at channel i and position u, H' u is the output feature map H u The eigenvector at position u, the H u The feature map output by the sixth layer of feature pyramid; (5) The feature map containing context information output by the cyclic cross attention module is deformed and correlated, and then input into the shared optical flow decoder to calculate the optical flow to obtain the initial optical flow; (6) The initial optical flow output in step (5) is refined by the context network and then optimized by the bilateral filter to obtain the final refined optical flow calculation result.