Efficient optical flow estimation method and device based on Mama

By combining the Mamba-based optical flow estimation method with the Mamba state-space model and differential autoregressive refinement module, the problem of insufficient robustness of optical flow estimation in fast-moving and weakly textured regions is solved, achieving efficient and real-time optical flow estimation, which is applicable to fields such as video surveillance, moving target detection, and autonomous driving.

CN120997251AActive Publication Date: 2025-11-21ZHEJIANG UNIV OF TECH

Patent Information

Application Number
CN202511512181.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2025-11-21
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Existing optical flow estimation methods are not robust enough in fast motion, significant illumination changes or weak texture regions, have high computational overhead, and are difficult to meet the real-time requirements of engineering. Furthermore, the limitations of the local receptive field of traditional convolutional neural networks and the bottleneck of the quadratic complexity of Transformers result in slow inference speed and high resource consumption.

Method used

An efficient optical flow estimation method based on Mamba is adopted. A single-scale high-dimensional feature with a fixed downsampling rate is extracted by a shared weight convolutional encoder. Intra-frame and cross-frame feature enhancement is performed by combining the Mamba state space model. A four-dimensional cost volume is constructed for global matching. Iterative refinement is performed by the differential Mamba autoregressive refinement module, and finally, high-precision optical flow results are output.

Benefits of technology

While maintaining accuracy, it significantly reduces computational resource consumption and inference latency, and improves estimation accuracy and robustness under large displacement, illumination changes and occlusion areas, making it suitable for high-resolution, long-sequence and resource-constrained real-world application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997251A_ABST
    Figure CN120997251A_ABST
Patent Text Reader

Abstract

The invention discloses an efficient optical flow estimation method and device based on Mama, and the method comprises the steps: carrying out the normalization and size alignment of two adjacent frames of images, and extracting the dense features of a fixed down-sampling rate through a shared weight convolution encoder; the two-frame features are sent to a multi-level feature enhancement module, an intra-frame modeling unit and a cross-frame interaction unit are cascaded and matched with channel reforming and residual error correction, and enhanced features are obtained; constructing a four-dimensional cost body on a low resolution, performing probability normalization along a target coordinate dimension, weighting a target coordinate grid according to a probability to obtain a corresponding coordinate, and subtracting the corresponding coordinate from a source coordinate to obtain an initial optical flow; and carrying out attention-guided space fusion on the initial optical flow and context and local correlation, sending the fused optical flow to a differential Mama-based autoregressive refinement module, carrying out iterative updating according to a small number of fixed steps, recovering to a target resolution through convex combination up-sampling, and outputting a final optical flow. According to the method, the optical flow field can be accurately estimated under the conditions of low complexity and low time delay.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of computer vision, and develops an efficient optical flow estimation method and device based on Mamba. BACKGROUND

[0002] Optical flow estimation is a basic problem in computer vision, aiming to describe scene motion information by analyzing pixel displacement between consecutive frames, and has important application value in video monitoring, moving target detection, automatic driving and robot navigation.

[0003] Traditional methods are mostly based on the assumptions of constant brightness and small displacement, and obtain the optical flow field through energy function optimization. Such methods are theoretically complete, but are often insufficient in robustness in fast motion, significant light changes or weak texture areas, and have large computational overhead, making it difficult to meet the real-time needs of engineering.

[0004] Deep learning has promoted the development of end-to-end optical flow methods. Convolutional neural network (CNN) based schemes are superior in efficiency, but are limited by the local receptive field of convolution; many methods rely on multiple iterations of refinement to achieve accuracy, which increases the number of parameters and time delay, and still has performance bottlenecks in large displacement and complex motion scenes. In contrast, the Transformer uses self-attention to achieve global modeling, which can better handle large displacement and complex motion, but its quadratic computational complexity results in slow inference speed, bloated models, high resource occupation, and limited deployment.

[0005] In summary, while maintaining estimation accuracy, reducing inference delay and computing power, and balancing global information modeling and efficient computation, are key contradictions that need to be solved in current optical flow research and application. Therefore, there is an urgent need for an efficient optical flow estimation scheme that can balance global dependence representation and linear / quasi-linear complexity to adapt to high-resolution, long-sequence and resource-constrained real application scenarios. SUMMARY

[0006] The present application overcomes the above-mentioned shortcomings of the prior art, and provides an efficient optical flow estimation method and device based on Mamba, which improves the high computational complexity and long inference delay of existing methods, significantly reduces the computational complexity and improves the inference speed while maintaining or even improving the accuracy.

[0007] To solve the above technical problems, the present application provides the following technical solutions: The first aspect of the present application provides an efficient optical flow estimation method based on Mamba, comprising the following steps: Collecting two adjacent frames of images to be processed and performing normalization and size alignment processing to form an input image pair with uniform resolution and dynamic range; The image pairs are respectively subjected to feature extraction, and a convolutional encoder with shared weights is used to output single-scale high-dimensional dense feature maps at a fixed down-sampling rate to balance the representation ability and computational complexity; The two frames of dense features are input into a multi-level feature enhancement module, which first aggregates long-range context within the frame, then interacts and fuses in the cross-frame dimension, and introduces a spatial neighborhood state fusion mechanism in the state updating process, so that the feature representation can combine sequence dependence and two-dimensional spatial context, and combine multi-layer perceptron (MLP) channel reorganization and residual correction, thereby obtaining global and cross-frame information supporting matching and estimation with linear or near-linear time complexity; A four-dimensional cost volume is constructed at the low-resolution feature level to represent the correspondence relationship from the source pixels to the target pixels, a softmax normalization is performed in the target coordinate dimension to obtain a matching probability distribution, and the distribution is used to perform weighted averaging on the two-dimensional coordinate grid of the target image to obtain the corresponding coordinates, which are then subtracted from the source coordinates to obtain the initial optical flow. One-time global matching obtains high-quality initial values and reduces the number of subsequent iterations and the overall computational amount; The initial optical flow and the local correlation of the target feature are jointly encoded into motion representation and sent into an attention-guided aggregator together with the context feature for spatial adaptive weighted fusion, highlighting the dynamically relevant areas and suppressing irrelevant interference to obtain aggregated features for refinement; The aggregated features are input into an autoregressive refinement module based on differential Mamba, which gradually outputs optical flow increments according to a predetermined small number of fixed iteration steps and superimposes them on the current estimate to obtain a more accurate optical flow field; The refined low-resolution optical flow is restored to the target resolution using convex combination upsampling, and the final result is reconstructed and the boundary consistency is processed in combination with the local continuity constraint; The final optical flow result is output for motion analysis, video understanding or other downstream visual tasks.

[0008] The feature extraction stage uses a shallow residual network with shared weights as an encoder to output single-scale high-dimensional features at a fixed down-sampling rate to balance the representation ability and computational complexity; The encoder is composed of stacked basic residual units and combined with normalization and nonlinear activation to stabilize training and inference; To facilitate transfer learning and multi-channel input, the encoder supports loading general pre-trained weights and adapting the first layer weights; Channel alignment and light fusion are set at the tail of the encoder to form a unified dimensional feature representation for subsequent enhancement and matching stages.

[0009] The multi-level feature enhancement module is composed of a plurality of intra-frame Mamba units and a cross-frame Mamba unit in cascade, and a light-weight MLP is introduced after each layer to perform channel reorganization and residual correction, so as to enhance the discriminability without significantly increasing the constant term overhead.

[0010] Global matching is realized in a global correlation manner on low-resolution features, a global correlation matrix is calculated after the spatial dimension of two frames of features is flattened, and is scaled by the square root of the number of channels, so as to obtain a cost volume representing the one-to-one correspondence relationship between source pixels and target pixels; The cost volume is subjected to softmax normalization in the target coordinate dimension to obtain a matching probability distribution, and the corresponding coordinates are obtained by weighted averaging of the target coordinate grid according to the distribution; The initial optical flow is obtained by subtracting the corresponding coordinates from the source coordinates, and bidirectional optical flow is output by simultaneously calculating the forward and reverse correlations when needed.

[0011] The autoregressive refinement module adopts a fixed small number of iterations to control the inference delay and the amount of calculation, and the number of iterations is set as a predetermined constant during deployment, and is updated by accumulating the optical flow increment output by the differential Mamba-based decoding unit at each step.

[0012] The up-sampling realizes the continuity reconstruction of sub-pixel accuracy by convex combination of neighborhood optical flows through learned local weights.

[0013] The training adopts an end-to-end joint optimization strategy to simultaneously improve the collaborative effect of the four stages of encoding, enhancement, matching and refinement, data enhancement includes brightness and contrast disturbance, random cropping and pyramid scaling, to enhance the generalization ability to large displacement, weak texture and repetitive texture scenes, and the loss design includes a reprojection consistency term, an occlusion and boundary robustness term and a smoothing regularization term.

[0014] The second aspect of the application relates to a Mamba-based efficient optical flow estimation device, comprising a memory and one or more processors, the memory stores executable code, and the one or more processors execute the executable code to implement the Mamba-based efficient optical flow estimation method of the application.

[0015] The third aspect of the application relates to a computer readable storage medium having a program stored thereon, which, when executed by a processor, implements the Mamba-based efficient optical flow estimation method of the application.

[0016] After the normalization and size alignment of the two adjacent frames of images are completed, the single-scale high-dimensional features of a fixed down-sampling rate are extracted through a shared weight convolutional encoder; the two frame features are input into a multi-level feature enhancement module for intra-frame and cross-frame information aggregation, and the intra-frame modeling unit and the cross-frame interaction unit are cascaded and cooperated with channel reorganization and residual correction to obtain enhanced features; a four-dimensional cost volume is constructed on a low resolution, and probability normalization is performed along the target coordinate dimension, the corresponding coordinates are obtained by weighting the target coordinate grid according to the probability, and the difference between the source coordinates and the corresponding coordinates is obtained, and the probabilistic global matching is performed to obtain the initial optical flow; the initial optical flow, the context features and the attention guided spatial fusion of the local correlation are sent into the autoregressive refinement module based on the differential Mamba for iterative refinement, and are restored to the target resolution through convex combination up-sampling, and the final optical flow result is output, so that low complexity, high throughput and low time delay are realized on the premise of maintaining the existing accuracy.

[0017] The innovation of the present application is: (1) The Mamba state space model is introduced into the optical flow estimation framework, which breaks through the local receptive field limitation of traditional convolutional neural networks and the secondary complexity bottleneck of Transformer, realizes global dependence modeling under nearly linear complexity, and significantly reduces the consumption of computing resources and the inference time delay.

[0018] (2) A spatial neighborhood state fusion mechanism is proposed in the state recursion process of the intra-frame and cross-frame Mamba units, which improves the consistency and discriminability of the representation under large displacement, illumination change and occlusion area.

[0019] (3) The autoregressive refinement structure based on the differential Mamba is designed, which can highlight the significant motion patterns and suppress redundant noise, effectively alleviate the state diffusion and numerical instability problems in the long sequence recursion process, and realize high-precision optical flow refinement in a small number of iteration steps.

[0020] (4) The attention mechanism is used to weight and aggregate the motion features, context features and historical states, realize spatial adaptive information fusion, highlight the dynamic related areas and suppress irrelevant interference.

[0021] The working principle of the present application is: (1) The Mamba state space model can establish long-range dependencies under linear complexity by selective scanning and recursive updating, so that each position can combine input features and hidden states when updating. This mechanism ensures that optical flow estimation can perceive the global and keep low computational overhead.

[0022] (2) The spatial neighborhood state fusion mechanism combines the hidden state of the current position and the state of the two-dimensional neighborhood when updating, so that the update not only depends on the time sequence recursion, but also combines the local context. This joint modeling method can reduce matching ambiguity and enhance feature stability in large displacement and occlusion scenes.

[0023] (3) The differential Mamba autoregressive refinement calculates two paths of states in parallel and performs differential fusion in each iteration to highlight significant motion patterns and suppress redundant noise. By gradually superimposing the flow increments, the model can converge to a higher precision result in a small number of fixed iteration steps.

[0024] (4) The attention mechanism generates weights pixel by pixel for motion features, context features, and historical states in the refinement stage, and then performs weighted summation after input dimension normalization, so as to adaptively highlight motion-related regions, reduce irrelevant interference, and provide stable input for subsequent iterations.

[0025] The advantages of the present application are: (1) An efficient optical flow estimation method based on Mamba is proposed, which can complete optical flow estimation with linear or near-linear complexity while maintaining the ability to model global dependencies, significantly reducing computational resource consumption and inference latency.

[0026] (2) Through multi-level feature enhancement of intra-frame and cross-frame Mamba units, as well as spatial neighborhood state fusion mechanism, the estimation accuracy and robustness in large displacement, weak texture, illumination change and occlusion scenes are effectively improved.

[0027] (3) The autoregressive refinement module based on differential Mamba combined with attention guided aggregator can complete optical flow refinement in a small number of iteration steps, balancing high precision and high real-time performance.

[0028] (4) The method has good generality and scalability, supports pre-trained weight loading and diversified data augmentation, and can be widely applied to visual task scenarios such as motion analysis, video understanding, autonomous driving and robot navigation. BRIEF DESCRIPTION OF DRAWINGS

[0029] Figure 1 is the overall flowchart of the method of the present application.

[0030] Figure 2 is the overall architecture of the model of the method of the present application.

[0031] Figure 3 is a structural schematic diagram of multi-level feature enhancement of the present application.

[0032] Figure 4 is an iteration process schematic diagram of the autoregressive refinement module of the present application.

[0033] Figure 5 is a schematic diagram of the device of the present application. DETAILED DESCRIPTION

[0034] In order to make the technical problems, technical solutions and advantages of the present application clearer, a kind of high-efficiency end-to-end optical flow estimation method based on Mamba will be described in detail below with the drawings.

[0035] Embodiment 1

[0036] In order to overcome the shortcomings of the existing optical flow estimation method in global information modeling, calculation complexity and precision improvement, the present embodiment proposes a high-efficiency optical flow estimation method based on Mamba. The Mamba architecture used has linear calculation complexity, which can greatly reduce the consumption of calculation resources and inference delay while realizing the ability of global information modeling. This method can effectively reduce the inference delay and calculation resource consumption while ensuring high-precision optical flow estimation, and is suitable for visual scenes with high real-time and precision requirements.

[0037] Figure 1 The overall flowchart of the high-efficiency optical flow estimation method based on Mamba of the present embodiment is shown, Figure 2 The model structure diagram of the present application is shown, which specifically includes the following steps: Step 101: feature extraction, normalizing and size aligning the adjacent two frames of images, and extracting dense feature maps with fixed down-sampling rate through a convolutional encoder with shared weights; Step 102: feature enhancement, sending the two frames of features into a multi-level state space enhancement module, first aggregating long-range context by intra-frame unit, then interacting and fusing by cross-frame unit, introducing spatial neighborhood state fusion mechanism in the state updating process, so that the features can combine sequence dependence and two-dimensional spatial context, and then completing channel reorganization and residual correction through a multi-layer perceptron to obtain enhanced features; Step 103: initial optical flow estimation, constructing a four-dimensional cost volume on a low resolution and applying exponential normalization along the target coordinate dimension to obtain a matching probability distribution, and then obtaining the corresponding coordinates by weighted averaging the target coordinate grid according to the distribution and differencing the source coordinates to obtain the initial optical flow; Step 104: autoregressive refinement, after the initial optical flow is fused with context and local correlation through attention-guided spatial weighting, it is sent into an autoregressive refinement module for a small number of fixed iterations, and is restored to the target resolution through convex combination up-sampling, and the final optical flow is output.

[0038] Step 101 specifically includes: (1) a convolutional encoder based on ResNet34 structure is used to extract features from the input two consecutive frames of images. The encoder parameters are shared between the two frames of images to ensure the consistency of the feature extraction process.

[0039] (2) The convolutional encoder adopts public dataset (ImageNet) pre-training weight for parameter initialization, and is further optimized in the end-to-end training process. The encoder is composed of multiple residual modules, which can effectively improve the feature expression ability and model generalization performance.

[0040] (3) After the above-mentioned encoder multi-layer convolution and down-sampling processing, a single-scale high-dimensional dense feature map is output, the height and width of which are 1 / 8 of the original input image, which is used to provide basic feature expression for subsequent feature enhancement and optical flow estimation.

[0041] Step 102 specifically comprises: (1) Figure 3 The overall flow of the multi-level feature enhancement module of the application is shown, which specifically comprises the following steps: first, the input two frames of high-dimensional features pass through the intra-frame Mamba unit, which further introduces a two-dimensional space scanning mechanism on the basis of the traditional state space recursion, realizes global modeling and context information aggregation in the spatial range of a single frame, and then the dynamic interaction and fusion of the two frames of features are realized through the cross-frame Mamba unit, so as to fully capture the spatio-temporal dependence relationship and cross-frame motion features; the MLP module is introduced after each cross-frame Mamba unit, so as to improve the nonlinear expression ability of the features. The above-mentioned units can be stacked in multiple layers, and finally output two frames of high expression force features enhanced by intra-frame and cross-frame, so as to provide rich and discriminative feature basis for subsequent optical flow estimation, realize efficient and accurate end-to-end pixel-level motion estimation.

[0042] (2) The application introduces a self-supervised feature enhancement mechanism of the Mamba state space model in the feature enhancement stage. Mamba is a kind of state space neural network structure which can efficiently model long-distance dependence of sequences with linear computational complexity. For the high-dimensional feature map output by the encoder for each frame (wherein ), this step first inputs it into the Mamba state space model, and uses a selective scanning strategy to sequentially perform forward and reverse state updates in the spatial dimension. The state update unit includes learnable parameters 、 、 、 , and the state recursion equation is as follows:

[0043] wherein, is the input feature at the current position, is the hidden state vector, is the output feature.

[0044] On this basis, the application proposes an improved intra-frame Mamba unit, which is different from the traditional recursive method which only depends on the state of the previous moment, and further introduces a spatial neighborhood state fusion mechanism in the state updating process, and after obtaining the original state variable , the state of the current position is fused with several neighbor states in the two-dimensional spatial neighborhood of the current position to obtain an enhanced structure-aware state , and the formula is as follows:

[0045] Among them, indicates the first neighbor index of the position , and is a learnable fusion weight. The aggregation process makes the state not only depend on the sequence information obtained by time recursion, but also integrate the context of its spatial neighborhood.

[0046] Finally, the output feature is given by the enhanced , and the formula is as follows:

[0047] In order to eliminate the directional bias of one-dimensional spatial scanning and make full use of the context from both sides of the sequence, this step performs forward and reverse state space processing on the same spatial sequence, respectively. The forward is used to aggregate the cumulative information before the current position, and the reverse is used to supplement the clues of the position after it. The two complement each other, form a symmetric and stable long-range dependence in a non-causal constraint space, and improve the consistency and discriminability in the boundary and occlusion area. The bidirectional aggregation relationship is as follows:

[0048] Finally, the output is denoted as , through the above process, the intra-frame Mamba unit can fully aggregate the global context information in the feature map space, and provide a high expressive feature basis for cross-frame feature interaction.

[0049] (3) The application further introduces a cross-frame interaction feature enhancement mechanism of the Mamba state space model in the feature enhancement stage, which is used to perform bidirectional interaction and fusion on the enhanced features of two frames under the premise of maintaining linear computational complexity, so as to depict the long-range space-time dependence between frames and improve the stability and discriminability of optical flow estimation. Let the enhanced features of two frames processed by the intra-frame Mamba unit be and . First, taking the first frame as the main branch, the dynamic response more sensitive to cross-frame alignment is extracted through depth separable convolution and nonlinear activation to obtain the intermediate representation, and the formula is as follows: ​

[0050] wherein, is the first frame enhanced feature, represents a depthwise separable convolution, is an activation function, is the main branch output feature.

[0051] Meanwhile, the second frame is taken as the modulation branch, and linear projection is used to obtain cross-frame modulation information that can be injected into state update, and the formula is as follows:

[0052] wherein, is the second frame enhanced feature, is a linear mapping operator, is the modulation branch output feature.

[0053] Subsequently, the two branch representations are concatenated in the channel dimension and linearly mapped to generate input-related state space parameters, and the formula is as follows:

[0054] wherein, represents the concatenation of the two in the channel dimension, is a linear mapping, , , respectively correspond to the gating, input mapping and readout parameters, which are all adaptive to the input.

[0055] In the selective scanning process, the , , and the input sequence unfolded in spatial order (derived from ) are sent to the state space update unit together to complete the hidden state update and readout, and the intermediate output is obtained, and the formula is as follows:

[0056] wherein, represents a state space update unit driven based on input-related parameters, is the output result. Similarly, the application also introduces spatial neighborhood state aggregation in the state update process of the cross-frame unit, so that the cross-frame dependency relationship can also consider the context of adjacent regions when modeling, further enhancing the robustness in large displacement and occlusion scenarios.

[0057] Consistent with intra-frame Mamba units, to alleviate the directional bias of unidirectional scanning and achieve temporal symmetry, this invention also employs bidirectional selective scanning in cross-frame Mamba units, performing forward and reverse processing on the same spatial sequence separately. The two outputs are then linearly mapped and summed in the channel dimension to form a cross-frame interactive representation.

[0058] (4) In this invention, an MLP is set after the cross-frame interaction features of the cross-frame Mamba unit to perform nonlinear mixing and residual correction of the source features and interaction features in the channel dimension, forming stable and highly expressive features for subsequent global matching and iterative refinement.

[0059] Specifically, the output of the cross-frame Mamba unit first undergoes linear merging and layer normalization to align the statistical distribution and suppress drift caused by dynamic interactions. Then, it is concatenated with the corresponding source features in the channel dimension before being fed into the MLP. The MLP employs a two-layer linear mapping and nonlinear activation and layer normalization structure: first, the number of channels after concatenation is expanded by a factor of four and GELU activation is applied; then, the number of channels is restored to match the source features, and layer normalization stabilizes the training and inference processes. To ensure smooth information flow and stable gradient propagation, the output of this unit is residually added to the source features, serving as the final enhancement result for this layer.

[0060] Step 103 specifically includes: (1) This invention constructs a global matching cost body. Based on the two frames of enhanced features output in step 2, they are denoted as query features. (From the first frame) and the matched features (From the second frame). For any source pixel location With the target pixel position Calculate the scaled dot product similarity to form a four-dimensional cost volume, using the following formula:

[0061] in, For feature channel dimension, The above equation represents the inner product over the channel dimension. That is, the similarity mapping set of each source pixel on the entire target image.

[0062] (2) This invention probabilistically transforms the cost body into a matching distribution. In the cost body... Target coordinate dimension Normalization is performed on the above to obtain the matching probability distribution of each source pixel, and the formula is as follows:

[0063] in, Indicates to fixed All normalized by the similarity of the source and target images, outputting a probability distribution summed to 1, which characterizes the matching likelihood of the source pixels at each location of the target image.

[0064] (3) The application calculates the corresponding coordinates and generates the initial optical flow according to the matching distribution. The matching distribution is weighted and averaged on the two-dimensional coordinate grid of the target image to obtain the corresponding coordinates of the source pixels in the target image , and then the initial optical flow field is obtained by subtracting the source coordinates, and the formula is as follows:

[0065] wherein, and are two-dimensional coordinate representations of the target and source images, is the weighted corresponding coordinate, is the initial optical flow output by this step, which is used as the input of the subsequent refinement stage.

[0066] Step 104 specifically includes: (1) Figure 4 The iterative flow diagram of the autoregressive refinement module of the application is shown, which specifically includes the following steps: first, motion features are constructed according to the initial optical flow. Under the guidance of the initial optical flow obtained in step 3, local correlation is calculated in the neighborhood of the second frame feature map, and the optical flow is jointly coded to obtain motion representation , and the formula is as follows:

[0067] wherein, represents that the local correlation is calculated in a window with a radius of , is the correlation feature obtained by convolution on , is the displacement feature obtained by convolution on , represents channel dimension splicing, is a combination of one layer convolution and nonlinear mapping. Thus, the , which can express local displacement and correlation, is provided as a basis for subsequent fusion.

[0068] ​​​​​(2) This invention employs attention-guided aggregation to achieve multi-source information fusion. In (1) obtained... Based on this, combined with the features output in step 2 and the current hidden state By adaptively assigning weights, motion-related regions are highlighted spatially, forming aggregated features for decoding. The formula is as follows:

[0069] in, This is a shallow convolutional and nonlinear mapping used to generate weights for the three inputs at each spatial location. Normalize over the input path dimension so that at the same position , In order to represent , and , This indicates element-wise multiplication. Thus, By organically unifying motion, context, and historical memory, a stable input is provided for iterative refinement.

[0070] (3) This invention uses Mamba to perform autoregressive refinement. As a temporal input, the hidden state is updated by the Mamba unit, and the optical flow increment is predicted. This prediction is then superimposed on the current optical flow to form a new estimate, as shown in the following formula:

[0071] in, Two-story Convolution mapping outputs optical flow increments, with the initial term taken as... . For improved differential Mamba, this unit performs two independent Mamba mappings in parallel during state updates:

[0072] Then, after normalization, differential fusion is performed to obtain the final hidden state:

[0073] in, For learnable weights, This represents the normalization function.

[0074] The design philosophy of Differential Mamba is to enable the hidden states to highlight significant motion patterns in the input through two-way parallel mapping and differential operations, while suppressing redundant or noisy features. This avoids the energy accumulation and state diffusion problems that occur in traditional single-path recursion when processing long sequences. The normalization process further stabilizes the numerical range, ensuring that the state distribution remains consistent across different positions and iteration steps, thereby improving the model's training convergence and inference robustness.

[0075] Therefore, the aggregation results are gradually corrected from the initial estimate through an autoregressive mechanism.

[0076] (4) This invention restores the resolution and outputs the result through convex combination upsampling. The refined low-resolution optical flow is upsampled through learnable weights using convex combination upsampling to restore it to the target resolution and output it as the final result. The formula is as follows:

[0077] in, Indicates in Convex combination upsampling based on predicted weights within the neighborhood (weights are normalized within the neighborhood). To refine the total number of iterations.

[0078] Thus, steps 101 to 104 form a continuous chain of "construction-fusion-refinement-restoration", outputting high-precision optical flow results.

[0079] Example 2

[0080] like Figure 5 As shown, this embodiment provides a high-efficiency optical flow estimation device based on Mamba, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement a high-efficiency optical flow estimation method based on Mamba according to Embodiment 1.

[0081] Example 3

[0082] This embodiment relates to a computer-readable storage medium storing a program that, when executed by a processor, implements an efficient optical flow estimation method based on Mamba as described in Embodiment 1.

[0083] The embodiments described in this specification are merely examples of implementations of the inventive concept. The scope of protection of this invention should not be considered as limited to the specific forms stated in the embodiments. The scope of protection of this invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.

Claims

1. A method for efficient optical flow estimation based on Mamba, characterized in that, The application relates to a method for estimating optical flow, comprising the following steps: Collecting two adjacent images to be processed and performing normalization and size alignment processing to form an input image pair with unified resolution and dynamic range; Respectively performing feature extraction on the image pair, and outputting a single-scale high-dimensional dense feature map under a fixed down-sampling rate by using a convolutional encoder with shared weights; Inputting the two-frame dense features into a multi-level feature enhancement module, first aggregating long-range context within the frame, then performing interactive fusion in the cross-frame dimension, and introducing a spatial neighborhood state fusion mechanism in the state updating process; Constructing a four-dimensional cost volume at the low-resolution feature level to represent the correspondence relationship from the source pixel to the target pixel, performing softmax normalization on the target coordinate dimension to obtain a matching probability distribution, and performing weighted averaging on the target coordinate grid according to the distribution to obtain the corresponding coordinates, and then performing difference between the corresponding coordinates and the source coordinates to obtain the initial optical flow; Encoding the initial optical flow and the local correlation of the target feature into motion representation, and inputting the motion representation and the context feature into an attention-guided aggregator for spatial adaptive weighted fusion to obtain aggregated features for refinement; Inputting the aggregated features into an autoregressive refinement module based on differential Mamba, and outputting the optical flow increment step by step according to a preset small number of fixed iteration steps and superimposing the optical flow increment on the current estimation to obtain a more accurate optical flow field; Restoring the refined low-resolution optical flow to the target resolution by convex combination up-sampling, and combining local continuity constraints to complete the final result reconstruction and boundary consistency processing; Outputting the final optical flow result.

2. The method of claim 1, wherein, The feature extraction comprises the following steps: using a shallow residual network with shared weights as an encoder to output a single-scale high-dimensional feature under a fixed down-sampling rate, and balancing the representation capability and the computational complexity.

3. The method of claim 1, wherein, The encoder is composed of a stack of basic residual units, and normalization and nonlinear activation are combined to stabilize the training and inference; In order to facilitate transfer learning and multi-channel input, the encoder supports loading general pre-trained weights and adapting the first layer weights; Channel alignment and light fusion are arranged at the tail of the encoder to form a unified dimensional feature representation for subsequent enhancement and matching stages.

4. The method of claim 1, wherein, The multi-level feature enhancement module is composed of a cascade of multi-layer intra-frame Mamba units and cross-frame Mamba units, and a light MLP is introduced after each layer to perform channel reordering and residual correction, so as to enhance the discriminability without significantly increasing the constant term overhead.

5. The method of claim 1, wherein, Global matching is realized in a global correlation manner on the low-resolution feature, the global correlation matrix is calculated after the spatial dimension of the two features is flattened, and the global correlation matrix is scaled by the square root of the channel number, so that a cost volume representing the correspondence relationship between the source pixel and the target pixel is obtained; The cost volume is subjected to softmax normalization in the target coordinate dimension to obtain a matching probability distribution, and the corresponding coordinates are obtained by weighted averaging on the target coordinate grid according to the distribution; The corresponding coordinates and the source coordinates are subtracted to obtain the initial optical flow, and bidirectional optical flow is output by simultaneously calculating the forward and reverse correlations when needed.

6. The method of claim 1, wherein, The autoregressive refinement module adopts a fixed small number of iterations to control the reasoning time delay and the amount of calculation, the iteration step number is set as a predetermined constant during deployment, and is updated by accumulating after each step of outputting the optical flow increment by the differential Mamba-based decoding unit.

7. The method of claim 1, wherein, The upsampling is realized by convex combination of the neighborhood optical flow through the learned local weights, realizing the continuity reconstruction of sub-pixel accuracy.

8. The method of claim 3, wherein, The training adopts an end-to-end joint optimization strategy, simultaneously improving the collaborative effect of the four stages of coding, enhancement, matching and refinement, the data enhancement includes brightness and contrast disturbance, random cropping and pyramid scaling, to enhance the generalization ability to large displacement, weak texture and repetitive texture scenes, and the loss design includes a reprojection consistency term, a robustness term for occlusion and boundary, and a smoothing regularization term.

9. A high efficiency optical flow estimation device based on Mamba, characterized in that, The memory stores executable code, and the one or more processors execute the executable code to implement the Mamba-based efficient optical flow estimation method in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The program is stored on the memory and executed by the processor to implement the Mamba-based efficient optical flow estimation method in any one of claims 1-8.

Citation Information

Patent Citations

  • Optical flow estimation method based on state space model, program, equipment and storage medium

    CN119559219A

  • Differential Mama-based adaptive background reconstruction hyperspectral anomaly detection method

    CN120544038A

Cited By

  • Optical flow estimation method and system fusing Mama and visual basis model knowledge

    CN121616625A

  • Optical flow estimation method and system fusing mamba with visual based model knowledge

    CN121616625B

  • Optical flow estimation method and device based on depth perception and global-local cooperation

    CN121962207A

  • Optical flow estimation method and device based on depth perception and global-local collaboration

    CN121962207B

  • Diffusion model space-time fusion method based on optical flow guidance

    CN122048667A