Optical flow estimation method and device, computer device and medium

CN121437552BActive Publication Date: 2026-09-22ATHENAEYES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511486631.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-09-22
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

[0003]然而,尽管PWC-Net等现有方法在多个基准数据集上表现良好,其在复杂真实场景下的应用仍存在若干固有局限:首先,多尺度特征融合机制较为刚性,缺乏对金字塔各层预测结果可靠性的动态评估

Benefits of technology

[0016]本发明实施例提供的光流估计方法、装置、计算机设备及存储介质,通过金字塔网络结构提取图像的多尺度特征,并基于多尺度特征,通过特征翘曲处理与代价体计算,得到各层级的初始光流;基于初始光流,通过不确定性估计分支计算各层级光流的不确定性图,不确定性图表征每个像素位置的光流预测可靠程度;采用当前层的不确定性图和上一层级融合后的不确定性图,计算当前层的融合权重,并利用融合权重对当前层光流和上一层级融合光流进行加权融合,得到当前层融合光流;基于融合权重和当前层不确定性图,更新当前层融合后的不确定性图;基于相邻金字塔层的光流梯度差异,计算运动梯度一致性损失,用于约束光流在跨层级间保持梯度方向一致;基于初始光流的端点误差损失、不确定性估计分支的不确定性损失和运动梯度一致性损失,构建总损失函数;使用总损失函数对光流估计网络进行训练,得到优化后的光流估计模型,并采用优化后的光流估计模型对待估计图像进行光流预测。实现采用引入不确定性估计分支动态评估各层级光流预测的可靠性,并据此加权融合多尺度结果,有效抑制了遮挡、无纹理区域的误差传递与累积;同时,通过运动梯度一致性约束损失强制相邻层间光流梯度对齐,保障了运动边界处的平滑性与物理合理性;两者协同作用,共同解决了传统方法在多尺度融合鲁棒性与跨层级一致性方面的不足,从而显著提升了光流估计在复杂场景下的整体精度、鲁棒性和可解释性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121437552B_ABST
    Figure CN121437552B_ABST
Patent Text Reader

Abstract

The application discloses an optical flow estimation method, device, equipment and medium, comprising: extracting multi-scale features of an image through a pyramid network structure, and obtaining initial optical flows of each level through feature warping processing and cost volume calculation; calculating uncertainty maps of the optical flows of each level based on the initial optical flows; calculating fusion weights of a current layer by using the uncertainty map of the current layer and a fused uncertainty map of a previous level, and performing weighted fusion to obtain a fused optical flow of the current layer; calculating motion gradient consistency loss based on optical flow gradient differences of adjacent pyramid layers; constructing a total loss function based on endpoint error loss, uncertainty loss and motion gradient consistency loss of the initial optical flow; training an optical flow estimation network by using the total loss function to obtain an optimized optical flow estimation model, and performing optical flow prediction on an image to be estimated. The application improves the overall precision, robustness and interpretability of optical flow estimation in complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and more particularly to an optical flow estimation method, apparatus, computer equipment, and medium. Background Technology

[0002] Optical flow estimation is one of the core tasks in computer vision, aiming to compute the motion vector of each pixel between adjacent frames in a video sequence. It is widely used in scenarios such as autonomous driving, video action recognition, object tracking, video interpolation, and enhancement. In recent years, deep learning-based optical flow estimation methods have made significant progress, with PWC-Net (Pyramid, Warping, and Cost Volume Network) and its derivatives becoming the mainstream architecture. These methods construct an image pyramid, warp features at different scales, and construct a cost volume to compute the matching cost. Finally, they recursively optimize optical flow prediction in a coarse-to-fine manner, achieving a good balance between accuracy and efficiency.

[0003] However, despite the good performance of existing methods such as PWC-Net on multiple benchmark datasets, their application in complex real-world scenes still has several inherent limitations: First, the multi-scale feature fusion mechanism is relatively rigid and lacks dynamic evaluation of the reliability of prediction results at each layer of the pyramid. Existing methods typically fuse optical flow from different layers with fixed or heuristic weights, making it difficult to handle situations where predictions at certain layers are significantly unreliable in areas such as occlusion, motion blur, and weak texture. This leads to the propagation and accumulation of errors between layers, affecting the accuracy of the final output. Second, there is a lack of explicit consistency constraints on optical flow predictions between different layers of the pyramid. Lower-level optical flow is often responsible for capturing large displacement motions but has lower accuracy, while higher-level optical flow is used to refine details but may deviate from the dominant motion direction of the lower layers. This inconsistency in cross-layer predictions can disrupt the overall smoothness and physical rationality of the optical flow field, especially at object edges and in fast-moving regions where estimation conflicts are likely to occur. Third, feature warping operations heavily depend on the quality of the optical flow predictions of the previous layer. If there are errors in the coarse-scale optical flow, feature warping based on it will amplify the errors and propagate them to subsequent layers, causing cost volume construction matching failures and leading to estimation degradation.

[0004] Therefore, existing technologies still lack the ability to explicitly model prediction uncertainty, fail to achieve confidence-based adaptive multi-scale fusion, and lack effective mechanisms to constrain cross-level prediction consistency, resulting in insufficient robustness and generalization ability in complex dynamic environments. A novel optical flow estimation method is urgently needed that can introduce an uncertainty-aware mechanism and impose cross-level consistency constraints to improve estimation accuracy and reliability in complex scenarios. Summary of the Invention

[0005] This invention provides an optical flow estimation method, apparatus, computer device, and storage medium to improve the overall accuracy, robustness, and interpretability of optical flow estimation in complex scenarios.

[0006] To address the aforementioned technical problems, embodiments of this application provide an optical flow estimation method, including: Multi-scale features of the image are extracted through a pyramid network structure, and based on the multi-scale features, the initial optical flow of each level is obtained through feature warping processing and cost volume calculation. Based on the initial optical flow, uncertainty maps of optical flow at each level are calculated through uncertainty estimation branches. The uncertainty maps characterize the reliability of optical flow prediction at each pixel location. Using the uncertainty map of the current layer and the uncertainty map of the previous layer after fusion, the fusion weight of the current layer is calculated, and the fusion weight is used to perform weighted fusion of the optical flow of the current layer and the fused optical flow of the previous layer to obtain the fused optical flow of the current layer. Based on the fusion weights and the uncertainty map of the current layer, update the uncertainty map of the current layer after fusion; based on the difference in optical flow gradient between adjacent pyramid layers, calculate the motion gradient consistency loss to constrain the optical flow to maintain consistent gradient direction across layers; Based on the endpoint error loss of the initial optical flow, the uncertainty loss of the uncertainty estimation branch, and the motion gradient consistency loss, a total loss function is constructed; The optical flow estimation network is trained using the total loss function to obtain an optimized optical flow estimation model, and the optimized optical flow estimation model is used to predict the optical flow of the image to be estimated.

[0007] Optionally, the calculation of the uncertainty map of optical flow at each level through the uncertainty estimation branch includes: After the cost volume of each layer of the pyramid network structure, a light quantum network consisting of convolutional layers and activation functions is connected; The light quantum network outputs an uncertainty map with the same resolution as the current layer optical flow, where a larger uncertainty value indicates that the optical flow prediction at that location is less reliable.

[0008] Optionally, calculating the fusion weights of the current layer and performing weighted fusion includes: Upsample the uncertainty graph after fusion at the previous level and subtract it from the uncertainty graph at the current level. The difference results are input into the learnable sensitivity parameter and the Sigmoid activation function to obtain the fusion weights of the current layer; Using the fusion weights, the current layer optical flow and the upsampled previous layer fused optical flow are weighted and summed to obtain the current layer fused optical flow.

[0009] Optionally, updating the uncertainty graph after fusion of the current layer includes: Upsample the uncertainty graph after fusion at the previous level; Compare the uncertainty map of the current layer with the uncertainty map of the previous layer after upsampling pixel by pixel, and take the minimum value as the uncertainty map after fusion of the current layer.

[0010] Optionally, the calculation of motion gradient consistency loss includes: Calculate the gradient map of the optical flow between two adjacent layers after upsampling and alignment; The differences between the gradient maps are calculated and dynamically weighted based on the optical flow amplitude; The motion gradient consistency loss is obtained by summing the weighted differences of all adjacent levels in the pyramid.

[0011] Optionally, the dynamic weighting based on optical flow amplitude includes: Based on the amplitude value of the optical flow after upsampling from the previous layer, the weighting coefficient is calculated using an exponential function. The smaller the optical flow amplitude, the smaller the weighting coefficient.

[0012] Optionally, the total loss function is a weighted sum of endpoint error loss, uncertainty loss, and motion gradient consistency loss, wherein the uncertainty loss is constructed based on the weighted error between the uncertainty map of each layer and the true optical flow.

[0013] To address the aforementioned technical problems, embodiments of this application also provide an optical flow estimation device, comprising: The feature extraction module is used to extract multi-scale features of the image through the pyramid network structure, and based on the multi-scale features, to obtain the initial optical flow of each level through feature warping processing and cost volume calculation. A reliable calculation module is used to calculate the uncertainty map of the optical flow at each level based on the initial optical flow through an uncertainty estimation branch. The uncertainty map represents the reliability of the optical flow prediction at each pixel location. The weighted fusion module is used to calculate the fusion weight of the current layer using the uncertainty graph of the current layer and the fused uncertainty graph of the previous layer, and to use the fusion weight to perform weighted fusion of the optical flow of the current layer and the fused optical flow of the previous layer to obtain the fused optical flow of the current layer. The fusion update module is used to update the uncertainty graph of the current layer after fusion based on the fusion weight and the uncertainty graph of the current layer; The consistency calculation module is used to calculate the motion gradient consistency loss based on the difference in optical flow gradient between adjacent pyramid layers, which is used to constrain the optical flow to maintain the consistent gradient direction across layers. The loss construction module is used to construct the total loss function based on the endpoint error loss of the initial optical flow, the uncertainty loss of the uncertainty estimation branch, and the motion gradient consistency loss; The optical flow prediction module is used to train the optical flow estimation network using the total loss function to obtain an optimized optical flow estimation model, and then use the optimized optical flow estimation model to predict the optical flow of the image to be estimated.

[0014] To address the aforementioned technical problems, this application also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the aforementioned optical flow estimation method.

[0015] To address the aforementioned technical problems, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the aforementioned optical flow estimation method.

[0016] The optical flow estimation method, apparatus, computer device, and storage medium provided in this invention extract multi-scale features of an image through a pyramid network structure. Based on these multi-scale features, initial optical flow for each level is obtained through feature warping processing and cost volume calculation. Based on the initial optical flow, uncertainty maps of optical flow for each level are calculated through an uncertainty estimation branch. The uncertainty map represents the reliability of optical flow prediction at each pixel location. The fusion weight of the current layer is calculated using the uncertainty map of the current layer and the fused uncertainty map of the previous layer. The fusion weight is then used to perform weighted fusion of the current layer's optical flow and the fused optical flow of the previous layer to obtain the fused optical flow of the current layer. The fused uncertainty map of the current layer is updated based on the fusion weight and the current layer's uncertainty map. The motion gradient consistency loss is calculated based on the difference in optical flow gradient between adjacent pyramid layers to constrain the optical flow to maintain consistent gradient directions across layers. A total loss function is constructed based on the endpoint error loss of the initial optical flow, the uncertainty loss of the uncertainty estimation branch, and the motion gradient consistency loss. The optical flow estimation network is trained using the total loss function to obtain an optimized optical flow estimation model. The optimized optical flow estimation model is then used to predict the optical flow of the image to be estimated. The method employs an uncertainty estimation branch to dynamically evaluate the reliability of optical flow predictions at each level, and then weights and fuses multi-scale results accordingly, effectively suppressing error propagation and accumulation in occluded and textureless regions. Simultaneously, it forces the alignment of optical flow gradients between adjacent layers through motion gradient consistency constraint loss, ensuring smoothness and physical rationality at motion boundaries. The synergistic effect of these two methods addresses the shortcomings of traditional methods in multi-scale fusion robustness and cross-level consistency, thereby significantly improving the overall accuracy, robustness, and interpretability of optical flow estimation in complex scenes. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of the optical flow estimation method of this application; Figure 3 This is a structural example of the optical flow estimation network after adding the feature fusion module guided by optical flow uncertainty in this application; Figure 4 This is a schematic diagram of one embodiment of the optical flow estimation device according to this application; Figure 5 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0020] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Please see Figure 1 ,like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0023] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc.

[0024] Terminal devices 101, 102, and 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), laptops, and desktop computers, etc.

[0025] Server 105 can be a server that provides various services, such as a backend server that supports the pages displayed on terminal devices 101, 102, and 103.

[0026] It should be noted that the optical flow estimation method provided in this application embodiment is executed by the server, and correspondingly, the optical flow estimation device is set in the server.

[0027] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included. The terminal devices 101, 102, and 103 in this embodiment can specifically correspond to application systems in actual production.

[0028] Please see Figure 2 , Figure 2 This embodiment illustrates an optical flow estimation method, which is then applied to... Figure 1 Taking the server-side as an example, the details are as follows: S201: Extract multi-scale features of the image through a pyramid network structure, and based on the multi-scale features, obtain the initial optical flow of each level through feature warping processing and cost volume calculation.

[0029] In one specific optional implementation, the step of "extracting multi-scale features of the image through a pyramid network structure, and obtaining the initial optical flow of each level based on the multi-scale features through feature warping processing and cost volume calculation" is specifically implemented through the following process: First, a multi-scale feature pyramid is constructed. A pair of consecutive input images (denoted as source image I1 and target image I2) are fed into a shared-weight feature extraction backbone network (such as VGG-16). This network contains multiple downsampling layers, naturally outputting a set of feature maps with decreasing resolution. For example, a pyramid with 6 layers (L=6) can be constructed, where the first layer is the original image resolution (e.g., H×W), and the resolutions of layers 2 through 6 are halved sequentially (e.g., H / 2×W / 2, H / 4×W / 4, ..., H / 32×W / 32). Thus, we obtain the multi-scale feature sets for both the source image I1 and the target image I2.

[0030] Subsequently, a coarse-to-fine strategy was adopted, starting from the coarsest 6th layer and recursively calculating the initial optical flow of each layer from top to bottom.

[0031] For the l-th level of the pyramid currently being processed (l ranges from 6 to 1): Feature Warping: Using the initial optical flow estimate f_{l+1} calculated from the previous layer (l+1 layer) and upsampled by 2, a backward warping operation is performed on the feature map c2 of the target image I2 in the current layer l. Physically, this means that based on the optical flow prediction from the previous layer, the features of I2 are "pulled back" to a position as aligned as possible with the features of I1, thereby reducing the motion offset between the two frames in the current layer and preparing for subsequent fine-grained calculations.

[0032] Constructing the Cost Volume: The feature map c1 of the source image I1 and the warped target feature map c_warped are concatenated along the channel dimension to form a comprehensive feature tensor. This tensor is then input into a lightweight multi-layer convolutional network (i.e., the "Cost Volume Network"). The core function of this network is to calculate the matching cost (or correlation) between c1 and c_warped at each pixel location and for each possible small displacement, and encode it into a three-dimensional cost volume. This cost volume is essentially a dense matching possibility space, clearly representing the matching confidence corresponding to different motion vectors.

[0033] Decoding the initial optical flow: The cost volume data obtained above is input into another small convolutional neural network (optical flow decoder). This decoder learns to regress the optimal, dense pixel-level motion vector from the complex matching cost information, i.e., outputs the initial optical flow estimate f of the current l-th layer. l The optical flow f l It is a residual correction and refinement of the coarse-grained optical flow f_{l+1} of the previous layer.

[0034] Finally, through iterative steps from layer 6 to layer 1, the initial optimized optical flow estimate is obtained at layer 1 (highest resolution).

[0035] Through the above-described "feature pyramid-warping-cost volume" implementation, this method offers the following key advantages compared to the traditional approach of directly calculating optical flow at full resolution: Significantly improved estimation capability for large displacement motions: The core advantage of the pyramid structure lies in converting large pixel displacements into small displacements at a coarse level. For example, a large motion of 64 pixels is only represented as a tiny motion of 2 pixels at layer 6 (downsampled by 32 times), which is easily captured and matched by standard convolutional kernels. This coarse-to-fine estimation framework effectively resolves the contradiction between the limited receptive field of deep networks and large displacement estimation, enabling the network to reliably capture fast, large-scale object motions.

[0036] A balance between computational efficiency and estimation accuracy is achieved: the computational complexity of cost volume calculation is quadratically related to the search range. Directly constructing a cost volume covering large displacements at the original resolution is computationally infeasible. This embodiment estimates coarse-grained motion of large displacements at a low-resolution level (small feature map), and then refines tiny residuals by constructing only a small cost volume at a high-resolution level (large feature map), greatly reducing the overall computational overhead and making it possible to achieve high efficiency while maintaining high accuracy.

[0037] The robustness and accuracy of optical flow estimation are enhanced: the feature warping operation applies the prediction results of the previous layer to the current layer, realizing cross-layer prediction propagation and motion alignment. This allows the current layer to focus only on the residual motion that was not interpreted by the previous layer when calculating the cost volume, greatly simplifying the matching task of the current layer and reducing its sensitivity to image noise, illumination changes, and weak textures. As a result, the overall accuracy and robustness of the final optical flow estimation are improved layer by layer and stably.

[0038] In summary, this step, through its ingenious multi-scale design and iterative optimization architecture, provides an effective way to address the fundamental challenge of "large displacement, high accuracy, and high efficiency" in optical flow estimation, and lays a solid foundation for further applying uncertainty guidance and consistency constraints.

[0039] S202: Based on the initial optical flow, the uncertainty map of the optical flow at each level is calculated through the uncertainty estimation branch. The uncertainty map represents the reliability of the optical flow prediction at each pixel location.

[0040] In one specific optional implementation, step S202, "calculating the uncertainty map of optical flow at each level based on the initial optical flow through uncertainty estimation branches," is specifically implemented through the following process: For the initial optical flow f of the l-th layer of the pyramid, which has already been calculated l The uncertainty estimation process is as follows: Branch Input and Structure: The uncertainty estimation branch exists as a lightweight subnetwork. Its input is the initial optical flow f calculated by the current layer. l The cost volume characteristics that the time depends on and f l The sub-network itself typically consists of two or more consecutive 3x3 convolutional layers, followed by a sigmoid activation function as the output layer. The entire branch is designed to be extremely lightweight, adding only a small amount of computational overhead.

[0041] Uncertainty graph calculation: The aforementioned sub-network encodes and processes the input features, ultimately outputting a graph that corresponds to the initial optical flow f of the current layer. l Uncertainty graph U with exactly the same resolution l The value of each pixel in this image ranges from [0,1], guaranteed by the Sigmoid function. Its value has a clear physical meaning: the closer the value is to 0, the higher the confidence level of the optical flow prediction at that location, and the more reliable it is; the closer the value is to 1, the higher the uncertainty of the prediction at that location, and the less reliable it is.

[0042] Mapping of uncertainty sources: High uncertainty (values ​​close to 1) typically occurs in the following regions: Occlusions: A pixel appears in only one frame and is not visible in another frame, resulting in an inherently blurred matching cost.

[0043] Weakly textured or repetitive textured regions (Textureless / Repeated Patterns): Lacking distinctive features, the matching cost cannot provide a unique solution, and there is matching ambiguity.

[0044] Motion boundaries: Motion is discontinuous at the edges of an object, belonging to ambiguous regions where a single pixel may correspond to multiple motion hypotheses.

[0045] Regions with drastic changes in illumination: The appearance changes are huge, leading to feature matching failure.

[0046] Supervision Signal (Training Phase): To enable the network to accurately predict uncertainty, a dedicated uncertainty loss, L_uncertainty, is introduced during the training phase for supervision. This loss function is designed to encourage a behavior where the uncertainty value approaches 1 (high uncertainty) where the optical flow prediction error is large, and approaches 0 (low uncertainty) where the prediction error is small. A common implementation is based on negative log-likelihood. The first term of this loss function means that if a pixel has a large prediction error but also high uncertainty, the penalty is smaller; the second term prevents the network from "lazying" by predicting all pixels as high uncertainty.

[0047] By introducing the aforementioned uncertainty estimation branch and generating an uncertainty graph, the method in this embodiment achieves the following core beneficial effects: This enables a quantitative perception of prediction reliability, providing crucial information for decision-making: the uncertainty map provides a pixel-by-pixel reliability map corresponding to the optical flow field. This fundamentally changes the traditional black-box model of optical flow estimation, which only provides predictions without indicating confidence levels. Users can clearly identify which parts of the prediction results are reliable and which parts are at risk (such as occluded boundaries or weakly textured areas), providing vital confidence information for subsequent advanced applications (such as decision-making systems for autonomous driving, robot vision navigation, and motion vector selection in video compression), thus improving the overall system's safety and robustness.

[0048] The intelligence and robustness of multi-scale fusion are significantly enhanced: the uncertainty map directly serves the subsequent uncertainty-guided fusion module. Fusion weights are dynamically generated based on the uncertainty difference between the current layer and the previous layer. This means that when the current layer's prediction at a certain location is more certain (lower uncertainty) than the previous layer's, the fusion algorithm assigns a higher weight to the current layer's prediction; conversely, it places more trust in the previous layer's correction results. This dynamic fusion strategy, based on "whoever is more reliable is chosen," effectively avoids error accumulation and propagation caused by incorrect corrections in low-confidence regions, thereby significantly improving the overall accuracy of the final optical flow field, especially in the aforementioned challenging regions.

[0049] Improved model interpretability and debugging efficiency: Uncertainty graphs, as a visual output, greatly enhance model interpretability. Developers can intuitively observe which scenarios and regions the model faces estimation difficulties, thus enabling targeted analysis of the root causes of problems (whether it's insufficient data, model capacity issues, or algorithmic defects), accelerating the iterative optimization and debugging process of the model.

[0050] In summary, the introduction of the uncertainty estimation branch not only enables the model to self-evaluate its prediction quality, but more importantly, it transforms this perception ability into a control signal that dynamically guides the optical flow generation process, thereby fundamentally improving the performance, reliability, and practicality of the optical flow estimation system.

[0051] In an optional implementation, step S202, calculating the uncertainty map of optical flow at each level through the uncertainty estimation branch, includes: After the cost volume of each layer of the pyramid network structure, a light quantum network consisting of convolutional layers and activation functions is connected; The uncertainty map is output by a light quantum network with the same resolution as the optical flow of the current layer. The larger the uncertainty value, the less reliable the optical flow prediction at that location.

[0052] In one specific optional implementation, the process of calculating the optical flow uncertainty map of each level through the uncertainty estimation branch in step S202 is as follows: The core of this step is to equip each layer of the pyramid network with a "reliability evaluator" for optical flow prediction. At each layer, after the main network completes the initial optical flow prediction for that layer through feature warping and cost volume calculation, the rich intermediate information generated during its calculation process—especially the cost volume features—is immediately captured and fed into a lightweight sub-network.

[0053] This lightweight quantum network is directly attached to the cost body module of the main network. Its structure is very ingenious, typically consisting of only two consecutive convolutional layers followed by a sigmoid activation function as the output layer. The entire sub-network is designed according to the principle of lightweightness, aiming to complete the critical task with minimal computational overhead.

[0054] This sub-network efficiently encodes and analyzes the input cost volume features, ultimately outputting an uncertainty map with the exact same resolution as the initial optical flow prediction of the current layer. Each pixel value in this map is a continuous numerical value between 0 and 1, with a clear physical meaning: the closer the value is to 0, the higher the confidence level of the optical flow prediction at that location, and the more reliable it is; the closer the value is to 1, the more unreliable the prediction at that location is, and there is a high degree of uncertainty. High uncertainty typically corresponds precisely to occluded areas, weakly textured areas, motion boundaries, and areas with drastic changes in illumination in the image.

[0055] By introducing the aforementioned uncertainty estimation branch and generating an uncertainty graph, the method in this embodiment achieves the following core beneficial effects: This approach achieves pixel-by-pixel reliability quantification of prediction results, providing crucial metadata for subsequent decision-making. Traditional optical flow estimation methods only output an optical flow field without specifying the confidence level of each prediction. This embodiment, through an uncertainty map, assigns a continuous confidence score to each pixel of the optical flow field. This fundamentally changes the output mode of optical flow estimation, transforming it from a "black box" point estimate into an interpretable prediction with accompanying uncertainty information. This reliability map is crucial for any downstream application; for example, autonomous driving systems can use it to ignore unreliable motion estimates in high-uncertainty regions, thereby making safer decisions.

[0056] Providing intelligent decision-making support for multi-scale fusion is crucial for achieving precise error control: one of the most important roles of the uncertainty map is to directly drive subsequent uncertainty-guided fusion modules. The fusion weights are calculated entirely based on the difference between the uncertainty maps of the current layer and the previous layer. This enables the system to make pixel-level intelligent decisions: at the current position, should it trust the new prediction from the current layer or place more trust in the historical estimate passed down from the previous layer? This dynamic fusion strategy based on "whoever is more reliable is trusted" fundamentally avoids the accumulation and propagation of errors caused by error correction in low-confidence regions, and is one of the core mechanisms for achieving high-precision optical flow estimation.

[0057] It enhances the model's inherent interpretability and debugging efficiency, accelerating the R&D process: the uncertainty graph, as a visual output, acts like a "model self-check report," clearly indicating which scenarios and regions the model faces perceptual difficulties. R&D personnel can intuitively observe problem areas (such as which occlusions are not correctly identified), enabling them to perform targeted data augmentation, model adjustments, or algorithm improvements, greatly improving the efficiency of model iteration and debugging, and shortening the development cycle.

[0058] S203: Using the uncertainty map of the current layer and the uncertainty map of the previous layer after fusion, calculate the fusion weight of the current layer, and use the fusion weight to perform weighted fusion of the optical flow of the current layer and the fused optical flow of the previous layer to obtain the fused optical flow of the current layer.

[0059] In a specific optional implementation, step S203, "using the uncertainty map of the current layer and the uncertainty map after fusion of the previous layer, calculating the fusion weight of the current layer, and using the fusion weight to perform weighted fusion of the optical flow of the current layer and the fused optical flow of the previous layer to obtain the fused optical flow of the current layer," is specifically implemented through the following process: This step is the core of the uncertainty-guided fusion module, and its purpose is to intelligently fuse optical flow predictions from two different sources: the newly calculated initial optical flow f from the current layer l. lThe fused optical flow Up(f_fused^{l+1}) is obtained by fusing and upsampling the previous layer l+1. The fusion is based on the reliability of the predictions of both layers.

[0060] Calculate the fusion weights: Input: Uncertainty graph U of the current layer l The uncertainty graph U_fused^{l+1} after fusion with the previous layer (already aligned with the current layer resolution through the upsampling Up(·) operation).

[0061] Procedure: First, calculate the difference between the two uncertainty plots: ΔU = Up(U_fused^{l+1}) - U l This difference ΔU is key: If ΔU>0, it means that the prediction uncertainty of the current layer l at this pixel is lower than that of the previous layer, that is, the prediction of the current layer is more reliable.

[0062] If ΔU<0, it means that the prediction uncertainty of the current layer l at this pixel is higher than that of the previous layer, that is, the prediction of the previous layer is more reliable.

[0063] Output: Input the difference into a learnable scaling parameter α and a sigmoid activation function to generate the final pixel-to-pixel fused weight map W. l The Sigmoid function constrains the weight values ​​to the range (0, 1). l The numerical value intuitively represents the predicted optical flow f for the current layer. l The level of trust.

[0064] Perform weighted fusion: Using the calculated fusion weight graph W l For the current layer optical flow f l The current layer's fused optical flow is obtained by performing a pixel-wise weighted summation of the upsampled fused optical flow Up(f_fused^{l+1}) from the previous layer. In regions where the reliability of both is similar, W l A value close to 0.5 achieves a smooth transition and blending.

[0065] Through the aforementioned uncertainty-guided dynamic fusion mechanism, the method in this embodiment, compared to the simple, fixed-strategy fusion method in traditional PWC-Net, brings the following key benefits: The most significant benefit is the achievement of adaptive and intelligent multi-scale information fusion, which effectively suppresses error accumulation. Traditional methods, when using upsampling of the optical flow from the rough layer to guide the fine layer, directly transmit and amplify the errors it carries. This embodiment addresses this by using fusion weight W... lThis system achieves pixel-level fine-grained control. It can automatically identify unreliable regions in the results passed from the previous layer (such as blurring errors caused by low resolution) and trust and adopt the higher-resolution correction results from the current layer in these regions. Simultaneously, it can also identify unreliable predictions caused by the limitations of the current layer itself (such as occlusion or weak texture) and resolutely retain the relatively more reliable coarse-grained estimates from the previous layer in these regions. This "selective adoption" mechanism fundamentally prevents the vicious cycle and accumulation of errors between pyramid levels, significantly improving the accuracy of the final prediction.

[0066] This enhances the robustness of predictions in challenging regions: occlusion, weak textures, and moving boundaries are inherently difficult areas for optical flow estimation and also regions with high uncertainty. This fusion strategy automatically marks these regions as "low-confidence" using an uncertainty map and relies more on contextual information from other levels for "correction" or "protection" during fusion, rather than blindly correcting errors. This makes the final optical flow field more stable and reasonable in these challenging regions, avoiding obvious outliers or boundary distortions.

[0067] Improved the rationality and interpretability of the optical flow estimation process: fusion of weight map W l It is itself a highly informative visualization. It clearly demonstrates the decision-making process of the network during fusion: at different locations in the image, does the network ultimately place more trust in the detailed corrections of the current layer or in the general outline of the previous layer? This provides developers with an intuitive window to analyze model behavior and diagnose the causes of failures, greatly enhancing the interpretability and transparency of the entire system.

[0068] In summary, this uncertainty-guided fusion step is not a simple weighted operation, but rather an intelligent decision-making unit that dynamically schedules and integrates the advantages of information flows at different scales, thereby outputting more reliable and accurate optical flow estimation results in complex scenarios.

[0069] In an optional implementation, step S203, calculating the fusion weights of the current layer and performing weighted fusion, includes: Upsample the uncertainty graph after fusion at the previous level and subtract it from the uncertainty graph at the current level. The difference results are input into the learnable sensitivity parameter and the Sigmoid activation function to obtain the fusion weights of the current layer; Using the fusion weights, the current layer optical flow and the upsampled fused optical flow of the previous layer are weighted and summed to obtain the current layer fused optical flow.

[0070] Specifically, the process of calculating the fusion weights and performing weighted fusion in step S203 is as follows: This step is the core of the uncertainty-guided fusion module. Its core idea is to dynamically decide which prediction to believe based on the reliability difference between the current layer and the previous layer.

[0071] First, two key inputs are obtained: one is the uncertainty map fused from the previous level (layer l+1), which represents the historical reliability assessment derived from information from all previous levels; the other is the initial uncertainty map just calculated for the current layer (layer l), which reflects the immediate reliability of the new round of predictions at this layer. The historical uncertainty map is upsampled to the resolution of the current layer to make it spatially perfectly aligned with the immediate uncertainty map.

[0072] Next, a reliability difference comparison is performed: the upsampled historical uncertainty map is subtracted from the current layer's instantaneous uncertainty map pixel by pixel. This difference is crucial: a negative result indicates that the current layer's prediction at that pixel is more reliable than the historical assessment; a positive result indicates that the historical assessment is more reliable than the current layer's prediction.

[0073] Next, this difference value is input into a decision unit consisting of a learnable sensitivity parameter and a sigmoid activation function. The sensitivity parameter, automatically optimized by the network, controls the system's sensitivity to reliability differences; the sigmoid function maps this difference value to a fusion weight between 0 and 1. This weight directly represents the level of confidence in the optical flow prediction of the current layer: the closer the weight is to 1, the more confident the current layer is; the closer the weight is to 0, the more confident the historical results are.

[0074] Finally, intelligent weighted fusion is performed: using the calculated fusion weights, the initial optical flow of the current layer and the upsampled fused optical flow of the previous layer are summed pixel by pixel. The final output is a new fused optical flow of the current layer that incorporates current detail corrections and historical coarse-grained information.

[0075] Through the aforementioned uncertainty-guided dynamic fusion decision-making mechanism, this embodiment achieves the following core beneficial effects: This system achieves pixel-level adaptive fusion, fundamentally suppressing error propagation. Traditional methods typically employ fixed, heuristic strategies (such as direct replacement) for inter-layer fusion, failing to address the complexities of different regions within an image. This embodiment achieves unprecedented pixel-level fine-grained control by calculating reliability differences and generating dynamic weights. The system automatically identifies and adopts more reliable corrections made by the current layer; simultaneously, it identifies less reliable predictions made by the current layer due to its own limitations and decisively rejects corrections, retaining the more reliable results from the previous layer. This "choosing the best and following it" mechanism precisely blocks the propagation and amplification of errors during the coarse-to-fine optimization process, providing a core guarantee for improving final accuracy.

[0076] This significantly enhances the robustness and reasonableness of predictions in challenging regions: in areas with high uncertainty such as occlusion and weak texture, network predictions are often unreliable. Guided by the uncertainty map, this fusion strategy automatically reduces the weight of the current layer's predictions in these regions, relying more on relatively stable historical predictions derived from coarse-grained context. This prevents the network from making erroneous "guessing" corrections in these difficult regions, resulting in a more stable final optical flow field in these areas, avoiding severe outliers, and significantly improving the reasonableness and usability of the overall output.

[0077] By endowing the model with human-like decision-making intelligence, the transparency and interpretability of the process are enhanced: the fused weight graph is essentially a "decision map," clearly showing how the network makes choices at each position in the image—whether to believe new evidence (the current layer) or past experience (the previous layer). This greatly enhances the transparency and interpretability of the model's internal decision-making process, providing developers with an intuitive and powerful tool for analyzing model behavior and locating the causes of failures. This is not merely a computational process, but a process that embodies intelligent decision-making.

[0078] S204: Update the uncertainty graph of the current layer after fusion based on the fusion weights and the uncertainty graph of the current layer.

[0079] In an optional implementation, step S204, updating the uncertainty graph after fusion of the current layer, includes: Upsample the uncertainty graph after fusion at the previous level; Compare the uncertainty map of the current layer with the uncertainty map of the previous layer after upsampling pixel by pixel, and take the minimum value as the uncertainty map after fusion of the current layer.

[0080] In one specific optional implementation, the process of updating the fused uncertainty map in step S204 is as follows: This step aims to maintain a continuously evolving record of uncertainty across pyramid levels, with the core task of tracking the most unreliable prediction level in history for each pixel location during the coarse-to-fine optimization process.

[0081] First, the fusion uncertainty map of the previous layer (i.e., layer l+1) that has already been fused is upsampled to make its spatial resolution consistent with that of the current layer l, thus achieving pixel-level alignment.

[0082] Subsequently, a crucial uncertainty minimization fusion operation is performed: the initial uncertainty map of the current layer, calculated by the uncertainty branch freshly, is compared pixel-by-pixel with the upsampled fused uncertainty map of the previous layer. For each pixel location in the image, the system selects the one with the smaller value from these two uncertainty maps.

[0083] Ultimately, the selected minimum value is determined as the new fusion uncertainty map for the current layer. This means that if a pixel was previously marked as high uncertainty at a lower (coarse) level, it will conservatively be recorded as "high uncertainty" even if it is predicted as low uncertainty at the current, finer level; conversely, if a pixel has performed reliably (low uncertainty) at all previous levels, it will remain in a low uncertainty state. This process effectively constructs a "chain of unreliability memories" spanning all scales.

[0084] Through the above-described update mechanism, the method of this embodiment brings the following core beneficial effects: The most critical and beneficial effect is the establishment of a persistent uncertainty memory to prevent the premature propagation of erroneous confidence. In the coarse-to-fine estimation process, erroneous predictions at the coarse level can propagate and affect the computation at the fine level. This step creates an "unreliability labeling" mechanism by permanently recording the worst (maximum) uncertainty value of each pixel across all past levels. This ensures that once a pixel is identified as unreliable at a certain scale (e.g., due to occlusion or blurring), this label is retained in all subsequent finer scales, constantly reminding the fusion module to handle this location with caution. This effectively prevents the system from becoming prematurely overconfident in later levels due to seemingly reliable local computations, avoiding the amplification of errors during refinement and greatly improving the robustness of the final result.

[0085] It provides a stable and reliable decision-making basis for uncertainty-guided fusion: fusion weight (W) l The calculation of the fusion uncertainty graph relies on comparing the current layer with the uncertainty of the previous layer. The accuracy of the updated fusion uncertainty graph in this step, representing the uncertainty of the previous layer, is crucial. Using a minimum-value strategy, this graph represents a historically validated and most conservative reliability assessment. The fusion decision (W) based on this is then made. l ∝ (U_fused^{l+1} - U l It becomes more reliable and stable because it is based on global optimal information rather than local instantaneous information, enabling the fusion process to make the safest choice under any circumstances.

[0086] The interpretability and credibility of the model output are enhanced: the final output fused uncertainty map is a complete map recording all "problem regions" throughout the entire pyramid optimization process. Users can not only know where the model is ultimately uncertain, but also be certain that these regions represent fundamental difficulties that exist from the beginning, rather than instantaneous judgments at the current layer. This greatly improves the credibility and practical value of the uncertainty map as a reliability indicator, providing extremely valuable safety information for downstream applications that make decisions based on optical flow (such as autonomous driving and robot navigation).

[0087] S205: Based on the difference in optical flow gradient between adjacent pyramid layers, calculate the motion gradient consistency loss to constrain the optical flow to maintain consistent gradient direction across layers.

[0088] In an optional implementation, step S205, calculating the motion gradient consistency loss, includes: Calculate the gradient map of the optical flow between two adjacent layers after upsampling and alignment; The differences between gradient maps are calculated and dynamically weighted based on optical flow amplitude; The motion gradient consistency loss is obtained by summing the weighted differences of all adjacent levels in the pyramid.

[0089] In one specific optional implementation, the process of calculating the motion gradient consistency loss in step S205 is as follows: This step aims to ensure the consistency and rationality of optical flow prediction in spatial structure by imposing cross-level constraints. First, two adjacent pyramid levels (e.g., level l and level l+1) are selected. Since the optical flow field resolutions of the two levels are different, the lower-resolution optical flow field of level l needs to be upsampled to perfectly align its spatial dimensions with those of level l+1. Next, the gradient maps of these two aligned optical flow fields are calculated separately. These gradient maps reflect the rate of change and direction of the optical flow field at each pixel location, clearly depicting the structural information of the motion boundary and the flow field.

[0090] Next, the pixel-by-pixel difference between the two gradient maps is calculated. This difference quantifies the degree of inconsistency in the optical flow structure of adjacent layers at corresponding locations. Subsequently, based on the amplitude information of the optical flow of the previous layer (layer l) after upsampling, the difference at each pixel location is dynamically weighted. Specifically, regions with small optical flow amplitudes (such as static backgrounds) are assigned smaller weights to reduce their contribution to the total loss; regions with large optical flow amplitudes (such as moving objects) are assigned larger weights to emphasize their consistency. Finally, the weighted differences of all adjacent layer pairs (such as l and l+1, l+1 and l+2, etc.) across all pixels are summed, and the resulting scalar value is the motion gradient consistency loss. This loss, as part of the overall training objective, directly guides the network to learn its ability to maintain prediction consistency across layers.

[0091] By introducing and calculating the motion gradient consistency loss described above, the method in this embodiment brings the following core benefits: This invention effectively resolves the structural conflicts in cross-level optical flow prediction and enhances physical plausibility: In traditional pyramid models, each level optimizes optical flow independently, lacking collaborative constraints and prone to directional inconsistencies in regions such as object edges. This invention explicitly injects the strong physical prior of "cross-level structural consistency" into the training process by forcibly requiring gradient maps of adjacent levels to remain similar. This ensures that the coarse-to-fine optimization process is harmonious and coherent, with the coarse motion profiles provided by lower levels consistent with the fine correction directions added by higher levels. The resulting optical flow field has greater physical meaning in its overall structure, avoiding anomalous predictions that violate motion continuity.

[0092] Significantly enhanced estimation accuracy and visual quality at motion boundaries: Motion boundaries are a core challenge in optical flow estimation, characterized by dramatic gradient changes and the potential for prediction discrepancies between different layers. This loss function directly compares and constrains the differences in gradient maps, forcing the network to reach cross-layer consensus in this critical region. Combined with a dynamic weighting mechanism for optical flow amplitude, this loss function pays particular attention to boundary regions where real motion occurs, effectively reducing blurring, jagged artifacts, or erroneous displacements at the boundaries. This results in clearer and more accurate outlines of moving objects, greatly improving the visual quality of the optical flow field and the usability of downstream tasks.

[0093] As an effective regularizer, the motion gradient consistency loss enhances the model's generalization ability: it is essentially a structural constraint that penalizes unstable predictions that "jump back and forth" between pyramid levels. By minimizing this loss, the network learns a more robust intrinsic representation to input variations, reducing the risk of overfitting to specific noise or patterns in the training data. This allows the trained model to exhibit stronger generalization performance and stability when facing unknown scenarios, especially new data containing complex motions and occlusions.

[0094] In one alternative implementation, dynamic weighting based on optical flow amplitude includes: Based on the amplitude value of the optical flow after upsampling from the previous layer, the weighting coefficient is calculated using an exponential function. The smaller the optical flow amplitude, the smaller the weighting coefficient.

[0095] In one specific optional implementation, the "dynamic weighting based on optical flow amplitude" step is achieved through the following process: This step is crucial in calculating the motion gradient consistency loss, aiming to impose differentiated constraints on regions with varying motion intensities. First, the fused optical flow field from the previous layer (layer l), after upsampling, is obtained and aligned with the resolution of the current layer (layer l+1). Next, the amplitude of the optical flow vector at each pixel location in this field is calculated, representing its motion velocity. Then, this amplitude value is input into an exponential function for calculation, outputting a weighting coefficient between 0 and 1. Finally, this weighting coefficient is used to scale the contribution of the current pixel to the consistency loss: the smaller the optical flow amplitude (indicating the region is nearly stationary or moving slowly), the smaller the calculated weighting coefficient, and the lighter its proportion in the consistency loss; conversely, the larger the optical flow amplitude (indicating intense motion in the region), the closer the weighting coefficient is to 1, and the gradient difference at that location will be sufficiently constrained. This mechanism enables adaptive and refined control of different motion regions in the image.

[0096] Through the aforementioned dynamic weighting mechanism based on optical flow rate, the method in this embodiment achieves the following key beneficial effects: This approach avoids excessive constraints on static and low-speed regions, improving the rationality of the loss function. In optical flow fields, large background areas and static objects typically exhibit motion vectors with minimal or zero amplitude. Imposing gradient consistency constraints of equal strength on these regions as on high-speed motion regions would not only introduce unnecessary optimization pressure but may also obscure the true motion boundaries. This weighted strategy significantly reduces the penalty intensity for these low-amplitude regions, allowing network training to focus more on regions with significant motion. This prevents the loss function from getting trapped in local optima and improves the model's practicality in real-world scenarios.

[0097] This method enhances the precise constraints on high-speed motion and motion boundaries, improving estimation quality: Large optical flow amplitudes typically correspond to rapidly moving objects or boundaries between different motion modes, which are challenging and critical areas for optical flow estimation. By assigning higher weight coefficients to these regions, this method ensures sufficient and effective constraints on gradient consistency in these important areas. This strongly guarantees that optical flow predictions at different levels maintain structural coherence and physical consistency at boundaries where objects are moving rapidly or where motion discontinuities exist, thus significantly improving the accuracy and visual quality of the final optical flow field in these challenging regions.

[0098] Enhanced stability and convergence of the training process: This dynamic weighting mechanism, as a built-in "attention" model, automatically balances the magnitude of gradient contributions from different regions in the loss function. It prevents large but insignificant gradients from low-amplitude static regions from overwhelming important gradient signals from key motion regions, providing the optimizer with a clearer and more effective learning direction, thereby accelerating the model's convergence process and contributing to the training of a more powerful and stable optical flow estimation model.

[0099] S206: Construct the total loss function based on the endpoint error loss of the initial optical flow, the uncertainty loss of the uncertainty estimation branch, and the motion gradient consistency loss.

[0100] In one alternative implementation, the total loss function is a weighted sum of endpoint error loss, uncertainty loss, and motion gradient consistency loss, wherein the uncertainty loss is constructed based on the weighted error between the uncertainty map of each layer and the true optical flow.

[0101] Specifically, the total loss function is constructed and calculated in the following manner: The total loss function L_total is a multi-objective optimization function, consisting of a weighted sum of three core loss terms, which provides comprehensive supervision signals to the network during training.

[0102] Endpoint Error Loss (L_EPE) - Accuracy Guarantee Item: Calculation method: This loss is the most fundamental and direct supervisory signal in optical flow estimation tasks. It calculates the optical flow predicted at each layer l of the network (usually the initial optical flow f). l Or, the average Euclidean distance (L2 norm) between the fused optical flow (f_fusedˡ) and the true optical flow (f_gtˡ) downsampled to the corresponding resolution. For a pyramid with L layers, the formula is: L_EPE = Σ_{l=1}^{L} γ^{Ll} · || f_predictedˡ - f_gtˡ ||2 where γ is a discount factor (usually set to 0.8) used to balance the contribution of different layers to the total loss, giving greater weight to the prediction error of higher resolution (lower layers).

[0103] Function: It directly minimizes the pixel-level deviation between the predicted value and the true value, and is the fundamental driving force for ensuring the absolute accuracy of optical flow estimation.

[0104] Uncertainty loss (L_uncertainty) - Reliability learning term: Calculation method: This loss term is constructed based on the formula described in the appendix, and its form is similar to a weighted mean squared error. For each level l, its uncertainty loss is: L_uncertaintyˡ = Σ_{i,j} [ 0.5 · U l (i,j)· || f l (i,j) - f_gtˡ(i,j) ||2² + log(U l (i,j)) ], where (i,j) are pixel coordinates, U l It is the uncertainty graph output by the uncertainty estimation branch.

[0105] Mechanism of action: The design of this loss function is extremely ingenious, with its two components mutually constraining each other: First item: Encourage the network to reduce prediction error (f) l For positions with larger values ​​of -f_gt), assign a higher uncertainty value U. l This reduces the overall penalty for that item. It's equivalent to telling the network, "If a pixel is difficult to predict, you can admit you're uncertain, and I won't punish you severely." Second: Prevent the network from transmitting all U... l All values ​​are set to very large values ​​to "cheat" and minimize the first term. It imposes a continuous, negative penalty on high uncertainty (because log(x) is negative when x<1), encouraging the network to minimize uncertainty values ​​where accurate prediction is possible.

[0106] Final term: Summation over all levels L_uncertainty = Σ_{l} L_uncertaintyˡ.

[0107] Motion gradient consistency loss (L_consist) - structural constraint term: Calculation method: This loss term forces that the optical flow predictions of adjacent pyramid levels (such as level l and level l+1) have similar gradient structures.

[0108] Total loss synthesis: The three loss terms are weighted by preset hyperparameters λ1 and λ2 to obtain the final total loss function: L_total = L_EPE + λ1 · L_uncertainty + λ2 · L_consist. By minimizing L_total using the gradient descent algorithm, the main network for optical flow prediction and the uncertainty estimation branch can be optimized simultaneously.

[0109] In this embodiment, through the carefully designed multi-component loss function described above, the following significant beneficial effects are achieved during the model training process: This approach achieves a balance between accuracy and robustness: Traditional single L_EPE loss drives the model to blindly pursue the average accuracy of all pixels, but at the cost of generating huge, difficult-to-optimize errors in high-uncertainty regions (occlusion, weak texture) and potentially leading to overfitting. The L_uncertainty term introduced in this embodiment allows the model to distinguish between "easy" and "difficult" samples. The model learns to be "skeptical" of difficult regions, thus avoiding overconfident but erroneous predictions in these areas and focusing more optimization effort on regions that can be reliably predicted. This mechanism greatly enhances the model's generalization ability and robustness in complex real-world scenes, making its output not only accurate but also "honest" and reliable.

[0110] By incorporating structured prior knowledge into the model, the physical plausibility is enhanced: the L_consist loss term explicitly injects the strong physical prior that the motion field should maintain structural consistency across scales into the training process. It does not directly constrain optical flow values, but rather constrains their gradient changes. This effectively avoids directional contradictions between predictions at different levels, ensuring that the coarse-to-fine optimization process produces a harmonious, smooth, and clearly defined optical flow field. This reduces outliers that violate the laws of physical motion in the results, improving visual quality and reliability for downstream tasks.

[0111] At the same time, a "win-win" situation of collaborative optimization was achieved: the three loss mechanisms did not work independently, but rather through deep collaboration. L_uncertainty helps the model identify unreliable regions, which are often the main source of error in L_EPE. Reducing the over-penalty for these regions allows L_consist to more effectively apply the correct structural constraints in them.

[0112] The smoothness constraint provided by L_consist, in turn, provides a more reasonable context for estimating L_uncertainty, since the uncertainty estimate of a physically unreasonable optical flow field is itself unreliable.

[0113] This synergistic effect guides the network to learn an intrinsically consistent representation, and the final trained model outperforms any model trained by a single loss or a simple combination of losses in all metrics, achieving a 1+1+1>3 effect.

[0114] In summary, this composite loss function is the key to the superior performance of this patented method. By integrating accuracy targets, uncertainty learning, and structural constraints, it successfully trains an optical flow estimation model that is not only more accurate but also smarter and more reliable.

[0115] S207: Train the optical flow estimation network using the total loss function to obtain the optimized optical flow estimation model, and use the optimized optical flow estimation model to predict the optical flow of the image to be estimated.

[0116] In one specific optional implementation, step S207, "training the optical flow estimation network using the total loss function to obtain an optimized optical flow estimation model, and using the optimized optical flow estimation model to predict optical flow in the image to be estimated," is specifically implemented through the following process: Training phase (model optimization): Data preparation: Use datasets containing a large number of consecutive frame images and their ground truth optical flow fields (such as Sintel, KITTI) as training samples. Each training sample is a triple (I1, I2, F_gt).

[0117] Forward propagation: The image pair (I1, I2) is input into the optical flow estimation network. The network executes all the aforementioned steps (S201-S206), namely, it sequentially goes through pyramid feature extraction, initial optical flow calculation, uncertainty estimation, uncertainty-guided fusion, etc., and finally outputs the predicted optical flow {f} for all levels of the pyramid. l}, fused optical flow {f_fusedˡ} and uncertainty graph {U l}

[0118] Loss Calculation: Based on the network output and the actual optical flow F_gt, the total loss function L_total is calculated. This function is a weighted sum of multiple losses: Endpoint error loss (L_EPE): Calculates the predicted optical flow f at each level. l The Euclidean distance between the optical flow and the upsampled true optical flow serves as the basis for optical flow accuracy monitoring.

[0119] Uncertainty loss (L_uncertainty): According to the formula ∑ (0.5 * U l ·||f l -f_gt||²+log(U l )) Calculation, forcing uncertainty graph U l It accurately reflects the distribution of prediction errors.

[0120] Motion gradient consistency loss (L_consist): Calculates the difference between optical flow gradients of adjacent layers and applies consistency constraints.

[0121] Finally, L_total = L_EPE + λ1 * L_uncertainty + λ2 * L_consist, where λ1 and λ2 are pre-defined hyperparameters used to balance the contributions of each loss.

[0122] Backpropagation and optimization: The gradient of L_total with respect to all network parameters (including the weights of the backbone feature extraction network, optical flow decoder, uncertainty estimation branch, etc.) is calculated using a gradient descent algorithm (such as Adam). Through backpropagation, the gradients are propagated back and the network parameters are updated, aiming to minimize the total loss. This process is iterated multiple times (epochs) across the entire training set until the model performance converges, resulting in a fully optimized optical flow estimation model.

[0123] Inference phase (optical flow prediction): Deployment Model: Deploy the optimized optical flow estimation model, which has been trained and has fixed parameters, into the actual application environment.

[0124] Make predictions: For any new, unseen consecutive image pair (I1_new, I2_new), simply input it into the model.

[0125] Forward computation: The model performs the same forward computation process as during training (but without calculating the loss). After a series of steps including pyramid processing, warping, cost volume calculation, and uncertainty-guided fusion, the network finally outputs the highest resolution (layer 1) fused optical flow field f_fused¹.

[0126] Output: The output f_fused¹ is the final, high-precision predicted optical flow map of the model for the input image pair, which can be directly used for subsequent tasks, such as target tracking, video frame interpolation, motion estimation in autonomous driving, etc.

[0127] Through the end-to-end training and inference process that integrates multiple innovation losses, the method in this embodiment brings the following core beneficial effects: This approach achieves collaborative optimization, resulting in a model with significantly improved overall performance. Traditional training methods rely solely on endpoint error loss (L_EPE), leading to a tendency for the model to learn an "average" solution, which struggles to perform well in challenging regions. This embodiment achieves collaborative optimization by combining L_uncertainty and L_consist with L_EPE to form the total loss function. This forces the model to simultaneously pursue high accuracy (L_EPE) while learning to assess its own uncertainty (L_uncertainty) and maintain logical consistency in predictions (L_consist). This multi-objective optimization process guides the network to learn more general and robust feature representations and inference capabilities, ultimately resulting in a model that outperforms models trained using traditional methods across various performance metrics.

[0128] This ensures the innovative modules function effectively in inference: the uncertainty estimation branch and consistency constraints are not merely auxiliary tools designed for training. The total loss function ensures these modules learn to perform their intended tasks: one produces an accurate uncertainty map, and the other effectively regulates the optical flow field structure. During inference, the uncertainty-guided fusion module relies on the high-quality uncertainty map to make correct decisions, while the effect of the consistency constraints is implicitly reflected in the network parameters, ensuring that the output naturally satisfies cross-level consistency. Therefore, the optimization results from the training phase are fully and seamlessly reflected in the final model's inference performance.

[0129] Meanwhile, this embodiment provides an end-to-end, efficient solution: the entire system—from input image to output optical flow—is a complete deep learning model that can be trained and deployed end-to-end. Users do not need to design complex post-processing logic or decision-making processes during inference. After one training, the resulting model can be directly used for efficient inference, while enjoying the benefits of high accuracy, high robustness, and good consistency, greatly enhancing its practical value in real-time applications (such as autonomous driving and robot navigation).

[0130] In summary, this training and prediction step represents the final integration and embodiment of the technical solution in this embodiment. Through a carefully designed joint loss function, it deeply integrates and optimizes multiple innovative modules, ultimately producing an optical flow estimation model that is not only more accurate but also more intelligent and reliable, providing a powerful end-to-end solution for addressing core challenges in computer vision.

[0131] In this embodiment, multi-scale features of the image are extracted through a pyramid network structure. Based on these multi-scale features, initial optical flow at each level is obtained through feature warping processing and cost volume calculation. Based on the initial optical flow, uncertainty maps of the optical flow at each level are calculated through an uncertainty estimation branch. The uncertainty map represents the reliability of optical flow prediction at each pixel location. The fusion weight of the current layer is calculated using the uncertainty map of the current layer and the fused uncertainty map of the previous layer. The fusion weight is then used to perform weighted fusion of the current layer's optical flow and the fused optical flow of the previous layer to obtain the fused optical flow of the current layer. The fused uncertainty map of the current layer is updated based on the fusion weight and the current layer's uncertainty map. The motion gradient consistency loss is calculated based on the difference in optical flow gradient between adjacent pyramid layers to constrain the optical flow to maintain consistent gradient directions across layers. A total loss function is constructed based on the endpoint error loss of the initial optical flow, the uncertainty loss of the uncertainty estimation branch, and the motion gradient consistency loss. The optical flow estimation network is trained using the total loss function to obtain an optimized optical flow estimation model. The optimized optical flow estimation model is then used to predict the optical flow of the image to be estimated. The method employs an uncertainty estimation branch to dynamically evaluate the reliability of optical flow predictions at each level, and then weights and fuses multi-scale results accordingly, effectively suppressing error propagation and accumulation in occluded and textureless regions. Simultaneously, it forces the alignment of optical flow gradients between adjacent layers through motion gradient consistency constraint loss, ensuring smoothness and physical rationality at motion boundaries. The synergistic effect of these two methods addresses the shortcomings of traditional methods in multi-scale fusion robustness and cross-level consistency, thereby significantly improving the overall accuracy, robustness, and interpretability of optical flow estimation in complex scenes.

[0132] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0133] Figure 4 A block diagram illustrating the principle of an optical flow estimation device corresponding to the optical flow estimation methods described in the above embodiments is shown. Figure 4 As shown, the optical flow estimation device includes a feature extraction module 31, a reliable calculation module 32, a weight fusion module 33, a fusion update module 34, a consistency calculation module 35, a loss construction module 36, and an optical flow prediction module 37. Detailed descriptions of each functional module are as follows: The feature extraction module 31 is used to extract multi-scale features of the image through the pyramid network structure, and based on the multi-scale features, to obtain the initial optical flow of each level through feature warping processing and cost volume calculation. The reliable calculation module 32 is used to calculate the uncertainty map of the optical flow at each level based on the initial optical flow through the uncertainty estimation branch. The uncertainty map represents the reliability of the optical flow prediction at each pixel position. The weighted fusion module 33 is used to calculate the fusion weight of the current layer using the uncertainty map of the current layer and the uncertainty map of the previous layer after fusion, and to use the fusion weight to perform weighted fusion of the optical flow of the current layer and the fused optical flow of the previous layer to obtain the fused optical flow of the current layer. The fusion update module 34 is used to update the uncertainty graph of the current layer after fusion based on the fusion weight and the uncertainty graph of the current layer; The consistency calculation module 35 is used to calculate the motion gradient consistency loss based on the optical flow gradient difference between adjacent pyramid layers, and is used to constrain the optical flow to maintain the gradient direction consistency across layers. The loss construction module 36 is used to construct a total loss function based on the endpoint error loss of the initial optical flow, the uncertainty loss of the uncertainty estimation branch, and the motion gradient consistency loss; The optical flow prediction module 37 is used to train the optical flow estimation network using the total loss function to obtain an optimized optical flow estimation model, and to use the optimized optical flow estimation model to predict the optical flow of the image to be estimated.

[0134] Specific limitations regarding the optical flow estimation device can be found in the limitations of the optical flow estimation method above, and will not be repeated here. Each module in the aforementioned optical flow estimation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in a computer device, or stored in software in the memory of a computer device, so that the processor can call and execute the corresponding operations of each module.

[0135] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 5 , Figure 5 This is a basic structural block diagram of the computer device in this embodiment.

[0136] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components connected to the memory 41, processor 42, and network interface 43 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0137] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0138] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or D-interface display memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as the program code of the optical flow estimation method. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.

[0139] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run program code stored in the memory 41 or process data, for example, to run program code for an optical flow estimation method.

[0140] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.

[0141] This application also provides another embodiment, namely, a computer-readable storage medium storing an interface display program that can be executed by at least one processor to cause the at least one processor to perform the steps of the optical flow estimation method described above.

[0142] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0143] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. An optical flow estimation method, characterized in that, include: Multi-scale features of the image are extracted through a pyramid network structure, and based on the multi-scale features, the initial optical flow of each level is obtained through feature warping processing and cost volume calculation. Based on the initial optical flow, uncertainty maps of optical flow at each level are calculated through uncertainty estimation branches. The uncertainty maps characterize the reliability of optical flow prediction at each pixel location. Using the uncertainty map of the current layer and the uncertainty map of the previous layer after fusion, the fusion weight of the current layer is calculated, and the fusion weight is used to perform weighted fusion of the optical flow of the current layer and the fused optical flow of the previous layer to obtain the fused optical flow of the current layer. Based on the fusion weights and the uncertainty map of the current layer, update the uncertainty map of the current layer after fusion; based on the difference in optical flow gradient between adjacent pyramid layers, calculate the motion gradient consistency loss to constrain the optical flow to maintain consistent gradient direction across layers; Based on the endpoint error loss of the initial optical flow, the uncertainty loss of the uncertainty estimation branch, and the motion gradient consistency loss, a total loss function is constructed; The optical flow estimation network is trained using the total loss function to obtain an optimized optical flow estimation model, and the optimized optical flow estimation model is used to predict the optical flow of the image to be estimated.

2. The optical flow estimation method as described in claim 1, characterized in that, The uncertainty map calculated by the uncertainty estimation branch for each level of optical flow includes: After the cost volume of each layer of the pyramid network structure, a light quantum network consisting of convolutional layers and activation functions is connected; The light quantum network outputs an uncertainty map with the same resolution as the current layer optical flow, where a larger uncertainty value indicates that the optical flow prediction at that location is less reliable.

3. The optical flow estimation method as described in claim 1, characterized in that, The calculation of the fusion weights for the current layer, and the weighted fusion of the optical flow of the current layer and the fused optical flow of the previous layer to obtain the fused optical flow of the current layer, includes: Upsample the uncertainty graph after fusion at the previous level and subtract it from the uncertainty graph at the current level. The difference results are input into the learnable sensitivity parameter and the Sigmoid activation function to obtain the fusion weights of the current layer; Using the fusion weights, the current layer optical flow and the upsampled previous layer fused optical flow are weighted and summed to obtain the current layer fused optical flow.

4. The optical flow estimation method as described in claim 1, characterized in that, The updated uncertainty graph after current layer fusion includes: Upsample the uncertainty graph after fusion at the previous level; Compare the uncertainty map of the current layer with the uncertainty map of the previous layer after upsampling pixel by pixel, and take the minimum value as the uncertainty map after fusion of the current layer.

5. The optical flow estimation method as described in claim 1, characterized in that, The calculation of motion gradient consistency loss includes: Calculate the gradient map of the optical flow between two adjacent layers after upsampling and alignment; The differences between the gradient maps are calculated and dynamically weighted based on the optical flow amplitude; The motion gradient consistency loss is obtained by summing the weighted differences of all adjacent levels in the pyramid.

6. The optical flow estimation method as described in claim 5, characterized in that, The dynamic weighting based on optical flow amplitude includes: Based on the amplitude value of the optical flow after upsampling from the previous layer, the weighting coefficient is calculated using an exponential function. The smaller the optical flow amplitude, the smaller the weighting coefficient.

7. The optical flow estimation method as described in claim 1, characterized in that, The total loss function is a weighted sum of endpoint error loss, uncertainty loss, and motion gradient consistency loss, wherein the uncertainty loss is constructed based on the weighted error between the uncertainty map of each layer and the real optical flow.

8. An optical flow estimation device, characterized in that, include: The feature extraction module is used to extract multi-scale features of the image through the pyramid network structure, and based on the multi-scale features, to obtain the initial optical flow of each level through feature warping processing and cost volume calculation. A reliable calculation module is used to calculate the uncertainty map of the optical flow at each level based on the initial optical flow through an uncertainty estimation branch. The uncertainty map represents the reliability of the optical flow prediction at each pixel location. The weighted fusion module is used to calculate the fusion weight of the current layer using the uncertainty graph of the current layer and the fused uncertainty graph of the previous layer, and to use the fusion weight to perform weighted fusion of the optical flow of the current layer and the fused optical flow of the previous layer to obtain the fused optical flow of the current layer. The fusion update module is used to update the uncertainty graph of the current layer after fusion based on the fusion weight and the uncertainty graph of the current layer; The consistency calculation module is used to calculate the motion gradient consistency loss based on the difference in optical flow gradient between adjacent pyramid layers, which is used to constrain the optical flow to maintain the consistent gradient direction across layers. The loss construction module is used to construct the total loss function based on the endpoint error loss of the initial optical flow, the uncertainty loss of the uncertainty estimation branch, and the motion gradient consistency loss; The optical flow prediction module is used to train the optical flow estimation network using the total loss function to obtain an optimized optical flow estimation model, and then use the optimized optical flow estimation model to predict the optical flow of the image to be estimated.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the optical flow estimation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the optical flow estimation method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Optical flow speed estimation method and system based on image scale invariance

    CN119359768A

  • Digital processing method and system for determination of optical flow

    US20100124361A1