An optical flow estimation method and system based on a lightweight model for optical flow estimation

CN120726098BActive Publication Date: 2026-09-01EEASY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510816291.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2026-09-01
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

通过运用多种方法,有效在保证光流估计精度情况下缩小模型的计算量,同时实现在端侧任务中实时预测,为后续应用提供准确的光流估计,解决现有技术无法实现在端侧中实时应用光流估计的问题

Benefits of technology

[0050]综上,本申请实施例首先在光流估计网络模型中改进轻量化卷积块,优化模型结构,实现了轻量化的光流估计模型,通过对输入图像进行标准化变化、特征图变换,以及结合前后图像的特征图进行局部相关性计算,实现特征点与特征点之间相关性的计算,提高对光流估计误差的容忍度,最后通过损失函数计算预测光流和真实标签计算损失值,优化光流估计网络模型的参数,提高光流估计的精度,有效减少模型的参数量和计算量,能够在端侧实现实时的光流估计预测,现有技术法无法在端侧中实现光流估计应用的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726098B_ABST
    Figure CN120726098B_ABST
Patent Text Reader

Abstract

This application relates to an optical flow estimation method and system based on a lightweight optical flow estimation model, belonging to the field of image processing technology. The method includes: firstly, improving the lightweight convolutional blocks in the optical flow estimation network model and optimizing the model structure to achieve a lightweight optical flow estimation model; secondly, calculating the correlation between feature points by standardizing the input image, transforming the feature maps, and combining the feature maps of the previous and subsequent images to improve the tolerance to optical flow estimation errors; and finally, calculating the predicted optical flow and the loss value using a loss function and the ground truth label to optimize the parameters of the optical flow estimation network model and improve the accuracy of optical flow estimation. Therefore, this embodiment optimizes the optical flow network structure, effectively reducing the number of model parameters and computational load, and enabling real-time prediction at the edge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an optical flow estimation method and system based on a lightweight optical flow estimation model. Background Technology

[0002] Optical flow refers to the instantaneous velocity of pixels moving on the imaging plane of a moving object in space. Optical flow estimation refers to estimating the optical flow by using pixel changes between adjacent frames in an image sequence. Optical flow estimation has been widely used in the field of image processing, such as motion detection, video stabilization, and rolling shutter distortion correction.

[0003] Current optical flow estimation methods can be broadly categorized into two types. The first is based on traditional methods, utilizing feature matching or phase estimation. Traditional methods are computationally intensive, inefficient, and have low accuracy, being highly susceptible to lighting and noise, making them difficult to implement in practical applications. The second type is based on deep learning methods. Deep learning has developed mature techniques for optical flow estimation, significantly outperforming traditional methods in both accuracy and speed, and exhibiting better environmental tolerance. However, deep learning-based optical flow estimation is primarily used on high-performance computing devices such as servers. On the edge, real-time prediction remains challenging due to limited computing power, which cannot meet the model's optical flow prediction requirements. Furthermore, existing models cannot avoid using the Grid_sample operator, which is not implemented on many edge chips, and CPU computation is slow. Therefore, current optical flow estimation methods cannot achieve real-time application on the edge. Summary of the Invention

[0004] This application provides an optical flow estimation method and system based on a lightweight optical flow estimation model. Firstly, by improving the model structure, the computational cost is reduced, enabling real-time optical flow estimation. Secondly, by combining the advantages of deep learning and traditional optical flow estimation, the accuracy of optical flow estimation is improved while simplifying the model structure. By employing multiple methods, the computational cost of the model is effectively reduced while maintaining the accuracy of optical flow estimation, and real-time prediction is achieved in edge-side tasks, providing accurate optical flow estimation for subsequent applications and solving the problem that existing technologies cannot achieve real-time application of optical flow estimation in edge-side tasks.

[0005] Firstly, this application provides an optical flow estimation method based on a lightweight optical flow estimation model, including:

[0006] A lightweight optical flow estimation network model is constructed using a dual-tower input optical flow network structure. A lightweight convolutional block with a Convblock structure is constructed using a row-column convolution design. The lightweight convolutional block is then converted into a two-dimensional convolutional layer through Kronecker product calculation.

[0007] The optical flow estimation network model performs convolution processing on the input sample image pairs, and performs absolute value normalization processing on the feature maps obtained by convolution to obtain a normalized feature map set. The sample image pairs include the previous frame image and the next frame image.

[0008] A two-dimensional Gaussian distribution transform is applied to the normalized feature maps of the later frames in the normalized feature map set to obtain the transformed feature maps;

[0009] Based on the previous frame standardized feature map in the standardized feature map set, and combined with the transformed feature map, local correlation analysis is performed to obtain a local correlation map;

[0010] Based on the local correlation map, feature splicing and convolutional regression optical flow are performed to obtain the predicted optical flow;

[0011] The loss value is calculated based on the predicted optical flow and the real labels corresponding to the sample images. The parameters of the optical flow estimation network model are then optimized to obtain the target lightweight model.

[0012] Optical flow estimation is performed on the input image to be processed using the target lightweight model to obtain the optical flow estimation result.

[0013] Optionally, a lightweight convolutional block with a Convblock structure is constructed using a row-column convolutional design. This lightweight convolutional block is then converted into a two-dimensional convolutional layer using Kronecker product calculation, including:

[0014] In the optical flow estimation network model, a row-column convolution design is adopted, based on Convblock(x) = conv k*1 (conv 1*k (x)) defines a lightweight convolutional block;

[0015] The lightweight convolutional block is converted into a k*k convolutional layer by Kronecker product calculation;

[0016] Where k is the kernel size and x is the input image of the layer.

[0017] Optionally, the feature maps obtained from convolution are subjected to absolute value normalization to obtain a normalized feature map set, including:

[0018] Using the feature map set of sample image pairs obtained by convolution as input, according to Calculate the forward result of AbsNorm;

[0019] Gradient backpropagation of LayerNorm is used, and the forward result is back-differentiated according to LayerNorm≈α*AbsNorm to obtain the standardized feature map set;

[0020] in, x is the input feature map, n is the feature channel size, AbsNorm is used to calculate the forward result, and LayerNorm is used to calculate the backward derivative. Optical flow is obtained by matching the correlation of each pixel one by one, and it is ensured that the intermediate feature vectors representing pixels have their own uniqueness to reflect the real situation.

[0021] Optionally, a two-dimensional Gaussian distribution transform is applied to the normalized feature maps of the later frames in the normalized feature map set to obtain transformed feature maps, including:

[0022] Based on the pixel information of the standardized feature map of the next frame in the standardized feature map set, and combined with the target center coordinates of the obtained optical flow estimation, Gaussian weight information is calculated using a two-dimensional Gaussian distribution.

[0023] Based on the Gaussian weight information, the normalized feature map of the next frame is weighted and transformed to obtain the transformed feature map.

[0024] Optionally, based on the pixel information of the standardized feature map of the later frame in the standardized feature map set, and combined with the target center coordinates obtained from the optical flow estimation, Gaussian weight information is calculated using a two-dimensional Gaussian distribution, including:

[0025] Using the pixel information of the normalized feature map of subsequent frames and the target center coordinates estimated by optical flow as input, according to Calculate Gaussian weight information;

[0026] Where, c = (c x ,c y ) represents the coordinates of the center of the Gaussian distribution, μ = (μ x ,μ y ) represents the target center coordinates for optical flow estimation, and σ represents the standard deviation of the Gaussian distribution.

[0027] Optionally, based on the Gaussian weight information, a weighted transformation is performed on the normalized feature map of the subsequent frame to obtain a transformed feature map, including:

[0028] Obtain the optical flow estimate flow from the model regression and the grid matrix G of the pixel positions in the normalized feature map of the next frame;

[0029] Using optical flow estimation flow, grid matrix G, and normalized feature map f2 of the next frame as input, according to Perform a Gaussian distribution transformation to obtain the transformation feature map.

[0030]

[0031] Wherein, the grid matrix G has dimensions H*W*2, G x Let G be the x-axis, and G be the x-axis. y Let G be the y-axis.

[0032] Optionally, based on the previous frame normalized feature map in the normalized feature map set, and combined with the transformed feature map, local correlation analysis is performed to obtain a local correlation map, including:

[0033] Based on the normalized feature map of the previous frame, a sliding window folding method is used to calculate local window features, which are used to characterize the tolerable optical flow error.

[0034] Based on the local window features and the transformed feature map, the similarity between the feature points of the previous frame image and the pixels in the same feature point and nearby area of ​​the subsequent frame image is analyzed to obtain a local correlation map.

[0035] Optionally, based on the normalized feature map of the previous frame, a sliding window folding method is used to calculate local window features, including:

[0036] Construct a sliding window k;

[0037] Within the sliding window k, according to f1 unfold =Unfold(f1, k), which uses the Unfold operator to slide and unfold the normalized feature map of the previous frame, and calculates the local window feature f1. unfold ;

[0038] Specifically, based on local window features and the transformed feature map, the similarity between feature points in the previous frame image and pixels in the same feature point and nearby region in the subsequent frame image is analyzed to obtain a local correlation map, including:

[0039] Using local window features f1 unfold and transformation feature map For input, according to Correlation calculations are performed between feature points to obtain the local correlation map corr; f1 is the normalized feature map of the previous frame image.

[0040] Optionally, the loss value is calculated based on the predicted optical flow and the true labels corresponding to the sample images, including:

[0041] Based on the model's predicted optical flow pred and real label flow gt As input, according to L1 = mean(abs(flow) pred -flow gt )) Calculate the loss L1.

[0042] Secondly, this application provides an optical flow estimation system based on a lightweight optical flow estimation model, comprising:

[0043] The model building module is used to construct a lightweight optical flow estimation network model using a dual-tower input optical flow network structure, and to construct a lightweight convolutional block with a Convblock structure using a row and column convolution design. The lightweight convolutional block is then converted into a two-dimensional convolutional layer through Kronecker product calculation.

[0044] The convolution module is used to perform convolution processing on the input sample image pairs through the optical flow estimation network model, and perform absolute value normalization processing on the feature maps obtained by convolution to obtain a normalized feature map set. The sample image pairs include the previous frame image and the next frame image.

[0045] The Gaussian transform module is used to transform the normalized feature maps of the later frame in the normalized feature map set using a two-dimensional Gaussian distribution transform to obtain transformed feature maps.

[0046] The local correlation analysis module is used to perform local correlation analysis based on the previous frame standardized feature map in the standardized feature map set and the transformed feature map to obtain a local correlation map.

[0047] The optical flow prediction module is used to perform feature stitching and convolutional regression optical flow based on the local correlation map to obtain the predicted optical flow.

[0048] The model optimization module is used to calculate the loss value based on the predicted optical flow and the real labels corresponding to the sample images, optimize the parameters of the optical flow estimation network model, and obtain the target lightweight model.

[0049] The optical flow estimation module is used to perform optical flow estimation on the input image to be processed using the target lightweight model, and obtain the optical flow estimation result.

[0050] In summary, the embodiments of this application first improve the lightweight convolutional block in the optical flow estimation network model and optimize the model structure to achieve a lightweight optical flow estimation model. By standardizing the input image, transforming the feature map, and combining the feature maps of the previous and subsequent images to calculate the local correlation, the correlation between feature points is calculated, improving the tolerance to optical flow estimation errors. Finally, the predicted optical flow is calculated using a loss function, and the loss value is calculated using the true label. This optimizes the parameters of the optical flow estimation network model, improves the accuracy of optical flow estimation, and effectively reduces the number of model parameters and computational load. It enables real-time optical flow estimation and prediction at the edge, addressing the problem that existing methods cannot achieve optical flow estimation applications at the edge. Attached Figure Description

[0051] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0052] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 A flowchart illustrating an optical flow estimation method based on a lightweight optical flow estimation model provided in this application embodiment;

[0054] Figure 2 This is a flowchart illustrating the steps of an optical flow estimation method based on a lightweight optical flow estimation model, provided in an optional embodiment of this application.

[0055] Figure 3 Example diagram of a traditional convolutional layer;

[0056] Figure 4 A feature extraction module structure diagram is provided as an optional example of this application;

[0057] Figure 5 Example diagrams of the modules of the optical flow estimation network provided as an optional example of this application;

[0058] Figure 6 Example diagram of row and column convolutional blocks provided as an optional example for this application;

[0059] Figure 7 Example diagram of the downsampling module provided as an optional example of this application;

[0060] Figure 8 This is an example diagram of the optical flow estimation module designed in this invention;

[0061] Figure 9 A structural block diagram of an optical flow estimation system based on a lightweight optical flow estimation model provided in this application embodiment;

[0062] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0064] To facilitate understanding of the embodiments of this application, further explanations and descriptions will be provided below in conjunction with the accompanying drawings and specific embodiments. These embodiments do not constitute a limitation on the embodiments of this application.

[0065] Figure 1 This is a flowchart illustrating an optical flow estimation method based on a lightweight optical flow estimation model, provided in an embodiment of this application. In specific implementations, this method can be applied to edge devices with limited computing power or to servers; this embodiment does not impose any limitations on this. Figure 1 As shown, the optical flow estimation method may specifically include the following steps:

[0066] Step 110: A lightweight optical flow estimation network model is constructed using a dual-tower input optical flow network structure, and a lightweight convolutional block with a Convblock structure is constructed using a row-column convolution design. The lightweight convolutional block is then converted into a two-dimensional convolutional layer through Kronecker product calculation.

[0067] In its specific implementation, this embodiment addresses the problems of existing technologies by constructing a lightweight optical flow estimation model based on deep learning, or an optical flow estimation network model. This lightweight model employs a dual-tower input optical flow network structure. To enable real-time prediction at the edge, it is typically necessary to reduce the number of model parameters and computational cost. Therefore, all modules of this model utilize minimal parameters and computational cost.

[0068] In practical implementation, convolutional blocks are an important component of the model backbone. Reducing the number of parameters in convolutional blocks can effectively reduce the overall number of parameters in the model. This embodiment uses row-column convolution as the standard, designs a lightweight convolutional block (Convblock), and uses Kronecker product calculation to convert the lightweight convolutional block into a two-dimensional convolutional layer, simulating the effect of two-dimensional convolution. Thus, while reducing the number of convolutional parameters, it achieves the same convolutional effect as traditional models.

[0069] Step 120: The input sample image pairs are convolved using the optical flow estimation network model, and the absolute value is standardized based on the feature maps obtained from the convolution to obtain a standardized feature map set.

[0070] The sample image pair includes a previous frame image and a subsequent frame image, and the standardized feature map set includes a previous frame standardized feature map corresponding to the previous frame image and a subsequent frame standardized feature map corresponding to the subsequent frame image.

[0071] In practical implementations, optical flow estimation typically uses paired sample images, such as previous and subsequent frame images with a temporal relationship. During the model training phase, the optical flow estimation network model first performs convolution processing on the input sample image pairs to obtain convolutional feature maps of the previous and subsequent frame images.

[0072] Then, the convolutional feature maps of the two images are normalized by absolute value. Absolute value normalization processes the input data through operators in deep learning (such as the Norm operator) to obtain normalized feature maps of the previous frame image and the next frame image, forming a normalized feature map set.

[0073] In the specific implementation, considering that in the current image field, existing technologies usually use BatchNorm as a standardization algorithm, but in the optical flow field, BatchNorm can actually bring disadvantages, mainly for two reasons: ① In this embodiment, optical flow is calculated by matching the correlation of each pixel one by one, which is different from other optical flow estimation methods based on deep learning. It is necessary to ensure that the intermediate feature vectors representing pixels have their own uniqueness; ② At present, optical flow data is basically derived from synthetic data, which cannot reflect the real situation. The weights obtained by BatchNorm are difficult to meet the needs of practical use.

[0074] Therefore, this embodiment innovatively updates the absolute value standardization process, using LayerNorm and AbsNorm to perform forward calculation and backward differentiation in absolute value calculation, thereby greatly reducing the computational load of the model while achieving virtually no loss.

[0075] Step 130: Apply a two-dimensional Gaussian distribution transform to the normalized feature maps of the later frame in the normalized feature map set to obtain the transformed feature maps.

[0076] The associated image is the next frame of the standardized feature map.

[0077] In this implementation, the feature map transformation method for edge-side optical flow estimation is updated by using a Gaussian distribution to transform the feature map. Specifically, in the optical flow estimation process, edge-side optical flow estimation mainly combines the feature point matching characteristics of the two images. That is, feature points in the previous frame image will only be similar to the same feature points and nearby pixels in the subsequent frame image. Therefore, a Gaussian distribution transformation is introduced to transform the feature map, processing the feature map of the subsequent frame image to obtain the transformed feature map. Thus, the characteristics of traditional feature matching are retained in the optical flow estimation process, while effectively reducing the computational load of the model and making effective use of the optical flow results predicted by the previous layer.

[0078] Step 140: Based on the previous frame standardized feature map in the standardized feature map set, and combined with the transformed feature map, perform local correlation analysis to obtain a local correlation map.

[0079] In practical implementation, traditional feature matching methods utilize feature correlation to calculate the correlation map between the feature maps of the previous and subsequent frames. This embodiment pre-transforms the feature map of the subsequent frame image, thus eliminating the need to calculate global correlation; only the correlation between feature points in the two feature maps needs to be calculated. Furthermore, to address the potential error in estimated optical flow, this embodiment utilizes the feature maps of the previous and subsequent frames to analyze local correlations and calculate a local feature correlation map, i.e., a local correlation map.

[0080] Step 150: Based on the local correlation map, perform feature stitching and convolutional regression optical flow to obtain the predicted optical flow.

[0081] In the specific implementation, the model performs feature stitching on the local correlation map. During feature stitching, two types of features can be combined, including the transformed feature map after the transformation of the subsequent frame image and the optical flow result predicted by the previous layer. By stitching different features together according to channels, and then passing a convolution, the optical flow is directly regressed to obtain the predicted optical flow.

[0082] Step 160: Calculate the loss value based on the predicted optical flow and the real labels corresponding to the sample images, optimize the parameters of the optical flow estimation network model, and obtain the target lightweight model.

[0083] Step 170: Optical flow estimation is performed on the input image to be processed using the target lightweight model to obtain the optical flow estimation result.

[0084] Steps 160-170 are described uniformly as follows:

[0085] In the specific implementation, the predicted optical flow is compared with the real labels, and a regression loss value, such as L1 loss, is calculated using a loss function. The model parameters are then optimized based on the loss value until a target lightweight model that meets the preset accuracy is obtained. Finally, the target lightweight model is used to estimate the optical flow of the input data, enabling real-time prediction in edge tasks and providing accurate optical flow estimation for subsequent applications.

[0086] As can be seen, this embodiment first improves the convolutional blocks of the optical flow estimation network model, effectively reducing the number of parameters and computational cost of the network. Through improved module design, the computational cost of the model is significantly reduced. Then, combining the advantages of deep learning and traditional optical flow estimation, improvements are made to the model structure, including absolute value standardization, feature map transformation, and local correlation calculation, thereby enhancing the accuracy of optical flow estimation. Compared to existing solutions, this embodiment is more tolerant of environmental noise, camera noise, etc., and can output optical flow more stably, achieving real-time optical flow estimation and providing effective optical flow maps for subsequent applications.

[0087] Reference Figure 2This illustration shows a flowchart of an optical flow estimation method based on a lightweight optical flow estimation model, provided in an optional embodiment of this application. The method may specifically include the following steps:

[0088] Step 210: A lightweight optical flow estimation network model is constructed using a dual-tower input optical flow network structure. A lightweight convolutional block with a Convblock structure is constructed using a row-column convolution design. The lightweight convolutional block is then converted into a two-dimensional convolutional layer through Kronecker product calculation.

[0089] In this specific implementation, a row-column convolution design is used to construct a lightweight convolutional block with a Convblock structure. The lightweight convolutional block is then converted into a two-dimensional convolutional layer using the Kronecker product calculation. Specifically, the following lightweight convolutional structure can be used: In the optical flow estimation network model, a row-column convolution design is adopted, based on Convblock(x) = conv k*1 (conv 1*k (x)) defines a lightweight convolutional block; the lightweight convolutional block is converted into a k*k convolutional layer by Kronecker product calculation; where k is the convolutional kernel size and x is the input image of the layer.

[0090] In related technologies, existing optical flow estimation schemes primarily use k-series optical flow estimation models with convolutional structures based on k-series methods. h *k w Traditional convolutions, such as Figure 3 As shown, its parameter quantity is k. h ×k w ×C in ×C out The computational complexity is ((C) in ×k h ×k w )+(C in ×k h ×k w -1)×C out ×W×H), k h ,k w C represents the kernel row and column size. in C is the input channel for the convolution kernel. out As shown in the output channel of the convolution kernel, the existing convolution block structure requires high computing power and has a large number of parameters, making it unsuitable for edge-side applications.

[0091] In response to the above problems, refer to Figure 4 The row and column convolutional block structure shown in this embodiment is a 1×k column convolutional conv constructed based on Convblock. 1*k and k×1 row convolution conv k*1Then, through Kronecker product calculation, the Convblock is transformed into a k*k convolutional layer, for example, a conv... 1*k and conv k*1 The convolutional kernel weights can be combined into a k*k weight matrix using the Kronecker product, allowing the Convblock design to have the same receptive field as a k*k convolution while reducing the number of model parameters. The number of parameters in this convolutional block is (k... h ×1+1×k w )×C in ×C out The computational complexity is ((C) in ×(k h ×1+1×k w ))+(C in ×(k h ×1+1×k w )-1)×C out ×W×H), obviously when k w ,k h When ≥3, (k h ×1+1×k w )>k h ×k w Therefore, the number of parameters and computational cost of using row and column convolutional blocks is less than that of traditional convolution.

[0092] Therefore, this embodiment optimizes and improves the convolutional blocks, allowing step-by-step convolutions to be merged into standard convolution operations during model deployment, ensuring compatibility with mainstream deep learning frameworks and reducing inference time. It offers the following advantages: ① Preserving the receptive field: The receptive range of two-dimensional convolution is achieved through step-by-step one-dimensional convolution, without sacrificing feature extraction capabilities; ② Lightweight design: The number of parameters and computational cost k decreases exponentially with increasing size, making it suitable for deployment on edge devices; ③ Engineering friendliness: It can be mathematically transformed to be compatible with standard convolution operations, facilitating model optimization and deployment.

[0093] Step 220: The input sample image pairs are convolved using the optical flow estimation network model, and the absolute value is standardized based on the feature maps obtained from the convolution to obtain a standardized feature map set.

[0094] The sample image pair includes a previous frame image and a subsequent frame image.

[0095] In the specific implementation, refer to Figure 5 and Figure 6 As shown, the optical flow estimation network model includes a feature extraction module ( Figure 5 The conv_module and downsampling module in the middle. Figure 5 downsample) and optical flow estimation module ( Figure 5(compute_flow in the context of [the module]). The feature extraction module conv_module can be found in [the relevant section]. Figure 6 The process shown extracts image features from the input image.

[0096] Specifically, when extracting image features, the aforementioned Conblock row and column convolution blocks are used for convolution processing to extract feature maps. For example, if the input image size is H×W×C (height*width*number of channels), first, a one-dimensional convolution is performed along the column direction (vertical direction) on each channel of the input image. For each pixel, this operation fuses the information of its k pixels above and below (receptive field is k×1). Then, a one-dimensional convolution is performed along the row direction (horizontal direction) on the output feature map of the column convolution. For each pixel, this operation fuses the information of its k pixels to the left and right (receptive field is 1×k). Finally, the receptive field is superimposed by the two convolutions to form k*k. Assuming the input channel is C and the output channel is C', the parameters of a normal k*k convolution are: C'×C×k×k, and the parameters of the Convblock row and column convolution are:

[0097] C'×C×k×1+C'×C×1×k=2C×C'×k. Obviously, when k=3, the number of parameters decreases from 9C×C' to 6C×C', which is a significant reduction. As k increases, the number of parameters will decrease further. In addition, the computational cost also decreases linearly with the increase of k, which significantly reduces the computational cost of the model.

[0098] Then, in this embodiment, after convolving the input image, the absolute value of the convolved feature map can be standardized to obtain the standardized feature maps corresponding to the previous and next frame images.

[0099] Optionally, the above-mentioned feature maps obtained based on convolution are subjected to absolute value normalization to obtain a normalized feature map set. Specifically, this may include: taking the feature map set of sample image pairs obtained by convolution as input, and according to... Calculate the forward result of AbsNorm; use the gradient backpropagation of LayerNorm, and perform backpropagation on the forward result according to LayerNorm≈α*AbsNorm to obtain the standardized feature map set; where... x is the input feature map, n is the feature channel size, AbsNorm is used to calculate the forward result, and LayerNorm is used to calculate the backward derivative. Optical flow is obtained by matching the correlation of each pixel one by one, and it is ensured that the intermediate feature vectors representing pixels have their own uniqueness to reflect the real situation.

[0100] In related technologies, norms are commonly used operators in deep learning. For example, in the image domain, traditional normalization operations mainly use the BatchNorm and LayerNorm operators. However, in the field of optical flow, the BatchNorm operator has limitations in its use, mainly for two reasons: First, optical flow is calculated by matching the correlation of each pixel one by one, unlike other deep learning-based optical flow estimation methods, which require ensuring that the intermediate feature vectors representing pixels have their own uniqueness. Second, current optical flow data is mostly derived from synthetic data, which cannot reflect the real situation, and the resulting weights are difficult to meet practical needs. LayerNorm requires calculating the mean and variance of features, which greatly increases the computational load of the model and is not suitable for use on the edge.

[0101] To address the aforementioned technical problems, this embodiment improves the absolute value normalization process by employing AbsNorm for forward computation and fitting the forward result to LayerNorm. Specifically, the feature maps x of the previous and subsequent frames are first used as input, and then... Determine the intermediate variable z of feature x after absolute value mean centering, where, This represents the mean of the absolute values ​​of all channels of the input feature, and z represents the feature minus its mean absolute value, completing the centering operation and preparing for the subsequent calculation of the standardized scaling factor. Then, according to... The decentralized data is normalized to obtain a standardized scaling factor, namely the AbsNorm forward result.

[0102] Considering that AbsNorm uses absolute value operations while LayerNorm is based on square operations, and their statistical properties differ, α is used as a difference compensation factor to compensate for this difference, making the forward output distribution of AbsNorm closer to that of LayerNorm. Preferably, experimental verification shows that when α = 0.86, the outputs of AbsNorm and LayerNorm are numerically consistent. Furthermore, since AbsNorm uses absolute value operations, there are discontinuities in the backpropagation gradient calculation, and convergence is less precise than with square values ​​near zero. Therefore, gradient backpropagation is used for the backpropagation gradient calculation. The final calculation uses AbsNorm forward and LayerNorm backward. In experiments, this method introduces virtually no loss and significantly reduces the computational cost of the model, while ensuring compatibility of the standardization effect.

[0103] For example, the feature extraction process for optical flow estimation from input data includes convolution and normalization. Assume the two images (sample images) for which optical flow needs to be calculated are named TA1 and TB1, with an input size of 640*480*3 (i.e., height*width*number of channels). TA1 and TB1 are processed by... Figure 7The downsampling module shown is used. In this module, the kernel size is 7*7, the number of output channels is 64, the stride is 2, and the padding value is 3. After this convolution operation, convolutional images with a data size of 320*240*64 are obtained, denoted as TA2 and TB2 respectively.

[0104] Then, TA2 and TB2 are respectively passed through Figure 6 The feature extraction module is shown. Within the `convblock`, the row and column convolutional blocks contain 1x3 (column convolution) and 3x1 (row convolution) kernels, with 64 output channels and a stride of 1. The output image size after this block is 320x240x64. The first convolution after the `convblock` has a 1x1 kernel, 256 output channels, and a stride of 1, resulting in an output size of 320x240x256. A second convolution with a 1x1 kernel, 64 output channels, and a stride of 1 further yields data of 320x240x64, denoted as TA3 and TB3 respectively.

[0105] TA3 and TB3 were respectively passed through Figure 7 The downsampling module shown, after undergoing this convolution operation, produces data of sizes 160*120*64, denoted as TA4 and TB4 respectively.

[0106] TA4 and TB4 were respectively passed through Figure 6 The feature extraction module is shown. Within the `convblock`, the row and column convolutional blocks have kernels of 1x3 and 3x1 respectively, with 64 output channels and a stride of 1. The output size after this row and column convolutional block is 160x120x64. The first convolution after the `convblock` has a 1x1 kernel, 256 output channels, and a stride of 1, resulting in an output size of 160x120x256. Then, a second convolution with a 1x1 kernel, 64 output channels, and a stride of 1 yields data of 160x120x64, denoted as TA5 and TB5 respectively.

[0107] TA5 and TB5 respectively through Figure 7 The downsampling module shown, after the convolution operation, produces data of sizes 80*60*64, denoted as TA6 and TB6 respectively.

[0108] TA6 and TB6 were respectively passed through Figure 6The feature extraction module is shown. Within the convblock, the row and column convolutional blocks have kernels of 1x3 and 3x1 respectively, with 64 output channels and a stride of 1. The output size after this row and column convolutional block is 80x60x64. The first convolution after the convblock has a 1x1 kernel, 256 output channels, and a stride of 1, resulting in an output size of 80x60x256. Then, a second convolution with a 1x1 kernel, 64 output channels, and a stride of 1 is performed, resulting in data sizes of 80x60x64, denoted as TA7 and TB7 respectively.

[0109] TA7 and TB7 are passed through the downsampling module. In this module, the kernel size is 7*7, the number of output channels is 64, the stride is 2, and the padding value is 3. After this convolution operation, the data sizes obtained are 40*30*64, which are denoted as TA8 and TB8 respectively.

[0110] TA8 and TB8 are respectively passed through a Figure 6 The feature extraction module is shown. Within the `convblock`, the row and column convolutional blocks have kernels of 1x3 and 3x1 respectively, with 64 output channels and a stride of 1. After passing through this row and column convolutional block, the output size is 40x30x64. The first convolution after the `convblock` has a 1x1 kernel, 256 output channels, and a stride of 1, resulting in an output size of 40x30x256. Then, a second convolution with a 1x1 kernel, 64 output channels, and a stride of 1 is performed, resulting in data sizes of 40x30x64, denoted as TA9 and TB9 respectively.

[0111] This completes the extraction of feature maps from the input image and performs absolute value standardization.

[0112] Step 230: Based on the pixel information of the standardized feature map of the next frame in the standardized feature map set, and combined with the target center coordinates of the obtained optical flow estimation, Gaussian weight information is calculated using a two-dimensional Gaussian distribution.

[0113] Step 240: Based on the Gaussian weight information, perform a weighted transformation on the normalized feature map of the next frame to obtain a transformed feature map.

[0114] Steps 230-240 are described uniformly as follows:

[0115] In related technologies, previous optical flow estimation based on deep learning models, after directly regressing the optical flow from the model, generally adopted a coarse-to-fine approach to refine the optical flow results. That is, it gradually upsamples from deep, low-resolution feature regression to shallow, high-resolution features. To better utilize the predicted optical flow results of the previous layer, Grid_sample is used to transform the feature map, thereby reducing the predicted optical flow error at each layer. However, Grid_sample is not supported on the edge. Therefore, this embodiment introduces Gaussian distribution transformation. By calculating the Gaussian distribution, the feature map of the subsequent frame image is transformed, which fully utilizes the matching characteristics of each feature point on the feature map of the previous and subsequent frames. That is, the feature points of the previous frame will only be similar to the same feature points and nearby pixels of the subsequent frame. The related subsequent frame image is introduced to achieve feature matching.

[0116] Specifically, this embodiment first analyzes the pixel information of the feature map of the next frame, including the grid matrix of the pixel positions in the feature map of the next frame (usually composed of x and y coordinates), which contains the pixel coordinates. The optical flow estimation results of the model in optical flow estimation are obtained. During optical flow estimation, features of the previous and next frames are extracted and matched through convolutional blocks and lightweight modules of the neural network, ultimately outputting a two-dimensional optical flow field. Each pixel corresponds to a two-dimensional vector, representing the coordinate offset of that pixel from the previous frame to the next frame, from which the target center coordinates are extracted.

[0117] Then, the pixel coordinates are superimposed with the optical flow coordinate offset to obtain the center position of the Gaussian distribution. Using this as input, Gaussian weight information is calculated based on the Gaussian function. This Gaussian weight information includes a two-dimensional Gaussian weight distribution map (referred to as the Gaussian distribution map). Each element in the Gaussian distribution map represents the weight coefficient of the corresponding pixel in the feature transformation. This distribution map is generated using the optical flow estimation results and pixel position information, and is used to perform a weighted transformation on the normalized feature map of the subsequent frame to obtain the transformed feature map. This replaces the Grid_sample operation, which is not supported by the edge, while also conforming to the principle of local correlation in feature matching.

[0118] In an optional embodiment, the above-mentioned calculation of Gaussian weight information based on the pixel information of the later frame standardized feature map in the standardized feature map set, combined with the obtained target center coordinates of optical flow estimation, and using a two-dimensional Gaussian distribution, includes: taking the pixel information of the later frame standardized feature map and the target center coordinates of optical flow estimation as input, and according to... Calculate the Gaussian weight information; where c = (c x ,c y ) represents the coordinates of the center of the Gaussian distribution, μ = (μ x ,μ y ) represents the target center coordinates for optical flow estimation, and σ represents the standard deviation of the Gaussian distribution.

[0119] In an optional embodiment, this embodiment performs a weighted transformation on the normalized feature map of the subsequent frame based on the Gaussian weight information to obtain a transformed feature map, including: obtaining the optical flow estimate flow regressed from the model and the grid matrix G of the pixel positions in the normalized feature map of the subsequent frame; using the optical flow estimate flow, the grid matrix G, and the normalized feature map of the subsequent frame f2 as input, according to... Perform a Gaussian distribution transformation to obtain the transformation feature map. Wherein, the grid matrix G has dimensions H*W*2, G x Let G be the x-axis, and G be the x-axis. y Let G be the y-axis.

[0120] In the specific implementation, the optical flow estimation flow is obtained by the model based on the feature maps of the previous and next frame images to regress the optical flow field, which represents the offset of each pixel in the previous frame image in the next frame image. It includes the offset of the optical flow field in the x direction flow(:,:,0) and the offset in the y direction flow(:,:,1), which mainly reflects the "target position" of the pixel in the previous frame in the next frame. The center of the Gaussian distribution is based on this to ensure that the weights are concentrated in the area near the position, which conforms to the local correlation principle of feature matching.

[0121] At this point, the center position of the Gaussian distribution (μ) x ,μ y ), that is (c x ,c y (μ) x =c x μ y =c y The Gaussian weights, μ, can be obtained by superimposing pixel coordinates and optical flow offsets, focusing the Gaussian weights near the "matching position" of the optical flow prediction. x =G x +flow(:,:,0), μ y =G y +flow(:,:,1).

[0122] When μ x =c x At that time, c x -μ x =0, essentially aligning the center of the Gaussian distribution strictly with the target position predicted by the optical flow, so that the weight matrix has a two-dimensional Gaussian distribution centered at that position. By explicitly introducing the position prior, the rationality of "local feature matching" in optical flow estimation is strengthened.

[0123] When calculating the Gaussian distribution plot, take This embodiment takes into account that Softmax itself is not sparsity enough and lacks positional information, and fails to combine the characteristics of traditional feature matching. Therefore, it does not use Softmax, but uses Gaussian distribution transformation feature map to avoid the weight dispersion problem caused by the lack of positional constraints of Softmax, which is more suitable for the efficiency and accuracy requirements of end-side optical flow tasks.

[0124] Therefore, this embodiment utilizes Gaussian distribution to drive feature map transformation, aligning the features of the subsequent frame image with the features of the previous frame image in position, thereby optimizing the accuracy of optical flow estimation.

[0125] Step 250: Based on the normalized feature map of the previous frame, calculate the local window features using a sliding window folding method.

[0126] The local window feature is used to characterize the tolerance for optical flow errors.

[0127] Step 260: Based on the local window features and the transformed feature map, analyze the similarity between the feature points of the previous frame image and the pixels in the same feature point and nearby area of ​​the subsequent frame image to obtain a local correlation map.

[0128] Steps 250-260 are described uniformly as follows:

[0129] In related technologies, the core of feature matching between consecutive frames is calculating the similarity between feature maps. In existing technologies, global correlation calculation requires traversing all feature point pairs in the entire image, which is computationally complex and susceptible to optical flow estimation errors.

[0130] In this embodiment, the feature map of the subsequent frame image has already been transformed using a Gaussian distribution. Therefore, it is not necessary to calculate global correlation; only the correlation between feature points in the feature maps needs to be calculated. The calculated local feature correlation map is then used to improve the tolerance for errors in the estimated optical flow. Specifically, the feature map of the previous frame image is first folded using an Unfold sliding window. Then, the folded feature map is combined with the transformed feature map of the subsequent frame image to calculate local correlation, resulting in a local feature correlation map.

[0131] Therefore, by introducing local correlation, the following can be achieved: 1. Local constraints on optical flow estimation: There are errors in the pixel motion offset of optical flow prediction, and global matching is prone to failure due to accumulated errors. The feature correlation of local regions (such as 3×3 windows) is more reliable; 2. Optimization of computational efficiency: There is no need to traverse the entire image. By focusing on local regions through sliding windows, the amount of computation is reduced.

[0132] Optionally, the above-mentioned calculation of local window features based on the standardized feature map using a sliding window folding method may specifically include: constructing a sliding window k; within the sliding window k, according to f1 unfold=Unfold(f1, k), which uses the Unfold operator to slide and unfold the normalized feature map of the previous frame, and calculates the local window feature f1. unfold Where f1 is the feature map of the previous frame image, i.e., the standardized feature map.

[0133] Optionally, the above-mentioned analysis of the similarity between feature points of the previous frame image and pixels in the same feature point and nearby region of the subsequent frame image based on local window features and the associated transformation feature map to obtain a local correlation map may include: using local window features f1 unfold and transformation feature map For input, according to Correlation calculations are performed between feature points to obtain a local correlation map (corr).

[0134] In the specific implementation, a sliding window k is first constructed (when k=3, it represents a 3*3 window). Then, the Unfold(k) sliding window folding operation is performed to divide the previous frame feature map f1 into overlapping local blocks according to the window size, and the corresponding local window features f are output. unfold For example, if the feature map f1 has a shape of C×H×W, then after the folding operation, the output shape can be C×k. 2 ×H'×W' (H'×W' is the number of blocks after the window slides). By decomposing the global feature map into multiple local feature blocks, each block corresponds to the feature vector of a k×k region in the original feature map, preparation is made for the calculation of local correlations.

[0135] In the specific implementation, the feature map is transformed. After warp transformation, the shape is similar to f unfold Since they are consistent, the similarity of local blocks at corresponding locations can be calculated. and f unfold For each corresponding local block, calculate the vector dot product or cosine similarity and output the local correlation graph (corr). The value at each position in the corr graph represents the feature similarity between the corresponding local regions of the previous frame and the next frame. High-value regions are possible matching points.

[0136] Therefore, this embodiment leverages the advantage of local correlation in optical flow errors. On one hand, it improves error robustness. Specifically, offset errors in optical flow estimation can lead to feature position misalignment during global matching, while the local window only focuses on a 3×3 region near the current pixel. Even with a 1-2 pixel error in optical flow, the true matching point can still be found within the local window. On the other hand, it balances computational efficiency and accuracy in edge-side vision tasks, compared to global correlation (complexity O(HW)). 2 Local dependencies are reduced to O(HWk) complexity by using Unfold. 2 When k 2Efficiency is significantly improved when HW is less than 1. In addition, the overlapping design of the sliding window can also ensure complete coverage of local areas and avoid feature omission.

[0137] For example, based on TA9 and TB9 obtained in the previous example, optical flow is obtained by calculating feature correlation matching. The main process includes:

[0138] TA9 and TB9 are processed through the optical flow estimation module, such as Figure 8 As shown, `prev_flow` is initialized to 0, denoted as `f1`, with a size of 40*30*2. Based on the calculations in the Gaussian transform formula above, the optical flow map is converted to a Gaussian distribution map, with an output of 40*30*1200, denoted as `G1`. The feature map of TB9 is then transformed using the Gaussian transform formula above, resulting in TB9', with a size of 40*30*64. Next, based on the local correlation calculation formula, a 3*3 local window correlation is calculated for TA9 and TB9', denoted as `C1`, with an output size of 40*30*9. `C1`, TB9', and `f1` are concatenated together by channel, resulting in an output size of 40*30*75. Finally, an optical flow is obtained through a convolution with a kernel of 3*3, a stride of 1, and 2 output channels, with an output data size of 40*30*2, denoted as `f2`.

[0139] F2 is passed through an upsampling module, and the output size is 80*60*2, which is denoted as F2.

[0140] TA7 and TB7 are processed through the optical flow estimation module, such as Figure 8 As shown, F2 is calculated according to a Gaussian distribution, transforming the optical flow map into a Gaussian distribution map, with an output size of 80*60*4800, denoted as G2. TB7 is then transformed using the Gaussian distribution transformation formula to obtain TB7', with a size of 80*60*64. Next, according to the local correlation formula, a 3*3 local window correlation is calculated for TA7 and TB7', denoted as C2, with an output size of 80*60*9. C2, TB7', and f2' are concatenated together by channel, resulting in an output size of 80*60*75. Finally, an optical flow is obtained through a convolution with a 3*3 kernel, a stride of 1, and 2 output channels, with an output data size of 80*60*2, denoted as f3.

[0141] The output size of f3 is 160*120*2, which is denoted as F3, after passing f3 through an upsampling module.

[0142] TA5 and TB5 are processed through the optical flow estimation module, such as Figure 8As shown, F3 is calculated according to a Gaussian distribution, transforming the optical flow map into a Gaussian distribution map, with an output size of 160*120*19200, denoted as G3. TB5 is then transformed using the Gaussian distribution transformation formula to obtain TB5', with a size of 160*120*64. Next, based on local correlation calculations, a 3*3 local window correlation is calculated for TA5 and TB5', denoted as C3, with an output size of 160*120*9. C3, TB5', and f3' are concatenated together by channel, resulting in an output size of 160*120*75. Finally, an optical flow is obtained through a convolution with a 3*3 kernel, a stride of 1, and 2 output channels, with an output data size of 160*120*2, denoted as f4.

[0143] The output size of f4 is 320*240*2, which is denoted as F4.

[0144] TA3 and TB3 are processed through the optical flow estimation module, such as Figure 8 As shown, F4 is calculated according to a Gaussian distribution, transforming the optical flow map into a Gaussian distribution map, with an output size of 320*240*76800, denoted as G4. TB3 undergoes feature map transformation according to the Gaussian distribution transformation formula to obtain TB3', with a size of 320*240*64. Then, based on local correlation calculations, a 3*3 local window correlation is calculated for TA3 and TB3', denoted as C4, with an output size of 320*240*9. C4, TB3', and f4' are concatenated together by channel, resulting in an output size of 320*240*75. Finally, an optical flow is obtained through a convolution with a 3*3 kernel, a stride of 1, and 2 output channels, with an output data size of 320*240*2, denoted as f5.

[0145] The output size of f5 is 640*480*2, which is passed through an upsampling module and denoted as F5.

[0146] Step 270: Based on the local correlation map, perform feature splicing and convolutional regression optical flow to obtain the predicted optical flow.

[0147] Step 280: Calculate the loss value based on the predicted optical flow and the real labels corresponding to the sample images, optimize the parameters of the optical flow estimation network model, and obtain the target lightweight model.

[0148] In a specific implementation, this embodiment can use L1 loss to optimize model parameters and calculate the loss value based on the predicted optical flow and the ground truth labels corresponding to the sample images. This can include: using the model's predicted optical flow... pred and real label flow gt As input, according to L1 = mean(abs(flow) pred -flow gt )) Calculate the loss L1.

[0149] In the specific implementation, during the model training phase, the model calculation process described above is executed. Specifically, during training, L1 loss is applied using supervised F2, F3, F4, F5, and the labels.

[0150] For example, in practical applications, only the optical flow graph output by F5 can be used.

[0151] Step 290: Optical flow estimation is performed on the input image to be processed using the target lightweight model to obtain the optical flow estimation result.

[0152] In summary, this application's embodiments address optical flow estimation for edge applications. First, the optical flow estimation network model is improved, including lightweight convolutional blocks and optimized model structure, significantly reducing computational complexity and achieving a lightweight optical flow estimation model for edge applications. Then, algorithm improvements are made in areas such as absolute value normalization of feature maps from consecutive frames, feature map transformation, and local correlation calculation. This effectively combines the advantages of deep learning and optical flow estimation, enabling the lightweight model to calculate correlations between feature points, improving tolerance to environmental noise (such as from cameras), and providing more stable optical flow output, thereby enhancing accuracy. Finally, the predicted optical flow and the ground truth label are calculated using a loss function to optimize the parameters of the optical flow estimation network model, further improving accuracy and addressing the limitation of existing methods in edge-based optical flow estimation applications.

[0153] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should know that the embodiments of this application are not limited to the described order of actions, because according to the embodiments of this application, some steps may be performed in other orders or simultaneously.

[0154] like Figure 9 As shown in the embodiments of this application, an optical flow estimation system 900 based on a lightweight optical flow estimation model is also provided, including:

[0155] The model building module 910 is used to build a lightweight optical flow estimation network model with a dual-tower input optical flow network structure, and to build a lightweight convolutional block with a Convblock structure using a row and column convolution design. The lightweight convolutional block is then converted into a two-dimensional convolutional layer through Kronecker product calculation.

[0156] The convolution module 920 is used to perform convolution processing on the input sample image pairs through the optical flow estimation network model, and perform absolute value normalization processing on the feature maps obtained by convolution to obtain a normalized feature map set. The sample image pairs include the previous frame image and the next frame image.

[0157] The Gaussian transform module 930 is used to transform the later frame normalized feature maps in the normalized feature map set using a two-dimensional Gaussian distribution transform to obtain transformed feature maps.

[0158] The local correlation analysis module 940 is used to perform local correlation analysis based on the previous frame standardized feature map in the standardized feature map set and the transformed feature map to obtain a local correlation map.

[0159] The optical flow prediction module 950 is used to perform feature stitching and convolutional regression optical flow based on the local correlation map to obtain the predicted optical flow.

[0160] The model optimization module 960 is used to calculate the loss value based on the predicted optical flow and the real labels corresponding to the sample images, optimize the parameters of the optical flow estimation network model, and obtain the target lightweight model.

[0161] The optical flow estimation module 970 is used to perform optical flow estimation on the input image to be processed using the target lightweight model, and obtain the optical flow estimation result.

[0162] Optional, Gaussian transform module, including:

[0163] The Gaussian weight calculation submodule is used to calculate Gaussian weight information based on the pixel information of the standardized feature map of the next frame in the standardized feature map set, combined with the target center coordinates of the obtained optical flow estimation, using a two-dimensional Gaussian distribution.

[0164] The Gaussian transform submodule is used to perform a weighted transform on the normalized feature map of the next frame based on the Gaussian weight information to obtain a transformed feature map.

[0165] Optional, local correlation analysis module, including:

[0166] The local window feature calculation submodule is used to calculate local window features based on the normalized feature map of the previous frame and by using a sliding window folding method. The local window features are used to characterize the tolerable optical flow error.

[0167] The local correlation map calculation submodule is used to analyze the similarity between the feature points of the previous frame image and the pixels in the same feature point and the nearby region of the subsequent frame image based on the local window features and the transformed feature map, and to obtain the local correlation map.

[0168] It should be noted that the optical flow estimation system based on the lightweight optical flow estimation model provided in this application can execute the optical flow estimation system method based on the lightweight optical flow estimation model provided in any embodiment of this application, and has the corresponding functions and beneficial effects of the execution method.

[0169] In a practical implementation, the aforementioned optical flow estimation system based on a lightweight optical flow estimation model can be integrated into a device. This allows the device to run an optical flow estimation network model improved by both modules and algorithms. As an electronic device, it performs optical flow estimation on the edge, providing optical flow estimation information to applications. This electronic device can consist of two or more physical entities, or it can consist of a single physical entity. For example, the electronic device can be a personal computer (PC), a computer, a server, etc. This application embodiment does not impose specific limitations in this regard.

[0170] like Figure 10 As shown, this application embodiment provides an electronic device, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114. The processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114. The memory 113 is used to store computer programs. When the processor 111 executes the program stored in the memory 113, it implements the steps of the optical flow estimation system method based on a lightweight optical flow estimation model provided in any of the aforementioned method embodiments. For example, the method may include the following steps: constructing a lightweight optical flow estimation network model using a dual-tower input optical flow network structure, and constructing a lightweight convolutional block with a Convblock structure using a row-column convolution design, converting the lightweight convolutional block into a two-dimensional convolutional layer through Kronecker product calculation; performing convolution processing on the input sample image pair through the optical flow estimation network model, and performing absolute value normalization processing on the feature map obtained by convolution to obtain a normalized feature map set, wherein the sample image pair includes a previous frame image and a subsequent frame image; performing transformation processing on the normalized feature map of the subsequent frame in the normalized feature map set using a two-dimensional Gaussian distribution transformation to obtain a transformed feature map; performing local correlation analysis based on the normalized feature map of the previous frame in the normalized feature map set, combined with the transformed feature map, to obtain a local correlation map; performing feature concatenation and convolution regression optical flow based on the local correlation map to obtain a predicted optical flow; calculating the loss value based on the predicted optical flow and the ground truth labels corresponding to the sample images, optimizing the parameters of the optical flow estimation network model to obtain a target lightweight model; and performing optical flow estimation on the input image to be processed through the target lightweight model to obtain an optical flow estimation result.

[0171] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the optical flow estimation method based on a lightweight optical flow estimation model as provided in any of the foregoing method embodiments.

[0172] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0173] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. An optical flow estimation method based on a lightweight optical flow estimation model, characterized in that, include: A lightweight optical flow estimation network model is constructed using a dual-tower input optical flow network structure. A lightweight convolutional block with a Convblock structure is constructed using a row-column convolution design. The lightweight convolutional block is then converted into a two-dimensional convolutional layer through Kronecker product calculation. The optical flow estimation network model performs convolution processing on the input sample image pairs, and performs absolute value normalization processing on the feature maps obtained by convolution to obtain a normalized feature map set. The sample image pairs include the previous frame image and the next frame image. A two-dimensional Gaussian distribution transform is applied to the normalized feature maps of the later frames in the normalized feature map set to obtain the transformed feature maps; Based on the previous frame standardized feature map in the standardized feature map set, and combined with the transformed feature map, local correlation analysis is performed to obtain a local correlation map; Based on the local correlation map, feature splicing and convolutional regression optical flow are performed to obtain the predicted optical flow; The loss value is calculated based on the predicted optical flow and the real labels corresponding to the sample images. The parameters of the optical flow estimation network model are then optimized to obtain the target lightweight model. Optical flow estimation is performed on the input image to be processed using the target lightweight model to obtain the optical flow estimation result.

2. The method according to claim 1, characterized in that, A lightweight convolutional block with a Convblock structure is constructed using a row-column convolutional design. This lightweight convolutional block is then converted into a two-dimensional convolutional layer using Kronecker product calculation, including: In the optical flow estimation network model, a row-column convolution design is adopted, based on Convblock(x) = conv k*1 (conv 1*k (x)) defines a lightweight convolutional block; The lightweight convolutional block is converted into a k*k convolutional layer by Kronecker product calculation; Where k is the kernel size and x is the input image of the layer.

3. The method according to claim 1, characterized in that, The feature maps obtained from convolution are subjected to absolute value normalization to obtain a normalized feature map set, including: Using the feature map set of sample image pairs obtained by convolution as input, according to Calculate the forward result of AbsNorm; Gradient backpropagation of LayerNorm is used, and the forward result is back-differentiated according to LayerNorm≈α*AbsNorm to obtain the standardized feature map set; in, x is the input feature map, n is the feature channel size, AbsNorm is used to calculate the forward result, and LayerNorm is used to calculate the backward derivative. Optical flow is obtained by matching the correlation of each pixel one by one, and it is ensured that the intermediate feature vectors representing pixels have their own uniqueness to reflect the real situation.

4. The method according to claim 1, characterized in that, The normalized feature maps of the later frames in the normalized feature map set are transformed using a two-dimensional Gaussian distribution transform to obtain transformed feature maps, including: Based on the pixel information of the standardized feature map of the next frame in the standardized feature map set, and combined with the target center coordinates of the obtained optical flow estimation, Gaussian weight information is calculated using a two-dimensional Gaussian distribution. Based on the Gaussian weight information, the normalized feature map of the next frame is weighted and transformed to obtain the transformed feature map.

5. The method according to claim 4, characterized in that, Based on the pixel information of the standardized feature map of the later frame in the standardized feature map set, and combined with the target center coordinates obtained from the optical flow estimation, Gaussian weight information is calculated using a two-dimensional Gaussian distribution, including: Using the pixel information of the normalized feature map of subsequent frames and the target center coordinates estimated by optical flow as input, according to Calculate Gaussian weight information; Where, c = (c x ,c y ) represents the coordinates of the center of the Gaussian distribution, μ = (μ x ,μ y ) represents the target center coordinates for optical flow estimation, and σ represents the standard deviation of the Gaussian distribution.

6. The method according to claim 4, characterized in that, Based on the Gaussian weight information, a weighted transformation is performed on the normalized feature map of the subsequent frame to obtain a transformed feature map, including: Obtain the optical flow estimate flow from the model regression and the grid matrix G of the pixel positions in the normalized feature map of the next frame; Using optical flow estimation flow, grid matrix G, and normalized feature map f2 of the next frame as input, according to Perform a Gaussian distribution transformation to obtain the transformation feature map. Wherein, the grid matrix G has dimensions H*W*2, G x Let G be the x-axis, and G be the x-axis. y Let G be the y-axis.

7. The method according to claim 1, characterized in that, Based on the previous frame normalized feature map in the normalized feature map set, and combined with the transformed feature map, local correlation analysis is performed to obtain a local correlation map, including: Based on the normalized feature map of the previous frame, a sliding window folding method is used to calculate local window features, which are used to characterize the tolerable optical flow error. Based on the local window features and the transformed feature map, the similarity between the feature points of the previous frame image and the pixels in the same feature point and nearby area of ​​the subsequent frame image is analyzed to obtain a local correlation map.

8. The method according to claim 7, characterized in that, Based on the normalized feature map of the previous frame, local window features are calculated using a sliding window folding method, including: Construct a sliding window k; Within the sliding window k, according to f1 unfold =Unfold(f1, k), which uses the Unfold operator to slide and unfold the normalized feature map of the previous frame, and calculates the local window feature f1. unfold ; Specifically, based on local window features and the transformed feature map, the similarity between feature points in the previous frame image and pixels in the same feature point and nearby region in the subsequent frame image is analyzed to obtain a local correlation map, including: Using local window features f1 unfold and transformation feature map For input, according to Perform correlation calculations between feature points to obtain the local correlation map corr; f1 is the normalized feature map of the previous frame image.

9. The method according to claim 1, characterized in that, The loss value is calculated based on the predicted optical flow and the ground truth labels corresponding to the sample images, including: Based on the model's predicted optical flow pred and real label flow gt As input, according to L1 = mean(abs(flow) pred -flow gt )) Calculate the loss L1.

10. An optical flow estimation system based on a lightweight model for optical flow estimation, characterized in that, include: The model building module is used to construct a lightweight optical flow estimation network model using a dual-tower input optical flow network structure, and to construct a lightweight convolutional block with a Convblock structure using a row and column convolution design. The lightweight convolutional block is then converted into a two-dimensional convolutional layer through Kronecker product calculation. The convolution module is used to perform convolution processing on the input sample image pairs through the optical flow estimation network model, and perform absolute value normalization processing on the feature maps obtained by convolution to obtain a normalized feature map set. The sample image pairs include the previous frame image and the next frame image. The Gaussian transform module is used to transform the normalized feature maps of the later frame in the normalized feature map set using a two-dimensional Gaussian distribution transform to obtain transformed feature maps. The local correlation analysis module is used to perform local correlation analysis based on the previous frame standardized feature map in the standardized feature map set and the transformed feature map to obtain a local correlation map. The optical flow prediction module is used to perform feature stitching and convolutional regression optical flow based on the local correlation map to obtain the predicted optical flow. The model optimization module is used to calculate the loss value based on the predicted optical flow and the real labels corresponding to the sample images, optimize the parameters of the optical flow estimation network model, and obtain the target lightweight model. The optical flow estimation module is used to perform optical flow estimation on the input image to be processed using the target lightweight model, and obtain the optical flow estimation result.

Citation Information

Patent Citations

  • Micro-expression recognition method and device based on multi-task learning and global cyclic convolution

    CN116030516A

  • Indoor monocular depth estimation method based on self-supervised deep learning

    CN117218174A