Dense Optical Flow Generation Method, Apparatus, Device, and Storage Medium

The dense optical flow is generated through block matching and local image matching, which solves the problem of long calculation time of traditional methods and achieves efficient and accurate estimation of dense optical flow.

CN117934558BActive Publication Date: 2025-07-04MOORE THREADS TECHNOLOGY (CHENGDU) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311326417.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-12
Publication Date
2025-07-04
Estimated Expiration
2043-10-12

AI Technical Summary

Technical Problem

The traditional dense optical flow estimation method requires a lot of iterations, and the calculation time is long, so it cannot achieve high calculation speed.

Method used

By block matching of the current image frame and the next image frame, the first motion estimation vector of the to-processed pixel points is obtained, and the second motion estimation vector is fitted based on the local image matching relationship, and dense optical flow information is finally generated.

Benefits of technology

It improves the accuracy of dense optical flow information, reduces computing time, and is suitable for hardware implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117934558B_ABST
    Figure CN117934558B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a method, apparatus, device, and storage medium for generating dense optical flow. Among them, the method includes: performing block matching on a current image frame and a next image frame to obtain a first motion estimation vector corresponding to each to-be-processed pixel point in the current image frame; the first motion estimation vector corresponding to the to-be-processed pixel point is determined by the block matching result of the unit image block where the to-be-processed pixel point is located; fitting a second motion estimation vector corresponding to each to-be-processed pixel point based on the matching relationship between the local image of each to-be-processed pixel point in the current image frame and the local image of the corresponding matching pixel point in the next image frame; the matching pixel point is determined based on the to-be-processed pixel point and the corresponding first motion estimation vector; generating dense optical flow information based on the first motion estimation vector and the second motion estimation vector corresponding to each to-be-processed pixel point.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to, but is not limited to, the field of image processing technologies, and particularly relates to a method, apparatus, device, and storage medium for generating dense optical flow. Background Art

[0002] In the field of computer vision, optical flow describes the motion trajectories of pixel points in an image or the correspondence of pixel points in a pair of images. Optical flow generally includes sparse optical flow and dense optical flow. Sparse optical flow generally describes significant feature points, while dense optical flow describes all pixel points of an image. In image processing tasks such as behavior recognition and motion prediction, optical flow plays a very important role as a motion feature. Therefore, it is particularly important to accurately estimate optical flow in the field of computer vision. Traditional dense optical flow estimation methods often require a large number of iterations and have a long calculation time, and cannot achieve a high calculation speed. Summary of the Invention

[0003] In view of this, embodiments of the present application at least provide a method, apparatus, device, and storage medium for generating dense optical flow.

[0004] The technical solutions of the embodiments of the present application are implemented as follows:

[0005] On the one hand, an embodiment of the present application provides a method for generating dense optical flow, the method including: performing block matching on a current image frame and a next image frame to obtain a first motion estimation vector corresponding to each to-be-processed pixel point in the current image frame; the first motion estimation vector corresponding to the to-be-processed pixel point is determined by the block matching result of the unit image block where the to-be-processed pixel point is located; fitting a second motion estimation vector corresponding to each to-be-processed pixel point based on the matching relationship between the local image of each to-be-processed pixel point in the current image frame and the local image of the corresponding matching pixel point in the next image frame; the matching pixel point is determined based on the to-be-processed pixel point and the corresponding first motion estimation vector; generating dense optical flow information based on the first motion estimation vector and the second motion estimation vector corresponding to each to-be-processed pixel point.

[0006] In some embodiments, the first motion estimation vector corresponding to the pixel point to be processed is the target block motion vector of the unit image block where the pixel point to be processed is located; the block matching of the current image frame and the next image frame to obtain the first motion estimation vector corresponding to each pixel point to be processed in the current image frame includes: for each preset image block in the current image frame, determining the image block pixel difference of the preset image block at different division scales; based on the image block pixel difference of the preset image block at different division scales, determining the target block motion vector of each unit image block in the preset image block.

[0007] In some embodiments, the different division scales at least include the first division scale corresponding to the preset image block. The determining the target block motion vector of each unit image block in the preset image block based on the image block pixel difference of the preset image block at different division scales includes: in response to the image block pixel difference at the first division scale being greater than or equal to the image block pixel difference at any other division scale, for each unit image block, determining the target block motion vector of the unit image block based on the block motion vector corresponding to the unit image block at each division scale.

[0008] In some embodiments, the different division scales at least include the first division scale corresponding to the preset image block. The determining the target block motion vector of each unit image block in the preset image block based on the image block pixel difference of the preset image block at different division scales includes: in response to the image block pixel difference at the first division scale being less than the image block pixel difference at each other division scale, determining the block motion vector corresponding to the preset image block as the target block motion vector of each unit image block in the preset image block.

[0009] In some embodiments, the determining the image block pixel difference of the preset image block at different division scales includes: for each division scale, respectively determining the image block pixel difference between the first sub-image block and the corresponding second sub-image block at the division scale; the first sub-image block is obtained by dividing the preset image block based on the division scale; the second sub-image block is the image block with the smallest image block pixel difference from the first sub-image block in the next image frame; taking the sum of the image block pixel differences corresponding to the first sub-image block at the division scale as the image block pixel difference of the preset image block at the division scale.

[0010] In some embodiments, determining the target block motion vector of the unit image block based on the block motion vectors corresponding to the unit image block at each of the division scales includes: generating, based on the block motion vectors corresponding to the unit image block at each of the division scales, the pixel difference of the image block corresponding to the unit image block at each of the division scales; and determining the block motion vector corresponding to the smallest pixel difference of the image block as the target block motion vector of the unit image block.

[0011] In some embodiments, the method further includes:

[0012] For each first sub-image block at each of the division scales, determining a search area corresponding to the first sub-image block in the next image frame; the first sub-image block is obtained by dividing the preset image block based on the division scale; within the search area corresponding to the first sub-image block, intercepting candidate image blocks with a sliding window corresponding to the size of the first sub-image block, and calculating the pixel difference of the image block between each candidate image block and the first sub-image block; determining the candidate image block corresponding to the smallest pixel difference of the image block as the second sub-image block corresponding to the first sub-image block; generating, based on the first sub-image block and the corresponding second sub-image block, the block motion vector corresponding to the first sub-image block at the division scale; wherein, the block motion vectors corresponding to the unit image blocks within the first sub-image block at the division scale are the same as the block motion vector corresponding to the first sub-image block.

[0013] In some embodiments, fitting the second motion estimation vector corresponding to each to-be-processed pixel point based on the matching relationship between the local image of each to-be-processed pixel point in the current image frame and the local image of the corresponding matching pixel point in the next image frame includes: for each to-be-processed pixel point, obtaining a plurality of neighboring pixel points corresponding to the to-be-processed pixel point in the current image frame, and obtaining a plurality of neighboring pixel points corresponding to the matching pixel point in the next image frame; and fitting the second motion estimation vector corresponding to each to-be-processed pixel point based on the matching relationship between the plurality of neighboring pixel points corresponding to the to-be-processed pixel point and the plurality of neighboring pixel points corresponding to the matching pixel point.

[0014] In some embodiments, fitting a second motion estimation vector corresponding to each of the to-be-processed pixel points based on the matching relationship between multiple neighboring pixel points corresponding to the to-be-processed pixel points and multiple neighboring pixel points corresponding to the matching pixel points includes: constructing a first matrix function based on the multiple neighboring pixel points corresponding to the to-be-processed pixel points; the first matrix function includes a first pixel function of each neighboring pixel point corresponding to the to-be-processed pixel point, and the first pixel function is a quadratic polynomial function characterizing the association relationship between the position coordinates of the neighboring pixel point in the current image frame and the gray value of the neighboring pixel point; constructing a second matrix function based on the multiple neighboring pixel points corresponding to the neighboring pixel points; the second matrix function includes a second pixel function of each neighboring pixel point corresponding to the neighboring pixel point, and the second pixel function is a quadratic polynomial function characterizing the association relationship between the position coordinates of the neighboring pixel point in the next image frame and the gray value of the neighboring pixel point; generating a second motion estimation vector corresponding to the to-be-processed pixel point based on the matching relationship between the first pixel function determined by the first matrix function and the second pixel function determined by the second matrix function.

[0015] In some embodiments, the quadratic polynomial coefficients of the pixel function include matrix coefficients and vector coefficients; generating a second motion estimation vector corresponding to the to-be-processed pixel point based on the matching relationship between the first pixel function determined by the first matrix function and the second pixel function determined by the second matrix function includes: determining a mean matrix coefficient based on the first matrix coefficient corresponding to the first pixel function and the second matrix coefficient corresponding to the second pixel function; determining a difference vector coefficient based on the first vector coefficient corresponding to the first pixel function and the second vector coefficient corresponding to the second pixel function; and determining the product of the inverse matrix corresponding to the mean matrix coefficient, a preset coefficient, and the difference vector coefficient as the second motion estimation vector corresponding to the to-be-processed pixel point.

[0016] In some embodiments, the dense optical flow information includes a target motion estimation vector corresponding to each of the to-be-processed pixel points; generating dense optical flow information based on the first motion estimation vector and the second motion estimation vector corresponding to each of the to-be-processed pixel points includes: for each of the to-be-processed pixel points, respectively determining a first calculated window pixel difference corresponding to the to-be-processed pixel point under the first motion estimation vector and a second calculated window pixel difference corresponding to the to-be-processed pixel point under the second motion estimation vector based on a preset pixel difference calculation window; and using the motion estimation vector corresponding to the minimum calculated window pixel difference among the first calculated window pixel difference and the second calculated window pixel difference as the target motion estimation vector corresponding to the to-be-processed pixel point.

[0017] On the other hand, an embodiment of the present application provides a dense optical flow generation device, including:

[0018] A first vector acquisition module, configured to perform block matching on a current image frame and a next image frame to obtain a first motion estimation vector corresponding to each to-be-processed pixel point in the current image frame; the first motion estimation vector corresponding to the to-be-processed pixel point is determined by a block matching result corresponding to a unit image block where the to-be-processed pixel point is located;

[0019] A second vector acquisition module, configured to fit a second motion estimation vector corresponding to each to-be-processed pixel point based on a matching relationship between a local image of each to-be-processed pixel point in the current image frame and a local image of a corresponding matching pixel point in the next image frame; the matching pixel point is determined based on the to-be-processed pixel point and the corresponding first motion estimation vector;

[0020] A dense optical flow generation module, configured to generate dense optical flow information based on the first motion estimation vector and the second motion estimation vector corresponding to each to-be-processed pixel point.

[0021] In another aspect, an embodiment of the present application provides a computer device, including a memory and a processor, where the memory stores a computer program that can be run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.

[0022] In still another aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements some or all of the steps in the above method.

[0023] In the embodiment of the present application, by separately obtaining a first motion estimation vector at the image block level of a to-be-processed pixel point and a second motion estimation vector at the pixel level of the to-be-processed pixel point, and generating final dense optical flow information based on the first motion estimation vector and the second motion estimation vector. In this way, the obtained dense optical flow information not only pays attention to the global information of the image but also pays attention to the local information of the image, thereby improving the accuracy of the dense optical flow information; at the same time, by using the first motion estimation vector at the image block level to determine the initial direction during the process of pixel-level motion estimation, compared with the dense optical flow calculation scheme of the related art, since it is not necessary to consider all the image information of the image frame, the multiple iteration processes are omitted, and the optical flow calculation time is reduced.

[0024] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the technical solution of the present application. Description of the Drawings

[0025] The drawings here are incorporated into the specification and constitute a part of this specification. These drawings show embodiments that conform to the present application and are used together with the specification to explain the technical solution of the present application.

[0026] Figure 1 Schematic diagram of the implementation process of a dense optical flow generation method provided by an embodiment of the present application Figure 1 ;

[0027] Figure 2 Schematic diagram of the implementation process of a dense optical flow generation method provided by an embodiment of the present application Figure 2 ;

[0028] Figure 3A Schematic diagram III of the implementation process of a dense optical flow generation method provided by an embodiment of the present application;

[0029] Figure 3B Schematic diagram of the relationship between neighborhood pixel points in the current image frame and the next image frame provided by an embodiment of the present application;

[0030] Figure 4 Schematic diagram of the implementation process of a dense optical flow generation method provided by an embodiment of the present application Figure 4 ;

[0031] Figure 5 Schematic diagram of the data flow of a dense optical flow calculation system provided by an embodiment of the present application;

[0032] Figure 6 Schematic diagram of the motion estimation of a unit image block provided by an embodiment of the present application;

[0033] Figure 7 Flowchart of the motion estimation of a unit image block in the refined division state provided by an embodiment of the present application;

[0034] Figure 8 Schematic diagram of the motion estimation vector of a dense optical flow provided by an embodiment of the present application;

[0035] Figure 9 Schematic diagram of a neighborhood point provided by an embodiment of the present application;

[0036] Figure 10 Schematic diagram of the composition structure of a dense optical flow generation device provided by an embodiment of the present application;

[0037] Figure 11 Schematic diagram of the hardware entity of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0038] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be construed as limitations on this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.

[0039] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict. The terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing this application and are not intended to limit this application.

[0041] An embodiment of this application provides a method for generating dense optical flow, which can be executed by a processor of a computer device. Among them, the computer device may refer to a device with data processing capabilities such as a server, a laptop computer, a tablet computer, a desktop computer, a smart TV, a set-top box, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant, a dedicated messaging device, a portable game device), etc.

[0042] Figure 1 Schematic diagram of the implementation process of a method for generating dense optical flow provided by an embodiment of this application Figure 1 , as Figure 1 shown, the method includes the following steps S101 to S103:

[0043] Step S101: Perform block matching on the current image frame and the next image frame to obtain a first motion estimation vector corresponding to each pixel point to be processed in the current image frame; the first motion estimation vector corresponding to the pixel point to be processed is determined by the block matching result of the unit image block where the pixel point to be processed is located.

[0044] In some embodiments, the above block matching method targets unit image blocks. For each unit image block in the current image frame, a block matching result corresponding to the unit image block is generated based on the block matching algorithm. That is, the target image block corresponding to each unit image block is determined in the next image frame, and then a first motion estimation vector corresponding to each unit image block is generated based on each unit image block and its corresponding target image block.

[0045] In the embodiments of the present application, the processor sets the first motion estimation vector of each to-be-processed pixel point in the unit image block to the first motion estimation vector corresponding to the unit image block; that is to say, for all to-be-processed pixel points in the same unit image block, their corresponding first motion estimation vectors are the same.

[0046] It can be understood that the first motion estimation vector is a tile-level motion estimation vector determined by considering the texture information of the unit image block where the pixel point is located.

[0047] Step S102: Fit a second motion estimation vector corresponding to each to-be-processed pixel point based on the matching relationship between the local image of each to-be-processed pixel point in the current image frame and the local image of the corresponding matching pixel point in the next image frame; the matching pixel point is determined based on the to-be-processed pixel point and the corresponding first motion estimation vector.

[0048] In some embodiments, for each to-be-processed pixel point, the position coordinate of the matching pixel point in the next image frame can be determined based on the position coordinate of the to-be-processed pixel point in the current image frame and the first motion estimation vector corresponding to the to-be-processed pixel point, and then the matching pixel point is obtained.

[0049] In the embodiments of the present application, considering that the object motion in adjacent frames is relatively "small", that is, for the to-be-processed pixel point, the image information of the local image where it is located is substantially the same as the local image of the matching pixel point in the next image frame. Therefore, for each to-be-processed pixel point, step S102 can fit the second motion estimation vector corresponding to the to-be-processed pixel point based on the matching relationship between the local image of the to-be-processed pixel point in the current image frame and the local image of the corresponding matching pixel point in the next image frame.

[0050] In some embodiments, the processor can fit the second motion estimation vector corresponding to the to-be-processed pixel point based on the optical flow method. Among them, the optical flow method can be any one of the following: Horn-Schunck optical flow method, Lucas-Kanade optical flow method, and Farneback optical flow method, etc.

[0051] It can be understood that the second motion estimation vector is a pixel-level motion estimation vector determined by considering the texture information of the local image where the pixel is located.

[0052] Step S103: Generate dense optical flow information based on the first motion estimation vector and the second motion estimation vector corresponding to each of the to-be-processed pixel points.

[0053] In the embodiment of the present application, for each to-be-processed pixel point, the difference in the image information corresponding to the first motion estimation vector and the difference in the image information corresponding to the second motion estimation vector are determined between the current image frame and the next image frame. The motion estimation vector with a smaller difference is used as the final motion estimation vector of the to-be-processed pixel point, and thus the dense optical flow information of the entire current image frame and the next image frame is obtained.

[0054] In some embodiments, based on the position coordinates of the to-be-processed pixel point in the current image frame and the first motion estimation vector, the matching position coordinates can be determined in the next image frame, and the pixel value at the matching position coordinates is obtained as the pixel value corresponding to the first motion estimation vector; based on the same method, the pixel value corresponding to the second motion estimation vector is determined, and the difference between the pixel value corresponding to the to-be-processed pixel point and the pixel value corresponding to the first motion estimation vector, and the difference between the pixel value corresponding to the to-be-processed pixel point and the pixel value corresponding to the second motion estimation vector are respectively determined; the motion estimation vector with a smaller difference is used as the final motion estimation vector of the to-be-processed pixel point.

[0055] In some embodiments, based on the position coordinates of the to-be-processed pixel point in the current image frame and the first motion estimation vector, the matching position coordinates can be determined in the next image frame, and the local image at the matching position coordinates is obtained as the local image corresponding to the first motion estimation vector; based on the same method, the local image corresponding to the second motion estimation vector is determined, and the difference between the local image corresponding to the to-be-processed pixel point and the local image corresponding to the first motion estimation vector, and the difference between the local image corresponding to the to-be-processed pixel point and the local image corresponding to the second motion estimation vector are respectively determined; the motion estimation vector with a smaller difference is used as the final motion estimation vector of the to-be-processed pixel point.

[0056] In other embodiments, the first motion estimation vector and the second motion estimation vector can also be directly weighted and averaged, and the fused motion estimation vector is used as the final motion estimation vector of the to-be-processed pixel point, and thus the dense optical flow information of the entire current image frame and the next image frame is obtained.

[0057] In the embodiments of the present application, by separately obtaining the first motion estimation vector of the pixel to be processed at the image block level and the second motion estimation vector of the pixel to be processed at the pixel level, and generating the final dense optical flow information based on the first motion estimation vector and the second motion estimation vector. In this way, the obtained dense optical flow information not only focuses on the global information of the image but also focuses on the local information of the image, thereby improving the accuracy of the dense optical flow information; at the same time, by using the first motion estimation vector at the image block level to determine the preliminary direction during the pixel-level motion estimation process, compared with the dense optical flow calculation scheme of the related art, since it is not necessary to consider all the image information of the image frame, multiple iteration processes are omitted, and the optical flow calculation time is reduced.

[0058] Figure 2 is an optional process schematic diagram of the dense optical flow generation method provided by the embodiments of the present application Figure 2 , and this method can be executed by the processor of the computer device. Based on Figure 1 , the first motion estimation vector corresponding to the pixel to be processed is the target block motion vector of the unit image block where the pixel to be processed is located; Figure 1 S101 in Figure 2 can be updated to S201 to S203, and will be described in combination with the steps shown in

[0059] Step S201: For each preset image block in the current image frame, determine the pixel difference of the preset image block at different division scales.

[0060] In some embodiments, it is necessary to divide the preset image block at different division scales, that is, divide the preset image block at at least two division scales, and then at least one first sub-image block corresponding to each division scale can be obtained.

[0061] Among them, the preset image block includes a plurality of the unit image blocks.

[0062] In the embodiments of the present application, step S101 can be executed by a video codec, and the above preset image block is the smallest processing unit of the video codec. Exemplarily, the size of the preset image block can be set to 32*32, that is, the preset image block is a square image block with a side length of 32 pixels. The above current image frame can be divided into a plurality of preset image blocks, and the processor needs to execute step S201 for each preset image block in the plurality of preset image blocks.

[0063] In some embodiments, in order to obtain more refined motion estimation data, the preset image block may be further divided to obtain a plurality of unit image blocks corresponding to the preset image block. In the current embodiment, by determining the target block motion vector corresponding to each unit image block, and then using the target block motion vector corresponding to the unit image block as the target block motion vector of each pixel to be processed in the unit image block.

[0064] It can be understood that the size of the unit image block is smaller than the size of the preset image block. In some embodiments, a preset image block may include 4 to the power of n unit image blocks, where n is an integer greater than or equal to 1. Exemplarily, when the size of the preset image block is 32*32, the size of the unit image block may be 16*16, and at this time, a preset image block may include 4 unit image blocks; when the size of the preset image block is 32*32, the size of the unit image block may be 8×8, and at this time, a preset image block may include 16 unit image blocks; when the size of the preset image block is 32*32, the size of the unit image block may be 4*4, and at this time, a preset image block may include 64 unit image blocks, and so on.

[0065] In some embodiments, the different division scales at least include a first division scale and a second division scale. The first sub-image block corresponding to the first division scale has the same size as the preset image block; the first sub-image block corresponding to the second division scale has the same size as the unit image block.

[0066] In the embodiments of the present application, at least one third division scale other than the first division scale and the second division scale is further provided. It can be understood that the size of the first sub-image block corresponding to the third division scale is smaller than the size of the first sub-image block corresponding to the first division scale, and at the same time, the size of the first sub-image block corresponding to the third division scale is larger than the size of the first sub-image block corresponding to the second division scale.

[0067] Exemplarily, when the size of the preset image block is 32*32, the size of the first sub-image block corresponding to the first division scale is also 32*32, that is, the preset image block is directly used as the first sub-image block corresponding to the first division scale; correspondingly, when the size of the unit image block is 8×8, the size of the first sub-image block corresponding to the second division scale is 8×8, that is, the preset image block is directly divided into 16 first sub-image blocks corresponding to the second division scale. It can be understood that the different division scales may further include a third division scale, and the size of the first sub-image block corresponding to the third division scale is 16*16, that is, the preset image block is directly divided into 4 first sub-image blocks corresponding to the third division scale.

[0068] In this embodiment, the pixel difference of the image block at the division scale is related to the motion estimation of each first sub-image block at the division scale.

[0069] It can be understood that for each division scale, block matching can be performed on each first sub-image block at the division scale in the next frame of the image respectively, and the minimum pixel difference of the image block corresponding to each first sub-image block in the block matching process can be obtained. The sum of the minimum pixel differences of the image blocks corresponding to all the first sub-image blocks at the division scale is used as the pixel difference of the preset image block at the division scale.

[0070] In some embodiments, the above-mentioned determination of the pixel difference of the preset image block at different division scales can be implemented through steps S2011 to S2012.

[0071] Step S2011: For each division scale, respectively determine the pixel difference of the image block between the first sub-image block at the division scale and the corresponding second sub-image block.

[0072] Wherein, the first sub-image block is obtained by dividing the preset image block based on the division scale; the second sub-image block is the image block with the smallest pixel difference of the image block between it and the first sub-image block in the next image frame.

[0073] In the embodiment of the present application, for the first sub-image block and the corresponding second sub-image block, the difference of the pixel values is taken for each pixel point, and the sum of the differences of the pixel values of each pixel point is used as the pixel difference of the image block between the first sub-image block and the corresponding second sub-image block.

[0074] Among them, the second sub-image block matching the first sub-image block can be determined in the next image frame by the block matching method. In some embodiments, the second sub-image block is the image block with the smallest pixel difference of the image block between it and the first sub-image block in the search area corresponding to the first sub-image block in the next image frame.

[0075] Step 2012: Use the sum of the pixel differences of the image blocks corresponding to the first sub-image blocks at the division scale as the pixel difference of the preset image block at the division scale.

[0076] In the embodiment of the present application, since the number of first sub-image blocks corresponding to different division scales is also different, here, for each preset image block in each division scale, by statistically analyzing the pixel differences of the image blocks between all the first sub-image blocks and the corresponding second sub-image blocks in the preset image block, it is possible to obtain the differences of block matching with different granularities while unifying each division scale into a comparison space of the same size, which is convenient for judging whether refined matching is required in subsequent steps.

[0077] Step S202: Determine the target block motion vector of each unit image block in the preset image block based on the pixel differences of the image blocks at different division scales of the preset image block.

[0078] In some embodiments, for different division scales, the preset image block can be divided into different numbers of first sub-image blocks. Correspondingly, for different division scales, the sizes of the corresponding first sub-image blocks are also different.

[0079] For each division scale, the preset image block can be divided into at least one first sub-image block based on this division scale. It can be understood that in the process of calculating the block motion vector corresponding to the first sub-image block, the smaller the size of the first sub-image block obtained based on the division scale, the more the obtained block motion vector focuses on the local information of the image; the larger the size of the first sub-image block obtained based on the division scale, the more the obtained block motion vector focuses on the overall information of the image.

[0080] Among them, the block motion vector corresponding to the unit image block in the first sub-image block at the division scale is the same as the block motion vector corresponding to the first sub-image block.

[0081] Based on this, in the embodiments of the present application, the preset image block is divided by different division scales, and the target block motion vector of the unit image block is determined based on the block motion vectors of the first sub-image blocks obtained at each division scale. In this way, the target block motion vector of the unit image block obtained not only focuses on the global information of the image but also can focus on the local information of the image, thereby improving the accuracy of the target block motion vector.

[0082] In some embodiments, the above step S202 can be implemented through step S2021.

[0083] Step S2021: In response to the pixel difference of the image block at the first division scale being greater than or equal to the pixel differences of the image blocks at any other division scale, for each unit image block, determine the target block motion vector of the unit image block based on the block motion vectors corresponding to the unit image block at each division scale.

[0084] Among them, the other division scales include all scales other than the first division scale. When the different division scales include the first division scale and the second division scale, the other division scale includes the second division scale; when the different division scales include the first division scale, the second division scale, and at least one third division scale, the other division scale includes the second division scale and at least one third division scale.

[0085] Here, it is necessary to determine whether the pixel difference of the image block at the first division scale is less than the pixel differences of the image blocks at each other division scale. In the case where the pixel difference of the image block at the first division scale is greater than or equal to the pixel difference of any other division scale, for each of the unit image blocks, based on the block motion vectors corresponding to the unit image block at each division scale, the target block motion vector of the unit image block is determined.

[0086] In the embodiments of the present application, the above can be achieved in the following manner: based on the block motion vectors corresponding to the unit image block at each division scale, generate the pixel differences of the image blocks corresponding to the unit image block at each division scale; determine the block motion vector corresponding to the smallest pixel difference of the image block as the target block motion vector of the unit image block.

[0087] Among them, for each unit image block in the current image frame, based on different division scales, the block motion vectors corresponding to the unit image block are also different.

[0088] In the embodiments of the present application, based on the unit image block and the block motion vectors at each division scale, the target image block corresponding to the unit image block can be determined in the next image frame, and the difference between the pixel values of each pixel point of the unit image block and the target image block is taken, and the sum of the differences of the pixel values of each pixel point is used as the pixel difference of the image block corresponding to the unit image block at each division scale.

[0089] Exemplarily, in the case where at least two division scales include the first division scale and the second division scale, the pixel difference of the image block corresponding to the first division scale can be determined based on the unit image block and the target image block at the first division scale; the pixel difference of the image block corresponding to the second division scale can be determined based on the unit image block and the target image block at the second division scale. In the case where the pixel difference of the image block corresponding to the first division scale is less than the pixel difference of the image block corresponding to the second division scale, the block motion vector of the unit image block at the first division scale is determined as the target block motion vector of the unit image block; in the case where the pixel difference of the image block corresponding to the first division scale is greater than the pixel difference of the image block corresponding to the second division scale, the block motion vector of the unit image block at the second division scale is determined as the target block motion vector of the unit image block.

[0090] In the embodiments of the present application, by comparing the block motion vectors of the unit image block at different division scales and determining the block motion vector corresponding to the smallest pixel difference of the image block as the target block motion vector of the unit image block, the global information and local information can be integrated to improve the accuracy of block matching.

[0091] In some embodiments, the above step S202 may be implemented through step S2022.

[0092] Step S2022, in response to the pixel difference of the image block at the first division scale being less than the pixel differences of the image blocks at each other division scale, determines the block motion vector corresponding to the preset image block as the target block motion vector of each unit image block in the preset image block.

[0093] In some embodiments, in response to the pixel difference of the image block at the first division scale being less than the pixel differences of the image blocks at each other division scale, it indicates that the motion estimation at the current first division scale is relatively better, and the target block motion vector of each unit image block in the preset image block can be set as the block motion vector corresponding to the preset image block.

[0094] In the embodiments of the present application, by determining whether the pixel difference of the image block at the first division scale is greater than or equal to the pixel difference of the image block at any other division scale, and only when the pixel difference of the image block at the first division scale is greater than or equal to the pixel difference of the image block at any other division scale, comparing the block motion vectors corresponding to the unit image block at each of the division scales, and then determining the target block motion vector of the unit image block, in this way, unnecessary computational amount can be avoided, and the acquisition efficiency of the first motion estimation vector is improved as a whole.

[0095] In some embodiments, the method further includes: for each first sub-image block at each of the division scales, determining a search area corresponding to the first sub-image block in the next image frame; the first sub-image block is obtained by dividing the preset image block based on the division scale; within the search area corresponding to the first sub-image block, intercepting candidate image blocks with a sliding window corresponding to the size of the first sub-image block, and calculating the pixel difference between each candidate image block and the first sub-image block; determining the candidate image block corresponding to the smallest pixel difference as the second sub-image block corresponding to the first sub-image block; generating the block motion vector corresponding to the first sub-image block at the division scale based on the first sub-image block and the corresponding second sub-image block; wherein, the block motion vector corresponding to the unit image block within the first sub-image block at the division scale is the same as the block motion vector corresponding to the first sub-image block.

[0096] In some embodiments, the search area corresponding to the first sub-image block is an area constructed with the position coordinates of the first sub-image block as the center. Exemplarily, taking the search area as a rectangular area as an example, when the position coordinates of the first sub-image block are (x, y) in the current image frame, the search area corresponding to the first sub-image block is a rectangular area from (x - m, y - n) to (x + m, y + n) in the next image frame.

[0097] In an embodiment of the present application, within the search area corresponding to the first sub-image block, the corresponding candidate image blocks are sequentially intercepted from the next image frame by means of the sliding window, and the difference in pixel values is taken pixel by pixel. The sum of the differences in pixel values of each pixel point in the sliding window obtained is used as the pixel difference between the candidate image block and the first sub-image block.

[0098] It can be understood that in the embodiment of the present application, the sliding step of the sliding window can be adaptively set based on the calculation accuracy, and the minimum setting is one pixel.

[0099] In an embodiment of the present application, after obtaining the pixel difference between each candidate image block in the search area and the first sub-image block, the candidate image block corresponding to the minimum pixel difference is determined as the second sub-image block corresponding to the first sub-image block.

[0100] Here, the vector from the position coordinates of the first sub-image block to the position coordinates of the second sub-image block is used as the block motion vector corresponding to the first sub-image block at this division scale.

[0101] In the above embodiment, the position coordinates of the image block can be the relative coordinates of the center point of the image block in the image frame, or the relative coordinates of the upper left corner point of the image block in the image frame. Of course, it can also be other relative coordinates representing the relative position relationship between the image block and the image frame. The present application does not make any limitations in this regard.

[0102] In an embodiment of the present application, by setting a search area and performing block matching on the first sub-image block in the form of a sliding window within the search area, the matching efficiency can be improved while ensuring the matching accuracy.

[0103] Figure 3A It is an optional flowchart three of the dense optical flow generation method provided by the embodiment of the present application, and this method can be executed by the processor of a computer device. Based on Figure 1 , Figure 1 S102 in can be updated to S301 to S302, and will be described in combination with the steps shown in Figure 3A .

[0104] Step S301: For each of the to-be-processed pixel points, obtain a plurality of neighboring pixel points corresponding to the to-be-processed pixel point in the current image frame, and obtain a plurality of neighboring pixel points corresponding to the matching pixel point in the next image frame.

[0105] In an embodiment of the present application, the plurality of neighboring pixel points corresponding to the to-be-processed pixel point are evenly distributed around the to-be-processed pixel point. Correspondingly, the plurality of neighboring pixel points corresponding to the matching pixel point are evenly distributed around the matching pixel point.

[0106] It can be understood that the number of multiple neighboring pixels corresponding to the pixel to be processed is the same as the number of multiple neighboring pixels corresponding to the matching pixel, and there is a one-to-one correspondence relationship.

[0107] Exemplarily, please refer to Figure 3B , which shows a schematic diagram of the relationship between neighboring pixels in a current image frame and a next image frame. Among them, the pixel P10 to be processed in the current image frame 31 matches the matching pixel P20 in the next image frame 32, and the matching pixel P20 can be determined based on the pixel P10 to be processed and the corresponding first motion estimation vector 33. The multiple neighboring pixels corresponding to the pixel P10 to be processed (6 in the figure, namely P11 to P16) are evenly distributed around the pixel P10 to be processed. At the same time, the multiple neighboring pixels corresponding to the matching pixel P20 (6 in the figure, namely P21 to P26) are evenly distributed around the matching pixel P20. At the same time, the neighboring pixel P11 in the current image frame 31 corresponds to the neighboring pixel P21 in the next image frame 32, the neighboring pixel P12 in the current image frame 31 corresponds to the neighboring pixel P22 in the next image frame 32, and so on. That is, there is a one-to-one correspondence relationship between the multiple neighboring pixels corresponding to the pixel to be processed and the multiple neighboring pixels corresponding to the matching pixel.

[0108] It can be understood that the selection method of the above multiple neighboring pixels can be adaptively configured based on specific implementation scenarios, and the specific positions of the neighboring pixels are not limited in the embodiments of the present application.

[0109] In this embodiment, by respectively obtaining multiple neighboring pixels around the pixel to be processed and multiple neighboring pixels around the matching pixel, not only can the image information of the local image of the current pixel be reflected by several neighboring pixels, but also, compared with the scheme that needs to perform multiple iterations in the process of matching the local image where the current pixel is located in the related art, the embodiments of the present application can reduce the calculation amount in the subsequent process of determining the second motion estimation vector.

[0110] In some embodiments, the number of neighboring pixels can be determined based on the type of optical flow method used in the implementation process. For example, in the case of fitting the second motion estimation vector corresponding to the pixel to be processed based on the Farneback optical flow method, at least 6 neighboring pixels are required.

[0111] Step S302: Based on the matching relationship between the multiple neighboring pixels corresponding to the pixel to be processed and the multiple neighboring pixels corresponding to the matching pixel, fit the second motion estimation vector corresponding to each pixel to be processed.

[0112] In some embodiments, based on a plurality of neighboring pixel points corresponding to the pixel point to be processed, the association relationship between the texture information and the position coordinates in the local image where the pixel point to be processed is located can be extracted; at the same time, based on a plurality of neighboring pixel points corresponding to the matching pixel point, the association relationship between the texture information and the position coordinates in the local image where the matching pixel point is located can be extracted; based on the matching relationship between the texture information in the local image where the pixel point to be processed is located and the texture information in the local image where the matching pixel point is located, a second motion estimation vector of the pixel point to be processed is generated.

[0113] Exemplarily, in the case of fitting the second motion estimation vector corresponding to the pixel point to be processed based on the Farneback optical flow method, a corresponding quadratic polynomial function can be constructed based on a plurality of neighboring pixel points corresponding to the pixel point to be processed, and a corresponding quadratic polynomial function can be constructed based on a plurality of neighboring pixel points corresponding to the matching pixel point; since both of these quadratic polynomial functions characterize the association relationship between the image information and the position coordinates, and the image information in the local image where the pixel point to be processed is located is associated with the image information in the local image where the matching pixel point is located by the second motion estimation vector, therefore, the second motion estimation vector of the pixel point to be processed can be fitted based on the two obtained quadratic polynomial functions.

[0114] In the embodiments of the present application, by respectively obtaining a plurality of neighboring pixel points around the pixel point to be processed and a plurality of neighboring pixel points around the matching pixel point, the image information of the local image of the current pixel point can be reflected by several neighboring pixel points; at the same time, compared with the solution in the related art that requires multiple iterations in the process of matching the local image where the current pixel point is located, in the embodiments of the present application, based on the matching relationship between the plurality of neighboring pixel points corresponding to the pixel point to be processed and the plurality of neighboring pixel points corresponding to the matching pixel point, the second motion estimation vector corresponding to each pixel point to be processed is fitted, avoiding the process of multiple iterative calculations, reducing the main computational amount of optical flow estimation, and shortening the calculation time.

[0115] In some embodiments, the above-mentioned matching relationship between the local image of each pixel point to be processed in the current image frame and the local image of the corresponding matching pixel point in the next image frame can be realized through steps S3021 to S3023, and the second motion estimation vector corresponding to each pixel point to be processed is fitted.

[0116] Step S3021: Based on a plurality of neighboring pixel points corresponding to the pixel point to be processed, a first matrix function is constructed.

[0117] Among them, the first matrix function includes a first pixel function for each neighborhood pixel corresponding to the pixel to be processed, and the first pixel function is a quadratic polynomial function representing the correlation between the position coordinates of the neighborhood pixel in the current image frame and the gray value of the neighborhood pixel.

[0118] In the embodiment of the present application, it is necessary to obtain the gray value and position coordinates corresponding to each of the neighborhood pixels corresponding to the pixel to be processed, and construct a quadratic polynomial function corresponding to each of the neighborhood pixels corresponding to the pixel to be processed as the first pixel function of the neighborhood pixel corresponding to the pixel to be processed, as shown in formula (1).

[0119] f(X)~X T AX + b T X + c Formula (1);

[0120] Among them, f(X) is the pixel value of the neighborhood pixel corresponding to the pixel to be processed, A is a 2×2 matrix, b is a 2×1 matrix, X is the coordinate (x, y) of the neighborhood pixel corresponding to the pixel to be processed, and c is a scalar. Based on this, formula (1) can be written as formula (2).

[0121]

[0122] Among them, A is a 2×2 matrix b is a 2×1 matrix c is r1.

[0123] It can be seen that in order to determine the correlation between the texture information and the position coordinates in the local image where the pixel to be processed is located, at least 6 neighborhood pixels corresponding to the pixel to be processed are required, and then 6 quadratic polynomial functions, that is, the first matrix function, are constructed. By solving the first matrix function, the first pixel function corresponding to the pixel to be processed is obtained.

[0124] Step S3022: Based on the multiple neighborhood pixels corresponding to the neighborhood pixel, construct a second matrix function.

[0125] Among them, the second matrix function includes a second pixel function for each neighborhood pixel corresponding to the neighborhood pixel, and the second pixel function is a quadratic polynomial function representing the correlation between the position coordinates of the neighborhood pixel in the next image frame and the gray value of the neighborhood pixel.

[0126] In the embodiment of the present application, it is necessary to obtain the gray value and position coordinates corresponding to each of the neighborhood pixels corresponding to the matching pixel, and construct a quadratic polynomial function corresponding to each of the neighborhood pixels corresponding to the matching pixel as the second pixel function of the neighborhood pixel corresponding to the matching pixel, as shown in formula (3).

[0127] f′(X′)~X′ T A′X′ + b′ T X′ + c′ Formula (3);

[0128] Wherein, f′(X′) is the pixel value of the neighborhood pixel corresponding to the matching pixel, A′ is a 2×2 matrix, b′ is a 2×1 matrix, X′ is the coordinate (x′, y′) of the neighborhood pixel corresponding to the matching pixel, and c′ is a scalar.

[0129] Similar to Formula (2), in order to determine the correlation between the texture information and the position coordinates in the local image where the matching pixel is located, at least 6 neighborhood pixels corresponding to the matching pixel (the same number as the neighborhood pixels corresponding to the pixel to be processed) are required, and then 6 quadratic polynomial functions, that is, the second matrix function, are constructed. By solving the second matrix function, the second pixel function corresponding to the matching pixel is obtained.

[0130] Step S3023, generate the second motion estimation vector corresponding to the pixel to be processed based on the matching relationship between the first pixel function determined by the first matrix function and the second pixel function determined by the second matrix function.

[0131] In the embodiments of the present application, the first pixel function characterizes the correlation between the current local image information and the position coordinates in the current image frame, and the second pixel function characterizes the correlation between the current local image information and the position coordinates in the next image frame; based on the fact that the current local image information in the current image frame is the same as the current local image information in the next image frame and there is a coordinate offset, a relational expression between the first pixel function and the second pixel function can be constructed, and then the coordinate offset, that is, the second motion estimation vector corresponding to the pixel to be processed, can be obtained.

[0132] X = X′ - d Formula (4);

[0133] Wherein, X is the coordinate of the pixel in the current image frame, X′ is the coordinate of the pixel in the next image frame, and d is the second motion estimation vector. The relational expression between the above first pixel function and the second pixel function can be expressed as Formula (5);

[0134] f(X) = f′(X′ - d) Formula (5);

[0135] In the embodiments of the present application, a first pixel function representing the association relationship between the current local image information and the position coordinates in the current image frame and a second pixel function representing the association relationship between the current local image information and the position coordinates in the next image frame are respectively constructed by the Farneback optical flow method, and then the coordinate offset is solved to obtain the second motion estimation vector corresponding to the pixel point to be processed. In this way, compared with the conventional Farneback optical flow method that requires multiple iterations for all pixel points in the local image, the main computational amount of optical flow estimation is reduced and the calculation time is shortened.

[0136] In some embodiments, the quadratic polynomial coefficients of the pixel function include matrix coefficients and vector coefficients; generating the second motion estimation vector corresponding to the pixel point to be processed based on the matching relationship between the first pixel function determined by the first matrix function and the second pixel function determined by the second matrix function includes: determining the mean matrix coefficient based on the first matrix coefficient corresponding to the first pixel function and the second matrix coefficient corresponding to the second pixel function; determining the difference vector coefficient based on the first vector coefficient corresponding to the first pixel function and the second vector coefficient corresponding to the second pixel function; and determining the product of the inverse matrix corresponding to the mean matrix coefficient, the preset coefficient, and the difference vector coefficient as the second motion estimation vector corresponding to the pixel point to be processed.

[0137] Among them, the matrix coefficient is A in the above formula (1), and the vector coefficient is b in the above formula (1). To facilitate understanding of the above method for determining the second motion estimation vector, please refer to the following derivation process: In the embodiments of the present application, by transforming formula (5), formula (6) can be obtained.

[0138] f′(X′ - d) = (X′ - d) T A′(X′ - d) + b′ T (X′ - d) + c′ = X T A′X + (b′ - 2A′d) T X + d T Ad - b′ T d + c′ formula (6);

[0139] Since the parameter A in the quadratic polynomial function represents the image information, therefore, A corresponding to the pixel point to be processed in the current image frame should be the same as A′ corresponding to the matching pixel point in the next image frame. The coordinate offset d can be expressed as formula (7);

[0140]

[0141] Among them, A″ can be the first matrix coefficient corresponding to the first pixel function or the second matrix coefficient corresponding to the second pixel function. To improve the accuracy of image information, the average value of the first matrix coefficient corresponding to the first pixel function and the second matrix coefficient corresponding to the second pixel function can be taken as the mean matrix coefficient. (b′ - b) is the difference between the first vector coefficient corresponding to the first pixel function and the second vector coefficient corresponding to the second pixel function, that is, the difference vector coefficient, and 1 / 2 is a preset coefficient; furthermore, the coordinate offset d can be expressed as the product of the inverse matrix corresponding to the mean matrix coefficient, the preset coefficient, and the difference vector coefficient.

[0142] Among them, considering that A may not be a singular matrix and its inverse cannot be directly calculated, formula (7) can be transformed to obtain formula (8), and then the motion estimation vector of the current pixel (i.e., the local optical flow offset) can be obtained.

[0143] d = (A″ T A) -1 A″ T *(b′ - b) / 2 Formula (8).

[0144] In the embodiments of the present application, the second motion estimation vector corresponding to the pixel to be processed can be quickly obtained through matrix operations, improving the calculation efficiency.

[0145] Figure 4 is an optional process schematic diagram of the dense optical flow generation method provided by the embodiments of the present application Figure 4 , and this method can be executed by the processor of a computer device. Based on Figure 1 , the dense optical flow information includes the target motion estimation vector corresponding to each pixel to be processed; Figure 1 S103 in Figure 4 can be updated to S401 to S402, and will be described in combination with the steps shown in

[0146] Step S401: For each pixel to be processed, based on a preset pixel difference calculation window, determine the first calculation window pixel difference corresponding to the pixel to be processed under the first motion estimation vector and the second calculation window pixel difference corresponding to the pixel to be processed under the second motion estimation vector, respectively.

[0147] In some embodiments, the pixel difference calculation window is a window centered on the pixel to be processed. Among them, the shape of the pixel difference calculation window can be rectangular, square or circular, and the embodiments of the present application do not limit this. Exemplarily, taking the shape of the pixel difference calculation window as a square as an example, the size of the pixel difference calculation window can be (2n + 1)×(2n + 1). Correspondingly, the pixel to be processed is at the center point of the corresponding pixel difference calculation window, that is, the relative coordinates in the pixel difference calculation window are (n + 1, n + 1).

[0148] In the embodiments of the present application, based on the first motion estimation vector and the pixel to be processed, the first target pixel corresponding to the pixel to be processed can be determined in the next image frame; a corresponding pixel difference calculation window is constructed based on the pixel to be processed in the current image frame, and a corresponding pixel difference calculation window is constructed based on the first target pixel in the next image frame; the pixel values of each pixel point in the pixel difference calculation window corresponding to the pixel to be processed and the pixel difference calculation window corresponding to the first target pixel are taken as differences pixel by pixel, and the sum of the differences of the pixel values of each pixel point in the obtained pixel difference calculation window is used as the first calculation window pixel difference corresponding to the pixel to be processed.

[0149] Exemplarily, taking the pixel difference calculation window as a square window with a size of (2n + 1)×(2n + 1) as an example, (2n + 1)×(2n + 1) pixel points are intercepted in the current image frame based on the pixel to be processed and the pixel difference calculation window. Correspondingly, (2n + 1)×(2n + 1) pixel points are intercepted in the next image frame based on the first target pixel and the pixel difference calculation window. After taking the differences of the pixel values pixel by pixel, the differences of the pixel values of (2n + 1)×(2n + 1) pixel points can be obtained, and the sum after accumulation is the first calculation window pixel difference corresponding to the pixel to be processed.

[0150] In some embodiments, the second calculation window pixel difference is determined based on the same calculation method as the first calculation window pixel difference. Based on the second motion estimation vector and the pixel to be processed, the second target pixel corresponding to the pixel to be processed can be determined in the next image frame; a corresponding pixel difference calculation window is constructed based on the pixel to be processed in the current image frame, and a corresponding pixel difference calculation window is constructed based on the second target pixel in the next image frame; the pixel values of each pixel point in the pixel difference calculation window corresponding to the pixel to be processed and the pixel difference calculation window corresponding to the second target pixel are taken as differences pixel by pixel, and the sum of the differences of the pixel values of each pixel point in the obtained pixel difference calculation window is used as the second calculation window pixel difference corresponding to the pixel to be processed.

[0151] Step S402: Use the motion estimation vector corresponding to the minimum calculated window pixel difference among the first calculated window pixel difference and the second calculated window pixel difference as the target motion estimation vector corresponding to the pixel point to be processed.

[0152] In some embodiments, when the first calculated window pixel difference is less than the second calculated window pixel difference, that is, the first calculated window pixel difference is the minimum calculated window pixel difference, use the first motion estimation vector as the target motion estimation vector corresponding to the pixel point to be processed.

[0153] In some embodiments, when the first calculated window pixel difference is greater than the second calculated window pixel difference, that is, the second calculated window pixel difference is the minimum calculated window pixel difference, use the second motion estimation vector as the target motion estimation vector corresponding to the pixel point to be processed.

[0154] In some embodiments, when the first calculated window pixel difference is equal to the second calculated window pixel difference, any one of the first motion estimation vector and the second motion estimation vector can be used as the target motion estimation vector corresponding to the pixel point to be processed.

[0155] In the embodiments of the present application, for each pixel point to be processed, calculate the first calculated window pixel difference corresponding to the pixel point to be processed under the first motion estimation vector and the second calculated window pixel difference corresponding to the pixel point to be processed under the second motion estimation vector, and determine the final target motion estimation vector of the pixel point to be processed through the comparison result of the two. In this way, the target motion estimation vectors of each pixel point to be processed obtained not only consider the global image information but also the local image information; at the same time, by comprehensively comparing the calculated window pixel differences of the two motion estimation vectors, the accuracy of the target motion estimation vector of each pixel point to be processed can be improved, thereby improving the accuracy of the dense optical flow.

[0156] The following describes the application of the dense optical flow generation method provided by the embodiments of the present application in an actual scenario.

[0157] In the field of video image processing, algorithms such as object tracking, image registration, high dynamic range imaging (HDR) synthesis, and multi-frame noise reduction all require the use of image dense optical flow. There are many calculation methods for image dense optical flow, and common ones include block matching, local binary approximation, optical flow algorithms based on energy minimization, and even deep learning network training schemes. However, the common point of these methods is that they require a large amount of computational effort.

[0158] In some embodiments, the above-mentioned calculation of dense optical flow of an image can be completed by a block matching method. The block matching method divides an image into multiple image blocks. In adjacent frames, a search is performed within a fixed or variable neighborhood of this image block, and the difference is calculated between the image block in the current frame and the image blocks within the neighborhood in the adjacent frame. The image block with the smallest difference within the neighborhood in the adjacent frame is selected, and the coordinate difference of its current image block is recorded, which is the optical flow vector of the current image block. The advantage of this method is that the computational complexity is relatively low, but the disadvantages are that it requires a large amount of computational effort and cannot guarantee the accuracy of the matching.

[0159] In some other embodiments, the above-mentioned calculation of dense optical flow of an image can also be completed by a local binary approximation algorithm. This method assumes that a local image block can be represented by a quadratic polynomial function, as shown in Equation (9).

[0160] f(x)~x T Ax + b T x + c Equation (9);

[0161] Where A is a 2×2 matrix, b is a 2×1 matrix, x is the coordinate (xy) of the image block, and c is a scalar. Based on this, Equation (9) can be written as Equation (10).

[0162]

[0163] Among them, Equation (10) has six unknowns (including r1 to r6). Therefore, to obtain this equation and thus represent the local image block, at least the information of six local image blocks is required. According to the assumption conditions of optical flow, the same image block has the same image information, and adjacent image blocks have similar image information. Therefore, the current local image block and six local image blocks in its surrounding neighborhood can be used as a group to fit the above equation. Therefore, Equation (10) can be transformed into a matrix function with (1, x, y, x 2 , y 2 , xy) as the basis, as shown in Equation (11).

[0164]

[0165] For the current local image block, the position coordinate difference between two adjacent frames of images can be expressed as Equation (12).

[0166] x = x + d(x) Equation (12);

[0167] Where x is the position coordinate and d(x) is the displacement; then, intermediate variables Δb(x) and h(x) can be obtained, which are determined by Equation (13) and Equation (14) respectively.

[0168] Δb(x) = b2(x) - b1(x) + A(x)d(x), Equation (13);

[0169] h(x) = A(x) T *Δb(x), Equation (14);

[0170] Finally, based on the intermediate variables Δb(x) and h(x), the optical flow field d out (x) can be solved, as shown in Equation (15).

[0171] d out (x) = G avg (x) -1 *h avg (x), Equation (15);

[0172] The above algorithm iterates until the number of iterations is reached. The advantage of binary approximation is that it does not depend on image texture information and can calculate motion estimation vectors for very small image blocks or even pixels. The disadvantage is that the number of binary iterations is relatively large, and due to the lack of image texture support, problems will occur in noise regions and fine texture regions.

[0173] Both block matching and the method based on solving local image equations have the problem of excessive computational complexity. For block matching, it is impossible to estimate more fine-grained pixel-level motion vectors; for the method of solving local image equations, it lacks image texture information, is sensitive to noise, and individual calculation units require recursive iteration, making it difficult to implement on a hardware pipeline. Due to the above problems, it is very difficult to achieve both real-time performance and good results in products for the dense optical flow method provided in the above embodiments.

[0174] In the embodiments of the present application, by reusing the intermediate results of the existing encoder module, the computational complexity of the initial motion estimation vector for dense optical flow calculation is reduced; at the same time, for local small blocks or small pixels, the method of image binary approximation is used to calculate fine motion estimation vectors. This method can reduce the main computational complexity of optical flow estimation by reusing modules, shorten the calculation time, and shorten the calculation path of image binary approximation, ensuring the timeliness of the hardware.

[0175] In the embodiments of the present application, by transmitting the motion estimation vector information of the codec to the dense optical flow calculation module, the dense optical flow calculation module performs image binary approximation calculation based on the motion estimation vector information of the encoder, and compares the calculated result with the motion estimation vector information provided by the codec by block subtraction, and selects the motion estimation vector with the smallest image block difference as the final result. Through this process, problems such as large computational complexity, poor real-time performance, and inability to be implemented in hardware for image dense optical flow calculation are solved.

[0176] Please refer to Figure 5, which shows a schematic data flow diagram of a dense optical flow calculation system. Among them, the embodiments of the present application mainly provide a video codec 510 and a dense optical flow calculation module 520. Among them, the video codec 510 is used to generate motion estimation vectors corresponding to two adjacent frames of images in units of image blocks; the dense optical flow calculation module 520 is used to perform local binary optical flow calculation based on the two adjacent frames of images and the corresponding motion estimation vectors to obtain per-pixel motion estimation vectors. The size of the unit image block can be set to 8×8.

[0177] In some embodiments, the minimum processing unit of the video codec is an image block of 32×32 (corresponding to the preset image block in the above embodiments), and the video codec can calculate motion estimation vectors at different scales, and then determine the motion estimation vectors corresponding to each unit image block in the 32×32 image block.

[0178] In the embodiments of the present application, in order to make the derived motion estimation vectors as accurate as possible, it is necessary to compare the motion estimation vectors of unit image blocks at the same position and different scale divisions. Since the motion estimation vector as the optical flow needs to consider the continuity of adjacent regions of the image, before comparing different scale divisions, it is necessary to compare whether to divide the 32×32 block.

[0179] Please refer to Figure 6 , which shows a schematic diagram of motion estimation of a unit image block. Here, the motion estimation vector of the image block in the embodiment corresponds to the block motion vector in the above embodiment.

[0180] Step S601, calculate the image block pixel differences of 32×32 image blocks at different scales.

[0181] As Figure 6 shown, the current embodiment includes three division scales of 32×32, 16×16, and 8×8 (which can be simply referred to as scales), and thus it is necessary to calculate the image block pixel differences at the 32×32 scale, the image block pixel differences at the 16×16 scale, and the image block pixel differences at the 8×8 scale respectively.

[0182] Among them, taking the 16×16 scale as an example, the image block pixel difference of the 32×32 image block at this scale is determined by accumulating the pixel differences of 4 16×16 image blocks. For any 16×16 image block, obtain the 16×16 image block in the current frame, and obtain multiple 16×16 image blocks within the window where the image block is located in the adjacent frame, calculate the sum of absolute differences (SAD) between the 16×16 image block in the current frame and each 16×16 image block within the window where the image block is located in the adjacent frame, and take the minimum sum of absolute differences as the pixel difference of the 16×16 image block.

[0183] Based on the same method, the pixel differences of image blocks at the 32×32 scale, the pixel differences of image blocks at the 16×16 scale, and the pixel differences of image blocks at the 8×8 scale can be obtained.

[0184] Step S602: Whether the pixel difference of the image block at the 32×32 scale is less than the pixel difference of the image block at the 8×8 scale and the pixel difference of the image block at the 16×16 scale.

[0185] In the embodiment of the present application, the magnitude relationships among the pixel differences of the image blocks at the 32×32 scale, the pixel differences of the image blocks at the 16×16 scale, and the pixel differences of the image blocks at the 8×8 scale are compared. When the pixel difference of the image block at the 32×32 scale is less than the pixel difference of the image block at the 16×16 scale and the pixel difference of the image block at the 32×32 scale is less than the pixel difference of the image block at the 8×8 scale, it is determined that no further refinement is required, and step S603 is executed; otherwise, it is determined that further refinement is required, and step S604 is executed.

[0186] Step S603: The motion estimation vector of each 8×8 unit image block in the 32×32 image block comes from the motion estimation vector of the 32×32 image block.

[0187] Step S604: The motion estimation vector of each 8×8 unit image block in the 32×32 image block is determined based on the motion estimation vectors of different scales corresponding to each 8×8 unit image block.

[0188] In the embodiment of the present application, when the 32×32 image block needs to be further refined, the 32×32 image block needs to be divided into 8×8 unit image blocks. Each unit image block corresponds to 3 motion estimation vectors: the motion estimation vector at the 8×8 scale, the motion estimation vector at the 16×16 scale, and the motion estimation vector at the 32×32 scale. At the same time, the pixel differences of the 8×8 image blocks under these three motion estimation vectors are compared, and the motion estimation vector of the scale with the smallest difference is selected as the motion estimation vector of the output 8×8 image block.

[0189] Please refer to Figure 7 , which shows a flowchart of the motion estimation of a unit image block in a refined division state.

[0190] Step S701: For each unit image block, determine the motion estimation vectors of the unit image block at different scales respectively.

[0191] Step S702: Determine the pixel difference of the image block corresponding to the motion estimation vector of the unit image block at each scale.

[0192] Among them, the motion estimation vector and the corresponding image block pixel difference of the unit image block at the 8×8 scale are obtained as follows: Obtain the sum of absolute differences (SAD) between the unit image block in the current frame and each 8×8 image block within the window where the unit image block is located in the adjacent frame, determine the image block corresponding to the minimum sum of absolute differences in the adjacent frame, and then generate the motion estimation vector of the unit image block at the 8×8 scale based on the position between the image block corresponding to the minimum sum of absolute differences in the adjacent frame and the unit image block in the current frame. Take the minimum sum of absolute differences as the image block pixel difference of the unit image block at the 8×8 scale.

[0193] Among them, the motion estimation vector and the corresponding image block pixel difference of the unit image block at the 16×16 scale are obtained as follows: Determine the 16×16 image block where the unit image block is located in the current frame, and obtain multiple 16×16 image blocks within the window where the 16×16 image block is located in the adjacent frame. Calculate the sum of absolute differences between the 16×16 image block in the current frame and each 16×16 image block within the window where the image block is located in the adjacent frame, determine the 16×16 image block corresponding to the minimum sum of absolute differences in the adjacent frame, and then generate the motion estimation vector of the unit image block at the 16×16 scale based on the position between the 16×16 image block corresponding to the minimum sum of absolute differences in the adjacent frame and the 16×16 image block in the current frame, that is, the motion estimation vector at the 16×16 scale. Use this motion estimation vector to determine the initial image block corresponding to the unit image block in the adjacent frame, and determine the sum of absolute differences between the unit image block and the initial image block as the image block pixel difference of the unit image block at the 16×16 scale.

[0194] Among them, based on the same calculation method as the 16×16 scale, the motion estimation vector and the image block pixel difference of the unit image block at the 32×32 scale can be obtained.

[0195] Step S703: Take the motion estimation vector corresponding to the minimum image block pixel difference as the motion estimation vector of the unit pixel block.

[0196] In this embodiment, compare the image block pixel differences of the unit image block at the 8×8 scale, 16×16 scale, and 32×32 scale respectively, and take the motion estimation vector corresponding to the minimum image block pixel difference as the motion estimation vector of the unit pixel block.

[0197] Based on the above embodiments, the motion estimation vector corresponding to each unit image block can be obtained. It can be understood that the motion estimation vector of each pixel point in the unit image block is the same as the motion estimation vector corresponding to the unit image block, and thus the motion estimation vector of each pixel point can be obtained.

[0198] Please refer to Figure 8 , which shows a schematic diagram of the motion estimation vector of dense optical flow. The dense optical flow calculation module is based on the motion estimation vector of the 8×8 image block output by the video codec, and on this basis, image binarization approximation is performed pixel by pixel. As Figure 8 shown in

[0199] , based on the coordinates (x1, y1) of image block 2, the following processing is performed on each pixel within image block 2: Figure 9 Select a matrix function for binomial calculation by taking the pixel to be solved in the reference image 820 as the center and selecting 6 neighboring points, as shown in formula (16). Exemplarily, the selection method of the 6 neighboring points can be in accordance with

[0200]

[0201] shown in the schematic diagram of the neighboring points.

[0202]

[0203] where f(x) is the gray value of the pixel to be solved, (x, y) is the coordinate of the pixel to be solved, and r1 - r6 are the coefficients of the binomial. Another form of this binomial is formula (17). A is a 2×2 matrix b is a 2×1 matrix

[0204] For the pixels of the same image information, its A is the same. Therefore, the A of the pixel points in the current image 810 is the same as the A of the pixel points of the same image information in the reference image 820. The pixel function in the reference image 820 can be expressed as formula (18):

[0205]

[0206] Therefore, the offset of the A of pixel d between the current image 810 and the reference image 820, the offset d can satisfy formula (19).

[0207]

[0208] Since A may not be a singular matrix and cannot be directly inverted, formula (20) can be transformed to obtain formula (12), and then the motion estimation vector (i.e., the local optical flow offset) of the current pixel corresponding to the pixel to be solved is obtained.

[0209] d = (A T A) -1 A T *(b - b1) / 2 Formula (20);

[0210] In some embodiments, for the calculation of the motion estimation vector of the current pixel, the embodiments of the present application determine the coordinates of the reference pixel corresponding to the current pixel in the reference image 820 based on the motion estimation vector corresponding to the unit image block where the current pixel is located. A 7×7 region centered on the current pixel in the current image is used as the calculation window, and a 7×7 region centered on the reference pixel in the reference image is used as the corresponding other calculation window. The two calculation windows are subtracted pixel by pixel, and the sum of the pixel differences is counted as the pixel difference of the first calculation window. The embodiments of the present application also determine the coordinates of the reference pixel corresponding to the current pixel in the reference image based on the motion estimation vector of the current pixel determined by formula (20). A 7×7 region centered on the current pixel in the current image is used as the calculation window, and a 7×7 region centered on the reference pixel in the reference image is used as the corresponding other calculation window. The two calculation windows are subtracted pixel by pixel, and the sum of the pixel differences is counted as the pixel difference of the second calculation window. The first calculation window pixel difference and the second calculation window pixel difference are compared, and the motion estimation vector with the smaller calculation window pixel difference is used as the target motion estimation vector of the current pixel.

[0211] The embodiments of the present application transmit the motion estimation vector of the video codec module to the dense optical flow calculation module. Based on the motion estimation vector of the video codec, image binomial equation fitting is performed in the neighborhood of the pixel to be solved by combining 6-pixel information, and the local pixel offset is calculated by the least squares method. Then, the motion estimation vector corresponding to the local pixel offset and the motion estimation vector of the video codec are compared for the pixel difference of the local image block, and the motion estimation vector with the smallest difference is selected as the target motion estimation vector of the current pixel.

[0212] Based on the above embodiments, by using the motion estimation vector of the video codec as the basis, the complexity of calculating the optical flow of the image by the dense optical flow field through pyramid calculation is reduced, and the optical flow calculation time is reduced. At the same time, based on the motion estimation vector of the video codec, local optical flow calculation based on image binary function fitting is performed, and local dense optical flow is obtained through matrix operation without multiple iterations, reducing the data flow path of hardware processing and ensuring the hardware realizability. At the same time, the motion estimation vector of the video codec and the optical flow of the image binary function fitting are compared, reducing the instability of the image binary function fitting method to noise and fine textures, and improving the accuracy of the dense optical flow.

[0213] Based on the foregoing embodiments, an embodiment of the present application provides a dense optical flow generation device. The device includes each unit included therein, as well as each module included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; during implementation, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0214] Figure 10 FIG. is a schematic structural diagram of a dense optical flow generation device provided by an embodiment of the present application, as Figure 10 shown, the dense optical flow generation device 1000 includes: a first vector acquisition module 1010, a second vector acquisition module 1020, and a dense optical flow generation module 1030, where:

[0215] The first vector acquisition module 1010 is configured to perform block matching on a current image frame and a next image frame to obtain a first motion estimation vector corresponding to each to-be-processed pixel point in the current image frame; the first motion estimation vector corresponding to the to-be-processed pixel point is determined by a block matching result corresponding to a unit image block where the to-be-processed pixel point is located;

[0216] The second vector acquisition module 1020 is configured to fit a second motion estimation vector corresponding to each to-be-processed pixel point based on a matching relationship between a local image of each to-be-processed pixel point in the current image frame and a local image of a corresponding matching pixel point in the next image frame; the matching pixel point is determined based on the to-be-processed pixel point and the corresponding first motion estimation vector;

[0217] The dense optical flow generation module 1030 is configured to generate dense optical flow information based on the first motion estimation vector and the second motion estimation vector corresponding to each to-be-processed pixel point.

[0218] In some embodiments, the first motion estimation vector corresponding to the to-be-processed pixel point is a target block motion vector of a unit image block where the to-be-processed pixel point is located; the first vector acquisition module 1010 is further configured to: for each preset image block in the current image frame, determine an image block pixel difference of the preset image block at different division scales; based on the image block pixel difference of the preset image block at different division scales, determine a target block motion vector of each unit image block in the preset image block.

[0219] In some embodiments, the different division scales at least include a first division scale corresponding to the preset image block. The first vector obtaining module 1010 is further configured to: in response to the pixel difference of the image block at the first division scale being greater than or equal to the pixel difference of the image block at any other division scale, for each unit image block, determine the target block motion vector of the unit image block based on the block motion vectors corresponding to the unit image block at each division scale.

[0220] In some embodiments, the different division scales at least include a first division scale corresponding to the preset image block. The first vector obtaining module 1010 is further configured to: in response to the pixel difference of the image block at the first division scale being less than the pixel difference of the image block at each other division scale, determine the block motion vector corresponding to the preset image block as the target block motion vector of each unit image block in the preset image block.

[0221] In some embodiments, the first vector obtaining module 1010 is further configured to: for each division scale, respectively determine the pixel difference of the image block between the first sub-image block at the division scale and the corresponding second sub-image block; the first sub-image block is obtained by dividing the preset image block based on the division scale; the second sub-image block is the image block with the smallest pixel difference between the first sub-image block and the next image frame; and use the sum of the pixel differences of the first sub-image block at the division scale as the pixel difference of the preset image block at the division scale.

[0222] In some embodiments, the first vector obtaining module 1010 is further configured to: generate the pixel difference of the image block corresponding to the unit image block at each division scale based on the block motion vectors corresponding to the unit image block at each division scale; and determine the block motion vector corresponding to the smallest pixel difference as the target block motion vector of the unit image block.

[0223] In some embodiments, the first vector acquisition module 1010 is further configured to: for each first sub-image block at each of the division scales, determine a search area corresponding to the first sub-image block in the next image frame; the first sub-image block is obtained by dividing the preset image block based on the division scale; within the search area corresponding to the first sub-image block, intercept candidate image blocks with a sliding window corresponding to the size of the first sub-image block, and calculate the pixel difference between each candidate image block and the first sub-image block; determine the candidate image block corresponding to the smallest pixel difference as the second sub-image block corresponding to the first sub-image block; based on the first sub-image block and the corresponding second sub-image block, generate a block motion vector corresponding to the first sub-image block at the division scale; wherein, the block motion vector corresponding to the unit image block within the first sub-image block at the division scale is the same as the block motion vector corresponding to the first sub-image block.

[0224] In some embodiments, the second vector acquisition module 1020 is further configured to: for each of the to-be-processed pixel points, obtain a plurality of neighboring pixel points corresponding to the to-be-processed pixel point in the current image frame, and obtain a plurality of neighboring pixel points corresponding to the matching pixel point in the next image frame; based on the matching relationship between the plurality of neighboring pixel points corresponding to the to-be-processed pixel point and the plurality of neighboring pixel points corresponding to the matching pixel point, fit a second motion estimation vector corresponding to each of the to-be-processed pixel points.

[0225] In some embodiments, the second vector acquisition module 1020 is further configured to: based on the plurality of neighboring pixel points corresponding to the to-be-processed pixel point, construct a first matrix function; the first matrix function includes a first pixel function of each neighboring pixel point corresponding to the to-be-processed pixel point, and the first pixel function is a quadratic polynomial function representing the correlation between the position coordinates of the neighboring pixel point in the current image frame and the gray value of the neighboring pixel point; based on the plurality of neighboring pixel points corresponding to the neighboring pixel point, construct a second matrix function; the second matrix function includes a second pixel function of each neighboring pixel point corresponding to the neighboring pixel point, and the second pixel function is a quadratic polynomial function representing the correlation between the position coordinates of the neighboring pixel point in the next image frame and the gray value of the neighboring pixel point; based on the matching relationship between the first pixel function determined by the first matrix function and the second pixel function determined by the second matrix function, generate a second motion estimation vector corresponding to the to-be-processed pixel point.

[0226] In some embodiments, the quadratic polynomial coefficients of the pixel function include matrix coefficients and vector coefficients; the second vector obtaining module 1020 is further configured to: determine a mean matrix coefficient based on the first matrix coefficient corresponding to the first pixel function and the second matrix coefficient corresponding to the second pixel function; determine a difference vector coefficient based on the first vector coefficient corresponding to the first pixel function and the second vector coefficient corresponding to the second pixel function; and determine the product of the inverse matrix corresponding to the mean matrix coefficient, a preset coefficient, and the difference vector coefficient as the second motion estimation vector corresponding to the to-be-processed pixel point.

[0227] In some embodiments, the dense optical flow information includes a target motion estimation vector corresponding to each to-be-processed pixel point; the dense optical flow generation module 1030 is further configured to: for each to-be-processed pixel point, determine a first calculated window pixel difference corresponding to the to-be-processed pixel point under the first motion estimation vector and a second calculated window pixel difference corresponding to the to-be-processed pixel point under the second motion estimation vector respectively based on a preset pixel difference calculation window; and use the motion estimation vector corresponding to the minimum calculated window pixel difference among the first calculated window pixel difference and the second calculated window pixel difference as the target motion estimation vector corresponding to the to-be-processed pixel point.

[0228] The description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to those of the method embodiments. In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the methods described in the above method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0229] It should be noted that in the embodiments of the present application, if the above-mentioned dense optical flow generation method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related art can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present application are not limited to any specific hardware, software, or firmware, or any arbitrary combination among hardware, software, and firmware.

[0230] An embodiment of the present application provides a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.

[0231] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements some or all of the steps in the above method. The computer-readable storage medium can be transient or non-transient.

[0232] An embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code runs in a computer device, the processor in the computer device executes to implement some or all of the steps in the above method.

[0233] An embodiment of the present application provides a computer program product. The computer program product includes a non-transient computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above method. The computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0234] It should be noted here that: the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities can be referred to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects to the method embodiments. For the technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0235] Figure 11 A schematic diagram of the hardware entity of a computer device provided by an embodiment of the present application is as Figure 11 shown. The hardware entity of the computer device 1100 includes: a processor 1101 and a memory 1102. Among them, the memory 1102 stores a computer program that can run on the processor 1101, and when the processor 1101 executes the program, it implements the steps in the method of any of the above embodiments.

[0236] The memory 1102 stores a computer program that can run on the processor. The memory 1102 is configured to store instructions and applications executable by the processor 1101, and can also cache data to be processed or already processed by the processor 1101 and each module in the computer device 1100 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0237] When the processor 1101 executes the program, it implements the steps of the dense optical flow generation method in any of the above. The processor 1101 generally controls the overall operation of the computer device 1100.

[0238] The embodiment of the present application provides a computer storage medium. The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the dense optical flow generation method in any of the above embodiments.

[0239] It should be pointed out here that the descriptions of the above storage medium and device embodiments are similar to those of the above method embodiments, and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0240] The above processor can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that other electronic devices implementing the functions of the above processor are also possible, and the embodiments of the present application do not make specific limitations.

[0241] The above computer storage medium / memory can be a read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; it can also be various terminals including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0242] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures, or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitude of the sequence numbers of the above steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The sequence numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0243] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including that element.

[0244] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.

[0245] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0246] In addition, in each embodiment of this application, the various functional units can all be integrated in one processing unit, or each unit can be separately used as one unit, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional units. Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage medium includes: removable storage devices, read-only memory (ROM), magnetic disks, or optical disks, etc., which can store program codes.

[0247] Alternatively, if the above-mentioned integrated units of this application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of this application. The foregoing storage medium includes: removable storage devices, ROM, magnetic disks, or optical disks, etc., which can store program codes.

[0248] As described above, this is only the implementation mode of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.

Claims

1. A method for generating dense optical flow, characterized in that, The method includes: Performing block matching on the current image frame and the next image frame to obtain a first motion estimation vector corresponding to each to-be-processed pixel point in the current image frame; the first motion estimation vector corresponding to the to-be-processed pixel point is determined by the block matching result of the unit image block where the to-be-processed pixel point is located; Fitting a second motion estimation vector corresponding to each to-be-processed pixel point based on the matching relationship between the local image of each to-be-processed pixel point in the current image frame and the local image of the corresponding matching pixel point in the next image frame; the matching pixel point is determined based on the to-be-processed pixel point and the corresponding first motion estimation vector, and the second motion estimation vector is a pixel-level motion estimation vector; Generating dense optical flow information based on the first motion estimation vector and the second motion estimation vector corresponding to each to-be-processed pixel point.

2. The method according to claim 1, characterized in that, The first motion estimation vector corresponding to the to-be-processed pixel point is the target block motion vector of the unit image block where the to-be-processed pixel point is located; The performing block matching on the current image frame and the next image frame to obtain a first motion estimation vector corresponding to each to-be-processed pixel point in the current image frame includes: For each preset image block in the current image frame, determining the image block pixel difference of the preset image block at different division scales; Based on the image block pixel difference of the preset image block at different division scales, determining the target block motion vector of each unit image block in the preset image block.

3. The method according to claim 2, wherein The different division scales at least include a first division scale corresponding to the preset image block. The determining the target block motion vector of each unit image block in the preset image block based on the image block pixel difference of the preset image block at different division scales includes: In response to the image block pixel difference at the first division scale being greater than or equal to the image block pixel difference at any other division scale, for each unit image block, determining the target block motion vector of the unit image block based on the block motion vector corresponding to the unit image block at each division scale.

4. The method according to claim 2, characterized in that The different division scales at least include a first division scale corresponding to the preset image block. The determining the target block motion vector of each unit image block in the preset image block based on the image block pixel difference of the preset image block at different division scales includes: In response to the image block pixel difference at the first division scale being less than the image block pixel difference at each other division scale, determining the block motion vector corresponding to the preset image block as the target block motion vector of each unit image block in the preset image block.

5. The method according to claim 2, wherein The determining the image block pixel difference of the preset image block at different division scales includes: For each division scale, respectively determining the image block pixel difference between the first sub-image block at the division scale and the corresponding second sub-image block; the first sub-image block is obtained by dividing the preset image block based on the division scale; the second sub-image block is the image block with the smallest image block pixel difference from the first sub-image block in the next image frame; Take the sum of the pixel differences of the image blocks corresponding to the first sub-image blocks at the division scale as the pixel difference of the preset image block at the division scale.

6. The method according to claim 3, wherein The determining the target block motion vector of the unit image block based on the block motion vectors corresponding to the unit image block at each division scale includes: Generate the pixel difference of the image block corresponding to the unit image block at each division scale based on the block motion vectors corresponding to the unit image block at each division scale; Determine the block motion vector corresponding to the smallest pixel difference of the image block as the target block motion vector of the unit image block.

7. The method according to claim 3 or 4, characterized in that, The method further includes: For each first sub-image block at each division scale, determine the search area corresponding to the first sub-image block in the next image frame; the first sub-image block is obtained by dividing the preset image block based on the division scale; Within the search area corresponding to the first sub-image block, intercept candidate image blocks with a sliding window corresponding to the size of the first sub-image block, and calculate the pixel difference of the image block between each candidate image block and the first sub-image block; Determine the candidate image block corresponding to the smallest pixel difference of the image block as the second sub-image block corresponding to the first sub-image block; Generate the block motion vector corresponding to the first sub-image block at the division scale based on the first sub-image block and the corresponding second sub-image block; wherein, the block motion vector corresponding to the unit image block within the first sub-image block at the division scale is the same as the block motion vector corresponding to the first sub-image block.

8. The method according to any one of claims 1 to 6, characterized in that, The fitting the second motion estimation vector corresponding to each to-be-processed pixel point based on the matching relationship between the local image of each to-be-processed pixel point in the current image frame and the local image of the corresponding matching pixel point in the next image frame includes: For each to-be-processed pixel point, obtain multiple neighboring pixel points corresponding to the to-be-processed pixel point in the current image frame, and obtain multiple neighboring pixel points corresponding to the matching pixel point in the next image frame; Fit the second motion estimation vector corresponding to each to-be-processed pixel point based on the matching relationship between the multiple neighboring pixel points corresponding to the to-be-processed pixel point and the multiple neighboring pixel points corresponding to the matching pixel point.

9. The method according to claim 8, wherein The fitting the second motion estimation vector corresponding to each to-be-processed pixel point based on the matching relationship between the multiple neighboring pixel points corresponding to the to-be-processed pixel point and the multiple neighboring pixel points corresponding to the matching pixel point includes: Construct a first matrix function based on the multiple neighboring pixel points corresponding to the to-be-processed pixel point; the first matrix function includes a first pixel function of each neighboring pixel point corresponding to the to-be-processed pixel point, and the first pixel function is a quadratic polynomial function characterizing the correlation relationship between the position coordinates of the neighboring pixel points in the current image frame and the gray values of the neighboring pixel points; Construct a second matrix function based on multiple neighboring pixel points corresponding to the neighboring pixel points; the second matrix function includes a second pixel function of each neighboring pixel point corresponding to the neighboring pixel points, and the second pixel function is a quadratic polynomial function representing the correlation between the position coordinates and the gray value of the neighboring pixel points in the next image frame. Generate a second motion estimation vector corresponding to the pixel point to be processed based on the matching relationship between the first pixel function determined by the first matrix function and the second pixel function determined by the second matrix function.

10. The method according to claim 9, wherein The quadratic polynomial coefficients of the pixel function include matrix coefficients and vector coefficients; generating the second motion estimation vector corresponding to the pixel point to be processed based on the matching relationship between the first pixel function determined by the first matrix function and the second pixel function determined by the second matrix function includes: Determine a mean matrix coefficient based on the first matrix coefficient corresponding to the first pixel function and the second matrix coefficient corresponding to the second pixel function. Determine a difference vector coefficient based on the first vector coefficient corresponding to the first pixel function and the second vector coefficient corresponding to the second pixel function. Determine the product of the inverse matrix corresponding to the mean matrix coefficient, a preset coefficient, and the difference vector coefficient as the second motion estimation vector corresponding to the pixel point to be processed.

11. The method according to any one of claims 1 to 6, characterized in that, The dense optical flow information includes a target motion estimation vector corresponding to each pixel point to be processed; generating dense optical flow information based on the first motion estimation vector and the second motion estimation vector corresponding to each pixel point to be processed includes: For each pixel point to be processed, based on a preset pixel difference calculation window, respectively determine a first calculation window pixel difference corresponding to the pixel point to be processed under the first motion estimation vector and a second calculation window pixel difference corresponding to the pixel point to be processed under the second motion estimation vector. Use the motion estimation vector corresponding to the minimum calculation window pixel difference among the first calculation window pixel difference and the second calculation window pixel difference as the target motion estimation vector corresponding to the pixel point to be processed.

12. A dense optical flow generation device, characterized in that, The device includes: A first vector acquisition module, configured to perform block matching on the current image frame and the next image frame to obtain a first motion estimation vector corresponding to each pixel point to be processed in the current image frame; the first motion estimation vector corresponding to the pixel point to be processed is determined by the block matching result of the unit image block where the pixel point to be processed is located. A second vector acquisition module, configured to fit a second motion estimation vector corresponding to each pixel point to be processed based on the matching relationship between the local image of each pixel point to be processed in the current image frame and the local image of the corresponding matching pixel point in the next image frame; the matching pixel point is determined based on the pixel point to be processed and the corresponding first motion estimation vector, and the second motion estimation vector is a pixel-level motion estimation vector. A dense optical flow generation module, configured to generate dense optical flow information based on the first motion estimation vector and the second motion estimation vector corresponding to each pixel point to be processed.

13. A computer device, comprising a memory and a processor, the memory storing a computer program that can run on the processor, characterized in that, When the processor executes the program, it implements the steps in the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps in the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Interpolated frame generating method and interpolated frame generating apparatus

    CN101197999A

  • Bidirectional motion estimating method and video frame rate up-converting method and system

    CN104219533A

  • Inter-frame prediction method, video coding method, electronic equipment and storage device

    CN111970516A