Optical flow calculation method and device, electronic equipment and storage medium

By randomly initializing the optical flow processing on the top layer image of the image pyramid, the problem of inaccurate optical flow calculation in large motion scenes is solved, and higher accuracy of optical flow results is achieved.

CN119991743AActive Publication Date: 2025-05-13VIVO MOBILE COMM CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510134948.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-13
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

The calculation results of existing optical flow calculation methods in large motion scenarios are inaccurate, mainly because the initial optical flow of the top layer of the image pyramid is set to all zeros, resulting in a small size of the search window when block matching, making it difficult to accurately find the most matching block.

Method used

Using a randomly initialized optical flow method, a set of random offsets is set to each pixel point of the top layer image of the image pyramid, used to indicate the pixel position of the reference frame, perform block matching processing, and then fine-tune layer by layer to obtain a fine-tune optical flow.

Benefits of technology

Through the randomly initialized optical flow method, it is possible to cover scenes with various motion amplitudes, improve the accuracy of optical flow results in large motion scenes, and avoid matching errors due to small search window size.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991743A_ABST
    Figure CN119991743A_ABST
Patent Text Reader

Abstract

The invention discloses an optical flow calculation method. The method comprises the following steps: constructing an N-layer first image pyramid of a target frame and an N-layer second image pyramid of a reference frame; configuring an initial optical flow of an Nth layer image of the first image pyramid; the initial optical flow comprises a group of offsets of each pixel point, and each random offset in each group of offsets is used for indicating a pixel position in the Nth layer image of the second image pyramid; performing block matching processing on the Nth-layer image of the first image pyramid and the second image pyramid according to the initial optical flow, determining a fine adjustment optical flow of the Nth-layer image of the first image pyramid according to a block matching processing result, and performing up-sampling on the fine adjustment optical flow to obtain an initial optical flow of the (N-1) th-layer image of the first image pyramid; and on the basis of the initial optical flow of the (N-1) th layer of image, calculating the fine adjustment optical flow of each layer of image below the Nth layer of image of the first image pyramid layer by layer until the fine adjustment optical flow of the first layer of image is obtained, and de-noising the fine adjustment optical flow to obtain a target optical flow.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image processing technology, and specifically relates to an optical flow calculation method, device, electronic device and storage medium. Background Art

[0002] Optical flow is an important technology in image processing. It is used to track moving objects in image sequences. Optical flow technology is widely used in photography, games and other fields. For example, in the field of photography, photographic equipment can automatically stabilize video images based on the optical flow of image sequences, reduce jitter during video recording, and make the videos taken by users smoother and clearer. In the field of games, compared with the motion vector information that comes with the game, optical flow can solve the movement of shadows, translucency and other scenes, making the interpolation effect of the game screen more coherent.

[0003] In the related art, a method combining a block matching algorithm and an image pyramid algorithm is used to calculate the optical flow of an image. First, an image pyramid of a target frame and an image pyramid of a reference frame are constructed, wherein the image with the lowest resolution in the image pyramid is at the top layer and the image with the original resolution is at the bottom layer; then, starting from the top layer of the image pyramid, the initial optical flow of the top layer is set to all zeros, the fine-tuned optical flow of the top layer is calculated according to the block matching method, and the fine-tuned optical flow of the top layer is transferred to the second top layer of the image pyramid as the initial optical flow of the second top layer; then, the initial optical flow of the second top layer is fine-tuned according to the block matching method to obtain the fine-tuned optical flow of the second top layer, and the fine-tuned optical flow of the second top layer is transferred to the next layer of the image pyramid as the initial optical flow of the next layer, and the calculation process of the second top layer is performed on the remaining layers of the image pyramid until the fine-tuned optical flow of the bottom layer of the image pyramid is calculated; finally, the fine-tuned optical flow of the bottom layer of the image pyramid is denoised to obtain the denoised optical flow, which is the optical flow of each pixel in the target frame.

[0004] However, for the top layer of the image pyramid, the initial optical flow of the top layer is all zero, that is, there will be no information, so that for each image block of the top layer image of the image pyramid of the target frame, the best matching block in the search window can only be found pixel by pixel in the top layer image of the image pyramid of the reference frame based on the search window. Since the size of the search window is usually set to be relatively small, the position of the best matching block may be wrong for scenes with large motion, and the fine-tuned optical flow of the top layer will be passed to the second top layer as the initial optical flow of the second top layer. Therefore, if the optical flow of the top layer is wrong, it will be more difficult for the subsequent layers to find the best matching block, which ultimately leads to inaccurate optical flow results calculated for scenes with large motion. Summary of the invention

[0005] The purpose of the embodiments of the present application is to provide an optical flow calculation method, device, electronic device and storage medium, which can improve the accuracy of the optical flow results calculated for scenes with large motion.

[0006] In a first aspect, an embodiment of the present application provides an optical flow calculation method, the method comprising:

[0007] Constructing a first image pyramid of the target frame and a second image pyramid of the reference frame; wherein the number of layers of the first image pyramid and the number of layers of the second image pyramid are both N, and N is an integer greater than 1;

[0008] Configuring an initial optical flow of an N-th layer image of the first image pyramid; wherein the initial optical flow of the N-th layer image of the first image pyramid comprises: a group of offsets for each pixel in the N-th layer image of the first image pyramid, each group of offsets comprising: at least two different random offsets, each of the random offsets being used to indicate a pixel position in the N-th layer image of the second image pyramid;

[0009] According to the initial optical flow of the Nth layer image of the first image pyramid, block matching processing is performed on the Nth layer image of the first image pyramid and the Nth layer image of the second image pyramid, and a fine-tuning optical flow of the Nth layer image of the first image pyramid is determined according to the block matching processing result, and the fine-tuning optical flow of the Nth layer image of the first image pyramid is up-sampled, and the up-sampling result is determined as the initial optical flow of the N-1th layer image of the first image pyramid; wherein the fine-tuning optical flow of the Nth layer image of the first image pyramid includes: a fine-tuning optical flow value of each pixel point in the Nth layer image of the first image pyramid;

[0010] Based on the initial optical flow of the N-1th layer image of the first image pyramid, calculate the fine-tuning optical flow of each layer image below the Nth layer image of the first image pyramid layer by layer until the fine-tuning optical flow of the first layer image of the first image pyramid is obtained;

[0011] The fine-tuned optical flow of the first layer image of the first image pyramid is denoised to obtain a target optical flow of the target frame.

[0012] In a second aspect, an embodiment of the present application provides an optical flow calculation device, the device comprising:

[0013] A construction module, used to construct a first image pyramid of a target frame and a second image pyramid of a reference frame; wherein the number of layers of the first image pyramid and the number of layers of the second image pyramid are both N, and N is an integer greater than 1;

[0014] a configuration module, configured to configure an initial optical flow of an N-th layer image of the first image pyramid; wherein the initial optical flow of the N-th layer image of the first image pyramid comprises: a set of offsets for each pixel in the N-th layer image of the first image pyramid, each set of offsets comprising: at least two different random offsets, each of the random offsets being used to indicate a pixel position in the N-th layer image of the second image pyramid;

[0015] a first processing module, configured to perform block matching processing on the Nth layer image of the first image pyramid and the Nth layer image of the second image pyramid according to the initial optical flow of the Nth layer image of the first image pyramid, determine the fine-tuning optical flow of the Nth layer image of the first image pyramid according to the block matching processing result, up-sample the fine-tuning optical flow of the Nth layer image of the first image pyramid, and determine the up-sampling result as the initial optical flow of the N-1th layer image of the first image pyramid; wherein the fine-tuning optical flow of the Nth layer image of the first image pyramid includes: the fine-tuning optical flow value of each pixel point in the Nth layer image of the first image pyramid;

[0016] a second processing module, configured to calculate, based on the initial optical flow of the N-1th layer image of the first image pyramid, the refined optical flow of each layer image below the Nth layer image of the first image pyramid layer by layer, until the refined optical flow of the first layer image of the first image pyramid is obtained;

[0017] The denoising module is used to perform denoising on the fine-tuned optical flow of the first layer image of the first image pyramid to obtain the target optical flow of the target frame.

[0018] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the optical flow calculation method described in the first aspect are implemented.

[0019] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the optical flow calculation method described in the first aspect are implemented.

[0020] In a fifth aspect, an embodiment of the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the optical flow calculation method as described in the first aspect.

[0021] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the optical flow calculation method as described in the first aspect.

[0022] In an embodiment of the present application, a first image pyramid of a target frame and a second image pyramid of a reference frame are constructed; wherein the number of layers of the first image pyramid and the number of layers of the second image pyramid are both N, and N is an integer greater than 1; an initial optical flow of an N-th layer image of the first image pyramid is configured; wherein the initial optical flow of the N-th layer image of the first image pyramid includes: a group of offsets for each pixel point in the N-th layer image of the first image pyramid, each group of offsets includes: at least two different random offsets, each random offset is used to indicate a pixel position in the N-th layer image of the second image pyramid; according to the initial optical flow of the N-th layer image of the first image pyramid, block matching is performed on the N-th layer image of the first image pyramid and the N-th layer image of the second image pyramid. The method comprises the following steps: performing block matching processing, determining a fine-tuning optical flow of an N-th layer image of a first image pyramid according to a block matching processing result, up-sampling the fine-tuning optical flow of the N-th layer image of the first image pyramid, and determining the up-sampling result as an initial optical flow of an N-1-th layer image of the first image pyramid; wherein the fine-tuning optical flow of the N-th layer image of the first image pyramid comprises: a fine-tuning optical flow value of each pixel point in the N-th layer image of the first image pyramid; based on the initial optical flow of the N-1-th layer image of the first image pyramid, calculating the fine-tuning optical flow of each layer image below the N-th layer image of the first image pyramid layer by layer until the fine-tuning optical flow of the first layer image of the first image pyramid is obtained; and performing denoising processing on the fine-tuning optical flow of the first layer image of the first image pyramid to obtain a target optical flow of a target frame.

[0023] It can be seen that compared with the related art in which the optical flow of the top layer of the image pyramid is initialized to all zeros, the initial optical flow based on all zeros cannot provide any information for the top layer, and can only find the most matching block in the neighborhood of the search window based on the search window, and the small-sized search box cannot cover large motion scenes. In the embodiment of the present application, the optical flow of the top layer of the image pyramid is randomly initialized. Since the randomly initialized optical flow contains a set of random offsets for each pixel point of the top layer image, and there is always a random offset within a certain range that is lucky enough to hit the most matching image block, further fine-tuning in the next layer can better estimate the final optical flow, so that the block matching processing of the top layer image is no longer limited by the size of the search window, and the accuracy of the optical flow results calculated for scenes with large motion can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is an example diagram of a block matching method provided by some embodiments of the present application;

[0025] Figure 2 is an example diagram of an image pyramid provided by some embodiments of the present application;

[0026] Figure 3A It is one of the example diagrams of the optical flow calculation method in the related art provided by some embodiments of the present application;

[0027] Figure 3B This is the second example diagram of the optical flow calculation method in the related art provided by some embodiments of the present application;

[0028] Figure 3C This is the third example diagram of the optical flow calculation method in the related art provided by some embodiments of the present application;

[0029] Figure 4 is a flowchart of an optical flow calculation method provided by some embodiments of the present application;

[0030] Figure 5 is an example diagram of the initial optical flow of the top layer image of the first image pyramid provided by some embodiments of the present application;

[0031] Fig. 6A is a flowchart of an implementation of step 403 provided in some embodiments of the present application;

[0032] Figure 6B It is one of the example diagrams of an implementation of step 403 provided in some embodiments of the present application;

[0033] Figure 6C This is a second example diagram of an implementation of step 403 provided in some embodiments of the present application;

[0034] Fig. 7A is a flowchart of an implementation of step 404 provided in some embodiments of the present application;

[0035] Figure 7B is a flowchart of an implementation of step 4041 provided in some embodiments of the present application;

[0036] Figure 7C is one of the example diagrams of an implementation of step 4041 provided in some embodiments of the present application;

[0037] Fig.7D This is a second example diagram of an implementation of step 4041 provided in some embodiments of the present application;

[0038] Fig. 7E This is a third example diagram of an implementation of step 4041 provided in some embodiments of the present application;

[0039] Figure 7FThis is a fourth example diagram of an implementation of step 4041 provided in some embodiments of the present application;

[0040] Figure 7G This is a fifth example diagram of an implementation of step 4041 provided in some embodiments of the present application;

[0041] Fig. 8A is a flowchart of a process for generating a target denoising model provided by some embodiments of the present application;

[0042] Figure 8B is an example diagram of an initial denoising model provided by some embodiments of the present application;

[0043] Figure 8C is an example diagram of a target denoising model provided by some embodiments of the present application;

[0044] Fig. 9 is a structural block diagram of an optical flow calculation device provided by some embodiments of the present application;

[0045] Fig.10 is a schematic diagram of the structure of an electronic device provided by some embodiments of the present application;

[0046] Fig.11 It is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of the present application. DETAILED DESCRIPTION

[0047] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0048] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0049] To facilitate understanding, some relevant concepts and application scenarios involved in the embodiments of the present application are first introduced.

[0050] 1. Related concepts

[0051] Optical flow is a vector with direction and length. The purpose of optical flow calculation is to solve the motion vector (or offset) of the corresponding pixel based on two consecutive frames.

[0052] Optical flow field, in space, motion can be described by a motion field, and on an image plane, the motion of an object is often reflected by the different grayscale distributions of different images in an image sequence, so the motion field in space is transferred to the image and represented as an optical flow field. The optical flow field is a two-dimensional vector field, which reflects the trend of grayscale changes at each point on the image, and can be regarded as the instantaneous velocity field generated by the movement of grayscale pixels on the image plane. The information it contains is the instantaneous motion velocity vector information of each image point. The purpose of studying the optical flow field is to approximate the motion field that cannot be directly obtained from the sequence image. Ideally, the optical flow field corresponds to the motion field.

[0053] Dense optical flow is an image registration method that performs pixel-by-pixel matching on an image or a specified area. It calculates the offset of all pixels on the image to form a dense optical flow field. Through this dense optical flow field, pixel-level image registration can be performed.

[0054] In contrast to dense optical flow, sparse optical flow does not calculate each pixel of the image point by point.

[0055] Block matching is a commonly used method in image denoising and motion estimation. By matching the query block with adjacent image blocks, the K blocks closest to the query block are found from these adjacent blocks. The so-called adjacent blocks are not necessarily adjacent in terms of absolute position, which can also lead to local search (local) and global search (non-local).

[0056] For the block matching process in optical flow calculation, such as Figure 1As shown, the target frame 10 and the reference frame 20, taking the calculation of the optical flow value of a pixel point 11 in the target frame 10 as an example, first, determine the image block 12 with the pixel point 11 as the center pixel point. Then, determine the search window 21 corresponding to the image block 12 in the reference frame 20, wherein the image block 12 can only be block matched within the search window 21. Then, the image block 12 moves up and down and left and right in the search window 21, and searches for the matching image block with the highest similarity to the image block 12 in the search window 21, wherein the similarity calculation between the two image blocks can adopt the square difference (SSD), the absolute value difference (SAD), the census transformation and other algorithms. For example, after comparison, it is found that the image block 22 in the search window 21 has the highest similarity with the image block 12, and the image block 22 is determined as the matching image block of the image block 12. Finally, calculate the offset between the image block 12 and the image block 22, and use the calculated offset as the optical flow value of the center pixel point 11 of the image 12. When the size of the search window 21 is smaller than the size of the reference frame 20, the above block matching process is a local search. When the size of the search window 21 is equal to the size of the reference frame 20, the above block matching process is a global search.

[0057] Image pyramid is a kind of multi-scale representation of images. It is an effective but conceptually simple structure for interpreting images at multiple resolutions. Take a 4-layer image pyramid as an example. Figure 2 As shown in the figure, the image pyramid of an image is a set of image resolutions that are gradually reduced in a pyramid shape (from bottom to top) and are derived from the same original image. It is obtained by downsampling in stages until a certain termination condition is reached. The layers of images are likened to a pyramid. The higher the level, the smaller the image size and the lower the resolution.

[0058] Re-parameterization is a technique in machine learning and deep learning that is used to change the way model parameters are expressed so that they can be optimized more efficiently or more stably. Re-parameterization represents a random variable (such as a latent vector) as a function of another random variable, so that it can be processed by standard optimization algorithms.

[0059] Structural reparameterization means first constructing a structure for training, and then converting the parameters into another set of parameters in the inference phase. In this way, a larger overhead can be used during training, but a smaller overhead can be used in the inference phase. It can also be understood that the reparameterized structure adds some parameters in the training phase that can be removed in the inference phase. The core idea of ​​structural reparameterization is that convolution operations are linear operations and are additive. In order to achieve structural reparameterization, a special multi-branch module needs to be designed, which only has linear operations inside, such as convolution and batch normalization (Batch Normalization, BN).

[0060] 2. Application Scenarios

[0061] Optical flow is an important technology in computer vision and image processing, which is used to track moving objects in image sequences. Currently, optical flow technology is widely used in the fields of photography and games, providing users with a more immersive experience.

[0062] In the related art, a method combining block matching algorithm and image pyramid algorithm is used to calculate the optical flow of an image. Figure 3A As shown, an image pyramid of the target frame and an image pyramid of the reference frame are constructed, wherein the image with the lowest resolution in the image pyramid is at the top layer, and the image with the original resolution is at the bottom layer, for example, the number of layers of the image pyramid is 4.

[0063] After that, starting from the top layer (i.e., the fourth layer) of the image pyramid, the initial optical flow of the fourth layer of the target frame's image pyramid is set to all zero. Since the initial optical flow of all zero cannot provide any information, for each image block of the fourth layer of the target frame's image pyramid, the best matching block within the search window can only be found pixel by pixel in the fourth layer of the reference frame's image pyramid based on the search window. For example, Figure 3B As shown, the fourth layer image 31 of the image pyramid of the target frame and the fourth layer image 32 of the image pyramid of the reference frame, take the calculation of the optical flow value of a pixel point 311 in the image 31 as an example, first, determine the image block 312 with the pixel point 311 as the center pixel point. Then, determine the search window 321 corresponding to the image block 312 in the image 32, wherein the image block 312 can only be block matched within the search window 321. Then, the image block 312 moves up, down, left, and right pixel by pixel in the search window 321, and searches for the matching image block with the highest similarity to the image block 312 in the search window 321, wherein the similarity calculation between the two image blocks can adopt algorithms such as square difference (SSD), absolute difference (SAD), census transformation, etc. For example, after comparison, it is found that the image block 322 in the search window 321 has the highest similarity with the image block 312, and the image block 322 is determined as the matching image block of the image block 312. Finally, the offset between the image block 312 and the image block 322 is calculated, and the calculated offset is used as the fine-tuning optical flow value of the central pixel 311 of the image 312. Similarly, the fine-tuning optical flow value of each pixel in the fourth layer image of the image pyramid of the target frame can be calculated, that is, the fine-tuning optical flow of the fourth layer image. After the fine-tuning optical flow of the fourth layer image of the image pyramid of the target frame is calculated, the fine-tuning optical flow of the fourth layer image is upsampled, and the upsampling result is passed to the third layer of the image pyramid of the target frame as the initial optical flow of the third layer image.

[0064] Afterwards, the initial optical flow of the third layer image is fine-tuned according to the block matching method to obtain the fine-tuned optical flow of the third layer image. Figure 3C As shown, the third layer image 41 of the image pyramid of the target frame and the third layer image 42 of the image pyramid of the reference frame, taking the optical flow value of a pixel point 411 in the fine-tuning image 41 as an example, determine the image block 412 with the pixel point 411 as the center pixel point. According to the initial optical flow value of the pixel point 411 in the initial optical flow of the image 41, determine the image block 422 in the image 42 corresponding to the initial optical flow value of the pixel point 411. Determine the search window 421 corresponding to the image block 412 in the image 42. The image block 412 moves up, down, left, and right pixel by pixel in the search window 421, and search for the matching image block with the highest similarity to the image block 412 in the search window 421, wherein the similarity between the two image blocks can be calculated using algorithms such as square difference (SSD), absolute difference (SAD), and census transformation. For example, after comparison, it is found that the similarity between the image block 423 and the image block 412 in the search window 421 is the highest. Since the similarity between the image block 423 and the image block 412 is greater than the similarity between the image block 422 and the image block 412, it is explained that the initial optical flow value of the pixel point 413 in the image 41 corresponding to the central pixel point 424 of the image block 423 is more suitable than the initial optical flow value of the pixel point 411 as the fine-tuning optical flow value of the pixel point 411. At this time, the initial optical flow value of the pixel point 413 is determined as the fine-tuning optical flow value of 411. Similarly, the fine-tuning optical flow value of each pixel point in the third layer image of the image pyramid of the target frame can be calculated, that is, the fine-tuning optical flow of the third layer image. After the fine-tuning optical flow of the third layer image of the image pyramid of the target frame is calculated, the fine-tuning optical flow of the third layer image is upsampled, and the upsampling result is transferred to the second layer of the image pyramid of the target frame as the initial optical flow of the second layer image.

[0065] For the second layer of the image pyramid of the target frame, the same calculation process as the third layer is performed to obtain the fine-tuned optical flow of the second layer image of the image pyramid of the target frame. The fine-tuned optical flow of the second layer image is upsampled, and the upsampled result is passed to the first layer of the image pyramid of the target frame as the initial optical flow of the first layer image. For the first layer of the image pyramid of the target frame, the same calculation process as the third layer is performed to obtain the fine-tuned optical flow of the first layer image of the image pyramid of the target frame.

[0066] Finally, the fine-tuned optical flow of the bottom layer (i.e., the first layer) of the image pyramid is denoised to obtain the denoised optical flow, which is the optical flow of each pixel in the target frame. The optical flow calculation method based on block matching is adopted. Since local similarity is not considered during block search, the best matching block is searched pixel by pixel. Therefore, the optical flow calculated based on the block matching method usually has relatively large noise. In order to solve this problem, the related technology uses a spatial iteration method to denoise the optical flow, that is, after calculating the optical flow, the offset of the current pixel point and its upper and left positions is compared, and these two offsets are applied to the image to be matched to determine whether the block matching is more accurate than the offset of the current position. If so, the optical flow value of the current position is directly used for filling. Alternatively, the related technology uses an average filtering method or a median filtering method.

[0067] Although the related technology can calculate the optical flow to a certain extent, it is necessary to consider the following: 1) The block matching method mentioned above needs to find the best matching image block within a certain search window, so the range selection of the search window is very important. In order to find the best matching image block, global search is theoretically a better choice, that is, the size of the search window is the same as the size of the image. However, in actual operation, if global search is selected, the amount of calculation will be very large and time-consuming; in order to achieve real-time, even if it is a local search, the size of the search window will be limited to a small size, but this will greatly reduce the optical flow effect, and the optical flow can only estimate scenes with very small movements; with the change of the number of pyramid layers, the optical flow effect will also fluctuate greatly. 2) The block matching method, because its principle is to find the best matching block, the final offset may be anywhere. In other words, the optical flow between adjacent pixels may vary greatly, which is specifically manifested in the large noise of the optical flow map. However, this phenomenon is inconsistent with the actual movement. In actual scenes, adjacent pixels usually have similar movement directions and movement amplitudes, so the actual optical flow should have local correlation, that is, the optical flow of adjacent pixels is smooth or slowly changing.

[0068] The related art has the following problems: 1) For the top layer of the image pyramid, the optical flow estimation method based on block matching uses all zeros as initialization. The initialization based on all zeros usually does not have any information in the current layer, and can only find the best matching block in the neighborhood of the search window based on the search window. However, since the size of the search window is usually set to be relatively small, the position of the best matching block for scenes with large motion may be wrong, and the fine-tuned optical flow of the top layer will be transmitted to the second top layer as the initial optical flow of the second top layer. Therefore, if the optical flow of the top layer is wrong, it will be more difficult for the subsequent layers to find the best matching block, which ultimately leads to inaccurate optical flow results calculated for scenes with large motion. Although the size of the search window can be increased, increasing the size of the search window will greatly increase the computational complexity, greatly increasing the amount of calculation. There is another problem with all-zero initialization: since the motion size gradually increases with the number of pyramid layers, the range of motion vectors of each layer is always different. When you want to reduce the amount of calculation by changing the number of pyramid layers, the difference in the optical flow results will be relatively large. 2) For the second top layer and the layers below the second top layer of the image pyramid, the block search of the optical flow is based on the search window, and the search is performed pixel by pixel within the search window. However, since the optical flow values ​​of the neighborhood positions are usually smooth and the offset difference relative to the center pixel is very small, it is a relatively redundant calculation to update the optical flow of the center pixel with the neighborhood optical flow. 3) The spatial iterative denoising method is widely used in the current optical flow calculation. Since each position depends on the offset with the upper and left positions, there will be relatively large problems in parallelization. In addition, since the calculation of optical flow is based on block matching, there is a certain prior information in it, and it is not positive or negative random salt and pepper noise. Therefore, average-based filtering methods such as mean filtering and Gaussian filtering cannot solve this type of problem; median filtering is also not suitable for optical flow. The reason is that for scenes with straight lines, the inaccurate matching and superposition of median filtering will cause the image straight lines to become distorted after the application of optical flow.

[0069] In order to solve the above technical problems, the embodiments of the present application provide an optical flow calculation method, device, electronic device and storage medium, so as to cover scenes with various motion amplitudes, reduce the amount of calculation in the optical flow calculation process, and improve the accuracy of the optical flow calculation results.

[0070] An optical flow calculation method provided in an embodiment of the present application is introduced below in conjunction with the accompanying drawings.

[0071] It should be noted that the optical flow calculation method provided in the embodiment of the present application is applicable to electronic devices. In practical applications, the electronic devices include but are not limited to: mobile terminals such as mobile phones, tablet computers, laptops, PDAs, and computer devices such as servers and desktops. The embodiment of the present application does not limit this.

[0072] Figure 4is a flowchart of an optical flow calculation method provided by some embodiments of the present application, such as Figure 4 As shown, the method at least includes the following steps: step 401, step 402, step 403, step 404 and step 405;

[0073] In step 401 , a first image pyramid of a target frame and a second image pyramid of a reference frame are constructed; wherein the number of layers of the first image pyramid and the number of layers of the second image pyramid are both N, and N is an integer greater than 1.

[0074] In an embodiment of the present application, the target frame and the reference frame may be two adjacent frames in an image sequence, wherein the image sequence may be a video captured by a camera device, or the image sequence may be a game screen image of a game application.

[0075] In the embodiment of the present application, the first image pyramid and the second image pyramid may both be Gaussian pyramids.

[0076] Exemplarily, taking the construction of the first image pyramid as an example, first, a series of Gaussian smoothing and downsampling operations are performed on the target frame to generate a set of image hierarchies with gradually decreasing resolutions; wherein, Gaussian smoothing is used to apply a Gaussian filter to the original image to generate a smoothed image; the Gaussian filter is a low-pass filter used to reduce high-frequency noise in the image; the downsampling operation is used to downsample the smoothed image, usually by halving the width and height of the image to obtain a lower resolution image. Afterwards, the above steps are repeated: the Gaussian smoothing and downsampling process is repeated on the downsampled image until the predetermined resolution level is reached, and each newly generated image is called a pyramid layer. For example, N=4, the following can be constructed: Figure 2 Image pyramid of the shown structure.

[0077] Exemplarily, N=4, the first image pyramid is a 4-layer pyramid, from top to bottom, it is the fourth layer image, the third layer image, the second layer image and the first layer image, wherein the fourth layer is the top layer of the pyramid (or called the first layer), and the first layer is the bottom layer of the pyramid. Similarly, the second image pyramid is also a 4-layer pyramid, from top to bottom, it is the fourth layer image, the third layer image, the second layer image and the first layer image, wherein the fourth layer is the top layer of the pyramid, and the first layer is the bottom layer of the pyramid.

[0078] In step 402, an initial optical flow of an N-th layer image of a first image pyramid is configured; wherein the initial optical flow of the N-th layer image of the first image pyramid includes: a group of offsets for each pixel in the N-th layer image of the first image pyramid, each group of offsets includes: at least two different random offsets, each random offset is used to indicate a pixel position in the N-th layer image of a second image pyramid.

[0079] In the embodiment of the present application, the Nth layer image of the first image pyramid refers to the image of the top layer of the first image pyramid, which has the smallest resolution.

[0080] On the one hand, considering that in the related art the optical flow of the Nth layer image of the first image pyramid is initialized to all zeros and cannot provide any information, only a pixel-by-pixel search can be performed within the search window, and the search calculation amount is large; on the other hand, considering that although the search calculation amount can be reduced by reducing the size of the search window, for scenes with large motion, a small search window will cause errors in the position of the most matching image block, making it difficult to maintain a balance between the performance (computational amount) and effect (computational results) of the block matching method.

[0081] In an embodiment of the present application, the optical flow of the top image of the image pyramid is initialized to a random value, and the block matching process of the top image of the image pyramid is guided by the randomized initial optical flow, so as to replace the block matching process of the top image of the image pyramid based on the search window in the related art. On the one hand, the optical flow calculation of the top image is no longer limited by the size of the search window, and can cover motion scenes of various amplitudes. For example, for large motion scenes, it is no longer necessary to use a large-size search window, and there will be no situation where the amplitude difference of the motion vector in different layers is too large. On the other hand, since the randomly initialized optical flow is random, it can always hit some of the most matching image blocks. Therefore, further fine-tuning of each layer in the next layer can eventually calculate the optical flow with higher accuracy.

[0082] In the embodiment of the present application, the random offset is usually in the form of (m, n). Taking a pixel as an example, if the position indicated by the random offset is on the right side of the pixel, the value of m is a positive value; if the position indicated by the random offset is on the left side of the pixel, the value of m is a negative value; if the position indicated by the random offset is on the upper side of the pixel, the value of n is a positive value; if the position indicated by the random offset is on the lower side of the pixel, the value of n is a negative value.

[0083] For example, Figure 5As shown, the top image 51 of the first image pyramid of the target frame and the top image 52 of the second image pyramid of the reference frame, take a pixel point 511 in the image 51 as an example, randomly initialize a set of offsets for the pixel point 511, for example, the set of offsets includes 9 random offsets, recorded as [(5,4), (3,2), (-3,3), (-1,1), (-3,0), (4,0), (-3,-3), (0,-3), (7,-3)], where the random offset (5,4) of the pixel point 511 indicates the pixel position 521 in the image 52, the random offset (3,2) of the pixel point 511 indicates the pixel position 522 in the image 52, and the random offset (5,4) of the pixel point 511 indicates the pixel position 521 in the image 52. The random offset (-3, 3) of pixel 511 indicates pixel position 523 in image 52, the random offset (-1, 1) of pixel 511 indicates pixel position 524 in image 52, the random offset (-3, 0) of pixel 511 indicates pixel position 525 in image 52, the random offset (4, 0) of pixel 511 indicates pixel position 526 in image 52, the random offset (-3, -3) of pixel 511 indicates pixel position 527 in image 52, the random offset (0, -3) of pixel 511 indicates pixel position 528 in image 52, and the random offset (7, -3) of pixel 511 indicates pixel position 529 in image 52. Similarly, the random offsets of other pixels in image 51 are taken in a similar way to that of pixel 511, but the values ​​are random and will not be described here.

[0084] In the embodiment of the present application, when the random offset is in the form of (m, n), the values ​​of m and n can satisfy [-8, 8], that is, m∈[-8, 8], n∈[-8, 8]. For example, if the number of layers of the image pyramid is 4, the motion range that can be covered is [-8*2 3 ,8*2 3 ], which can cover most scenarios.

[0085] In the embodiment of the present application, considering that the motion vector of each pixel position in the top-level image of the image pyramid should be uniform, and should not be dominated by small motion or large motion, when configuring the initial optical flow of the N-th layer image of the first image pyramid, the initial optical flow of the N-th layer image of the first image pyramid obeys a uniform distribution, that is, all random offsets in the initial optical flow are uniformly distributed.

[0086] Exemplarily, the random offset is in the form of (m, n), m∈[-8, 8], n∈[-8, 8], and the probability of the 17 values ​​[-8, 8] appearing in the initial optical flow of the Nth layer image of the first image pyramid is the same, that is, uniformly distributed.

[0087] In step 403, according to the initial optical flow of the Nth layer image of the first image pyramid, block matching processing is performed on the Nth layer image of the first image pyramid and the Nth layer image of the second image pyramid, and the fine-tuning optical flow of the Nth layer image of the first image pyramid is determined according to the block matching processing result, and the fine-tuning optical flow of the Nth layer image of the first image pyramid is up-sampled, and the up-sampling result is determined as the initial optical flow of the N-1th layer image of the first image pyramid; wherein the fine-tuning optical flow of the Nth layer image of the first image pyramid includes: the fine-tuning optical flow value of each pixel point in the Nth layer image of the first image pyramid.

[0088] In the embodiment of the present application, according to the initial optical flow of the Nth layer image of the first image pyramid, in the process of performing block matching processing on the Nth layer image of the first image pyramid and the Nth layer image of the second image pyramid, the search window is no longer used, but the above-mentioned initial optical flow is used.

[0089] In some embodiments, Fig. 6A As shown, the above step 403 may include the following steps: step 4031, step 4032, step 4033, step 4034 and step 4035;

[0090] In step 4031, for each pixel point P in the Nth layer image of the first image pyramid ij , read P from the initial optical flow of the Nth layer image of the first image pyramid ij The corresponding set of offsets R ij ; Wherein, 1≤i≤L1, 1≤j≤W1, L1 and W1 are the length and width values ​​of the Nth layer image of the first image pyramid respectively.

[0091] Exemplarily, N=4, the first image pyramid and the second image pyramid are both 4-layer pyramids. The pyramids are, from top to bottom, the fourth layer image, the third layer image, the second layer image and the first layer image, wherein the fourth layer is the top layer of the pyramid and the first layer is the bottom layer of the pyramid.

[0092] Starting from the top layer (i.e., the fourth layer) of the image pyramid, for example, Figure 5As shown, the top image 51 of the first image pyramid of the target frame and the top image 52 of the second image pyramid of the reference frame, take a pixel point 511 in the image 51 as an example, obtain the offset of the pixel point 511 from the initial optical flow of the top image 51 of the first image pyramid, for example, the offset includes 9 random offsets, recorded as [(5,4), (3,2), (-3,3), (-1,1), (-3,0), (4,0), (-3,-3), (0,-3), (7,-3)], wherein the random offset (5,4) of the pixel point 511 indicates the pixel position 521 in the image 52, and the random offset (3,2) of the pixel point 511 indicates the pixel position 521 in the image 52. 22, the random offset (-3, 3) of pixel 511 indicates pixel position 523 in image 52, the random offset (-1, 1) of pixel 511 indicates pixel position 524 in image 52, the random offset (-3, 0) of pixel 511 indicates pixel position 525 in image 52, the random offset (4, 0) of pixel 511 indicates pixel position 526 in image 52, the random offset (-3, -3) of pixel 511 indicates pixel position 527 in image 52, the random offset (0, -3) of pixel 511 indicates pixel position 528 in image 52, and the random offset (7, -3) of pixel 511 indicates pixel position 529 in image 52.

[0093] In step 4032, the image block B in the Nth layer image of the first image pyramid is determined ij , where B ij P ij The image block with the center pixel as the center pixel.

[0094] For example, Figure 5 For example, a pixel 511 of the image 51 in Figure 6B As shown, an image block 512 with pixel 511 as the center pixel in the image 51 is determined.

[0095] It should be noted that the size of the image block in the block matching process can be set according to the actual situation. Figure 6B The following description only takes an image block of 3×3 size as an example.

[0096] In step 4033, a candidate image block in the Nth layer image of the second image pyramid is determined, wherein the candidate image block is a block having a pixel size of R ij Each random offset in is an image block with the center pixel.

[0097] For example, Figure 5 For example, a pixel 511 of the image 51 in Figure 6BAs shown, an image block 512 with pixel 511 as the center pixel in image 51 is determined. Figure 6C As shown, in image 52, determine the alternative image block 531 with pixel 521 as the center pixel, the alternative image block 532 with pixel 522 as the center pixel, the alternative image block 533 with pixel 523 as the center pixel, the alternative image block 534 with pixel 524 as the center pixel, the alternative image block 535 with pixel 525 as the center pixel, the alternative image block 536 with pixel 526 as the center pixel, the alternative image block 537 with pixel 527 as the center pixel, the alternative image block 538 with pixel 528 as the center pixel, and the alternative image block 539 with pixel 529 as the center pixel.

[0098] In step 4034, calculate B ij The similarity between each candidate image block and the candidate image block with the highest similarity is determined as B ij The matching image patches.

[0099] For example, Figure 5 Take the pixel 511 of the image 51 in as an example, Figure 6C As shown, the similarity between image block 512 and candidate image block 531, the similarity between image block 512 and candidate image block 532, the similarity between image block 512 and candidate image block 533, the similarity between image block 512 and candidate image block 534, the similarity between image block 512 and candidate image block 535, the similarity between image block 512 and candidate image block 536, the similarity between image block 512 and candidate image block 537, the similarity between image block 512 and candidate image block 538, and the similarity between image block 512 and candidate image block 539 are calculated. According to the 9 similarities obtained by calculation, the image block with the highest similarity is selected, for example, the image block with the highest similarity is candidate image block 536.

[0100] In step 4035, calculate B ij With B ij The offset between the matching image blocks is calculated and the calculated offset is determined as P ij Fine-tuned optical flow value; among them, all P ij The refined optical flow values ​​constitute the refined optical flow of the Nth layer image of the first image pyramid.

[0101] Exemplarily, since the similarity between the candidate image block 536 and the image block 512 is the highest, the offset between the candidate image block 536 and the image block 512 is calculated, and the calculated offset is determined as the fine-tuning optical flow value of the pixel point 511. Similarly, the fine-tuning optical flow values ​​of other pixel points in the image 51 can be calculated.

[0102] It can be seen that in the embodiment of the present application, random initialization is used for the layer with the smallest resolution of the image pyramid. The motivation for doing so is that random initialization will not be limited by the size of the search window and can better cover scenes with various motion amplitudes. This initialization can always hit some of the most matching blocks, and further fine-tuning in the next layer can eventually estimate the final motion vector relatively well. Due to the application of random initialization, there is no need to use a relatively large search box for large motion scenes, and there will be no situation where the amplitude difference of the motion vector in different layers is too large.

[0103] In step 404, based on the initial optical flow of the N-1th layer image of the first image pyramid, the refined optical flow of each layer image below the Nth layer image of the first image pyramid is calculated layer by layer until the refined optical flow of the first layer image of the first image pyramid is obtained.

[0104] In the embodiment of the present application, when N=2, the N-1th layer image of the first image pyramid is the first layer image. When N>2, it is necessary to first calculate the fine-tuning optical flow of the N-1th layer image based on the initial optical flow of the N-1th layer image of the first image pyramid, and then calculate the fine-tuning optical flow of the N-2th layer image, until the fine-tuning optical flow of the first layer image of the first image pyramid is calculated.

[0105] In some embodiments, Fig. 7A As shown, the above step 404 may include the following steps: step 4041, step 4042 and step 4043;

[0106] In step 4041, based on the search window, block matching processing is performed on the N-1th layer image of the first image pyramid and the N-1th layer image of the second image pyramid, and the initial optical flow of the N-1th layer image of the first image pyramid is fine-tuned according to the block matching processing result to obtain the fine-tuned optical flow of the N-1th layer image of the first image pyramid.

[0107] In the embodiment of the present application, when N=2, it is only necessary to perform the above step 4041 to obtain the fine-tuned optical flow of the first layer image of the first image pyramid.

[0108] In step 4042, when N>2, the refined optical flow of the N-1th layer image of the first image pyramid is upsampled, and the upsampling result is determined as the initial optical flow of the N-2th layer image of the first image pyramid.

[0109] In the embodiment of the present application, since the resolution of the N-1th layer image is lower than the resolution of the N-2th layer image, it is necessary to upsample the fine-tuned optical flow of the N-1th layer image of the first image pyramid to the same size as the N-2th layer image. At this time, the upsampling result is determined as the initial optical flow of the N-2th layer image of the first image pyramid.

[0110] In step 4043, the same processing operations as those of the N-1th layer image of the first image pyramid are sequentially performed on the N-2th layer image to the first layer image of the first image pyramid until a refined optical flow of the first layer image of the first image pyramid is obtained.

[0111] It can be seen that in the embodiment of the present application, the optical flow of other layers other than the top layer of the first image pyramid can be calculated based on the search window. Since the block matching technology of the search window is relatively mature, and the optical flow result of the top layer always has a random offset within a certain range that is lucky enough to hit the most matching image block, the accurate optical flow value of the target frame can be finally calculated.

[0112] For other layers besides the top layer of the image pyramid, considering that in the related art, when performing dense pixel-by-pixel search within the search window, the optical flow values ​​of the neighborhood positions are usually smooth, the offset difference relative to the center pixel is very small, and after several layers of pyramid iterations, the optical flow has gradually approached the actual motion. Therefore, it is a relatively redundant calculation to update the optical flow of the center pixel with the neighborhood optical flow.

[0113] In response to the above problems, in an embodiment of the present application, during the block matching process of the N-1th layer image of the first image pyramid and the N-1th layer image of the second image pyramid, a sparse search is performed within the search window to update the optical flow value of the center pixel with an offset at a farther position. At the same time, since the optical flow of the top image of the image pyramid is randomly initialized, this sparse search can effectively find the more appropriate offset to update the current optical flow value, thereby achieving fine-tuning of the optical flow.

[0114] In some embodiments, Figure 7B As shown, the above step 4041 may include the following steps: step 40411, step 40412, step 40413, step 40414, step 40415, step 40416 and step 40417;

[0115] In step 40411, for each pixel point Q in the N-1th layer image of the first image pyramid st , read Q from the initial optical flow of the N-1th layer image of the first image pyramid stThe initial optical flow value of ; wherein, 1≤s≤L2, 1≤t≤W2, L2, W2 are the length value and width value of the N-1th layer image of the first image pyramid respectively.

[0116] Exemplarily, N=4, the first image pyramid and the second image pyramid are both 4-layer pyramids. The pyramids are, from top to bottom, the fourth layer image, the third layer image, the second layer image and the first layer image, wherein the fourth layer is the top layer of the pyramid and the first layer is the bottom layer of the pyramid. The N-1th layer image of the first image pyramid is the third layer image of the first image pyramid.

[0117] Exemplarily, the initial optical flow of the third layer image of the first image pyramid is fine-tuned according to the block matching method to obtain the fine-tuned optical flow of the third layer image of the first image pyramid. Figure 7C As shown, the third layer image 61 of the image pyramid of the target frame and the third layer image 62 of the image pyramid of the reference frame, taking the optical flow value of a pixel point 611 in the fine-tuning image 61 as an example, obtain the initial optical flow value of the pixel point 611 in the initial optical flow of the image 61.

[0118] In step 40412, the image block A in the N-1th layer image of the first image pyramid is determined st , where A st Q st The image block with the center pixel as the center pixel.

[0119] For example, Figure 7C For example, a pixel 611 in the image 61 is Fig.7D As shown, an image block 612 with pixel 611 as the center pixel in the image 61 is determined.

[0120] In step 40413, the search window Z in the N-1th layer image of the second image pyramid is determined. st , where Z st A st The corresponding search window.

[0121] For example, Figure 7C For example, a pixel 611 in the image 61 is Fig. 7E As shown, a search window 621 corresponding to the image block 612 in the image 62 is determined.

[0122] It should be noted that, since a sparse search method is used when performing block matching, the size of the search window 621 can be set slightly larger than the size of the search window in the related art.

[0123] In step 40414, according to Q st The initial optical flow value C1 determines Zst The image block G in st , where G st A st The corresponding image block.

[0124] For example, Figure 7C For example, a pixel 611 in the image 61 is Figure 7F As shown, an image block 622 corresponding to the initial optical flow value of the pixel point 611 in the image 62 is determined.

[0125] In step 40415, configure A st In Z st The search step length is H, where H is an integer greater than 1.

[0126] In the embodiment of the present application, the value of the search step length H is set to an integer greater than 1 to achieve a sparse search within the search window. For example, H=2.

[0127] In step 40416, according to H, in Z st Search for matching image block Y in st , where Y st For A st The image blocks with the highest similarity.

[0128] For example, H=2, such as Figure 7G As shown, the image block 612 moves up, down, left and right in the search window 621 with a step size of 2 to perform block search, and finds the matching image block with the highest similarity to the image block 612 in the search window 621. For example, the matching image block with the highest similarity is image block 624.

[0129] In step 40417, determine A st With Y st Is the similarity X1 between them greater than A? st With G st If X1>X2, then determine the initial optical flow value C2 in the initial optical flow of the N-1th layer of the first image pyramid, and determine C2 as Q st The fine-tuned optical flow value, where C2 is Y st The initial optical flow value corresponding to the central pixel of st The initial optical flow value C1 is determined as Q st Fine-tuned optical flow value; among them, all Q st The refined optical flow values ​​constitute the refined optical flow of the N-1th layer image of the first image pyramid.

[0130] Exemplarily, if the similarity between image block 612 and image block 624 is greater than the similarity between image block 612 and image block 622, it means that the initial optical flow value of pixel 613 in image 61 corresponding to central pixel 623 of image block 624 is more suitable as the fine-tuning optical flow value of pixel 611 than the initial optical flow value of pixel 611. In this case, the initial optical flow value of pixel 613 is determined as the fine-tuning optical flow value of 611. If the similarity between image block 612 and image block 624 is less than the similarity between image block 612 and image block 622, it means that the initial optical flow value of pixel 611 is the most suitable optical flow value. The initial optical flow value of pixel 611 is determined as the fine-tuning optical flow value of pixel 611. Similarly, the fine-tuning optical flow values ​​of other pixels in the third layer image of the image pyramid of the target frame can be calculated, that is, the fine-tuning optical flow of the third layer image. After calculating the fine-tuned optical flow of the third layer image of the image pyramid of the target frame, up-sample the fine-tuned optical flow of the third layer image, and transfer the up-sampled result to the second layer of the image pyramid of the target frame as the initial optical flow of the second layer image.

[0131] In the embodiment of the present application, the same calculation process as the third layer is performed on the second layer of the image pyramid of the target frame to obtain the fine-tuned optical flow of the second layer image of the image pyramid of the target frame. The fine-tuned optical flow of the second layer image is transferred to the first layer of the image pyramid of the target frame as the initial optical flow of the first layer image. The same calculation process as the third layer is performed on the first layer of the image pyramid of the target frame to obtain the fine-tuned optical flow of the first layer image of the image pyramid of the target frame.

[0132] It can be seen that in the embodiment of the present application, for other layers besides the top layer of the image pyramid, block matching can be performed in a larger search space through sparse search, thereby reducing the amount of calculation while ensuring the accuracy of block matching.

[0133] In step 405, the fine-tuned optical flow of the first layer image of the first image pyramid is denoised to obtain the target optical flow of the target frame.

[0134] Taking into account that the optical flow obtained based on the block matching method has large noise, in order to improve the quality of optical flow estimation, in an embodiment of the present application, the fine-tuned optical flow of the first layer image of the first image pyramid is denoised to obtain the target optical flow of the target frame.

[0135] Considering that large convolution kernel filtering has the following advantages: 1) Larger receptive field: Large convolution kernels can cover a larger input area, thereby capturing more global information in one convolution operation. 2) Reduced network depth: Large convolution kernels can obtain more global information in one operation, reducing the number of layers required, thereby reducing network depth, helping to improve training stability and reduce the problem of gradient vanishing. 3) Smooth output: Large convolution kernels contain more adjacent pixel information, and have a more significant smoothing effect on feature maps, which helps to reduce sensitivity to local noise and enhance the capture of global patterns. Therefore, in an embodiment of the present application, a denoising model based on a large convolution kernel can be used to denoise the fine-tuned optical flow of the first layer image of the first image pyramid to obtain the target optical flow of the target frame.

[0136] In some embodiments, in order to improve the denoising efficiency, the above step 405 may include the following steps:

[0137] Step 4051;

[0138] In step 4051, the fine-tuned optical flow of the first layer image of the first image pyramid and the first layer image of the second image pyramid are input into the re-parameterized target denoising model for denoising, and the target optical flow of the target frame is output; wherein the target denoising model includes a type of convolution kernel and the size of the convolution kernel is greater than a preset size value.

[0139] In the embodiment of the present application, a neural network is used to learn the filter kernel. In order to enable the optical flow to be applied in real time in videos or games, the efficiency requirements will be relatively high. Since a large convolution kernel filter is used, a network model with too many layers cannot be used. In order to ensure that the number of layers is small and the denoising effect is good, a re-parameterization technology is used to design the network. Specifically, the training model is decoupled from the inference model. Since convolutions are all linear operations, multiple parallel convolution kernels can be merged into a single large-core convolution. Since the number of layers of the network is relatively small, after the training is completed, the filtering operation can be directly applied in the form of a lookup table, which can further improve the efficiency of the optical flow calculation. Among them, the basic principle of re-parameterization calculation is that convolution is a linear operation, so the addition of multiple linear operations is still a linear operation.

[0140] In some embodiments, Fig. 8A As shown, the generation process of the above target denoising model may include the following steps:

[0141] Step 801, step 802, step 803 and step 804;

[0142] In step 801, a training set is obtained; wherein the training set includes multiple groups of sample data, each group of sample data includes: a first sample image, a second sample image, a noisy optical flow of the first sample image and a true optical flow of the first sample image, and the second sample image is a reference image of the first sample image.

[0143] In the embodiment of the present application, in order to ensure the training effect of the model, the training set may include a large amount of sample data.

[0144] In the embodiment of the present application, the real optical flow of the first sample image may be manually annotated or may be derived from a data set publicly available on the Internet.

[0145] In step 802, a backbone network is constructed; wherein the backbone network includes at least two groups of convolution kernels, and each group of convolution kernels includes a plurality of convolution kernels of different sizes connected in parallel.

[0146] For example, Figure 8B As shown in the figure, the backbone network includes: two groups of convolution kernels and two activation layers, each group of convolution kernels includes: a 7×7 convolution kernel, a 5×5 convolution kernel, a 3×3 convolution kernel and a 1×1 convolution kernel.

[0147] In an embodiment of the present application, the backbone network includes at least two groups of convolution kernels, each group of convolution kernels includes multiple parallel-connected convolution kernels of different sizes, so that more comprehensive prior information of the training data can be learned during the model training process.

[0148] In step 803, the noisy optical flows of the second sample image and the first sample image in the sample data are input into the backbone network for denoising, and the predicted image and the denoised optical flow of the predicted image are output; the loss value is calculated according to the first sample image, the real optical flow of the first sample image, the predicted image and the denoised optical flow of the predicted image, and the network parameters of the backbone network are updated according to the loss value; the above training process is repeated using different groups of sample data until the model converges to obtain an initial denoising model.

[0149] For example, the loss function in the model training process can adopt the following formula (1):

[0150] loss=α*L1(flow, flowgt)+β*L2(I2, warp(I1, flow)) (1)

[0151] Wherein, loss represents the loss function, α and β are weight coefficients, for example, α=0.5, β=0.5, flow represents the denoised optical flow of the predicted image, flowgt represents the true optical flow of the first sample image, I1 represents the predicted image, I2 represents the first sample image, L1 represents the absolute value difference function, and L2 represents the square difference function.

[0152] In step 804, the initial denoising model is re-parameterized to generate a target denoising model.

[0153] For example, Figure 8B As shown in , all convolution kernels in each group of convolution kernels are jump-connected to the 7×7 convolution kernel. Since a smaller convolution kernel can be expanded into a larger convolution kernel by expanding the edge with 0, the sum of multiple convolution kernels of different sizes can be modeled into a large convolution kernel, as shown in Figure 8C As shown. Based on this principle, the model in the training phase and the model in the inference phase can be decoupled. In the training phase, more parallel branches are used. Figure 8B In the training phase, more parallel branches can be used to learn more information; in the inference phase, Figure 8C This is the network structure after the merger. This structure has no parallel branches and is more memory-friendly, so it can be reasoned more efficiently. Since there are only two layers of 7×7 convolution kernels, a lookup table can be used to further improve the efficiency of convolution operations.

[0154] It can be seen that in the embodiments of the present application, through the re-parameterized design, the training and reasoning of the model can be decoupled, so that both the training effect and the reasoning speed can be improved.

[0155] As can be seen from the above embodiment, in this embodiment, a first image pyramid of a target frame and a second image pyramid of a reference frame are constructed; wherein the number of layers of the first image pyramid and the number of layers of the second image pyramid are both N, and N is an integer greater than 1; an initial optical flow of an N-th layer image of the first image pyramid is configured; wherein the initial optical flow of the N-th layer image of the first image pyramid includes: a group of offsets for each pixel point in the N-th layer image of the first image pyramid, each group of offsets includes: at least two different random offsets, each random offset is used to indicate a pixel position in the N-th layer image of the second image pyramid; according to the initial optical flow of the N-th layer image of the first image pyramid, the N-th layer image of the first image pyramid and the N-th layer image of the second image pyramid are subjected to optical flow comparison. A block matching process is performed, and a fine-tuning optical flow of an N-th layer image of a first image pyramid is determined according to a result of the block matching process, and the fine-tuning optical flow of the N-th layer image of the first image pyramid is upsampled, and the upsampling result is determined as an initial optical flow of an N-1-th layer image of the first image pyramid; wherein the fine-tuning optical flow of the N-th layer image of the first image pyramid comprises: a fine-tuning optical flow value of each pixel point in the N-th layer image of the first image pyramid; based on the initial optical flow of the N-1-th layer image of the first image pyramid, the fine-tuning optical flow of each layer image below the N-th layer image of the first image pyramid is calculated layer by layer until the fine-tuning optical flow of the first layer image of the first image pyramid is obtained; and the fine-tuning optical flow of the first layer image of the first image pyramid is denoised to obtain a target optical flow of a target frame.

[0156] It can be seen that compared with the related art in which the optical flow of the top layer of the image pyramid is initialized to all zeros, the initial optical flow based on all zeros cannot provide any information for the top layer, and can only find the most matching block in the neighborhood of the search window based on the search window, and the small-sized search box cannot cover large motion scenes. In the embodiment of the present application, the optical flow of the top layer of the image pyramid is randomly initialized. Since the randomly initialized optical flow contains a set of random offsets for each pixel point of the top layer image, and there is always a random offset within a certain range that is lucky enough to hit the most matching image block, further fine-tuning in the next layer can better estimate the final optical flow, so that the block matching processing of the top layer image is no longer limited by the size of the search window, and the accuracy of the optical flow results calculated for scenes with large motion can be improved.

[0157] In summary, the beneficial effects of the embodiments of the present application are manifested in the following aspects: 1) Better optical flow effect: By adopting random initialization, scenes of various motion amplitudes can be better covered, and through sparse search, block matching can be performed in a larger search space, which can greatly improve the accuracy of block matching. In the denoising stage, by adding more parallel branches in the training stage, the network can be helped to learn more information. 2) Faster running speed: The sparse search method can greatly improve the accuracy of block matching without increasing the search space. In the denoising stage, by merging the parallel branches of the network, the speed of reasoning can be significantly improved, and the convolution operation can be further performed by using a lookup table, which can further improve the operation speed. 3) Easier parallel programming: No similar spatial iteration scheme is adopted, which is very convenient for parallelizing the program, so there is more room to further improve the optical flow effect. In general, the embodiments of the present application can have a better balance between the optical flow effect and the optical flow operation efficiency, and can make the optical flow real-time application in videos and games, and improve the picture effect of videos or games.

[0158] In addition, the optical flow calculation method provided in the embodiment of the present application can also be applied to the following scenarios: 1) Application in more scenarios: Since image alignment is an important step in the multi-frame image alignment task, a similar method can be used to calculate the optical flow in multi-frame high dynamic range imaging, multi-frame super-resolution and multi-frame denoising, and then the optical flow is used to align the image. 2) Binocular disparity calculation: Since the calculation principle of binocular disparity is basically similar to that of optical flow, the only difference is that disparity calculation is a one-dimensional match, while optical flow calculation is a two-dimensional match. Therefore, this optical flow solution can be directly applied to the binocular disparity estimation task.

[0159] The optical flow calculation method provided in the embodiment of the present application can be executed by an optical flow calculation device. In the embodiment of the present application, an optical flow calculation device executing the optical flow calculation method is taken as an example to illustrate the optical flow calculation device provided in the embodiment of the present application.

[0160] Fig. 9 is a structural block diagram of an optical flow calculation device provided in an embodiment of the present application, such as Fig. 9 As shown, the optical flow calculation device 900 may include: a construction module 901, a configuration module 902, a first processing module 903, a second processing module 904 and a denoising module 905;

[0161] The construction module 901 is used to construct a first image pyramid of a target frame and a second image pyramid of a reference frame; wherein the number of layers of the first image pyramid and the number of layers of the second image pyramid are both N, and N is an integer greater than 1;

[0162] The configuration module 902 is configured to configure an initial optical flow of the Nth layer image of the first image pyramid; wherein the initial optical flow of the Nth layer image of the first image pyramid comprises: a group of offsets for each pixel in the Nth layer image of the first image pyramid, each group of offsets comprises: at least two different random offsets, each of the random offsets is used to indicate a pixel position in the Nth layer image of the second image pyramid;

[0163] The first processing module 903 is configured to perform block matching processing on the Nth layer image of the first image pyramid and the Nth layer image of the second image pyramid according to the initial optical flow of the Nth layer image of the first image pyramid, determine the fine-tuning optical flow of the Nth layer image of the first image pyramid according to the block matching processing result, and up-sample the fine-tuning optical flow of the Nth layer image of the first image pyramid, and determine the up-sampling result as the initial optical flow of the N-1th layer image of the first image pyramid; wherein the fine-tuning optical flow of the Nth layer image of the first image pyramid includes: the fine-tuning optical flow value of each pixel point in the Nth layer image of the first image pyramid;

[0164] The second processing module 904 is configured to calculate, based on the initial optical flow of the N-1th layer image of the first image pyramid, the refined optical flow of each layer image below the Nth layer image of the first image pyramid layer by layer until the refined optical flow of the first layer image of the first image pyramid is obtained;

[0165] The denoising module 905 is used to perform denoising on the fine-tuned optical flow of the first layer image of the first image pyramid to obtain the target optical flow of the target frame.

[0166] As can be seen from the above embodiment, in this embodiment, a first image pyramid of a target frame and a second image pyramid of a reference frame are constructed; wherein the number of layers of the first image pyramid and the number of layers of the second image pyramid are both N, and N is an integer greater than 1; an initial optical flow of an N-th layer image of the first image pyramid is configured; wherein the initial optical flow of the N-th layer image of the first image pyramid includes: a group of offsets for each pixel point in the N-th layer image of the first image pyramid, each group of offsets includes: at least two different random offsets, each random offset is used to indicate a pixel position in the N-th layer image of the second image pyramid; according to the initial optical flow of the N-th layer image of the first image pyramid, the N-th layer image of the first image pyramid and the N-th layer image of the second image pyramid are subjected to optical flow comparison. A block matching process is performed, and a fine-tuning optical flow of an N-th layer image of a first image pyramid is determined according to a result of the block matching process, and the fine-tuning optical flow of the N-th layer image of the first image pyramid is upsampled, and the upsampling result is determined as an initial optical flow of an N-1-th layer image of the first image pyramid; wherein the fine-tuning optical flow of the N-th layer image of the first image pyramid comprises: a fine-tuning optical flow value of each pixel point in the N-th layer image of the first image pyramid; based on the initial optical flow of the N-1-th layer image of the first image pyramid, the fine-tuning optical flow of each layer image below the N-th layer image of the first image pyramid is calculated layer by layer until the fine-tuning optical flow of the first layer image of the first image pyramid is obtained; and the fine-tuning optical flow of the first layer image of the first image pyramid is denoised to obtain a target optical flow of a target frame.

[0167] It can be seen that compared with the related art in which the optical flow of the top layer of the image pyramid is initialized to all zeros, the initial optical flow based on all zeros cannot provide any information for the top layer, and can only find the most matching block in the neighborhood of the search window based on the search window, and the small-sized search box cannot cover large motion scenes. In the embodiment of the present application, the optical flow of the top layer of the image pyramid is randomly initialized. Since the randomly initialized optical flow contains a set of random offsets for each pixel point of the top layer image, and there is always a random offset within a certain range that is lucky enough to hit the most matching image block, further fine-tuning in the next layer can better estimate the final optical flow, so that the block matching processing of the top layer image is no longer limited by the size of the search window, and the accuracy of the optical flow results calculated for scenes with large motion can be improved.

[0168] Optionally, as an embodiment, the first processing module 903 may be specifically configured to process each pixel point P in the Nth layer image of the first image pyramid. ij , read the P from the initial optical flow of the Nth layer image of the first image pyramid ij The corresponding set of offsets R ij; Wherein, 1≤i≤L1, 1≤j≤W1, L1, W1 are respectively the length value and the width value of the Nth layer image of the first image pyramid; Determine the image block B in the Nth layer image of the first image pyramid ij , wherein the B ij For the P ij An image block with the R as the central pixel point; determining a candidate image block in the Nth layer image of the second image pyramid, wherein the candidate image block is an image block with the R ij Each random offset in the image block is a central pixel; calculate the B ij The similarity between each candidate image block and the candidate image block with the highest similarity is determined as the B ij The matching image block of B ij With the B ij The offset between the matching image blocks, and the calculated offset is determined as the P ij Fine-tuned optical flow value; wherein all of the P ij The refined optical flow values ​​constitute the refined optical flow of the Nth layer image of the first image pyramid.

[0169] Optionally, as an embodiment, the initial optical flow of the Nth layer image of the first image pyramid may obey a uniform distribution.

[0170] Optionally, as an embodiment, the second processing module 904 can be specifically used to perform block matching processing on the N-1th layer image of the first image pyramid and the N-1th layer image of the second image pyramid based on the search window, and fine-tune the initial optical flow of the N-1th layer image of the first image pyramid according to the block matching processing result to obtain the fine-tuned optical flow of the N-1th layer image of the first image pyramid; when N>2, upsample the fine-tuned optical flow of the N-1th layer image of the first image pyramid, and determine the upsampling result as the initial optical flow of the N-2th layer image of the first image pyramid; and perform the same processing operations as the N-1th layer image of the first image pyramid on the N-2th layer image to the first layer image of the first image pyramid in sequence until the fine-tuned optical flow of the first layer image of the first image pyramid is obtained.

[0171] Optionally, as an embodiment, the second processing module 904 may be specifically configured to process each pixel point Q in the N-1th layer image of the first image pyramid. st , read the Q from the initial optical flow of the N-1th layer image of the first image pyramid stThe initial optical flow value C1 of the image pyramid is: 1≤s≤L2, 1≤t≤W2, L2 and W2 are respectively the length and width of the image at the N-1th layer of the first image pyramid; determine the image block A in the image at the N-1th layer of the first image pyramid st , wherein the A st For the Q st The image block is the center pixel point; determine the search window Z in the N-1th layer image of the second image pyramid st , wherein the Z st For the A st The corresponding search window; according to the Q st The initial optical flow value C1 determines the Z st The image block G in st , wherein the G st For the A st The corresponding image block; configure the A st In the Z st The search step length H in Z is H, where H is an integer greater than 1; according to H, in Z st Search for matching image block Y in st , wherein the Y st For the A st The image block with the highest similarity between them; determine the A st With the Y st Is the similarity X1 between them greater than that of A st With the G st If X1>X2, then determine the initial optical flow value C2 in the initial optical flow of the N-1th layer of the first image pyramid, and determine C2 as the Q st The fine-tuned optical flow value of st The initial optical flow value corresponding to the central pixel of st The initial optical flow value C1 is determined as the Q st Fine-tuned optical flow value; wherein all of the Q st The refined optical flow values ​​constitute the refined optical flow of the N-1th layer image of the first image pyramid.

[0172] Optionally, as an embodiment, the denoising module 905 can be specifically used to input the fine-tuned optical flow of the first layer image of the first image pyramid and the first layer image of the second image pyramid into a re-parameterized target denoising model for denoising, and output the target optical flow of the target frame; wherein the target denoising model includes a type of convolution kernel and the size of the convolution kernel is greater than a preset size value.

[0173] Optionally, as an embodiment, the optical flow calculation device 900 may further include: a training module;

[0174] The training module is used to obtain a training set; wherein the training set includes multiple groups of sample data, each group of sample data includes: a first sample image, a second sample image, a noisy optical flow of the first sample image and a real optical flow of the first sample image, and the second sample image is a reference image of the first sample image; construct a backbone network; wherein the backbone network includes at least two groups of convolution kernels, and each group of convolution kernels includes a plurality of convolution kernels of different sizes connected in parallel; the noisy optical flows of the second sample image and the first sample image in the sample data are input into the backbone network for denoising, and a predicted image and a denoised optical flow of the predicted image are output; a loss value is calculated based on the first sample image, the real optical flow of the first sample image, the predicted image and the denoised optical flow of the predicted image, and the network parameters of the backbone network are updated based on the loss value; the above training process is repeatedly performed using different groups of the sample data until the model converges to obtain an initial denoising model; the initial denoising model is reparameterized to generate the target denoising model.

[0175] The optical flow calculation device in the embodiment of the present application can be an electronic device, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or it can be other devices other than a terminal. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (Augmented Reality, AR) / virtual reality (Virtual Reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (Ultra-Mobile Personal Computer, UMPC), a netbook, or a personal digital assistant (Personal Digital Assistant, PDA), etc., and can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (Personal Computer, PC), a television (Television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.

[0176] The optical flow calculation device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an IOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0177] The optical flow calculation device provided in the embodiment of the present application can realize Figure 4 , Fig. 6A , Fig. 7A , Figure 7B and Fig. 8A To avoid repetition, the various processes implemented by any of the method embodiments described in the present invention will not be described in detail here.

[0178] Alternatively, if Fig.10 As shown, an embodiment of the present application further provides an electronic device 1000, including a processor 1001 and a memory 1002, wherein the memory 1002 stores a program or instruction that can be executed on the processor 1001, and when the program or instruction is executed by the processor 1001, each step of the above-mentioned optical flow calculation method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0179] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0180] Fig.11 It is a schematic diagram of the hardware structure of an electronic device provided in various embodiments of the present application.

[0181] The electronic device 1100 includes but is not limited to: a radio frequency unit 1101, a network module 1102, an audio output unit 1103, an input unit 1104, a sensor 1105, a display unit 1106, a user input unit 1107, an interface unit 1108, a memory 1109 and a processor 1110 and other components.

[0182] Those skilled in the art will appreciate that the electronic device 1100 may also include a power source (such as a battery) for supplying power to each component, and the power source may be logically connected to the processor 1110 through a power management system, thereby implementing functions such as managing charging, discharging, and power consumption management through the power management system. Fig.11 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be described in detail here.

[0183] In some embodiments, the processor 1110 is configured to construct a first image pyramid of a target frame and a second image pyramid of a reference frame; wherein the number of layers of the first image pyramid and the number of layers of the second image pyramid are both N, and N is an integer greater than 1; configure an initial optical flow of an N-th layer image of the first image pyramid; wherein the initial optical flow of the N-th layer image of the first image pyramid comprises: a set of offsets for each pixel point in the N-th layer image of the first image pyramid, each set of offsets comprises: at least two different random offsets, each of the random offsets is used to indicate a pixel position in the N-th layer image of the second image pyramid; according to the initial optical flow of the N-th layer image of the first image pyramid, the N-th layer image of the first image pyramid and the N-th layer image of the second image pyramid are respectively The method comprises the steps of: performing block matching processing on the image, determining a fine-tuning optical flow of an N-th layer image of the first image pyramid according to a result of the block matching processing, upsampling the fine-tuning optical flow of the N-th layer image of the first image pyramid, and determining the upsampling result as an initial optical flow of an N-1-th layer image of the first image pyramid; wherein the fine-tuning optical flow of the N-th layer image of the first image pyramid comprises: a fine-tuning optical flow value of each pixel point in the N-th layer image of the first image pyramid; calculating the fine-tuning optical flow of each layer image below the N-th layer image of the first image pyramid layer by layer based on the initial optical flow of the N-1-th layer image of the first image pyramid, until the fine-tuning optical flow of the first layer image of the first image pyramid is obtained; and performing denoising processing on the fine-tuning optical flow of the first layer image of the first image pyramid to obtain a target optical flow of the target frame.

[0184] It can be seen that compared with the related art in which the optical flow of the top layer of the image pyramid is initialized to all zeros, the initial optical flow based on all zeros cannot provide any information for the top layer, and can only find the most matching block in the neighborhood of the search window based on the search window, and the small-sized search box cannot cover large motion scenes. In the embodiment of the present application, the optical flow of the top layer of the image pyramid is randomly initialized. Since the randomly initialized optical flow contains a set of random offsets for each pixel point of the top layer image, and there is always a random offset within a certain range that is lucky enough to hit the most matching image block, further fine-tuning in the next layer can better estimate the final optical flow, so that the block matching processing of the top layer image is no longer limited by the size of the search window, and the accuracy of the optical flow results calculated for scenes with large motion can be improved.

[0185] Optionally, as an embodiment, the processor 1110 is specifically configured to, for each pixel point P in the Nth layer image of the first image pyramid, ij , read the P from the initial optical flow of the Nth layer image of the first image pyramid ij The corresponding set of offsets R ij; Wherein, 1≤i≤L1, 1≤j≤W1, L1, W1 are respectively the length value and the width value of the Nth layer image of the first image pyramid; Determine the image block B in the Nth layer image of the first image pyramid ij , wherein the B ij For the P ij An image block with the R as the central pixel point; determining a candidate image block in the Nth layer image of the second image pyramid, wherein the candidate image block is an image block with the R ij Each random offset in the image block is a central pixel; calculate the B ij The similarity between each candidate image block and the candidate image block with the highest similarity is determined as the B ij The matching image block of B ij With the B ij The offset between the matching image blocks, and the calculated offset is determined as the P ij Fine-tuned optical flow value; wherein all of the P ij The refined optical flow values ​​constitute the refined optical flow of the Nth layer image of the first image pyramid.

[0186] Optionally, as an embodiment, the initial optical flow of the Nth layer image of the first image pyramid obeys a uniform distribution.

[0187] Optionally, as an embodiment, the processor 1110 is specifically used to perform block matching processing on the N-1th layer image of the first image pyramid and the N-1th layer image of the second image pyramid based on the search window, and fine-tune the initial optical flow of the N-1th layer image of the first image pyramid according to the block matching processing result to obtain the fine-tuned optical flow of the N-1th layer image of the first image pyramid; when N>2, upsample the fine-tuned optical flow of the N-1th layer image of the first image pyramid, and determine the upsampling result as the initial optical flow of the N-2th layer image of the first image pyramid; and perform the same processing operations as the N-1th layer image of the first image pyramid on the N-2th layer image to the first layer image of the first image pyramid in sequence until the fine-tuned optical flow of the first layer image of the first image pyramid is obtained.

[0188] Optionally, as an embodiment, the processor 1110 is specifically configured to, for each pixel point Q in the N-1th layer image of the first image pyramid, st , read the Q from the initial optical flow of the N-1th layer image of the first image pyramid stThe initial optical flow value C1 of the image pyramid is: 1≤s≤L2, 1≤t≤W2, L2 and W2 are respectively the length and width of the image at the N-1th layer of the first image pyramid; determine the image block A in the image at the N-1th layer of the first image pyramid st , wherein the A st For the Q st The image block is the center pixel point; determine the search window Z in the N-1th layer image of the second image pyramid st , wherein the Z st For the A st The corresponding search window; according to the Q st The initial optical flow value C1 determines the Z st The image block G in st , wherein the G st For the A st The corresponding image block; configure the A st In the Z st The search step length H in Z is H, where H is an integer greater than 1; according to H, in Z st Search for matching image block Y in st , wherein the Y st For the A st The image block with the highest similarity between them; determine the A st With the Y st Is the similarity X1 between them greater than that of A st With the G st If X1>X2, then determine the initial optical flow value C2 in the initial optical flow of the N-1th layer of the first image pyramid, and determine C2 as the Q st The fine-tuned optical flow value of st The initial optical flow value corresponding to the central pixel of st The initial optical flow value C1 is determined as the Q st Fine-tuned optical flow value; wherein all of the Q st The refined optical flow values ​​constitute the refined optical flow of the N-1th layer image of the first image pyramid.

[0189] Optionally, as an embodiment, the processor 1110 is specifically used to input the fine-tuned optical flow of the first layer image of the first image pyramid and the first layer image of the second image pyramid into a re-parameterized target denoising model for denoising, and output the target optical flow of the target frame; wherein the target denoising model includes a type of convolution kernel and the size of the convolution kernel is greater than a preset size value.

[0190] Optionally, as an embodiment, the processor 1110 is also used to obtain a training set; wherein the training set includes multiple groups of sample data, each group of sample data includes: a first sample image, a second sample image, a noisy optical flow of the first sample image and a real optical flow of the first sample image, and the second sample image is a reference image of the first sample image; construct a backbone network; wherein the backbone network includes at least two groups of convolution kernels, and each group of convolution kernels includes multiple parallel-connected convolution kernels of different sizes; the noisy optical flows of the second sample image and the first sample image in the sample data are input into the backbone network for denoising, and a predicted image and a denoised optical flow of the predicted image are output; a loss value is calculated based on the first sample image, the real optical flow of the first sample image, the predicted image and the denoised optical flow of the predicted image, and the network parameters of the backbone network are updated based on the loss value; the above training process is repeatedly performed using different groups of the sample data until the model converges to obtain an initial denoising model; the initial denoising model is reparameterized to generate the target denoising model.

[0191] It should be understood that in the embodiment of the present application, the input unit 1104 may include a graphics processor (Graphics Processing Unit, GPU) 11041 and a microphone 11042, and the graphics processor 11041 processes the image data of the static picture or video obtained by the image capture device (such as a camera) in the video capture mode or the image capture mode. The display unit 1106 may include a display panel 11061, and the display panel 11061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1107 includes a touch panel 11071 and at least one of other input devices 11072. The touch panel 11071 is also called a touch screen. The touch panel 11071 may include two parts: a touch detection device and a touch controller. Other input devices 11072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.

[0192] The memory 1109 can be used to store software programs and various data. The memory 1109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, an application program or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory 1109 may include a volatile memory or a non-volatile memory, or the memory 1109 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM) and a direct memory bus random access memory (DRRAM). The memory 1109 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0193] The processor 1110 may include one or more processing units; optionally, the processor 1110 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It is understandable that the modem processor may not be integrated into the processor 1110.

[0194] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned optical flow calculation method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0195] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0196] The embodiment of the present application also provides a chip, the chip includes a processor and a communication interface, the communication interface is coupled to the processor, the processor is used to run a program or instruction, implement the various processes of the above-mentioned optical flow calculation method embodiment, and can achieve the same technical effect. To avoid repetition, it is not repeated here. It should be understood that the chip mentioned in the embodiment of the present application can also be called a system-level chip, a system chip, a chip system, or a system-on-chip chip, etc.

[0197] The embodiment of the present application also provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned optical flow calculation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0198] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0199] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for enabling a terminal (such as a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0200] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. An optical flow calculation method, characterized in that: The method comprises: Constructing a first image pyramid of the target frame and a second image pyramid of the reference frame; wherein the number of layers of the first image pyramid and the number of layers of the second image pyramid are both N, and N is an integer greater than 1; Configuring an initial optical flow of an N-th layer image of the first image pyramid; wherein the initial optical flow of the N-th layer image of the first image pyramid comprises: a group of offsets for each pixel in the N-th layer image of the first image pyramid, each group of offsets comprising: at least two different random offsets, each of the random offsets being used to indicate a pixel position in the N-th layer image of the second image pyramid; According to the initial optical flow of the Nth layer image of the first image pyramid, block matching processing is performed on the Nth layer image of the first image pyramid and the Nth layer image of the second image pyramid, and a fine-tuning optical flow of the Nth layer image of the first image pyramid is determined according to the block matching processing result, and the fine-tuning optical flow of the Nth layer image of the first image pyramid is up-sampled, and the up-sampling result is determined as the initial optical flow of the N-1th layer image of the first image pyramid; wherein the fine-tuning optical flow of the Nth layer image of the first image pyramid includes: a fine-tuning optical flow value of each pixel point in the Nth layer image of the first image pyramid; Based on the initial optical flow of the N-1th layer image of the first image pyramid, calculate the fine-tuning optical flow of each layer image below the Nth layer image of the first image pyramid layer by layer until the fine-tuning optical flow of the first layer image of the first image pyramid is obtained; The fine-tuned optical flow of the first layer image of the first image pyramid is denoised to obtain a target optical flow of the target frame.

2. The method according to claim 1, characterized in that The method of performing block matching processing on the N-th layer image of the first image pyramid and the N-th layer image of the second image pyramid according to the initial optical flow of the N-th layer image of the first image pyramid, and determining the fine-tuning optical flow of the N-th layer image of the first image pyramid according to the block matching processing result includes: For each pixel point P in the Nth layer image of the first image pyramid ij , read the P from the initial optical flow of the Nth layer image of the first image pyramid ij The corresponding set of offsets R ij ; Wherein, 1≤i≤L1, 1≤j≤W1, L1 and W1 are the length and width of the Nth layer image of the first image pyramid respectively; Determine an image block B in the Nth layer image of the first image pyramid ij , wherein the B ij For the P ij The image block is the center pixel; Determine a candidate image block in the Nth layer image of the second image pyramid, wherein the candidate image block is a ij Each random offset in is an image block with the center pixel point; Calculate the B ij The similarity between each candidate image block and the candidate image block with the highest similarity is determined as the B ij The matching image patches; Calculate the B ij With the B ij The offset between the matching image blocks, and the calculated offset is determined as the P ij Fine-tuned optical flow value; wherein all of the P ij The refined optical flow values ​​constitute the refined optical flow of the Nth layer image of the first image pyramid.

3. The method according to claim 1 or 2, characterized in that: The initial optical flow of the Nth layer image of the first image pyramid obeys uniform distribution.

4. The method according to claim 1, characterized in that The step of calculating the fine-tuning optical flow of each layer of the image below the Nth layer of the first image pyramid layer by layer based on the initial optical flow of the N-1th layer of the image of the first image pyramid, until the fine-tuning optical flow of the first layer of the image of the first image pyramid is obtained, includes: Based on the search window, block matching is performed on the N-1th layer image of the first image pyramid and the N-1th layer image of the second image pyramid, and the initial optical flow of the N-1th layer image of the first image pyramid is fine-tuned according to the block matching result to obtain the fine-tuned optical flow of the N-1th layer image of the first image pyramid; when N>2, the fine-tuned optical flow of the N-1th layer image of the first image pyramid is up-sampled, and the up-sampling result is determined as the initial optical flow of the N-2th layer image of the first image pyramid; The same processing operations as those for the N-1th layer image of the first image pyramid are sequentially performed on the N-2th layer image to the first layer image of the first image pyramid until a fine-tuned optical flow of the first layer image of the first image pyramid is obtained.

5. The method according to claim 4, characterized in that The method of performing block matching processing on the N-1th layer image of the first image pyramid and the N-1th layer image of the second image pyramid based on the search window, and fine-tuning the initial optical flow of the N-1th layer image of the first image pyramid according to the block matching processing result to obtain the fine-tuned optical flow of the N-1th layer image of the first image pyramid includes: For each pixel point Q in the N-1th layer image of the first image pyramid st , read the Q from the initial optical flow of the N-1th layer image of the first image pyramid st The initial optical flow value C1; wherein, 1≤s≤L2, 1≤t≤W2, L2, W2 are the length value and width value of the N-1th layer image of the first image pyramid respectively; Determine the image block A in the N-1th layer image of the first image pyramid st , wherein the A st For the Q st The image block is the center pixel; Determine the search window Z in the N-1th layer image of the second image pyramid st , wherein the Z st For the A st The corresponding search window; According to the Q st The initial optical flow value C1 determines the Z st The image block G in st , wherein the G st For the A st The corresponding image block; Configure the A st In the Z st The search step length H in , where H is an integer greater than 1; According to the H, in the Z st Search for matching image block Y in st , wherein the Y st For the A st The image blocks with the highest similarity between them; Determine the A st With the Y st Is the similarity X1 between them greater than that of A st With the G st The similarity between them is X2; If X1>X2, then determine the initial optical flow value C2 in the initial optical flow of the N-1th layer of the first image pyramid, and determine C2 as the Q st The fine-tuned optical flow value of st The initial optical flow value corresponding to the central pixel of ; If X1≤X2, then the Q st The initial optical flow value C1 is determined as the Q st Fine-tuned optical flow value; wherein all the Q st The refined optical flow values ​​constitute the refined optical flow of the N-1th layer image of the first image pyramid.

6. The method according to claim 1, characterized in that The denoising process is performed on the fine-tuned optical flow of the first layer image of the first image pyramid to obtain the target optical flow of the target frame, including: The fine-tuned optical flow of the first layer image of the first image pyramid and the first layer image of the second image pyramid are input into the re-parameterized target denoising model for denoising, and the target optical flow of the target frame is output; wherein the target denoising model includes a type of convolution kernel and the size of the convolution kernel is greater than a preset size value.

7. The method according to claim 6, characterized in that Before the step of inputting the fine-tuned optical flow of the first layer image of the first image pyramid and the first layer image of the second image pyramid into the re-parameterized target denoising model for denoising and outputting the target optical flow of the target frame, the method further includes: Acquire a training set; wherein the training set includes multiple groups of sample data, each group of sample data includes: a first sample image, a second sample image, a noisy optical flow of the first sample image, and a true optical flow of the first sample image, and the second sample image is a reference image of the first sample image; Constructing a backbone network; wherein the backbone network includes at least two groups of convolution kernels, each group of convolution kernels includes a plurality of convolution kernels of different sizes connected in parallel; Inputting the noisy optical flows of the second sample image and the first sample image in the sample data into the backbone network for denoising, and outputting a predicted image and a denoised optical flow of the predicted image; calculating a loss value according to the first sample image, the real optical flow of the first sample image, the predicted image, and the denoised optical flow of the predicted image, and updating the network parameters of the backbone network according to the loss value; repeatedly executing the above training process using different groups of the sample data until the model converges to obtain an initial denoising model; The initial denoising model is re-parameterized to generate the target denoising model.

8. An optical flow calculation device, characterized in that: The device comprises: A construction module, used to construct a first image pyramid of a target frame and a second image pyramid of a reference frame; wherein the number of layers of the first image pyramid and the number of layers of the second image pyramid are both N, and N is an integer greater than 1; a configuration module, configured to configure an initial optical flow of an N-th layer image of the first image pyramid; wherein the initial optical flow of the N-th layer image of the first image pyramid comprises: a set of offsets for each pixel in the N-th layer image of the first image pyramid, each set of offsets comprising: at least two different random offsets, each of the random offsets being used to indicate a pixel position in the N-th layer image of the second image pyramid; a first processing module, configured to perform block matching processing on the Nth layer image of the first image pyramid and the Nth layer image of the second image pyramid according to the initial optical flow of the Nth layer image of the first image pyramid, determine the fine-tuning optical flow of the Nth layer image of the first image pyramid according to the block matching processing result, up-sample the fine-tuning optical flow of the Nth layer image of the first image pyramid, and determine the up-sampling result as the initial optical flow of the N-1th layer image of the first image pyramid; wherein the fine-tuning optical flow of the Nth layer image of the first image pyramid includes: the fine-tuning optical flow value of each pixel point in the Nth layer image of the first image pyramid; a second processing module, configured to calculate, based on the initial optical flow of the N-1th layer image of the first image pyramid, the refined optical flow of each layer image below the Nth layer image of the first image pyramid layer by layer, until the refined optical flow of the first layer image of the first image pyramid is obtained; The denoising module is used to perform denoising on the fine-tuned optical flow of the first layer image of the first image pyramid to obtain the target optical flow of the target frame.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the optical flow calculation method according to any one of claims 1 to 7 are implemented.

10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the optical flow calculation method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Effective estimation method for large displacement optical flows

    CN105809712A

  • Motion area detection method based on pyramid optical flow method

    CN115564804A

  • WasSAR image offset estimation method and device for three-dimensional information extraction

    CN119205919A

  • Method and apparatus for generating video intermediate frame

    US20230403371A1