Optical Flow Estimation Method, Computer Program Product, Storage Medium and Electronic Device
By combining the images and gyroscope data collected by the camera, and gyroscope data, and fusion of the gyroscope domain and temporary optical flow, the existing optical flow estimation method has solved the problem of low accuracy in special scenarios, achieving higher optical flow estimation accuracy.
Patent Information
- Application Number
- CN202111164952.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-09-30
AI Technical Summary
The existing optical flow estimation methods are difficult to meet the rich texture and similar lighting conditions of the image content in rainy days, foggy days, nights and other scenarios, resulting in low optical flow accuracy.
By acquiring the first and second images acquired by the same camera at different moments, as well as the first and second gyroscope data and the second gyroscope data during image acquisition, the gyroscope domain is calculated and fused with the temporary optical flow to improve the optical flow estimation accuracy.
This method can effectively respond to the challenge of optical flow estimation in special scenarios and significantly improve the accuracy of optical flow estimation, especially in the case of blurred background.
Smart Images

Figure CN114119678B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to an optical flow estimation method, a computer program product, a storage medium, and an electronic device. Background Art
[0002] Optical flow estimation is a fundamental and important computer vision task, and has been widely applied to applications such as object tracking, visual mapping, and image alignment. Existing optical flow estimation methods rely heavily on image content, usually requiring the images used for optical flow estimation to contain rich texture information and similar illumination conditions. However, for images acquired in scenarios such as rainy days, foggy days, and nights, the above requirements are often difficult to meet, resulting in low-precision optical flow estimated by existing methods for these scenarios. Summary of the Invention
[0003] The purpose of the embodiments of the present application is to provide an optical flow estimation method, a computer program product, a storage medium, and an electronic device to improve the above technical problems.
[0004] To achieve the above purpose, the present application provides the following technical solutions:
[0005] In a first aspect, the embodiments of the present application provide an optical flow estimation method, including: obtaining a first image, a second image, first gyroscope data, and second gyroscope data; wherein, the first image and the second image are images acquired by the same camera at different times, the first gyroscope data is the data acquired by the gyroscope during the acquisition of the first image, and the second gyroscope data is the data acquired by the gyroscope during the acquisition of the second image; calculating a gyro domain according to the first gyroscope data and the second gyroscope data, where the gyro domain is a two-dimensional motion field between the first image and the second image; estimating an instantaneous optical flow between the first image and the second image according to the first image and the second image, and fusing the instantaneous optical flow and the gyro domain to obtain the optical flow between the first image and the second image.
[0006] The gyro domain in the above method is calculated according to gyroscope data. Since gyroscope data is not affected by image content and gyroscope data is acquired during image acquisition, the calculated gyro domain can effectively reflect the background motion in the first image and the second image, while the instantaneous optical flow estimated according to the first image and the second image can better reflect the motion of the foreground and moving objects in the first image and the second image. Therefore, fusing the two can significantly improve the accuracy of optical flow estimation.
[0007] In particular, for images collected in scenarios such as rainy days, foggy days, and nights, the background is often blurred, and it is difficult to process using optical flow estimation methods based on image content. However, in the above method, gyroscope data is used for good estimation. Therefore, this method can effectively cope with the challenges brought by special scenarios.
[0008] In one implementation of the first aspect, the first gyroscope data is the data collected by the gyroscope during the exposure of the first image, and the second gyroscope data is the data collected by the gyroscope during the exposure of the second image; wherein, the camera capturing an image includes two stages: exposure and post-processing.
[0009] Since the image has been generated during the exposure stage, and the post-processing stage is only optimizing the quality of the generated image, there is no corresponding relationship between the gyroscope data and the image in the post-processing stage. Therefore, when calculating the gyro domain, only the gyroscope data during the exposure stage can be used to improve the calculation accuracy.
[0010] In one implementation of the first aspect, calculating the gyro domain according to the first gyroscope data and the second gyroscope data includes: calculating a rotation matrix between the first image and the second image according to the first gyroscope data and the second gyroscope data; calculating a homography matrix corresponding to the rotation matrix according to the rotation matrix and the internal parameters of the camera; calculating the gyro domain according to the homography matrix.
[0011] According to the gyroscope data, a rotation matrix in three-dimensional space can be calculated. Then, according to the internal parameters of the camera, the rotation matrix can be transformed into a homography matrix in two-dimensional space. Applying the homography matrix to the pixel coordinates can calculate the gyro domain. From the perspective of three-dimensional space, the gyro domain represents the rotation of the camera. From the perspective of two-dimensional space, the gyro domain represents the background movement in the image.
[0012] In an implementation of the first aspect, the camera adopts a rolling shutter. Calculating the rotation matrix between the first image and the second image based on the first gyroscope data and the second gyroscope data includes: calculating n rotation matrices between the first image and the second image according to n sets of data included in the first gyroscope data and n sets of data included in the second gyroscope data; where n is an integer greater than 1, each set of data is collected at different times, and the i-th rotation matrix is the rotation matrix between the i-th part of the first image in the exposure order and the i-th part of the second image in the exposure order, and i is any integer from 1 to n; calculating the homography matrix corresponding to the rotation matrix based on the rotation matrix and the internal parameters of the camera includes: calculating n homography matrices corresponding to the n rotation matrices according to the n rotation matrices and the internal parameters of the camera; calculating the gyro domain based on the homography matrix includes: calculating n partial gyro domains according to the n homography matrices and splicing the n partial gyro domains into the gyro domain; where the i-th partial gyro domain is the two-dimensional motion field between the i-th part of the first image in the exposure order and the i-th part of the second image in the exposure order, and i is any integer from 1 to n.
[0013] In an implementation of the first aspect, calculating n rotation matrices between the first image and the second image according to n sets of data included in the first gyroscope data and n sets of data included in the second gyroscope data includes: calculating n corresponding temporary rotation matrices M 1 ~M n according to the n sets of data included in the first gyroscope data, and calculating n corresponding temporary rotation matrices M n+1 ~M 2n according to the n sets of data included in the second gyroscope data; i traverses the integers from 1 to n, and calculates the i-th rotation matrix between the first image and the second image according to n + 1 temporary rotation matrices M i ~M n+i . After the traversal is completed, the n rotation matrices are obtained.
[0014] The above two implementation manners give possible calculation methods of the gyro domain when the camera adopts a rolling shutter. Among them, n rotation matrices (n > 1) are calculated in total, which is an effective approximation of the image generation method under the rolling shutter.
[0015] In an implementation of the first aspect, the camera uses a global shutter. Calculating the rotation matrix between the first image and the second image based on the first gyroscope data and the second gyroscope data includes: calculating a rotation matrix between the overall first image and the overall second image according to n sets of data included in the first gyroscope data and n sets of data included in the second gyroscope data; where n is an integer greater than 1, and each set of data is collected at different times.
[0016] In an implementation of the first aspect, calculating a rotation matrix between the overall first image and the overall second image according to n sets of data included in the first gyroscope data and n sets of data included in the second gyroscope data includes: calculating n corresponding temporary rotation matrices M 1 ~M n according to the n sets of data included in the first gyroscope data, and calculating n corresponding temporary rotation matrices M n+1 ~M 2n according to the n sets of data included in the second gyroscope data; calculating the rotation matrix according to the 2n temporary rotation matrices M 1 ~M 2n .
[0017] The above two implementations give possible calculation methods in the gyroscope domain when the camera uses a global shutter. Among them, only one rotation matrix is calculated, which conforms to the image generation method under the global shutter.
[0018] In an implementation of the first aspect, estimating the temporary optical flow between the first image and the second image and fusing the temporary optical flow with the gyroscope domain to obtain the optical flow between the first image and the second image includes: using a neural network model to estimate the temporary optical flow according to the first image and the second image; fusing the temporary optical flow with the gyroscope domain to obtain the optical flow between the first image and the second image.
[0019] In the above implementation, a neural network model is used to estimate the temporary optical flow. Since the gyroscope domain has already well estimated the background optical flow (the gyroscope domain itself can also be regarded as an optical flow estimated from gyroscope data), the neural network model can be more focused on estimating the optical flow of the foreground and moving objects, which is equivalent to narrowing the range of optical flow estimation that needs to be performed. Therefore, it is beneficial to improve the optical flow estimation accuracy of the model.
[0020] In an implementation of the first aspect, the fusion of the temporary optical flow and the gyroscope domain includes: using the neural network model to fuse the temporary optical flow and the gyroscope domain; the neural network model includes m consecutively connected optical flow estimation modules, where m is an integer greater than 1, and the k-th optical flow estimation module performs the following steps: extracting the k-th level feature of the first image according to the (k - 1)-th level feature of the first image, and extracting the k-th level feature of the second image according to the (k - 1)-th level feature of the second image, and downsampling the (k - 1)-th level feature when extracting the k-th level feature; fusing the k-th level gyroscope domain and the (k + 1)-th level temporary optical flow to obtain the k-th level fused optical flow, where the k-th level gyroscope domain is obtained by downsampling the (k - 1)-th level gyroscope domain; calculating the k-th level temporary optical flow according to the k-th level feature of the first image, the k-th level feature of the second image, and the k-th level fused optical flow, and upsampling the k-th level fused optical flow when calculating the k-th level temporary optical flow; where k is any integer from 1 to m, the 0-th level feature of the first image is the first image, the 0-th level feature of the second image is the second image, the 0-th level gyroscope domain is the gyroscope domain, the m-th level gyroscope domain is 0, the (m + 1)-th level temporary optical flow is 0, and the optical flow between the first image and the second image is obtained from the 1st level temporary optical flow.
[0021] In the above implementation, m consecutively connected optical flow estimation modules are used to perform feature extraction, optical flow fusion, and optical flow estimation at multiple scales. Each optical flow estimation module further optimizes the optical flow output by the previous optical flow estimation module, thereby achieving coarse-to-fine optical flow estimation, which is beneficial to improving the optical flow estimation result.
[0022] In an implementation of the first aspect, the fusion of the k-th level gyro domain and the (k + 1)-th level temporary optical flow to obtain the k-th level fused optical flow includes one of the following three methods: Using the first optical flow fusion unit in the k-th optical flow estimation module to fuse the k-th level gyro domain and the (k + 1)-th level temporary optical flow to obtain the k-th level fused optical flow, where the first optical flow fusion unit includes at least one convolutional layer; Using the weight prediction unit in the k-th optical flow estimation module to predict the k-th level weight map according to the k-th level features of the first image and the k-th level features of the second image, and using the k-th level weight map to perform weighted fusion on the k-th level gyro domain and the (k + 1)-th level temporary optical flow to obtain the k-th level fused optical flow; where the pixel value in the k-th level weight map represents the fusion weight, and the weight prediction unit includes at least one convolutional layer; Using the first optical flow fusion unit in the k-th optical flow estimation module to fuse the k-th level gyro domain and the (k + 1)-th level temporary optical flow to obtain the k-th level temporary fused optical flow, using the weight prediction unit in the k-th optical flow estimation module to predict the k-th level weight map according to the k-th level features of the first image and the k-th level features of the second image, and using the k-th level weight map to perform weighted fusion on the k-th level gyro domain and the k-th level temporary fused optical flow to obtain the k-th level fused optical flow; where both the first optical flow fusion unit and the weight prediction unit include at least one convolutional layer, and the pixel value in the k-th level weight map represents the fusion weight.
[0023] Among the above three optical flow fusion methods, the first one belongs to implicit fusion, that is, using a network (the first optical flow fusion unit) to learn how to fuse the gyro domain and the temporary optical flow; the second one belongs to explicit fusion, that is, fusing the gyro domain and the temporary optical flow according to a clear indication information (weight map); the third one uses both implicit fusion and explicit fusion, which is beneficial to obtaining a more accurate optical flow estimation result.
[0024] In a second aspect, an embodiment of the present application provides an optical flow estimation device, including: a data acquisition component for acquiring a first image, a second image, first gyroscope data, and second gyroscope data; where the first image and the second image are images acquired by the same camera at different times, the first gyroscope data is the data acquired by the gyroscope during the acquisition of the first image, and the second gyroscope data is the data acquired by the gyroscope during the acquisition of the second image; a gyro domain calculation component for calculating a gyro domain according to the first gyroscope data and the second gyroscope data, where the gyro domain is a two-dimensional motion field between the first image and the second image; an optical flow estimation component for estimating the temporary optical flow between the first image and the second image according to the first image and the second image, and fusing the temporary optical flow and the gyro domain to obtain the optical flow between the first image and the second image.
[0025] In a third aspect, an embodiment of the present application provides a computer program product, including computer program instructions. When the computer program instructions are read and run by a processor, the method provided in the first aspect or any possible implementation manner of the first aspect is executed.
[0026] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are read and run by a processor, the method provided in the first aspect or any possible implementation manner of the first aspect is executed.
[0027] In a fifth aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor. Computer program instructions are stored in the memory. When the computer program instructions are read and run by the processor, the method provided in the first aspect or any possible implementation manner of the first aspect is executed. Description of the Drawings
[0028] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative efforts.
[0029] Figure 1 Shows the flow of an optical flow estimation method provided by an embodiment of the present application;
[0030] Figure 2 Shows a data acquisition method provided by an embodiment of the present application;
[0031] Figure 3 Shows a calculation method in the gyro domain when using a rolling shutter;
[0032] Figure 4 Shows Figure 3 a calculation method of the rotation matrix in
[0033] Figure 5 Shows a calculation method in the gyro domain when using a global shutter;
[0034] Figure 6 Shows Figure 5 a calculation method of the rotation matrix in
[0035] Figure 7 Shows the structure of a neural network model provided by an embodiment of the present application;
[0036] Figure 8 Shows Figure 7 The structure of the optical flow fusion submodule in the neural network model;
[0037] Figure 9 The structure of an optical flow estimation device provided by an embodiment of the present application is shown;
[0038] Figure 10 The structure of an electronic device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0039] In recent years, research on computer vision, deep learning, machine learning, image processing, image recognition and other technologies based on artificial intelligence has made important progress. Artificial Intelligence (AI) is an emerging science and technology that studies and develops theories, methods, technologies and application systems for simulating and extending human intelligence. Artificial intelligence is a comprehensive discipline involving many types of technologies such as chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, neural networks, etc. Computer vision, as an important branch of artificial intelligence, specifically allows machines to recognize the world. Computer vision technology usually includes face recognition, liveness detection, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, target detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, text recognition, video processing, video content recognition, behavior recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous positioning and map construction, computational photography, robot navigation and positioning and other technologies. With the research and advancement of artificial intelligence technology, this technology has been applied in many fields, such as security, urban management, traffic management, building management, park management, facial access, facial attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile phone imaging, cloud services, smart homes, wearable devices, unmanned driving, automatic driving, smart medical care, facial payment, facial unlocking, fingerprint unlocking, identity verification, smart screens, smart TVs, cameras, mobile Internet, live streaming, beauty, makeup, medical beauty, and smart temperature measurement.
[0040] The optical flow estimation method in the embodiment of the present application generally belongs to the category of computer vision technology. The method uses gyroscope data to improve the accuracy of optical flow estimation, especially in scenes that are difficult to handle with existing methods such as rainy days, foggy days, and nights.
[0041] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.
[0042] The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element qualified by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.
[0043] The terms "first", "second", etc. are used only to distinguish one entity or operation from another entity or operation, and should not be construed as indicating or implying relative importance, nor should it be construed as requiring or implying any actual relationship or order between these entities or operations.
[0044] Figure 1 The flowchart of an optical flow estimation method provided by an embodiment of the present application is shown. The method can, but is not limited to, be executed by Figure 10 the electronic device shown. Referring to Figure 1 , the method includes:
[0045] S110: Obtain a first image, a second image, first gyroscope data, and second gyroscope data.
[0046] Step S110 is a step of obtaining the data required for estimating the optical flow. The data required to be obtained includes two categories. One category is the image data (first image, second image) collected by the camera, and the other category is the gyroscope data (first gyroscope data, second gyroscope data) collected by the gyroscope.
[0047] It should be noted here that Figure 1 the camera and the gyroscope in the method should be installed on the same device, which may be referred to as device A, but the device executing the method is not necessarily device A and may also be device B. For example, image data and gyroscope data are obtained in real time from the camera and gyroscope of a mobile phone (device A), and optical flow estimation is performed locally on device A; or for another example, after the camera and gyroscope of the mobile phone (device A) generate image data and gyroscope data, these data are exported to a PC (device B), and optical flow estimation is performed on device B, and so on.
[0048] The first image and the second image in step S110 are images collected by the same camera at different times. The optical flow to be estimated is the optical flow between the first image and the second image. Of course, according to requirements, the optical flow from the first image to the second image or the optical flow from the second image to the first image can be estimated.
[0049] The optical flow between the so-called first image and second image is in the form of a two-dimensional vector field. Each vector in this vector field (including two components in the x and y directions) reflects the gray-scale change rate at the corresponding position in the image. Generally, the change in gray-scale in the image is caused by the movement of pixels on the image plane. Therefore, to a certain extent, this optical flow can reflect the pixel movement between the first image and the second image.
[0050] The image captured by the camera should be understood as the image generated by the camera's image sensor at any time, rather than just the image generated after the user issues a capture instruction. For example, the video frames generated after the user clicks the "Start Recording" button in the camera APP belong to the images captured by the camera; another example is that the images seen by the user on the preview interface of the camera APP also belong to the images captured by the camera.
[0051] The first image and the second image are conceptually universal. For example, a large number of video frames are included in the video captured by the camera. Any two adjacent frames can be selected, with one frame as the first image and the other as the second image. However, no matter which two frames are selected, the method for estimating the optical flow is similar.
[0052] When elaborating later, mainly the case where the first image and the second image meet the following conditions is taken as an example, and other cases can be analyzed similarly:
[0053] (1) The content of the first image and the second image is for the same scene
[0054] If the content of the first image and the second image is not for the same scene, then it doesn't make much sense to estimate the optical flow between them. If the acquisition time interval between the first image and the second image is not large (for example, two adjacent frames in a video, two consecutive photos taken, etc.), it is easy to meet condition (1).
[0055] (2) The first image and the second image have the same size
[0056] Since the first image and the second image are captured by the same camera, condition (2) is easy to meet.
[0057] The first gyroscope data in step S110 is the data collected by the gyroscope during the acquisition of the first image, and the second gyroscope data is the data collected by the gyroscope during the acquisition of the second image. Thus, the first image and the first gyroscope data are corresponding, and the second image and the second gyroscope data are corresponding.
[0058] The acquisition of each image does not occur instantaneously and takes a certain period of time, while gyroscope data can be considered to be acquired instantaneously. Each time a set of gyroscope data is acquired, the gyroscope data can be acquired at a fixed frequency. In the embodiments of the present application, during the acquisition of one image, at least one set of gyroscope data can be acquired. Each image generates corresponding timestamps during acquisition, including the start timestamp and the end timestamp of the acquisition, and each set of gyroscope data also corresponds to a timestamp. Thus, the correspondence between the image and the gyroscope data can be determined based on the mutual relationship between the timestamps. The time on these timestamps can be the kernel time of the system (referring to the operating system of the device where the camera and gyroscope are located).
[0059] Referring to Figure 2 , I a represents the first image, and the start timestamp of its acquisition is t a S, and the end timestamp of the acquisition is t a E. During the period from t a S to t a E, the gyroscope acquired a total of 14 sets of data, denoted as g a (1) to g a (14); I b represents the second image, and the start timestamp of its acquisition is t b S, and the end timestamp of the acquisition is t b E. During the period from t b S to t b E, the gyroscope acquired a total of 14 sets of data, denoted as g b (1) to g b (14). Among them, t a E and t b S may also be implemented as the same timestamp. In this case, each image only corresponds to one timestamp representing the start of the acquisition.
[0060] However, it should be noted that some implementation methods will use all the data acquired by the gyroscope during image acquisition for optical flow estimation. For Figure 2 example, g a (1) to g a (14) are used as the first gyroscope data, and g b (1) to g b (14) are used as the second gyroscope data. While some other implementation methods will only use part of the data acquired by the gyroscope during image acquisition for optical flow estimation. For Figure 2 example, g a (1) to g a (6) are used as the first gyroscope data, and g b (1) to g b(6) As the second gyroscope data, up to g a (7) to g a (14), and g b (7) to g b (14) is discarded. Here, "discarded" can either mean deletion or simply not participating in subsequent calculations.
[0061] The following focuses on the latter case. Figure 2 In, the camera captures I a during the process t a S to t a E is divided into two stages, namely t a S to t a M and t a M to t a E. Among them, the former stage is the exposure stage, and the latter stage is the post-processing stage. The two stages are separated by the exposure end time t a M. The original image of I a is generated by the image sensor of the camera during the exposure stage. However, the image quality at this time is poor and still needs to be post-processed by the image signal processor (ISP) of the camera to obtain a better-quality I a .
[0062] It can be seen that the post-processing stage is only optimizing the quality of the original image of I a . Therefore, the gyroscope data g a (7) to g a (14) collected in this stage has no corresponding relationship with I a . If these data are used for gyroscope domain calculations in step S120, it may instead lead to a decrease in calculation accuracy. Therefore, it is better to discard them in step S110 and only retain the gyroscope data g a (1) to g a (6) as the first gyroscope data.
[0063] It should be noted that to obtain g a (1) to g a (6), one way is as described above, first obtain g a (1) to g a (14), and then filter out the g a (7) to g a (14) among them. Another way is to directly obtain only g a (1) to g a (6) because the duration of the exposure stage is fixed, and the number of groups of the corresponding gyroscope data is also fixed, for example, 6 groups. Thus, with t aTaking 6 groups of gyroscope data with S as the starting point of time, g can be obtained. a (1) to g a (6).
[0064] For the case of I b it can also be analyzed similarly to I a and will not be repeated here. In the following description, it is mainly taken as an example that the obtained gyroscope data only includes the data during the exposure stage.
[0065] The gyroscope data used in the embodiments of this application at least includes the angular velocity information in three directions. Of course, it does not exclude that the gyroscope can also collect other information. In addition, the timestamp corresponding to each group of gyroscope data can also be regarded as a part of the gyroscope data.
[0066] Next, taking a mobile phone as an example, a possible data acquisition process is introduced. For other electronic devices, the situation is similar:
[0067] First, select a mobile phone without optical image stabilization (or the function can be turned off through technical means) and supporting root. Gyroscope data can always reflect the true movement trajectory of the device (referring to the device where the camera and gyroscope are located), but the optical image stabilization function will cause the images obtained from the camera not to reflect the true movement trajectory of the device, which is not conducive to optical flow estimation. Therefore, this function needs to be blocked. And supporting the root function is to obtain the highest permissions of the mobile phone in order to normally access the data collected by the sensors (referring to the camera and gyroscope).
[0068] Then, install and run a customized library in the mobile phone. The code of this customized library at least implements the following functions:
[0069] First, obtain the required data (images and gyroscope data) from the system's Hardware Abstraction Layer (HAL for short). Obtaining data directly from the HAL layer has higher accuracy compared to obtaining data from some upper-layer applications.
[0070] Second, export the obtained data through a data transmission protocol. Depending on the location for optical flow estimation, it can be exported to the local memory of the mobile phone or to an external device, etc.
[0071] Note that the logic of obtaining images and their corresponding gyroscope data using timestamp information introduced above can be implemented in the customized library.
[0072] Finally, preprocess the exported data, including operations such as formatting and deleting useless data, to obtain data convenient for optical flow estimation. This step may or may not be executed by a custom library. For example, if the data has been exported to an external device, the preprocessing can be performed on the external device.
[0073] S120: Calculate the gyro domain based on the first gyroscope data and the second gyroscope data.
[0074] The gyro domain is a two-dimensional motion field between the first image and the second image, which can to a certain extent reflect the pixel motion between the first image and the second image. The size of the gyro domain is the same as that of the first image and the second image, and each pixel in it is a motion vector (including two components in the x and y directions). It is not difficult to see that the gyro domain can also be regarded as an optical flow calculated based on gyroscope data (in the prior art, optical flow is generally calculated based on images).
[0075] For the optical flow estimation problem, the image content can generally be divided into three parts: background, foreground, and moving objects.
[0076] The background can be considered as the area composed of objects that are relatively far from the camera in the image. The background motion (referring to the motion of pixels in the background) is consistent with the motion of the camera. When elaborating on step S110, it has been mentioned that the gyroscope data can reflect the true motion trajectory of the device, and naturally can also reflect the motion trajectory of the camera installed on the device. And according to step S110, the gyroscope data are all collected during image acquisition. Therefore, the calculated gyro domain can effectively reflect the background motion in the first image and the second image. The calculation of the gyro domain is not affected by the image content, but depends on the special hardware of the gyroscope. Therefore, its estimation of the background motion can reach a very high accuracy.
[0077] The foreground can be considered as the area composed of objects that are relatively close to the camera in the image. Although theoretically the foreground motion (referring to the motion of pixels in the foreground) is also consistent with the motion of the camera, the gyroscope data cannot describe the foreground motion very well. The reason is that the gyroscope can only record the rotation information of the device and cannot record the translation information of the device. Since the first image and the second image are collected at different times, during this period, the device is likely to have a certain translation, which will cause a parallax between the first image and the second image. The parallax can be basically ignored in the background of the image, but is more obvious in the foreground of the image.
[0078] However, according to statistics, when mobile devices such as mobile phones perform image acquisition, 90% of the generated motion is rotation and only 10% is translation. Therefore, in addition to being able to well describe the background motion, the gyroscope domain can also describe the foreground motion to a certain extent, and the remaining part can be left for step S130 to solve.
[0079] A moving object refers to an object that can move autonomously in an image, and its motion trajectory has no necessary relationship with the motion trajectory of the camera. Naturally, the gyroscope domain cannot effectively describe it.
[0080] Regarding how to calculate the gyroscope domain, it will be elaborated in detail later and will not be expanded here for the time being.
[0081] S130: Estimate the temporary optical flow between the first image and the second image, and fuse the temporary optical flow with the gyroscope domain to obtain the optical flow between the first image and the second image.
[0082] The temporary optical flow is an intermediate optical flow estimation result. As pointed out when elaborating step S120, the gyroscope domain can also be regarded as an optical flow. Therefore, there is no obstacle to the fusion of the gyroscope domain and the temporary optical flow. What is obtained after fusion is the final optical flow estimation result, that is, the optical flow between the first image and the second image.
[0083] For the estimation of the temporary optical flow, a deep learning-based method can be used, or a traditional optical flow estimation method can be used. In the following, the method of using a neural network model for optical flow estimation (which belongs to the deep learning-based method) will be mainly introduced.
[0084] For optical flow fusion, there are also various methods, including: directly fusing using a neural network model, which is called implicit fusion; fusing using a weight map calculated in a certain way (where the pixel values represent the fusion weights), which is called explicit fusion; and a method that combines implicit fusion and explicit fusion. Specific examples of these three fusion methods will be given later.
[0085] Regarding step S130, the following issues also need to be explained:
[0086] (1) Estimating the temporary optical flow will at least use the first image and the second image, and in some implementation methods, the gyroscope domain may also be used.
[0087] (2) The process of estimating the temporary optical flow and optical flow fusion does not necessarily only be executed once, and may need to be executed multiple times to obtain the final optical flow estimation result.
[0088] (3) The size of the finally obtained optical flow is the same as that of the first image and the second image, but the size of the temporary optical flow is not necessarily the same as that of the first image and the second image, that is, estimating the temporary optical flow may be performed at different scales.
[0089] (4) It is possible that the two stages of estimating the temporary optical flow and fusing the optical flow are unified into a neural network model, that is, the model will not only learn how to estimate the temporary optical flow, but also learn how to fuse the temporary optical flow with the gyroscope domain (that is, the entire step S130 is completed by the neural network model). Of course, it is not excluded that only the neural network is used to estimate the temporary optical flow, while the optical flow fusion does not use the neural network, or the neural network is not used to estimate the temporary optical flow, while the optical flow fusion uses the neural network.
[0090] The corresponding content for the above questions (1) to (4) can be found in the examples in the following text, and will not be elaborated here for the time being.
[0091] When elaborating on step S120, it was mentioned that the gyroscope domain can effectively reflect the background motion in the first image and the second image, but there are still deficiencies in the description of the foreground motion and moving objects. The temporary optical flow estimated based on the first image and the second image can better reflect the motion of the foreground and moving objects in the first image and the second image. Taking the case of using a neural network model to estimate the temporary optical flow as an example, since the gyroscope domain has already well estimated the background optical flow (as mentioned before, the gyroscope domain itself can also be regarded as an optical flow), the neural network model can focus more on estimating the optical flow of the foreground and moving objects, which is equivalent to narrowing the range of optical flow estimation that needs to be performed. Therefore, it is beneficial to improve the estimation accuracy of the optical flow of the foreground and moving objects by the model. Thus, fusing the temporary optical flow and the gyroscope domain in step S130 can make up for each other's strengths and weaknesses, achieving better estimation effects in the background, foreground, and moving object regions, and significantly improving the optical flow estimation accuracy.
[0092] In particular, for images collected in scenarios such as rainy days, foggy days, and nights, the background is often blurred, and it is difficult for the optical flow estimation method based on image content to handle. In the above optical flow estimation method, the gyroscope data is used to well estimate the background optical flow, and the optical flow estimation of the foreground and moving objects is also improved. Therefore, this method can effectively handle the images collected in these special scenarios.
[0093] Regarding the optical flow finally estimated in step S130, its uses are not limited. For example, it can be used to align the first image to the second image (or vice versa), and can be used to track the target in the video sequence where the first image and the second image are located, and so on. Since the estimation accuracy of the optical flow has been improved, the execution effects of these optical flow-based tasks will naturally be improved accordingly.
[0094] Next, based on the above embodiments, the possible calculation methods of the gyroscope domain in step S120 will be continued to be introduced:
[0095] In some implementation manners, the gyroscope domain can be calculated according to the following steps:
[0096] Step A: Calculate the rotation matrix between the first image and the second image based on the first gyroscope data and the second gyroscope data.
[0097] Step B: Calculate the homography matrix corresponding to the rotation matrix according to the rotation matrix and the internal parameters of the camera.
[0098] Step C: Calculate the gyro domain according to the homography matrix.
[0099] Among them, the rotation matrix in the three-dimensional space can be calculated based on the gyroscope data, and then the rotation matrix can be converted into the homography matrix in the two-dimensional space according to the internal parameters of the camera. Then, by applying the homography matrix to the pixel coordinates, the gyro domain can be calculated. From the calculation process of the gyro domain, it can be seen that in the three-dimensional space, the gyro domain represents the rotation of the camera, and in the two-dimensional space, the gyro domain represents the background movement in the image.
[0100] It should be noted that the homography matrix calculated according to this method only contains rotation information and does not contain translation information (because the gyroscope cannot provide translation information). Therefore, it is not excluded that in some alternative solutions, translation information is obtained through other sensors and incorporated into the calculation of the homography matrix.
[0101] The shutters currently used in cameras are mainly divided into two types. One is the rolling shutter, and the other is the global shutter. For a camera using a rolling shutter, the image is generally exposed row by row (or column by column), that is, the generation time of each row of pixels in the image is different, which is also called the rolling shutter effect. For example, when taking a photo on a high-speed train, the telegraph poles in the photo appear tilted, which is the influence of the rolling shutter effect. For a camera using a global shutter, the image is exposed as a whole, that is, the generation time of each row of pixels in the image is the same. Currently, the global shutter is higher in cost and technical implementation difficulty than the rolling shutter. Therefore, the cameras of the vast majority of devices use the rolling shutter. For the cases where the camera uses a rolling shutter and a global shutter, there will be some differences in the implementation of the above Steps A to C. The following is a specific description:
[0102] Using a rolling shutter (Steps A1 to C1 are respectively an implementation method of Steps A to C)
[0103] Step A1: Calculate n rotation matrices between the first image and the second image according to the n groups of data included in the first gyroscope data and the n groups of data included in the second gyroscope data.
[0104] Among them, n is an integer greater than 1. For example, for Figure 2, n can be taken as 6. If i is any integer from 1 to n, the i-th rotation matrix among the n rotation matrices is: the rotation matrix between the i-th part of the first image in the exposure order and the i-th part of the second image in the exposure order. For example, for the case of n = 6, if the image size is 720×480 (width×height) with progressive exposure, the image can be equally divided into 6 parts in the row direction, each part having a size of 720×80, and numbered from 1 to 6 in the order from top to bottom.
[0105] Assume that the sizes of the first image and the second image are both W×H (width×height). Strictly speaking, due to the influence of the rolling shutter effect, there is a rotation matrix between each row of the first image and the corresponding row in the second image, that is, there are a total of W different rotation matrices. However, in most cases, W rotation matrices are not calculated. On the one hand, there are not so many sets of gyroscope data to support such calculations (n < W or even n << W). On the other hand, even if there is enough gyroscope data, the amount of computation required to calculate W rotation matrices is too large. Therefore, in step A1, only n rotation matrices are calculated. Although the calculation accuracy is reduced compared to W rotation matrices, it can still effectively reflect the image generation method when using a rolling shutter.
[0106] Refer to Figure 3 , the first image I a is divided into n parts I a (1) to I a (n) according to the exposure order, and the second image I b is also divided into n parts I b (1) to I b (n) according to the exposure order. The first gyroscope data g a includes a total of n groups, which are g a (1) to g a (n) respectively. The second gyroscope data g b also includes n groups, which are g b (1) to g b (n) respectively. According to g a (1) to g a (n) and g b (1) to g b (n), the rotation matrix R a between I b (1) and I 1 (1) can be calculated, the rotation matrix R a between I b (2) and I 2 (2),..., the rotation matrix R a between I b (n) and I n .
[0107] In some implementations, the n rotation matrices are calculated as follows:
[0108] Step A11: Calculate the corresponding n temporary rotation matrices M 1 ~M n based on the n sets of data included in the first gyroscope data, and calculate the corresponding n temporary rotation matrices M n+1 ~M 2n .
[0109] Referring to Figure 4 , without loss of generality, taking the gyroscope data g a (1) as an example, g a (1) includes the angular velocity information in three directions. Multiply the angular velocity information in these three directions by a sampling time interval (the time stamp t a (2) of g a (2) minus the time stamp t a (1) of g a (1)) to obtain the corresponding three rotation vectors. Substitute these three rotation vectors into the Rodriguez formula to calculate M 1 . It can be considered that M 1 describes the rotation information of the device during the period from t a (1) to t a (2). For the corresponding temporary rotation matrices M a (2) to g a (n), g b (1) to g b (n), namely M 2 ~M n , M n+1 ~M 2n , the calculation methods are similar and will not be elaborated repeatedly. It should be noted that for g a (n), if it is not the last set of gyroscope data corresponding to I a , for example, in Figure 2 , there is g a (6) followed by g a (7), but g a (7) is discarded, then the angular velocity information in g a (6) should be multiplied by t a (7)-t a (6), rather than t b (1)-t a (6).
[0110] Step A12: Let i iterate over the integers from 1 to n. According to the n + 1 temporary rotation matrices M i ~M n+i, calculate the i-th rotation matrix between the first image and the second image. After the traversal is completed, n rotation matrices are obtained.
[0111] For example, when i = 1, multiply the n + 1 temporary rotation matrices M 1 ~M n+1 to obtain the first rotation matrix R 1 between the first image and the second image. It should be noted that R 1 is not only related to g a (1), g b (1), but also needs to combine g a (1)~g a (n), g b (1) to accurately calculate. For the rotation matrices R 2 ~R n , the calculation method is similar and will not be repeated.
[0112] Figure 4 The right side of Figure 4 shows the calculation process described in step A12. At Figure 4 the leftmost side, each part of the image in the exposure order and the corresponding set of gyroscope data are written together.
[0113] In some alternative solutions, if n is relatively large, only a part of the gyroscope data can be selected for calculating the rotation matrix to save the amount of computation. For example, when n = 10, only the 1st, 3rd, 5th, 7th, and 9th groups of data can be selected to calculate 5 rotation matrices. At this time, the first image and the second image will only be divided into 5 parts.
[0114] Step B1: Calculate n homography matrices corresponding to the n rotation matrices according to the internal parameters of the camera.
[0115] Assume that the internal parameter matrix formed by the internal parameters of the camera is K. Then the homography matrix H 1 corresponding to R 1 = K × R 1 × K -1 , and H 1 can also be considered as the homography matrix between the i-th part of the first image in the exposure order and the i-th part of the second image in the exposure order. For the homography matrices H 2 ~H n corresponding to R 2 ~R n , the calculation method is similar and will not be repeated. The final calculation result is as Figure 3 shown.
[0116] Step C1: Calculate n partial gyrodomains according to the n homography matrices and splice the n partial gyrodomains into a gyrodomain.
[0117] Among them, the i-th partial gyroscopic domain is the two-dimensional motion field between the i-th part of the first image in the exposure order and the i-th part of the second image in the exposure order, where i is any integer from 1 to n. Due to the curtain rolling effect, each homography matrix can only correspond to the rotation information of a part of the image, so the obtained gyroscopic domain is only a part of the entire gyroscopic domain, thus it is called a partial gyroscopic domain. Finally, the complete gyroscopic domain can be obtained only after splicing the partial gyroscopic domains.
[0118] For example, H 1 The corresponding partial gyroscopic domain is G ab (1), and the calculation method of G ab (1) is as follows: For any coordinate (x, y) in the first part of the first image in the exposure order, multiply it by H 1 , and a new coordinate (x’, y’) will be obtained. Subtracting the original coordinate from this new coordinate can obtain the motion vector (u, v) = (x’ - x, y’ - y) at the coordinate (x, y). Traversing all the coordinates in the first part of the first image in the exposure order can obtain G ab (1). Note that in the calculation process, the image content of the first image is not used. For the partial gyroscopic domains G 2 ~H n corresponding to H ab (2)~G ab (n), the calculation methods are similar and will not be elaborated repeatedly. Finally, the obtained G ab is as Figure 3 shown.
[0119] Adopt a global shutter (Steps A2 to C2 are respectively one implementation manner of Steps A to C)
[0120] Step A2: Calculate a rotation matrix between the whole of the first image and the whole of the second image according to the n groups of data included in the first gyroscope data and the n groups of data included in the second gyroscope data.
[0121] Among them, n is an integer greater than 1. For example, for Figure 2 , n = 6 can be taken. The difference between Step A2 and Step A1 is that Step A1 needs to calculate n rotation matrices, while Step A2 only needs to calculate one rotation matrix R, because when adopting a global shutter, the rotation information of the entire image can be represented by one rotation, as Figure 5 shown.
[0122] In some implementation manners, the calculation method of this rotation matrix is as follows:
[0123] Step A21: Calculate the corresponding n temporary rotation matrices M according to the n groups of data included in the first gyroscope data1 to M n and calculating corresponding n temporary rotation matrices M according to n groups of data included in the second gyroscope data n+1 to M 2n .
[0124] Step A21 is the same as step A11 and will not be elaborated again.
[0125] Step A22: Calculating the rotation matrix between the whole of the first image and the whole of the second image according to 2n temporary rotation matrices M 1 to M 2n .
[0126] For example, multiplying 2n temporary rotation matrices M 1 to M 2n can obtain the rotation matrix R between the whole of the first image and the whole of the second image.
[0127] Figure 6 Illustrates the calculation process described in steps A11 and A12.
[0128] Step B2: Calculating the homography matrix corresponding to the rotation matrix according to the rotation matrix and the internal parameters of the camera.
[0129] Step B2 is similar to step B1. The difference is that only one rotation matrix is calculated in step A2, so only one corresponding homography matrix H needs to be calculated in step B2. H can also be considered as the homography matrix between the whole of the first image and the whole of the second image. The calculation result is as Figure 5 shown.
[0130] Step C2: Calculating the gyro domain according to the homography matrix.
[0131] Step C2 is similar to step C1. The difference is that only one homography matrix is calculated in step B2, so the complete gyro domain G ab can be directly calculated in step C2. The calculation result is as Figure 5 shown.
[0132] In some alternative solutions, if n is relatively large, only a part of the gyroscope data can also be selected for calculating the rotation matrix R to save the amount of computation. For example, when n = 10, the first, third, fifth, seventh, and ninth groups of data can also be selected to calculate the rotation matrix R.
[0133] Next, on the basis of the above embodiments, a method for implementing step S130 using a neural network model will be continued to be introduced:
[0134] Figure 7 Illustrates the structure of a neural network model provided by an embodiment of the present application. Refer toFigure 7 , when viewed horizontally, the neural network model includes m (m is an integer greater than 1, Figure 7 where m = 4) successively connected optical flow estimation modules, and each optical flow estimation module has the functions of feature extraction, optical flow fusion, and optical flow estimation; when viewed vertically, the neural network model includes an encoding network and a decoding network (the structure related to the left side and the gyro domain may or may not be part of the neural network model). The encoding network is mainly used to extract multi-scale features, and the decoding network is mainly used for multi-scale optical flow estimation and fusion. When elaborating below, it is mainly elaborated according to the optical flow estimation module.
[0135] Taking the k-th (k is any integer from 1 to m) optical flow estimation module as an example, it mainly includes three sub-modules:
[0136] Feature extraction sub-module ( Figure 7 the E module in ): used to extract the k-th level feature of the first image according to the (k - 1)-th level feature of the first image and, according to the (k - 1)-th level feature of the second image extract the k-th level feature of the second image
[0137] Among them, when the E module extracts it performs downsampling on , so, in Figure 7 it can be seen that the size of is smaller than and The situation is similar. The E module can be implemented as a convolutional module, which includes at least one convolutional layer, and of course may also contain other structures.
[0138] In particular,
[0139] Optical flow fusion sub-module ( Figure 7 the SGF module in ): used to fuse the k-th level gyro domain and the (k + 1)-th level temporary optical flow
[0140] Among them, is obtained by downsampling , the downsampling process of the gyro domain is as Figure 7As shown in the leftmost column, DOWN represents the downsampling module. The DOWN module can be part of the neural network model (i.e., perform downsampling during the operation of the neural network model) or an independent part (i.e., input the result obtained after prior downsampling into the neural network model).
[0141] In particular, (where 0 refers to a two-dimensional motion field of all zeros), is not shown in Figure 7 and (where 0 refers to an optical flow of all zeros), is not shown in Figure 7 At the same time, since and so Therefore, the SGF module in the m-th optical flow estimation module can be omitted, as shown in the optical flow estimation module 4 of Figure 7 .
[0142] It should be noted that when introducing the function of the SGF module above, it is not mentioned that its input includes and However, in Figure 7 , the input of the SGF module includes these two pieces of information. The reason is that in different implementation methods of the SGF module, whether these two pieces of information are to be used as input is optional. For details, see the introduction to the internal structure of the SGF module later.
[0143] Decoding sub-module (the D module in Figure 7 ): Used to calculate the k-th level temporary optical flow based on the k-th level feature of the first image the k-th level feature of the second image and the k-th level fused optical flow
[0144] Among them, when the D module calculates , it upsamples . The upsampling can be implemented through structures such as deconvolution layers. Of course, the D module also contains other structures. For example, these structures can predict and the residual of and and then use to calculate
[0145] In particular, the optical flow V between the first image and the second image ab can be obtained from , for example, after upsampling according to requirements, as shown in the UP module of Figure 7 . If upsampling is not required, the UP module can also be omitted.
[0146] Briefly summarize the structure of the above neural network model. This model simultaneously implements optical flow estimation and optical flow fusion in step S130. The model uses m sequentially connected optical flow estimation modules to perform feature extraction, optical flow fusion, and optical flow estimation at multiple scales. Each optical flow estimation module further optimizes the optical flow output by the previous optical flow estimation module, thus realizing coarse-to-fine optical flow estimation, which is beneficial to improving the optical flow estimation result. It should be understood that Figure 7 Only the structure related to optical flow estimation in the neural network model is shown, and it does not exclude that the neural network model may also include other structures.
[0147] Figure 8 Shows the possible structure that the SGF module may adopt. Refer to Figure 8 , the SGF module may include the following three units:
[0148] The first optical flow fusion unit: used to fuse the k-th level gyro domain and the (k + 1)-th level temporary optical flow to obtain the k-th level temporary fusion optical flow
[0149] Among them, the first optical flow fusion unit can be implemented as a convolution module, which includes at least one convolutional layer. Of course, it may also include other structures.
[0150] The weight prediction unit: predicts the k-th level weight map according to the k-th level features of the first image and the k-th level features of the second image
[0151] Among them, each pixel value represents a fusion weight, which is used to fuse and vectors at the same position. For example, the fusion weight can take a value between [0, 255].
[0152] Optionally, before inputting into the weight prediction unit, it can also be first warped (for the case of estimating the optical flow from I to I a to I b If it is to estimate the optical flow from I b to I a the optical flow, then should be warped). The so-called warping is an operation to align to as shown in Figure 8 .
[0153] Optionally, in addition to and In addition to, the weight prediction unit may also include other inputs, such as and / or
[0154] The weight prediction unit can be implemented as a convolutional module, which includes at least one convolutional layer and may also contain other structures.
[0155] Second optical flow fusion unit: used to utilize the k-th level weight map to perform weighted fusion on the k-th level gyro domain and the k-th level temporary fusion optical flow to obtain the k-th level fused optical flow
[0156] wherein, if the pixel value in represents the corresponding fusion weight, then a possible fusion formula (the second optical flow fusion unit implements this formula) is as follows:
[0157]
[0158] The symbol ⊙ in the formula represents pixel-by-pixel multiplication between matrices, and in the formula has been normalized to ensure that each of its pixel values is within the interval [0, 1]. For convenience, still use to represent the normalized weight map. Intuitively understanding this formula, (before normalization) the darker part (corresponding to the background area in the image) fuses more because the optical flow in this area has been accurately given by the gyro domain, while the lighter part (corresponding to the foreground area or moving object in the image) fuses more because the optical flow in this area cannot be accurately given by the gyro domain and needs to rely more on the optical flow estimated from the image.
[0159] It can be understood that if the pixel value in represents the corresponding fusion weight, then the above formula needs to be adjusted accordingly.
[0160] Figure 8 There are also some variations in the SGF module in, such as:
[0161] Variation 1: Remove the weight prediction unit and the second optical flow fusion unit, that is, directly use the first optical flow fusion unit to fuse and to obtain and use as Output. This method is the implicit fusion mentioned above, that is, using a network (the first optical flow fusion unit) to learn how to fuse the gyroscope domain and the temporary optical flow.
[0162] It should be noted that since the weight prediction unit is removed, the input of the SGF module no longer includes and
[0163] Variant 2: Remove the first optical flow fusion unit, that is, use the weight prediction unit. First, according to and predict Then use the second optical flow fusion unit. According to for and perform weighted fusion to obtain This method is the explicit fusion mentioned above, that is, fuse the gyroscope domain and the temporary optical flow according to a clear indication information (weight map).
[0164] As for Figure 8 the SGF module in
[0165] uses both implicit fusion and explicit fusion, which is beneficial to obtaining a more accurate optical flow estimation result.
[0166] The training of the neural network model can be divided into two methods: supervised and unsupervised. Supervised training requires data annotation. However, since annotating optical flow data is difficult and time-consuming, the former is mainly trained on synthetic data (for example, videos generated by computer vision algorithms), and its application range is relatively narrow. While the latter can be trained on actual data (for example, real-shot videos), and its application range is relatively wide. Therefore, here we mainly introduce the method of unsupervised training.
[0167] When performing unsupervised training, the following two loss functions can be but are not limited to being set:
[0168] 1. Photo loss
[0169] Assume that V ab is the optical flow from I a to I b (I a to I b in the training are two images in the training set). Then V ab can be used to warp I a to obtain warp(I a ), and then calculate the representation of warp(I a ) and I bThe image loss of the difference (formally, it can be L1 loss or L2 loss). Setting this loss is intended to reduce the difference between warp(I a ) and I b .
[0170] 2. Smooth loss
[0171] This loss calculates a smoothness index within V ab (formally, it can be TV loss). Setting this loss is intended to make the estimated optical flow relatively smooth. Since the movement of actual objects generally has a certain integrity, correspondingly, the distribution of vectors in the optical flow also has a certain regularity. If the vectors in the optical flow are chaotic, it will surely not conform to the actual situation.
[0172] The above two losses can be weighted as the total loss. According to the gradient of the total loss function, the network parameters can be updated until the training termination condition (for example, the model converges) is reached. In some implementation manners, only the image loss can be calculated without calculating the smooth loss.
[0173] Figure 9 Fig. shows the functional block diagram of the optical flow estimation device 200 provided by an embodiment of the present application. Referring to Figure 9 , the optical flow estimation device 200 includes:
[0174] A data acquisition component 210, configured to acquire a first image, a second image, first gyroscope data, and second gyroscope data; wherein, the first image and the second image are images acquired by the same camera at different times, the first gyroscope data is the data acquired by the gyroscope during the acquisition of the first image, and the second gyroscope data is the data acquired by the gyroscope during the acquisition of the second image;
[0175] A gyroscope domain calculation component 220, configured to calculate a gyroscope domain according to the first gyroscope data and the second gyroscope data, where the gyroscope domain is a two-dimensional motion field between the first image and the second image;
[0176] An optical flow estimation component 230, configured to estimate the temporary optical flow between the first image and the second image according to the first image and the second image, and fuse the temporary optical flow and the gyroscope domain to obtain the optical flow between the first image and the second image.
[0177] In an implementation manner of the optical flow estimation device 200, the first gyroscope data is the data acquired by the gyroscope during the exposure of the first image, and the second gyroscope data is the data acquired by the gyroscope during the exposure of the second image; wherein, the camera's acquisition of an image includes two stages: exposure and post-processing.
[0178] In an implementation of the optical flow estimation device 200, the gyroscopic domain calculation component 220 calculates the gyroscopic domain based on the first gyroscope data and the second gyroscope data, including: calculating a rotation matrix between the first image and the second image according to the first gyroscope data and the second gyroscope data; calculating a homography matrix corresponding to the rotation matrix according to the rotation matrix and the internal parameters of the camera; and calculating the gyroscopic domain according to the homography matrix.
[0179] In an implementation of the optical flow estimation device 200, the camera uses a rolling shutter. The gyroscopic domain calculation component 220 calculates a rotation matrix between the first image and the second image according to the first gyroscope data and the second gyroscope data, including: calculating n rotation matrices between the first image and the second image according to n sets of data included in the first gyroscope data and n sets of data included in the second gyroscope data; where n is an integer greater than 1, each set of data is collected at different times, and the i-th rotation matrix is the rotation matrix between the i-th part of the first image in the exposure order and the i-th part of the second image in the exposure order, and i is any integer from 1 to n; the gyroscopic domain calculation component 220 calculates a homography matrix corresponding to the rotation matrix according to the rotation matrix and the internal parameters of the camera, including: calculating n homography matrices corresponding to the n rotation matrices according to the n rotation matrices and the internal parameters of the camera; the gyroscopic domain calculation component 220 calculates the gyroscopic domain according to the homography matrix, including: calculating n partial gyroscopic domains according to the n homography matrices, and splicing the n partial gyroscopic domains into the gyroscopic domain; where the i-th partial gyroscopic domain is the two-dimensional motion field between the i-th part of the first image in the exposure order and the i-th part of the second image in the exposure order, and i is any integer from 1 to n.
[0180] In an implementation of the optical flow estimation device 200, the gyroscopic domain calculation component 220 calculates n rotation matrices between the first image and the second image according to n sets of data included in the first gyroscope data and n sets of data included in the second gyroscope data, including: calculating n corresponding temporary rotation matrices M 1 ~M n according to the n sets of data included in the first gyroscope data, and, calculating n corresponding temporary rotation matrices M n+1 ~M 2n according to the n sets of data included in the second gyroscope data; i traverses the integers from 1 to n, and according to n + 1 temporary rotation matrices M i ~M n+i , calculates the i-th rotation matrix between the first image and the second image. After the traversal is completed, the n rotation matrices are obtained.
[0181] In an implementation of the optical flow estimation device 200, the camera uses a global shutter. The gyroscope domain calculation component 220 calculates a rotation matrix between the first image and the second image according to the first gyroscope data and the second gyroscope data, including: calculating a rotation matrix between the overall first image and the overall second image according to n sets of data included in the first gyroscope data and n sets of data included in the second gyroscope data; where n is an integer greater than 1, and each set of data is collected at different times.
[0182] In an implementation of the optical flow estimation device 200, the gyroscope domain calculation component 220 calculates a rotation matrix between the overall first image and the overall second image according to n sets of data included in the first gyroscope data and n sets of data included in the second gyroscope data, including: calculating corresponding n temporary rotation matrices M 1 ~M n according to the n sets of data included in the first gyroscope data, and calculating corresponding n temporary rotation matrices M n+1 ~M 2n according to the n sets of data included in the second gyroscope data; calculating the rotation matrix according to the 2n temporary rotation matrices M 1 ~M 2n .
[0183] In an implementation of the optical flow estimation device 200, the optical flow estimation component 230 estimates the temporary optical flow between the first image and the second image, and fuses the temporary optical flow and the gyroscope domain to obtain the optical flow between the first image and the second image, including: using a neural network model to estimate the temporary optical flow according to the first image and the second image; fusing the temporary optical flow and the gyroscope domain to obtain the optical flow between the first image and the second image.
[0184] In one implementation of the optical flow estimation device 200, the optical flow estimation component 230 fuses the temporary optical flow and the gyroscope domain, including: fusing the temporary optical flow and the gyroscope domain using the neural network model; the neural network model includes m sequentially connected optical flow estimation modules, where m is an integer greater than 1, and the k-th optical flow estimation module performs the following steps: extracting the k-th level feature of the first image based on the (k - 1)-th level feature of the first image, and extracting the k-th level feature of the second image based on the (k - 1)-th level feature of the second image, and downsampling the (k - 1)-th level feature when extracting the k-th level feature; fusing the k-th level gyroscope domain and the (k + 1)-th level temporary optical flow to obtain the k-th level fused optical flow, where the k-th level gyroscope domain is obtained by downsampling the (k - 1)-th level gyroscope domain; calculating the k-th level temporary optical flow based on the k-th level feature of the first image, the k-th level feature of the second image, and the k-th level fused optical flow, and upsampling the k-th level fused optical flow when calculating the k-th level temporary optical flow; where k is any integer from 1 to m, the 0-th level feature of the first image is the first image, the 0-th level feature of the second image is the second image, the 0-th level gyroscope domain is the gyroscope domain, the m-th level gyroscope domain is 0, the (m + 1)-th level temporary optical flow is 0, and the optical flow between the first image and the second image is obtained by upsampling the 1st level temporary optical flow.
[0185] In an implementation of the optical flow estimation device 200, the k-th optical flow estimation module fuses the k-th gyroscope domain and the (k + 1)-th temporary optical flow to obtain the k-th fused optical flow, including one of the following three methods: using the first optical flow fusion unit in the k-th optical flow estimation module to fuse the k-th gyroscope domain and the (k + 1)-th temporary optical flow to obtain the k-th fused optical flow, where the first optical flow fusion unit includes at least one convolutional layer; using the weight prediction unit in the k-th optical flow estimation module to predict the k-th weight map according to the k-th features of the first image and the k-th features of the second image, and using the k-th weight map to perform weighted fusion on the k-th gyroscope domain and the (k + 1)-th temporary optical flow to obtain the k-th fused optical flow; wherein, the pixel value in the k-th weight map represents the fusion weight, and the weight prediction unit includes at least one convolutional layer; using the first optical flow fusion unit in the k-th optical flow estimation module to fuse the k-th gyroscope domain and the (k + 1)-th temporary optical flow to obtain the k-th temporary fused optical flow, using the weight prediction unit in the k-th optical flow estimation module to predict the k-th weight map according to the k-th features of the first image and the k-th features of the second image, and using the k-th weight map to perform weighted fusion on the k-th gyroscope domain and the k-th temporary fused optical flow to obtain the k-th fused optical flow; wherein, both the first optical flow fusion unit and the weight prediction unit include at least one convolutional layer, and the pixel value in the k-th weight map represents the fusion weight.
[0186] The optical flow estimation device 200 provided by the embodiments of the present application, its implementation principle and the generated technical effects have been introduced in the foregoing method embodiments. For a brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding content in the method embodiments.
[0187] Figure 10 Shows a possible structure of the electronic device 300 provided by the embodiments of the present application. Refer to Figure 10 , the electronic device 300 includes (solid line boxes): a processor 310, a memory 320, and a communication interface 350. These components are interconnected and communicate with each other through a communication bus 360 and / or other forms of connection mechanisms (not shown).
[0188] Among them, the processor 310 includes one or more, which may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 310 may be a general-purpose processor, including a central processing unit (CPU), a microcontroller unit (MCU), a network processor (NP), or other conventional processors; it may also be a dedicated processor, including a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Moreover, when there are multiple processor 310s, a part of them may be general-purpose processors, and another part may be dedicated processors.
[0189] The memory 320 includes one or more, which may be, but is not limited to, random access memory (RAM), read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0190] The processor 310 and other possible components can access the memory 320, read and / or write data therein. In particular, one or more computer program instructions may be stored in the memory 320, and the processor 310 can read and run these computer program instructions to implement the optical flow estimation method provided in the embodiments of the present application.
[0191] The communication interface 350 includes one or more interfaces that can be used to communicate directly or indirectly with other devices for data interaction. The communication interface 350 may include interfaces for wired and / or wireless communication. If there is no need to communicate with other devices, the electronic device 300 may not be provided with the communication interface 350.
[0192] It can be understood that Figure 10 The structure shown is only schematic, and the electronic device 300 may also include more or fewer components than those shown Figure 10 in the figure, or have a different configuration from that shown Figure 10 in the figure:
[0193] In some implementation manners, the electronic device 300 further includes (dashed boxes): a camera 330 and a gyroscope 340, which are interconnected and communicate with other components introduced above through a communication bus 360 and / or other forms of connection mechanisms (not shown).
[0194] Among them, the camera 330 may be one or more. For example, it may be a wide-angle camera, a telephoto camera, etc. The camera 330 is used to collect image or video data, including a first image and a second image required for optical flow estimation. The first image and the second image may be collected by the same camera at different times.
[0195] The gyroscope 340 may be a three-axis gyroscope, a six-axis gyroscope, etc. The gyroscope 340 is used to collect gyroscope data, including a first gyroscope data and a second gyroscope data required for optical flow estimation. The gyroscope data at least includes angular velocity information.
[0196] Figure 10 The components shown in the figure may be implemented by hardware, software, or a combination thereof. The electronic device 300 may be a physical device, such as a mobile phone, a camera, a digital camera, a tablet computer, a laptop computer, a PC, a drone, a wearable device, a robot, a server, etc., or may be a virtual device, such as a virtual machine, a virtualization container, etc. Moreover, the electronic device 300 is not limited to a single device, and may also be a combination of multiple devices or a cluster composed of a large number of devices.
[0197] The embodiment of the present application further provides a computer-readable storage medium, on which computer program instructions are stored. When these computer program instructions are read and run by a processor, the optical flow estimation method provided by the embodiment of the present application is executed. For example, the computer-readable storage medium may be implemented as Figure 10 the memory 320 in the electronic device 300 shown in the figure.
[0198] An embodiment of the present application further provides a computer program product, which includes computer program instructions. When these computer program instructions are read and executed by a processor, they perform the optical flow estimation method provided by the embodiment of the present application.
[0199] The above are only embodiments of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. An optical flow estimation method, characterized in that, it includes: Obtain a first image, a second image, first gyroscope data, and second gyroscope data; wherein, the first image and the second image are images collected by the same camera at different times, the first gyroscope data is the data collected by the gyroscope during the acquisition of the first image, and the second gyroscope data is the data collected by the gyroscope during the acquisition of the second image; Calculate a gyroscopic domain according to the first gyroscope data and the second gyroscope data, where the gyroscopic domain is a two-dimensional motion field between the first image and the second image, and the gyroscopic domain is the optical flow calculated according to the first gyroscope data and the second gyroscope data; Estimate the temporary optical flow between the first image and the second image according to the first image and the second image, and fuse the temporary optical flow and the gyroscopic domain to obtain the optical flow between the first image and the second image.
2. The optical flow estimation method according to claim 1, characterized in that, the first gyroscope data is the data collected by the gyroscope during the exposure of the first image, and the second gyroscope data is the data collected by the gyroscope during the exposure of the second image; wherein, the camera image acquisition includes two stages: exposure and post-processing.
3. The optical flow estimation method according to claim 1 or 2, characterized in that, the calculating the gyroscopic domain according to the first gyroscope data and the second gyroscope data includes: Calculate the rotation matrix between the first image and the second image according to the first gyroscope data and the second gyroscope data; Calculate the homography matrix corresponding to the rotation matrix according to the rotation matrix and the internal parameters of the camera; Calculate the gyroscopic domain according to the homography matrix.
4. The optical flow estimation method according to claim 3, characterized in that, the camera uses a rolling shutter, and the calculating the rotation matrix between the first image and the second image according to the first gyroscope data and the second gyroscope data includes: Calculate n rotation matrices between the first image and the second image according to n sets of data included in the first gyroscope data and n sets of data included in the second gyroscope data; where n is an integer greater than 1, each set of data is collected at different times, and the i-th rotation matrix is the rotation matrix between the i-th part of the first image in the exposure order and the i-th part of the second image in the exposure order, and i is any integer from 1 to n; the calculating the homography matrix corresponding to the rotation matrix according to the rotation matrix and the internal parameters of the camera includes: Calculate n homography matrices corresponding to the n rotation matrices according to the n rotation matrices and the internal parameters of the camera; the calculating the gyroscopic domain according to the homography matrix includes: Calculate n partial gyrodomains according to the n homography matrices, and splice the n partial gyrodomains into the gyrodomain; wherein, the i-th partial gyrodomain is a two-dimensional motion field between the i-th part of the first image in the exposure order and the i-th part of the second image in the exposure order, and i is any integer from 1 to n.
5. The optical flow estimation method according to claim 4, wherein, the calculating of n rotation matrices between the first image and the second image according to the n sets of data included in the first gyroscope data and the n sets of data included in the second gyroscope data includes: Calculate the corresponding n temporary rotation matrices M based on the n groups of data included in the first gyroscope data 1 ~M n , and calculate the corresponding n temporary rotation matrices M based on the n groups of data included in the second gyroscope data n+1 ~M 2n ; i traverses the integers from 1 to n, and calculates the i-th rotation matrix between the first image and the second image according to the n+1 temporary rotation matrices M i ~M n+i , and after the traversal is completed, the n rotation matrices are obtained.
6. The optical flow estimation method according to claim 3, wherein, the camera adopts a global shutter, and the calculating of the rotation matrix between the first image and the second image according to the first gyroscope data and the second gyroscope data includes: Calculate a rotation matrix between the whole of the first image and the whole of the second image according to the n sets of data included in the first gyroscope data and the n sets of data included in the second gyroscope data; wherein, n is an integer greater than 1, and each set of data is collected at different times.
7. The optical flow estimation method according to claim 6, wherein, the calculating of a rotation matrix between the whole of the first image and the whole of the second image according to the n sets of data included in the first gyroscope data and the n sets of data included in the second gyroscope data includes: Calculate the corresponding n temporary rotation matrices M according to the n groups of data included in the first gyroscope data 1 ~M n , and calculate the corresponding n temporary rotation matrices M according to the n groups of data included in the second gyroscope data n+1 ~M 2n ; Calculate the rotation matrix according to 2n temporary rotation matrices M 1 ~M 2n 8. The optical flow estimation method according to claim 1, wherein, the estimating of the temporary optical flow between the first image and the second image and the fusing of the temporary optical flow and the gyrodomain to obtain the optical flow between the first image and the second image includes: Estimate the temporary optical flow according to the first image and the second image by using a neural network model; Fuse the temporary optical flow and the gyrodomain to obtain the optical flow between the first image and the second image.
9. The optical flow estimation method according to claim 8, wherein, the fusing of the temporary optical flow and the gyrodomain includes: fusing the temporary optical flow and the gyrodomain by using the neural network model; The neural network model includes m successively connected optical flow estimation modules, m is an integer greater than 1, and the k-th optical flow estimation module performs the following steps: Extract the k-th level feature of the first image according to the k-1-th level feature of the first image, and extract the k-th level feature of the second image according to the k-1-th level feature of the second image. When extracting the k-th level feature, downsampling is performed on the k-1-th level feature; Fuse the k-th level gyrodomain and the k+1-th level temporary optical flow to obtain the k-th level fused optical flow, and the k-th level gyrodomain is obtained by downsampling the k-1-th level gyrodomain; Calculate the k-th level temporary optical flow according to the k-th level feature of the first image, the k-th level feature of the second image and the k-th level fused optical flow. When calculating the k-th level temporary optical flow, upsampling is performed on the k-th level fused optical flow; Wherein, k is any integer from 1 to m, the 0th-level feature of the first image is the first image, the 0th-level feature of the second image is the second image, the 0th-level gyro domain is the gyro domain, the mth-level gyro domain is 0, the (m + 1)th-level temporary optical flow is 0, and the optical flow between the first image and the second image is obtained from the 1st-level temporary optical flow.
10. The optical flow estimation method according to claim 9, characterized in that the fusing the kth-level gyro domain and the (k + 1)th-level temporary optical flow to obtain the kth-level fused optical flow includes one of the following three methods: Using the first optical flow fusion unit in the kth optical flow estimation module to fuse the kth-level gyro domain and the (k + 1)th-level temporary optical flow to obtain the kth-level fused optical flow, and the first optical flow fusion unit includes at least one convolutional layer; Using the weight prediction unit in the kth optical flow estimation module to predict the kth-level weight map according to the kth-level features of the first image and the kth-level features of the second image, and using the kth-level weight map to perform weighted fusion on the kth-level gyro domain and the (k + 1)th-level temporary optical flow to obtain the kth-level fused optical flow; wherein, the pixel value in the kth-level weight map represents the fusion weight, and the weight prediction unit includes at least one convolutional layer; Using the first optical flow fusion unit in the kth optical flow estimation module to fuse the kth-level gyro domain and the (k + 1)th-level temporary optical flow to obtain the kth-level temporary fused optical flow, using the weight prediction unit in the kth optical flow estimation module to predict the kth-level weight map according to the kth-level features of the first image and the kth-level features of the second image, and using the kth-level weight map to perform weighted fusion on the kth-level gyro domain and the kth-level temporary fused optical flow to obtain the kth-level fused optical flow; wherein, both the first optical flow fusion unit and the weight prediction unit include at least one convolutional layer, and the pixel value in the kth-level weight map represents the fusion weight.
11. A computer program product, characterized in that it includes computer program instructions, and when the computer program instructions are read and run by a processor, the method described in any one of claims 1-10 is executed.
12. A computer-readable storage medium, characterized in that computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are read and run by a processor, the method described in any one of claims 1-10 is executed.
13. An electronic device, characterized in that it includes a memory and a processor, computer program instructions are stored in the memory, and when the computer program instructions are read and run by the processor, the method described in any one of claims 1-10 is executed.
14. The electronic device according to claim 13, characterized in that the device further includes a camera and a gyroscope, the camera is used to collect a first image and a second image, and the gyroscope is used to collect first gyroscope data and second gyroscope data.
Citation Information
Patent Citations
Video anti-shake method and device integrated with gyroscope
CN109618091A
Optical flow estimation using a neural network and egomotion optimization
US10262224B1