Image registration method, surface normal vector reconstruction method, system and electronic device
The method improves image alignment precision and reduces computational complexity by using edge detection and a high-frequency loss function in a two-step optical flow estimation process, addressing the limitations of traditional and neural network-based methods for dynamic objects.
Patent Information
- Application Number
- CN202310325418.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-29
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-03-29
AI Technical Summary
In the prior art, the image registration method based on the optical flow estimation calculation method has poor accuracy when processing dynamic objects and has high computational complexity, which cannot meet the real-time requirements.
Using an image registration method based on an optical flow estimation algorithm, by acquiring edge images of multiple video frames, a loss fraction containing a high frequency information change rate is used to combine the first and second optical flow estimation results to adjust the target frame to achieve accurate registration.
It improves the optical flow estimation accuracy of dynamic objects, reduces the amount of calculation, and realizes real-time and accurate registration between video frames.
Smart Images

Figure CN116309755B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technologies, and in particular, to an image registration method, a surface normal vector reconstruction method, an image registration system, a surface normal vector reconstruction system, an electronic device, and a storage medium. Background Art
[0002] In many entertainment or industrial image applications, it may be necessary to use an optical flow estimation algorithm to register different images. For example, for the case where the target object is an object that may have movement changes in a short time, such as a human body that will have slight movements such as breathing in a short time, additional calculations are required for the video frames of the target object to achieve registration. Specifically, for example, an optical flow estimation algorithm can be used to estimate the optical flow between frames, so as to achieve rough registration between video frames.
[0003] However, since this algorithm assumes that the optical flow in all regions of the picture is globally smooth and has no mutation points, and performs smooth interpolation processing on the optical flow of the entire picture. In fact, the global optical flow between frames is very likely not to be smooth. For example, if the target object is an irregular three-dimensional object or a human body, its movement is likely to cause regions that are not visible in the previous frame among multiple captured video frames to appear in the subsequent frame. As a result, it is impossible to accurately calculate the optical flow of this part of the region, and thus the accuracy of image registration is poor.
[0004] To solve the problem of registration accuracy, in recent years, an optical flow estimation algorithm based on a neural network has emerged, but this method has a very high computational complexity. For the same amount of tasks, the hardware resources consumed may be hundreds of times that of the traditional optical flow estimation algorithm, and it cannot be applied to scenarios with high real-time requirements and large amounts of tasks. Summary of the Invention
[0005] In order to at least partially solve the above problems existing in the prior art, an image registration method, a surface normal vector reconstruction method, an image registration system, a surface normal vector reconstruction system, an electronic device, and a storage medium are provided.
[0006] According to one aspect of the present application, an image registration method is provided, including:
[0007] Obtain a plurality of video frames;
[0008] Based on an optical flow estimation algorithm, determine a first optical flow estimation result from a target frame to a reference frame among the plurality of video frames;
[0009] Perform edge detection on the target frame and the reference frame respectively to obtain an edge image of the target frame and an edge image of the reference frame;
[0010] Using a second loss function and based on an optical flow estimation algorithm, determine a second optical flow estimation result from the edge image of the target frame to the edge image of the reference frame, where the second loss function includes a fraction of the loss regarding the change rate of high-frequency information in the edge image; and
[0011] Based on the first optical flow estimation result and the second optical flow estimation result, adjust the target frame to obtain an adjusted video frame.
[0012] Exemplarily, represent the second loss function E2(u) using the following formula:
[0013]
[0014] where P1(x,y) represents the pixel value of the pixel (x, y) in the edge image of the target frame, P2(x,y) represents the pixel value of the pixel (x, y) in the edge image of the reference frame, u2[x,y] represents the optical flow displacement matrix of the second optical flow estimation result of the pixel (x, y), and β and γ respectively represent adjustment coefficients.
[0015] Exemplarily, the method further includes:
[0016] Determine an occlusion area and a non-occlusion area of the target frame relative to the reference frame at least based on the difference between the first optical flow estimation result and the second optical flow estimation result;
[0017] Adjusting the target frame based on the first optical flow estimation result and the second optical flow estimation result to obtain an adjusted video frame includes:
[0018] For each first pixel in the non-occlusion area of the target frame, directly adjust the first pixel according to the first optical flow estimation result of the first pixel;
[0019] For each second pixel in the occlusion area of the target frame, adjust the second pixel according to the first optical flow estimation result of other pixels in the target frame.
[0020] Exemplarily, adjusting the second pixel according to the first optical flow estimation result of other pixels in the target frame includes:
[0021] For each second pixel in the occlusion area of the target frame, adjust the second pixel according to the first optical flow estimation result of at least one first pixel closest to the second pixel in the occlusion area of the target frame.
[0022] Exemplarily, adjusting the second pixel according to the first optical flow estimation result of at least one first pixel closest to the second pixel in the occlusion area of the target frame includes:
[0023] Determine multiple first pixels closest to the second pixel in the occlusion area of the target frame;
[0024] According to the distance between each determined first pixel and the second pixel, perform inverse distance weighted averaging on the first optical flow estimation results of the determined multiple first pixels to determine the optical flow displacement of the second pixel; and
[0025] Adjust the second pixel according to the optical flow displacement of the second pixel.
[0026] Exemplarily, determine the occluded area and the non-occluded area of the target frame relative to the reference frame at least based on the difference between the first optical flow estimation result and the second optical flow estimation result, including:
[0027] Determine the area where the pixel (x, y) that satisfies the following formula is located as the occluded area of the target frame relative to the reference frame:
[0028]
[0029] where u1[x,y] represents the optical flow displacement matrix of the first optical flow estimation result of the pixel (x, y), u2[x,y] represents the optical flow displacement matrix of the second optical flow estimation result of the pixel (x, y), η represents an adjustment coefficient, η > 0, and T represents an optical flow threshold; and
[0030] Determine the area outside the occluded area as the non-occluded area.
[0031] Exemplarily, the method further includes:
[0032] Respectively determine at least some frames in the multiple video frames as pending frames;
[0033] For each pending frame, perform the following operations on each target frame of the pending frame,
[0034] Based on the optical flow estimation algorithm, determine the first optical flow estimation result from the target frame to the pending frame;
[0035] Perform edge detection on the target frame and the pending frame respectively to obtain the edge image of the target frame and the edge image of the pending frame respectively;
[0036] Use the second loss function and based on the optical flow estimation algorithm, determine the second optical flow estimation result from the edge image of the target frame to the edge image of the pending frame;
[0037] Determine the occluded area of the target frame relative to the pending frame at least based on the difference between the first optical flow estimation result from the target frame to the pending frame and the second optical flow estimation result from the target frame to the pending frame;
[0038] Calculate the area of the occluded area of the target frame relative to the pending frame;
[0039] Calculate the sum of the areas of the occluded regions of each target frame with respect to the pending frame; and
[0040] Compare the sums of the areas calculated for each pending frame respectively, and determine the pending frame corresponding to the smallest sum of the areas as the reference frame.
[0041] Exemplarily, the method further includes:
[0042] Determine a middle frame among the multiple video frames as the reference frame.
[0043] Exemplarily, the optical flow estimation algorithm includes a two-frame differential optical flow estimation algorithm and a dense inverse search optical flow estimation algorithm.
[0044] According to another aspect of the present application, there is also provided a surface normal vector reconstruction method, including:
[0045] Use the above image registration method to register multiple video frames of the target object;
[0046] Use the adjusted video frames to reconstruct the surface normal vector of the target object.
[0047] According to another aspect of the present application, there is also provided an image registration system, including:
[0048] An acquisition module for acquiring multiple video frames;
[0049] A first determination module for determining a first optical flow estimation result from the target frame to the reference frame among the multiple video frames based on the optical flow estimation algorithm;
[0050] An edge detection module for respectively performing edge detection on the target frame and the reference frame to respectively obtain an edge image of the target frame and an edge image of the reference frame;
[0051] A second determination module for determining a second optical flow estimation result from the edge image of the target frame to the edge image of the reference frame by using a second loss function and based on the optical flow estimation algorithm, where the second loss function includes a fraction of the loss regarding the change rate of the high-frequency information in the edge image; and
[0052] An adjustment module for adjusting the target frame based on the first optical flow estimation result and the second optical flow estimation result to obtain adjusted video frames.
[0053] According to another aspect of the present application, there is also provided a surface normal vector reconstruction system, including:
[0054] A registration module for registering multiple video frames of the target object by using the above image registration method; and
[0055] A reconstruction module for reconstructing the surface normal vector of the target object by using the adjusted video frames.
[0056] According to another aspect of the present application, there is also provided an electronic device, including a processor and a memory. Among them, computer program instructions are stored in the memory, and when the computer program instructions are run by the processor, they are used to execute the above-mentioned image registration method and / or the above-mentioned surface normal vector reconstruction method.
[0057] According to another aspect of the present application, there is also provided a storage medium, on which program instructions are stored, and when the program instructions are run, they are used to execute the above-mentioned image registration method and / or the above-mentioned surface normal vector reconstruction method.
[0058] In the above solution, by obtaining the first optical flow estimation result from each target frame to the reference frame in multiple video frames, and obtaining the second optical flow estimation result between the edge images of the target frame and the reference frame respectively, finally, based on the results of the two optical flow estimations, the accurate registration of each target frame to the reference frame is realized. Moreover, since the second loss function used in the process of determining the second optical flow estimation result includes a fraction of the loss regarding the change rate of the high-frequency information in the two edge images, the optical flow for each detailed area between each target frame and the reference frame can be accurately estimated. In particular, the accuracy of optical flow estimation when shooting dynamic objects at close range can be improved, and the accuracy of optical flow estimation for areas with sudden optical flow between video frames can be increased. And, the computational complexity of this solution is close to that of traditional optical flow estimation algorithms, so the computational complexity is small and the processing speed is fast. Therefore, the real-time and accurate registration of multiple video frames can be effectively realized.
[0059] A series of simplified concepts are introduced in the summary of the invention, which will be further described in detail in the detailed implementation section. The content part of the present application does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the protection scope of the claimed technical solution.
[0060] The advantages and features of the present application will be described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] The following drawings of the present application are hereby incorporated as part of the present application for understanding the present application. The embodiments of the present application are shown in the drawings and their descriptions are used to explain the principles of the present application. In the drawings,
[0062] Figure 1 A schematic flowchart of an image registration method according to an embodiment of the present application is shown;
[0063] Figure 2 A schematic diagram showing the principle of determining the first optical flow estimation result according to an embodiment of the present application is shown;
[0064] Figure 3A schematic diagram showing the principle of determining a second optical flow estimation result according to an embodiment of the present application;
[0065] Figure 4 A schematic flowchart showing a surface normal vector reconstruction method according to an embodiment of the present application;
[0066] Figure 5 A schematic block diagram showing an image registration system according to an embodiment of the present application; and
[0067] Figure 6 A schematic block diagram showing a surface normal vector reconstruction system according to an embodiment of the present application; and
[0068] Figure 7 A schematic block diagram showing an electronic device according to an embodiment of the present application. Detailed implementation manners
[0069] In the following description, a large number of details are provided to enable a thorough understanding of the present application. However, those skilled in the art can understand that the following description only exemplarily shows the preferred embodiments of the present application, and the present application can be implemented without one or more of such details. In addition, in order to avoid confusion with the present application, some well-known technical features in the art are not described in detail.
[0070] To at least partially solve the above technical problems, according to a first aspect of the present application, an image registration method is provided. This image registration method can accurately calculate the optical flow between multiple video frames of a target object, and then achieve precise registration of multiple video frames based on the optical flow estimation result. The computational amount of this solution is small and the processing efficiency is high, so that real-time registration between video frames can be effectively achieved.
[0071] Figure 1 A schematic flowchart showing an image registration method 100 according to an embodiment of the present application. As Figure 1 shown, the image registration method 100 includes the following steps.
[0072] Step S110, obtain multiple video frames.
[0073] According to an embodiment of the present application, the multiple video frames can be video frames of a target object collected by using any existing or future-developed suitable image acquisition device. The target object can be any suitable object. For example, it can be a target human body or animal, or other three-dimensional objects.
[0074] According to an embodiment of the present application, the multiple video frames may be acquired by the same image acquisition device or by different image acquisition devices. The multiple video frames may be consecutive video frames, for example, multiple images of a target object consecutively acquired by an image acquisition device. The multiple video frames may also be non-consecutive video frames, for example, video frames that meet preset time-domain requirements and are selected from consecutive video frames. In a specific embodiment, the target object is a target human body or a three-dimensional object, and the multiple video frames may be a set of consecutive video frames of the target object taken at close range by an image acquisition device under a light source environment where the light source irradiates the surface of the target object from different preset angles at different times. It can be understood that in this case, in the multiple video frames, the position of the target object in the image is approximately the same, but the brightness of each corresponding region is different. The multiple video frames obtained by this solution can facilitate registration using the optical flow estimation method.
[0075] According to an embodiment of the present application, each of the multiple video frames may be an image of any size or resolution, or an image that meets preset resolution requirements. The multiple video frames may be black-and-white images or color images. The requirements for the multiple video frames may be set based on actual processing needs, the hardware of the image acquisition device, etc. The multiple video frames may be the original images directly acquired by the image acquisition device or the images after preprocessing operations on the original images. The preprocessing operations may include operations such as geometric transformation and filtering of the original images to facilitate subsequent processing.
[0076] Step S130, based on the optical flow estimation algorithm, determine the first optical flow estimation result from the target frame to the reference frame among the multiple video frames.
[0077] It can be understood that the reference frame and the target frame among the multiple video frames are relative. The reference frame may be a certain video frame selected from the multiple video frames, and the other video frames except the reference frame among the multiple video frames may all be called target frames. For example, the multiple video frames are 3 consecutive video frames, the reference frame is the selected second frame, then the first frame and the third frame both belong to the target frames.
[0078] According to the embodiments of the present application, any suitable method can be used to determine a reference frame. Exemplarily, the method 100 may further include determining a reference frame from multiple video frames. In one example, any one of the multiple video frames can be determined as the reference frame. In another example, a specific frame among the multiple video frames can also be determined as the reference frame. Optionally, a video frame that meets the timing requirement among the multiple video frames can be determined as the reference frame. For example, the method 100 may include step S121 of determining a middle frame among the multiple video frames as the reference frame. For example, if the number of multiple video frames is 7, the 4th frame sorted according to the shooting timing can be determined as the reference frame. It can be understood that the movement of the target object is usually continuous and stable. Therefore, the method of using the middlemost video frame as the reference frame for optical flow estimation calculation and subsequent registration of the target frame has less calculation amount and relatively more accurate registration results. Alternatively, a video frame that meets the preset image quality requirement can also be used as the reference frame. For example, a target recognition algorithm can be used to identify the area where the target object is located, and the proportion of the target object area in the image can be calculated. Then, the video frame with the highest proportion of the target object among the multiple video frames can be used as the reference frame. Alternatively, multiple video frames that meet the preset timing requirement can be used as reference frames in sequence, and the optical flow from each corresponding target frame to the reference frame can be calculated one by one. And finally, the final reference frame can be selected according to the optical flow estimation result. The specific implementation of this example will be described later and will not be elaborated here. Of course, other suitable methods can also be used to determine the reference frame.
[0079] According to the embodiments of the present application, any existing or future-developed suitable optical flow estimation algorithm can be used to determine the optical flow between each target frame and the determined reference frame. Exemplarily and non-limitingly, step S130 can use any one of the two-frame differential optical flow estimation algorithm (Lucas Kanade optical flow algorithm, abbreviated as L-K optical flow algorithm) and the dense inverse search optical flow estimation algorithm (Dense Inverse Search-based method, abbreviated as DIS optical flow algorithm) to calculate the first optical flow estimation result between the target frame and the reference frame. Optionally, the sparse optical flow estimation between each target frame and the reference frame can be calculated using the L-K optical flow algorithm, and then the dense optical flow estimation between each target frame and the reference frame can be obtained by combining the interpolation algorithm. Alternatively, the DIS optical flow algorithm can also be used to directly obtain the dense optical flow estimation between each target frame and the reference frame. For the latter, since the DIS optical flow algorithm is an algorithm that maximally balances the optical flow quality and calculation time, the quality of the first optical flow estimation result obtained by adopting this scheme is good, and the calculation amount is small and the calculation efficiency is high. Thus, it can ensure real-time acquisition of a high-precision first optical flow estimation result, which is helpful for realizing real-time registration of multiple video frames.
[0080] The following takes the direct acquisition of the dense optical flow estimation between each target frame and the reference frame using the DIS optical flow algorithm as an example to describe the process of obtaining the optical flow estimation result. Figure 2 FIG. shows a schematic diagram of the principle of determining the first optical flow estimation result according to an embodiment of the present application. As Figure 2 shown, it can be understood that if the target frame and the reference frame are respectively denoted as I(a) and I(b), and the optical flow between the target frame and the reference frame is denoted as O(a, b), then the following relationship can exist for each pixel point in the target frame and the reference frame: I(a) + O(a, b) = I(b). That is, using the optical flow between two video frames, the target frame can be converted into the reference frame.
[0081] It is assumed that the pixel point P'2(x, y) at the pixel position (x, y) in the reference frame may be moved from the pixel point P′1((x, y) - u1[x, y]) at the pixel position (x, y) - u1[x, y] on the target frame. Among them, u1[x, y] can be regarded as the optical flow displacement matrix of the pixel point. The optical flow displacement matrix between the target frame and the reference frame can be solved using the following first loss function.
[0082]
[0083] It can be understood that the above first loss function formula includes two parts. The first part represents the error magnitude of each pixel point in the target frame after moving according to the optical flow displacement matrix u1[x, y] and matching the corresponding pixel point in the reference frame. The second part represents the smoothness inside the optical flow. Assuming that all pixels move in an approximate displacement in one direction, it can be considered that the inside of the optical flow is smooth, which is more in line with the natural law. The α in the above formula represents a regulation coefficient, and α is a fixed value, which can be arbitrarily set according to actual processing requirements. The above loss function setting method of the DIS optical flow algorithm can balance the optical flow smoothness and the matching degree after image deformation.
[0084] The above method can be used to sequentially calculate the optical flow displacement matrix between each target frame and the reference frame in multiple video frames to obtain the first optical flow estimation result.
[0085] Step S150, perform edge detection on the target frame and the reference frame respectively to obtain the edge image of the target frame and the edge image of the reference frame respectively.
[0086] According to the embodiments of the present application, any existing or future-developed edge detection algorithm can be used to obtain the edge images of the target frame and the reference frame. Exemplarily, any one of the edge detection algorithms such as the Roberts Cross operator, Prewitt operator, Sobel operator, Kirsch operator, and compass operator based on the first derivative method can be used to implement the edge detection in this step. Preferably, any one of the edge detection algorithms such as the Canny operator, Laplacian operator, and Marr-Hildreth operator based on the second derivative method can also be used to implement the edge detection in this step.
[0087] In a specific embodiment, the Canny algorithm can be used to perform edge detection on the reference frame and each target frame, and the edge images of the reference frame and each target frame are obtained respectively. Exemplarily, the edge images of the reference frame and the target frame can both be grayscale images or black-and-white images with the color pixel information removed. The feature information of the color gradient of the corresponding video frame can be presented in this edge image.
[0088] Step S170: Using the second loss function and based on the optical flow estimation algorithm, determine the second optical flow estimation result from the edge image of the target frame to the edge image of the reference frame. Wherein, the second loss function includes a fraction of the loss regarding the change rate of the high-frequency information in the edge image.
[0089] In this step, the same optical flow estimation algorithm as in step S130 can be used to calculate the optical flow between the edge image of the target frame and the edge image of the reference frame, and obtain the second optical flow estimation result. For example, if the DIS optical flow algorithm is used in step S130 to obtain the first optical flow estimation result, then this step can also use the DIS optical flow algorithm to obtain the second optical flow estimation result between the edge image of each target frame and the edge image of the reference frame. The second optical flow estimation result obtained by adopting this scheme has good quality, small calculation amount, and high calculation efficiency, so as to ensure real-time acquisition of a high-precision second optical flow estimation result, and further contribute to the real-time registration of multiple video frames.
[0090] Different from the first optical flow estimation result, the second optical flow estimation result is the optical flow between edge images. Figure 3 The schematic diagram shows the principle of determining the second optical flow estimation result according to an embodiment of the present application. As Figure 3 shown, and Figure 2The process of determining the first optical flow estimation result is similar. The edge images of each target frame and the reference frame can be denoted as I’(a) and I’(b) respectively. Denote the optical flow from the edge image of each target frame to the edge image of the reference frame as O’(a, b). Then, there can also be a relationship for each pixel point in the edge images of the target frame and the reference frame: I’(a) + O’(a, b) = I’(b). That is, using the optical flow between the edge images of two video frames, the edge image of the target frame can be converted into the edge image of the reference frame. Using the edge images of two video frames to estimate the optical flow again can greatly improve the robustness of the optical flow estimation algorithm and enhance the smoothness of the optical flow estimation under changing lighting conditions.
[0091] Similarly, taking the example of obtaining the dense optical flow estimation result from the edge image of each target frame to the edge image of the reference frame using the DIS optical flow algorithm. Similar to the method for determining the first optical flow estimation result, it can be assumed that each pixel point P2(x, y) at the pixel position (x, y) in the edge image of the reference frame can move from the pixel point P1((x, y) + u2[x, y]) at the pixel position (x, y) + u2[x, y] on the edge image of the target frame. Here, u2[x, y] can be regarded as the optical flow displacement matrix of the pixel points. The optical flow displacement matrix from the edge image of each target frame to the edge image of the reference frame can be solved using the second loss function.
[0092] It should be noted that the second loss function adopted in this step can be different from the above first loss function. In addition to including the loss fraction representing the error magnitude of each pixel point in the edge image of the target frame matching each corresponding pixel point in the edge image of the reference frame and the loss fraction representing the smoothness within the optical flow, it can also include the loss fraction regarding the change rate of the high-frequency information in the two edge images. In other words, during the process of estimating the optical flow for the edge images of the target frame and the reference frame respectively, the matching of the detailed textures in each region of the edge images is also fully considered.
[0093] For example, in the case where the target object is a person, multiple high-definition video frames of the target person all include detailed texture information such as the wrinkles and hair in the target person's hair and clothing. Or, in multiple video frames, there may also be slight changes in the left-right shaking of the target person's head. For example, the exposed parts of the ears in the reference frame and the target frame are not exactly the same. During the process of registering multiple video frames using the image registration method of the present application, the second optical flow estimation result between the edge image of the target frame and the edge image of the reference frame can be obtained in this step. Since the loss fraction of the change rate of the high-frequency information in the edge image is added to the second loss function of the optical flow estimation, it is possible to accurately estimate the dense optical flow between the edge features of the target frame and the edge features of the reference frame. Specifically, for example, it is possible to accurately calculate the optical flow of the hair, the wrinkles and hair in the clothing of the target person in the video frame, as well as the exposed parts of the facial features of the person. This can overcome the shortcoming of existing optical flow estimation algorithms that only estimate globally smooth optical flow, and can accurately estimate the non-smooth optical flow between two video frames, such as the optical flow of the exposed parts or occluded parts of the facial features between the front and rear frames, so that the accuracy of the optical flow estimation for various video frames can be ensured to be better, and the computational amount of the optical flow estimation is also smaller.
[0094] Step S190: Based on the first optical flow estimation result and the second optical flow estimation result, adjust the target frame to obtain an adjusted video frame.
[0095] As described above, the first optical flow estimation result can be the optical flow estimation result from each target frame to the reference frame obtained by using a traditional optical flow estimation algorithm. The second optical flow estimation result can be the optical flow estimation result between the edge image of each target frame and the edge image of the reference frame obtained by using the optimized optical flow estimation algorithm of the present application. It can be understood that for the same target frame and reference frame, the first optical flow estimation result and the second optical flow estimation result are different. For example, the optical flow displacement matrices obtained by the two optical flow estimations are different. The first optical flow estimation result can be the globally smoothed optical flow displacement matrix, while the second optical flow estimation result is the optical flow displacement matrix considering the frame-to-frame mutation region. In this step, the target frame can be adjusted by integrating the two optical flow estimation results.
[0096] It can be understood that for multiple video frames of a target object captured within a short period of time, the optical flow mutation region between two adjacent video frames of the target object is relatively small, or is merely a partial region in the entire image. Exemplarily, in this step, the optical flow mutation region between frames can be determined based on the difference between the results of two optical flow estimations. Then, the non-mutation regions outside the optical flow mutation region in the target frame can be adjusted using the first optical flow estimation result to align these regions with the corresponding regions in the reference frame. For the pixels located in the optical flow mutation region, other suitable methods can be adopted to adjust these pixels so as to align the optical flow mutation region with the corresponding region in the reference frame. Exemplarily but not restrictively, the optical flow mutation region in the target frame can be adjusted based on the first optical flow estimation result of the pixels in the non-mutation region near the optical flow mutation region. Finally, the effect that each target frame in multiple video frames is precisely aligned with the reference frame can be achieved, thereby enabling precise registration between multiple video frames.
[0097] It can be understood that if the optical flow between the target frame and the reference frame is estimated using a traditional optical flow estimation algorithm, there are still obvious differences between the target frame and the reference frame in the multiple video frames finally registered based on the optical flow, especially the optical flow mutation region in the image simply cannot achieve the registration effect. For example, in multiple video frames of a target person, there may be differences in the degree to which the target person's mouth opens, the degree to which the eyes open, the fitting of hair strands, or the texture of clothing. By superimposing and displaying multiple video frames registered using the traditional optical flow estimation algorithm, the dynamic changes in these optical flow mutation regions can be clearly observed, and the registration effect is poor. However, for the multiple registered video frames obtained by registering based on the results of two optical flow estimations in the embodiments of the present application, all image regions including those within the optical flow mutation region can be precisely aligned with the reference frame. By superimposing and displaying the multiple registered video frames, an effect of precise alignment in both the overall and detailed parts can be presented. For example, in the superimposed video, in addition to presenting the changes in light, there will be no micro-movement effect such as the target person's mouth opening.
[0098] In addition, the surface normal vector of the target object can be reconstructed by using the image registration method in the embodiments of the present application to obtain the external contour shape of the target object. In the prior art, photometric stereo measurement is a relatively simple method for obtaining the surface normal vector of a target object, which can be realized without using expensive photographing equipment. However, the current photometric stereo measurement method is usually only applicable to obtaining the surface normal vector of a target object that is absolutely stationary, and it is impossible to achieve accurate registration for a target object that still has dynamic changes within a short period of time, so the surface normal vector of the target object cannot be reconstructed. According to the embodiments of the present application, the above image registration method can be loaded into the image registration process of photometric stereo measurement to effectively reconstruct the surface normal vector of the target object, and both the reconstruction efficiency and accuracy are relatively high. Thus, the applicable range of the photometric stereo measurement method can be expanded.
[0099] In the above solution, by obtaining the first optical flow estimation result from each target frame to the reference frame in multiple video frames, and obtaining the second optical flow estimation result between the edge images of the target frame and the reference frame respectively, finally, based on the results of the two optical flow estimations, the precise registration of each target frame to the reference frame is realized. Moreover, since the second loss function used in the process of determining the second optical flow estimation result includes a fraction of the loss regarding the change rate of the high-frequency information in the two edge images, the optical flow for each detailed area between each target frame and the reference frame can be accurately estimated. In particular, the optical flow estimation accuracy for dynamically photographed objects at close range can be improved, and the accuracy of optical flow estimation for areas with sudden optical flow between video frames can be increased. And the computational complexity of this solution is close to that of traditional optical flow estimation algorithms, so the computational complexity is small and the processing speed is fast. Therefore, the real-time and precise registration of multiple video frames can be effectively achieved.
[0100] It can be understood that the above steps S110 to S190 can be executed in a variety of suitable execution orders. Although the steps S110 to S190 shown in Figure 1 are executed sequentially, some of these steps can also be executed in other suitable execution orders. For example, step S150 can also be executed before step S130 or can be executed simultaneously with step S130. The present application does not limit this.
[0101] Exemplarily, in step S170, the second loss function E2(u) is represented by the following formula:
[0102]
[0103] Among them, P1(x, y) represents the pixel value of the pixel (x, y) in the edge image of the target frame, and P2(x, y) represents the pixel value of the pixel (x, y) in the edge image of the reference frame. u2[x, y] represents the optical flow displacement matrix of the second optical flow estimation result of the pixel (x, y). Both β and γ are adjustment coefficients, usually fixed values, and can be arbitrarily set according to actual processing requirements. Their values can be equal or unequal.
[0104] As described above, the DIS optical flow algorithm can be used to obtain the dense optical flow estimation result from the edge image of each target frame to the edge image of the reference frame. It can be understood that in the formula of the above loss function, the first loss fraction can represent the error magnitude of matching the corresponding pixel points in the edge image of the reference frame after each pixel point in the edge image of the target frame is moved according to the optical flow displacement matrix u2[x, y]. The second loss fraction is the fraction of the loss regarding the change rate of the high-frequency information in the edge image, which represents the matching error magnitude of the change rates of the high-frequency information in the two edge images. The third loss fraction represents the smoothness inside the optical flow. The smaller the value of the third loss fraction, the smoother the optical flow, and the larger the value, the less smooth the optical flow.
[0105] It can be understood that in the above second loss function, on the basis that the quality of the matched image and the smoothness of the optical flow are both good through the first loss fraction and the third loss fraction, the second loss fraction is also used to set a penalty term to control the change rate of matching the high-frequency information in the two edge images. That is, for the detailed parts in the pictures of the target frame and the reference frame, such as human skin or clothes, through the second loss fraction in the above second loss function, not only can it be ensured that the color information of each pixel between video frames in the calculated optical flow can be accurately matched, but also the way of color change of the pixels can be accurately matched. Thus, an accurate optical flow estimation result can be obtained.
[0106] The setting method of the above second loss function is relatively simple, with a small amount of calculation and a fast processing speed. Using the above loss function, the optical flow displacement matrix of each pixel from the edge image of the target frame to the edge image of the reference frame can be calculated relatively quickly and accurately. Furthermore, real-time and accurate registration of multiple video frames can be achieved, and the user experience is good.
[0107] Exemplarily, according to the embodiments of the present application, the reference frame among multiple video frames may not be determined first, but multiple video frames are traversed, and each video frame is used as a pending reference frame for separate processing, and then the reference frame can be determined according to the processing results. Specifically, the method 100 may further include steps S121 to step S124.
[0108] Step S121: Determine each frame in at least some of the multiple video frames as a pending frame. This pending frame can be a pending reference frame.
[0109] Step S122: For each pending frame, perform the following steps S122.1 to S122.5 for each target frame of this pending frame.
[0110] Step S122.1: Based on the optical flow estimation algorithm, determine the first optical flow estimation result from this target frame to this pending frame. Step S122.2: Perform edge detection on the target frame and the pending frame respectively to obtain the edge image of the target frame and the edge image of the pending frame respectively. Step S122.3: Use the second loss function and based on the optical flow estimation algorithm, determine the second optical flow estimation result from the edge image of the target frame to the edge image of the pending frame.
[0111] It can be understood that steps S122.1 to S122.3 are respectively similar to steps S130, S150, and S170. Those of ordinary skill in the art can understand the implementation manners of these steps and will not be elaborated herein. For example, if the target object is a target human face, the multiple video frames can be multiple images of the target human face captured successively within a short period of time, such as 10 video frames captured successively. During the capture process, there may be slight micro-movement phenomena of the human face, such as slight head swaying. For example, among these 10 video frames, first take the 5th frame as the pending frame, and successively take the other 9 video frames as target frames to perform the above steps S122.1 to S122.3. The first optical flow estimation results between these 9 target frames and the 5th frame can be obtained, and correspondingly, the second optical flow estimation results between the edge images of these 9 target frames and their respective edge images to the edge image of the 5th frame can be obtained.
[0112] Step S122.4. Determine the occlusion area of the target frame relative to the pending frame based at least on the difference between the first optical flow estimation result from the target frame to the pending frame and the second optical flow estimation result from the target frame to the pending frame. It can be understood that the occlusion area of the target frame relative to the pending frame can be a pixel area that exists in the target frame but does not exist in the pending frame. It can represent the optical flow mutation area between the target frame and the pending frame. For example, among the 10 video frames of the above-mentioned target face, the 5th frame is a frontal face image, while the face in the first frame is slightly rotated to the left, so that the exposed area of the right ear of the target face in the first frame is larger than the exposed area of the right ear of the target face in the 5th frame, that is, a partial area of the right ear of the target face in the 5th frame is occluded. Then the occlusion area of the first frame relative to the 5th frame can be approximately the area where the exposed area of the right ear of the target face in the first frame is more exposed than in the 5th frame. The occlusion area can be calculated based on the difference between the first optical flow estimation result calculated for these two video frames and the second optical flow estimation result calculated for the edge images of these two video frames. Any suitable analysis method can be used to analyze the difference between the two optical flow estimation results to determine the occlusion area.
[0113] Exemplarily, taking the use of the DIS optical flow algorithm for optical flow estimation as an example, the first optical flow estimation result can be the globally smoothed optical flow displacement matrix, and the second optical flow estimation result is the optical flow displacement matrix considering the inter-frame mutation area. In this step, the occlusion area can be directly determined based on the difference matrix of these two optical flow displacement matrices. For example, a first threshold can be set, and the area formed by the pixels corresponding to the determinant value of the difference matrix of the two images being greater than the first threshold can be determined as the occlusion area. Alternatively, an index value representing the optical flow smoothness of each pixel position in the image can be calculated based on the two optical flow displacement matrices, and the area formed by the pixels with poor optical flow smoothness can be determined as the occlusion area of the target frame relative to the pending frame according to the magnitude of the index value.
[0114] In a specific example, step S122.4 may include step S122.41.
[0115] Step S122.41. Determine the area where the pixels (x, y) in the target frame that satisfy the following formula are located as the occlusion area of the target frame relative to the pending frame:
[0116]
[0117] Where u1′[x,y] represents the optical flow displacement matrix of the first optical flow estimation result of the pixel (x, y). u2′[x,y] represents the optical flow displacement matrix of the second optical flow estimation result of the pixel (x, y). η’ represents the adjustment coefficient, and η’ can be set to any value greater than zero according to actual processing requirements. T’ represents the optical flow threshold, and this optical flow threshold can also be an empirical value set according to actual needs.
[0118] It can be understood that in the above formula, |u1′[x,y] - u2′[x,y]| represents the determinant value of the difference matrix between the two optical flow displacement matrices obtained from the two optical flow estimations. It can represent the difference magnitude of the two optical flow estimation results for each pixel in the video frame. It represents the magnitude of the sum of the change rates of the two optical flow displacements of each pixel in the statistics, which can represent the internal smoothness of the optical flow of the pixel to a certain extent. The larger the value, the less smooth the optical flow of the pixel. Through the above formula, the pixels with large differences in the two optical flow estimation results and unsmooth optical flow in the video frame can be screened out, and thus the occlusion area of the target frame relative to the pending frame can be accurately determined.
[0119] Step S122.5, calculate the area of the occlusion area of the target frame relative to the pending frame. After determining the occlusion area of each target frame relative to the corresponding pending frame through the above steps, any suitable method can be used to calculate the area of the occlusion area. Exemplarily but not restrictively, the number of pixels included in the occlusion area can be used as the area of the occlusion area. Alternatively, other methods can also be used to determine the area of the occlusion area. For example, the area of the minimum bounding rectangle of the occlusion area can be used as the area of the occlusion area, etc.
[0120] Step S123, calculate the sum of the areas of the occlusion areas of each target frame of the pending frame relative to the pending frame. In the embodiment where the 5th frame among the above 10 video frames is the pending frame, the areas of the occlusion areas of the remaining 9 video frames relative to the 5th frame can be calculated using the above step S122, and the areas of 9 occlusion areas can be obtained. In this step, the sum of the areas of these 9 occlusion areas can be calculated.
[0121] Step S124, compare the sums of the areas calculated for each pending frame respectively, and determine the pending frame corresponding to the smallest sum of the areas as the reference frame. In the embodiment where the above multiple video frames are 10 video frames of a target face, the first sum of the areas calculated with the 1st frame as the pending frame, the second sum of the areas calculated with the 2nd frame as the pending frame... the tenth sum of the areas calculated with the 10th frame as the pending frame can be calculated respectively. The minimum value among these 10 sums of the areas can be obtained, and the pending frame corresponding to the minimum value is determined as the reference frame. For example, if the fifth sum of the areas is the minimum value among the 10 sums of the areas, then the 5th frame can be determined as the reference frame. Those of ordinary skill in the art can understand the implementation manner and various extended manners of this solution, which will not be elaborated herein.
[0122] In the above scheme, multiple video frames can be traversed, and each video frame can be used as a pending frame. The first optical flow estimation result between the target frame of each pending frame and the pending frame and the second optical flow estimation result between the edge maps of the two are determined respectively. Then, the occlusion area and the occlusion area area of the target frame relative to the pending frame are determined based on the difference between the two optical flow estimation results, and then the area sum of the occlusion areas of all target frames relative to their pending frames is calculated. By comparing the size of the area sum corresponding to each pending frame, the pending frame with the smallest area sum is determined as the reference frame. Through this scheme, the optimal reference frame among multiple video frames can be accurately determined, and the target frame can be registered based on the optimal reference frame in subsequent steps. For example, for multiple video frames of a face, if the multiple video frames are respectively the face video frames taken when the left face gradually turns to the right face, the video frame corresponding to the front face can be determined as the reference frame through the above scheme for determining the reference frame. Therefore, the scheme not only helps to improve the accuracy of image registration, but also helps to ensure the better visual effect of each video frame after registration, so that the user experience is better.
[0123] In another example, in order to save the amount of calculation, the above step S121 can also use the video frames that meet the preset timing requirements of the multiple video frames as pending frames. The video frames that meet the preset timing requirements can be video frames captured at the middle moment in the multiple video frames. For example, if the multiple video frames are 11 video frames, then in step S121, each video frame in the 6th ± 2nd video frame can also be used as a pending frame, that is, the 4th frame to the 8th frame can be used as the pending frames to perform the above steps S122 to S124.
[0124] Since the dynamic changes of the target object in multiple video frames shot successively are often stable and continuous, the pending frame corresponding to the minimum area is usually located near the middle frame. This solution does not need to calculate each video frame as a pending frame, which can effectively reduce the amount of calculation and improve the processing speed while ensuring accuracy.
[0125] Exemplarily, the method 100 further includes step S180.
[0126] Step S180, at least based on the difference between the first optical flow estimation result and the second optical flow estimation result, determine the occlusion area and the non-occlusion area of the target frame relative to the reference frame. The occlusion area of the target frame relative to the reference frame can be determined by a method similar to the aforementioned step S122.4, which will not be repeated here. The non-occlusion area can be other areas in the target frame except the occlusion area. It can be understood that the occlusion area can be an area where the optical flow changes suddenly between two video frames, while the non-occlusion area can be an area where the optical flow is relatively smooth.
[0127] In one example, step S180 determines the occluded area and non-occluded area of the target frame relative to the reference frame based at least on the difference between the first optical flow estimation result and the second optical flow estimation result, including step S180.1 and step S180.2.
[0128] In step S180.1, the area where the pixel (x, y) satisfying the following formula is located is determined as the occluded area of the target frame relative to the reference frame: Where u1[x,y] represents the optical flow displacement matrix of the first optical flow estimation result of the pixel (x, y), u2[x,y] represents the optical flow displacement matrix of the second optical flow estimation result of the pixel (x, y), η represents an adjustment coefficient, which can be set to any value greater than zero according to actual processing requirements. T represents the optical flow threshold. This step is similar to the solution of the foregoing step S122.41, and those of ordinary skill in the art can understand the implementation manner of this step according to its description, and will not be elaborated herein. In step S180.2, the area outside the occluded area in the target frame is determined as the non-occluded area, and this non-occluded area can be an area where the optical flow in the target frame relative to the reference frame is relatively smooth.
[0129] It can be understood that through the above solution, the pixels with large differences between the two optical flow estimation results and uneven optical flow in the target frame can be screened out, and thus the occluded area of the target frame relative to the reference frame can be accurately determined. In other words, the area of optical flow mutation in the target frame relative to the reference frame can be determined more accurately. Furthermore, the area where the optical flow in the target frame relative to the reference frame is relatively smooth can be accurately determined. In addition, the computational amount of this solution is also small, and the execution logic is relatively simple.
[0130] Step S190 adjusts the target frame based on the first optical flow estimation result and the second optical flow estimation result to obtain an adjusted video frame, including step S191 and step S192.
[0131] In step S191, for each first pixel in the non-occluded area of the target frame, the first pixel is directly adjusted according to the first optical flow estimation result of the first pixel. The pixels located in the non-occluded area can be called first pixels. As described above, the non-occluded area can be an area where the optical flow is relatively smooth. Based on the traditional optical flow estimation algorithm, the optical flow of this area can be determined more accurately. Therefore, each pixel in this area can be directly registered to the corresponding pixel in the reference frame based on the first optical flow estimation result. For example, each first pixel in the target frame can be adjusted based on the optical flow displacement matrix u1[x,y] in the foregoing example. The position coordinates of each first pixel can be substituted into u1[x,y] to obtain the optical flow displacement of each first pixel, so that each first pixel can be moved by the corresponding displacement in the target frame.
[0132] Step S192: For each second pixel in the occluded region of the target frame, adjust the second pixel according to the first optical flow estimation result of other pixels in the target frame. Each pixel located in the occluded region can be called a second pixel. Since the occluded region is an optical flow mutation region, it is impossible to accurately determine the corresponding position of this pixel in the reference frame. Therefore, the optical flow displacement can be determined based on the first optical flow estimation result of other pixels outside this pixel for registration with the reference frame. Optionally, the optical flow displacement can be determined based on the first optical flow estimation result of the first pixels near the second pixel, and then the position of this pixel can be adjusted based on this displacement. Alternatively, the optical flow displacement can also be determined based on the first optical flow estimation result of other second pixels near the second pixel, and then the position of this pixel in the target frame can be adjusted. For example, the optical flow displacements of the first optical flow estimation results of all second pixels can be averaged, and each second pixel can be adjusted based on the average displacement. Of course, other suitable methods can also be used to adjust the second pixel.
[0133] In the above solution, the occluded region and non-occluded region of the target frame relative to the reference frame can be determined first based on the difference between the first optical flow estimation result and the second optical flow estimation result. Then, different methods are used to adjust the pixels in these two regions respectively to register the target frame with the reference frame. For each pixel in the region with relatively smooth optical flow represented by the non-occluded region, the first optical flow estimation result is directly used for adjustment. For the region with optical flow mutation represented by the non-occluded region, the first optical flow estimation result of other pixels is used for adjustment. In this solution, the registration accuracy of the target frame with the reference frame is relatively high, and the calculation amount is also small.
[0134] Exemplarily, step S192 of adjusting the second pixel according to the first optical flow estimation result of other pixels in the target frame includes step S192.1.
[0135] Step S192.1: For each second pixel in the occluded region of the target frame, adjust the second pixel according to the first optical flow estimation result of at least one first pixel in the occluded region of the target frame that is closest to the second pixel. Optionally, the second pixel can be directly adjusted based on the first optical flow estimation result of the first pixel closest to the second pixel. Alternatively, the second pixel can also be adjusted based on the first optical flow estimation result of a specific pixel among the multiple pixels closest to the second pixel. This specific pixel is, for example, the pixel with the pixel value closest to that of the second pixel. Alternatively, the second pixel can also be adjusted based on the first optical flow estimation results of the multiple pixels closest to the second pixel. For example, the optical flow displacements of the first optical flow estimation results of the 10 pixels closest to the second pixel can be weighted and averaged, and the weighted average optical flow displacement is used as the optical flow displacement of this second pixel for adjustment.
[0136] It can be understood that since the optical flow of each first pixel located in the non-occluded area is relatively smooth, the optical flow estimation of these pixels is relatively accurate. Since the dynamic changes of the target object are stable and continuous, for two adjacent pixels in the target frame, their positions in the reference frame are also likely to be adjacent. Therefore, based on the relatively accurate optical flow estimation result of the pixel closest to the second pixel, the optical flow of each second pixel can be determined relatively more accurately. Furthermore, by adjusting each second pixel based on the determined optical flow, the occluded area in the target frame can also be accurately aligned with the reference frame, thereby significantly improving the accuracy of image registration. Moreover, this solution is simple in calculation and small in computational amount, so the processing efficiency is also relatively high.
[0137] Exemplarily, step S192.1 adjusts the second pixel according to the first optical flow estimation results of at least one first pixel in the occluded area of the target frame that is closest to the second pixel, and includes step S192.11, step S192.12, and step S192.13.
[0138] Step S192.11 determines a plurality of first pixels in the occluded area of the target frame that are closest to the second pixel. For example, the plurality of pixels is 10 pixels, and in this step, 10 first pixels closest to each second pixel can be determined.
[0139] Step S192.12 performs distance reciprocal weighted averaging on the first optical flow estimation results of the determined plurality of first pixels according to the distances between each determined first pixel and the second pixel, and determines the optical flow displacement of the second pixel. For example, the optical flow displacements of the first optical flow estimation results of 10 first pixels can be subjected to distance reciprocal weighted averaging. That is, among the 10 first pixels, the closer the first pixel is to the second pixel, the greater the weight of its optical flow displacement, and the farther the first pixel is from the second pixel, the smaller the weight of its optical flow displacement. Thus, the accuracy of the determined optical flow displacement of the second pixel can be guaranteed to a relatively high degree.
[0140] Step S192.13 adjusts the second pixel according to the optical flow displacement of the second pixel. For example, in the target frame, the second pixel can be moved to a new position according to the optical flow displacement.
[0141] The above method for adjusting the second pixel is more accurate and reasonable, and can further improve the accuracy of image registration. Moreover, the computational amount is also small, so the processing efficiency is also relatively high.
[0142] In the prior art, the implementation cost of some methods for reconstructing the surface normal vector of a target object using relatively complex devices is extremely high. Some others register multiple video frames of the target object captured successively within a short period of time to reconstruct the surface normal vector of the target object. Taking photometric stereo as an example, an image acquisition device and a group of programmatically controllable point light sources can be used. During the process of photographing the target object, each point light source is sequentially lit, and multiple video frames with different brightnesses of the target object under different light sources are obtained. Then, the multiple video frames are registered to obtain the surface normal vector of the target object. However, photometric stereo often requires the target object to be absolutely stationary during the photographing process, otherwise it will be difficult to achieve registration and thus impossible to reconstruct the surface normal vector of the target object. Therefore, in the prior art, this method is mostly used to reconstruct the surface normal vector of static objects, rather than dynamic objects such as the human body. This is because dynamic objects cannot remain absolutely stationary even within a short period of time. Therefore, the applicable range of the photometric stereo method in the prior art for reconstructing the surface normal vector of a target object is relatively narrow.
[0143] According to the second aspect of the present application, a method for reconstructing a surface normal vector is further provided. Figure 4 The schematic flowchart of a method 400 for reconstructing a surface normal vector according to an embodiment of the present application is shown. As shown, the reconstruction method 400 includes step S410 and step S420.
[0144] In step S410, the above image registration method 100 is used to register multiple video frames of the target object. Exemplarily, the multiple video frames can be captured by using the same video frame acquisition device. For example, the above image registration method can be used to register consecutive video frames of the target human face captured by the same camera at different light source positions successively.
[0145] In step S420, the surface normal vector of the target object is reconstructed by using the adjusted video frames. For example, based on the consecutive video frames of the registered target human face, the surface normal vector of the target human face can be reconstructed by using photometric stereo.
[0146] It can be understood that since the image registration method according to the embodiment of the present application can achieve the optical flow estimation accuracy of video frames of various target objects including dynamic objects, the registration accuracy of multiple video frames of dynamic objects can be improved. In particular, the applicable range of surface normal vector reconstruction methods such as photometric stereo measurement methods can be significantly increased. In addition, the computational amount of this method is also small, and the reconstruction effect is also high.
[0147] According to the third aspect of the present application, an image registration system is further provided. Figure 5The schematic block diagram of an image registration system 500 according to an embodiment of the present application is shown. As shown in the figure, the image registration system 500 includes an acquisition module 510, a first determination module 520, an edge detection module 530, a second determination module 540, and an adjustment module 550.
[0148] The acquisition module 510 is configured to acquire a plurality of video frames.
[0149] The first determination module 520 is configured to determine a first optical flow estimation result from a target frame to a reference frame among the plurality of video frames based on an optical flow estimation algorithm.
[0150] The edge detection module 530 is configured to perform edge detection on the target frame and the reference frame respectively to obtain an edge video frame of the target frame and an edge video frame of the reference frame respectively.
[0151] The second determination module 540 is configured to determine a second optical flow estimation result from the edge video frame of the target frame to the edge video frame of the reference frame by using a second loss function and based on an optical flow estimation algorithm, where the second loss function includes a fraction of the loss regarding the change rate of high-frequency information in the edge video frame.
[0152] The adjustment module 550 is configured to adjust the target frame based on the first optical flow estimation result and the second optical flow estimation result to obtain an adjusted video frame.
[0153] According to the fourth aspect of the present application, a surface normal vector reconstruction system is further provided. Figure 6 The schematic block diagram of a surface normal vector reconstruction system 600 according to an embodiment of the present application is shown. As shown in the figure, the reconstruction system 600 includes a registration module 610 and a reconstruction module 620.
[0154] The registration module 610 is configured to register a plurality of video frames of a target object by using the above image registration method 100. Exemplarily, the plurality of video frames may be acquired by using the same video frame acquisition device.
[0155] The reconstruction module 620 is configured to reconstruct the surface normal vector of the target object by using the adjusted video frame.
[0156] According to the fifth aspect of the present application, an electronic device is further provided. Figure 7 The schematic block diagram of an electronic device 700 according to an embodiment of the present application is shown. As shown in the figure, the electronic device 700 includes a processor 710 and a memory 720. Among them, computer program instructions are stored in the memory 720, and when the computer program instructions are run by the processor 710, they are used for the above image registration method 100 and / or the above surface normal vector reconstruction method 400.
[0157] According to a sixth aspect of the present application, a storage medium is further provided, on which program instructions are stored, and the program instructions are used to execute the above image registration method 100 and / or the above surface normal vector reconstruction method 400 when running.
[0158] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0159] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.
[0160] Similarly, it should be understood that, in order to streamline the present application and help understand one or more of the various inventive aspects, in the description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the method of the present application should not be construed as reflecting the intention that the claimed present application requires more features than those expressly recited in each claim. Rather, as reflected by the corresponding claims, the inventive point lies in being able to solve the corresponding technical problems with fewer features than all the features of a single disclosed embodiment. Therefore, the claims following the specific implementation manner are hereby expressly incorporated into the specific implementation manner, where each claim itself serves as a separate embodiment of the present application.
[0161] Those skilled in the art can understand that, except for mutual exclusion between features, any combination can be adopted for all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.
[0162] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some modules in the image registration system or the surface normal vector reconstruction system according to the embodiments of the present application. The present application can also be implemented as a device program (for example, a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0163] It should be noted that the above embodiments illustrate the present application rather than limit the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.
[0164] As described above, it is only the specific implementation manner of the present application or the description of the specific implementation manner, and the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all should be covered by the protection scope of the present application. The protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. An image registration method, comprising: Obtaining a plurality of video frames; Based on an optical flow estimation algorithm, determining a first optical flow estimation result from a target frame to a reference frame among the plurality of video frames; Performing edge detection on the target frame and the reference frame respectively to obtain an edge image of the target frame and an edge image of the reference frame respectively; Using a second loss function and based on the optical flow estimation algorithm, determining a second optical flow estimation result from the edge image of the target frame to the edge image of the reference frame, wherein the second loss function includes a fraction of the loss regarding the change rate of high-frequency information in the edge image; And Based on the first optical flow estimation result and the second optical flow estimation result, adjusting the target frame to obtain an adjusted video frame.
2. The image registration method according to claim 1, wherein, The second loss function E2(i) is represented by the following formula: Wherein, P1(x,y) represents the pixel value of the pixel (x, y) in the edge image of the target frame, P2(x,y) represents the pixel value of the pixel (x, y) in the edge image of the reference frame, u2[x,y] represents the optical flow displacement matrix of the second optical flow estimation result of the pixel (x, y), and β and γ respectively represent adjustment coefficients.
3. The image registration method according to claim 1 or 2, wherein, The method further includes: Determining at least an occlusion area and a non-occlusion area of the target frame relative to the reference frame based on a difference between the first optical flow estimation result and the second optical flow estimation result; The adjusting the target frame based on the first optical flow estimation result and the second optical flow estimation result to obtain an adjusted video frame includes: For each first pixel in the non-occlusion area of the target frame, directly adjusting the first pixel according to the first optical flow estimation result of the first pixel; For each second pixel in the occlusion area of the target frame, adjusting the second pixel according to the first optical flow estimation result of other pixels in the target frame.
4. The image registration method according to claim 3, wherein, The adjusting the second pixel according to the first optical flow estimation result of other pixels in the target frame includes: For each second pixel in the occlusion area of the target frame, adjusting the second pixel according to the first optical flow estimation result of at least one first pixel closest to the second pixel in the occlusion area of the target frame.
5. The image registration method according to claim 4, wherein, The adjusting the second pixel according to the first optical flow estimation result of at least one first pixel closest to the second pixel in the occlusion area of the target frame includes: Determining a plurality of first pixels closest to the second pixel in the occlusion area of the target frame; According to the distance between each determined first pixel and the second pixel, performing distance reciprocal weighted averaging on the first optical flow estimation results of the determined plurality of first pixels to determine the optical flow displacement of the second pixel; and Adjusting the second pixel according to the optical flow displacement of the second pixel.
6. The image registration method according to claim 3, wherein, The determining at least an occlusion area and a non-occlusion area of the target frame relative to the reference frame based on a difference between the first optical flow estimation result and the second optical flow estimation result includes: Determining the area where the pixel (x, y) satisfying the following formula is located as the occlusion area of the target frame relative to the reference frame: Among them, u1[x, y] represents the optical flow displacement matrix of the first optical flow estimation result of the pixel (x, y), u2[x, y] represents the optical flow displacement matrix of the second optical flow estimation result of the pixel (x, y), η represents an adjustment coefficient, η > 0, T represents an optical flow threshold; and Determine the area outside the occlusion area as the non-occlusion area.
7. The image registration method according to claim 1 or 2, wherein, The method further includes: Determine at least some frames in the multiple video frames as pending frames respectively; For each pending frame, perform the following operations for each target frame of the pending frame, Based on the optical flow estimation algorithm, determine the first optical flow estimation result from the target frame to the pending frame; Perform edge detection on the target frame and the pending frame respectively to obtain the edge image of the target frame and the edge image of the pending frame respectively; Using the second loss function and based on the optical flow estimation algorithm, determine the second optical flow estimation result from the edge image of the target frame to the edge image of the pending frame; Determine the occlusion area of the target frame relative to the pending frame based at least on the difference between the first optical flow estimation result from the target frame to the pending frame and the second optical flow estimation result from the target frame to the pending frame; Calculate the area of the occlusion area of the target frame relative to the pending frame; Calculate the sum of the areas of the occlusion areas of each target frame of the pending frame relative to the pending frame; and Compare the sums of the areas calculated for each pending frame respectively, and determine the pending frame corresponding to the smallest sum of the areas as the reference frame.
8. The image registration method according to claim 1 or 2, wherein The method further includes: Determine the middle frame in the multiple video frames as the reference frame.
9. The image registration method according to claim 1 or 2, wherein, The optical flow estimation algorithm includes a two-frame differential optical flow estimation algorithm and a dense inverse search optical flow estimation algorithm.
10. A surface normal vector reconstruction method, including: Using the image registration method according to any one of claims 1 to 9, register multiple video frames of a target object; Reconstruct the surface normal vector of the target object using the adjusted video frames.
11. An image registration system, including: An acquisition module, configured to acquire multiple video frames; A first determination module, configured to determine the first optical flow estimation result from the target frame to the reference frame in the multiple video frames based on the optical flow estimation algorithm; An edge detection module, configured to perform edge detection on the target frame and the reference frame respectively to obtain the edge image of the target frame and the edge image of the reference frame respectively; A second determination module, configured to use the second loss function and based on the optical flow estimation algorithm, determine the second optical flow estimation result from the edge image of the target frame to the edge image of the reference frame, where the second loss function includes a fraction of the loss regarding the change rate of the high-frequency information in the edge image; And An adjustment module, configured to adjust the target frame based on the first optical flow estimation result and the second optical flow estimation result to obtain an adjusted video frame.
12. A surface normal vector reconstruction system, including: A registration module, configured to use the image registration method according to any one of claims 1 to 9 to register multiple video frames of a target object; And A reconstruction module, configured to reconstruct the surface normal vector of the target object using the adjusted video frames.
13. An electronic device, comprising a processor and a memory, wherein, The computer program instructions are stored in the memory, and when the computer program instructions are run by the processor, they are used to execute the image registration method according to any one of claims 1 to 9 and / or the surface normal vector reconstruction method according to claim 10.
14. A storage medium, on which program instructions are stored, and when the program instructions are run, they are used to execute the image registration method according to any one of claims 1 to 9 and / or the surface normal vector reconstruction method according to claim 10.
Citation Information
Patent Citations
Video blind denoising method and device based on deep learning
CN111539879A
Automatic object recognition method and system thereof, shopping device and storage medium
US20190303650A1