Image processing method, device, electronic device and storage medium
By improving the grid optical flow algorithm for image photometric error, constructing the objective function and optimization matrix for iterative optimization, the problems of low alignment accuracy and poor real-time performance in mobile phone video shooting are solved, the risk of mobile phone heating is reduced, and the efficiency of video image alignment is improved.
Patent Information
- Application Number
- CN202111612276.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2041-12-27
AI Technical Summary
Existing video image alignment methods have problems such as low alignment accuracy, poor real-time performance, and easy heating of the phone in mobile phone video shooting.
A grid optical flow algorithm based on image photometric error is adopted. By constructing the objective function and optimization matrix, the grid vertex pairs are iteratively optimized, the derivation method of calculating the image gradient derivative is reduced, the photometric error process is optimized, the amount of calculation is reduced, and the real-time performance and alignment accuracy are improved.
While ensuring the accuracy of video image alignment, it significantly reduces processing time and the probability of mobile phone heating, and improves the real-time performance and computing speed of the algorithm.
Smart Images

Figure CN114331816B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and more specifically, to an image processing method, device, electronic device, and storage medium. Background Art
[0002] When shooting mobile videos, to obtain clear, stable, and realistic images, it is often necessary to align the video images during the recording process. However, existing video image alignment methods either have low alignment accuracy, poor real-time performance, and can easily cause the phone to overheat. Summary of the Invention
[0003] The present disclosure provides an image processing method, apparatus, electronic device, and storage medium to at least solve at least one problem in the above-mentioned related technologies.
[0004] According to a first aspect of an embodiment of the present disclosure, there is provided an image processing method, comprising: acquiring a first image and a second image in a video, wherein the first image is a previous frame image of the second image, the first image and the second image are respectively divided into a plurality of first grid areas and a plurality of second grid areas corresponding to each other, the corresponding first grid areas and the corresponding second grid areas forming a grid area pair, and the corresponding first grid vertices in the first grid areas and the second grid vertices in the second grid areas forming a grid vertex pair; constructing an objective function for each grid vertex pair, wherein the objective function characterizes a displacement error between the second grid vertex and the first grid vertex in the current grid vertex pair; determining an optimization matrix for the current grid vertex pair using at least one first grid area near the current grid vertex pair; iteratively optimizing the objective function using the optimization matrix to obtain a displacement adjustment amount of the second grid vertex in the current grid vertex pair relative to the corresponding first grid vertex; and adjusting the position of the second grid vertex in each grid vertex pair relative to the corresponding first grid vertex by the corresponding displacement adjustment amount to obtain a target second image aligned with the first image.
[0005] Optionally, before the first image and the second image are respectively divided into a plurality of first grid areas and a plurality of second grid areas corresponding one to one, the method includes: removing a first dynamic image area in the first image and a second dynamic image area in the second image, respectively, the first dynamic image area corresponding to the second dynamic image area; the first image and the second image are respectively divided into a plurality of first grid areas and a plurality of second grid areas corresponding one to one, including: an area in the first image other than the first dynamic image area is divided into a plurality of first grid areas; an area in the second image other than the second dynamic image area is divided into a plurality of second grid areas, the plurality of second grid areas corresponding one to one to the plurality of first grid areas.
[0006] Optionally, constructing the objective function includes: constructing the objective function based on photometric error information and deformation error information, wherein the photometric error information represents the photometric error between the sampling pixel points in at least one first grid area near the current grid vertex pair and the sampling pixel points in the corresponding second grid area, and the deformation error information represents the deformation error between the first grid vertex and the second grid vertex in the current grid vertex pair.
[0007] Optionally, the position of the sampling pixel point in each first grid area is obtained by interpolating all first grid vertices in each first grid area.
[0008] Optionally, the iteratively optimizing the objective function through the optimization matrix to obtain the displacement adjustment amount of the second mesh vertex in the current mesh vertex pair relative to the corresponding first mesh vertex includes: obtaining the current displacement adjustment amount of the second mesh vertex in the current mesh vertex pair undergoing iterative optimization; calculating, based on the current displacement adjustment amount, an update amount of the current displacement adjustment amount of the second mesh vertex in the current mesh vertex pair undergoing iteration through the optimization matrix; adjusting the update amount on the current displacement adjustment amount to obtain an updated displacement adjustment amount; using the updated displacement adjustment amount as the current displacement adjustment amount, repeating the steps of calculating the update amount and obtaining the updated displacement adjustment amount until an iteration end condition is met, thereby obtaining the displacement adjustment amount of the second mesh vertex in the current mesh vertex pair relative to the corresponding first mesh vertex.
[0009] Optionally, the step of repeatedly executing the step of calculating the update amount and the step of obtaining the updated displacement adjustment amount until an iteration end condition is reached includes: when the value of the update amount is less than or equal to a preset threshold, determining that the iteration end condition is reached.
[0010] Optionally, adjusting the position of the second mesh vertex in each mesh vertex pair relative to the corresponding first mesh vertex by a corresponding displacement adjustment amount to obtain a target image aligned with the first image includes:
[0011] adding the position coordinates of the first mesh vertex in each mesh vertex pair to the corresponding displacement adjustment amount to obtain a pre-adjusted position of the second mesh vertex in each mesh vertex pair in the second image;
[0012] The position of each second mesh vertex is adjusted to the pre-adjusted position to obtain a target image aligned with the first image.
[0013] According to a second aspect of an embodiment of the present disclosure, an image processing apparatus is provided, comprising: an image acquisition unit, configured to: acquire a first image and a second image in a video, wherein the first image is a previous frame image of the second image, the first image and the second image are respectively divided into a plurality of first grid areas and a plurality of second grid areas corresponding to each other, the corresponding first grid areas and the second grid areas constitute a grid area pair, and the corresponding first grid vertex in the first grid area and the second grid vertex in the second grid area constitute a grid vertex pair; an objective function construction unit, configured to: construct an objective function for each of the grid vertex pairs, wherein the objective function table The method includes: determining a displacement error between a second mesh vertex and a first mesh vertex in a current mesh vertex pair; an optimization matrix determining unit configured to determine an optimization matrix for the current mesh vertex pair using at least one first mesh region near the current mesh vertex pair; a displacement adjustment amount acquiring unit configured to iteratively optimize the objective function using the optimization matrix to obtain a displacement adjustment amount of the second mesh vertex in the current mesh vertex pair relative to the corresponding first mesh vertex; and a position adjusting unit configured to adjust the position of the second mesh vertex in each mesh vertex pair relative to the corresponding first mesh vertex by the corresponding displacement adjustment amount to obtain a target second image aligned with the first image.
[0014] Optionally, the image processing unit also includes an image removal unit, which can be configured to: remove the first dynamic image area in the first image and the second dynamic image area in the second image before the first image and the second image are respectively divided into a plurality of first grid areas and a plurality of second grid areas corresponding to each other, and the first dynamic image area corresponds to the second dynamic image area; the image acquisition unit can be configured to: divide the area of the first image other than the first dynamic image area into a plurality of first grid areas; and divide the area of the second image other than the second dynamic image area into a plurality of second grid areas, and the plurality of second grid areas correspond to the plurality of first grid areas one-to-one.
[0015] Optionally, the objective function construction unit can be configured to: construct the objective function based on photometric error information and deformation error information, the photometric error information represents the photometric error between the sampling pixel points in at least one first grid area near the current grid vertex pair and the sampling pixel points in the corresponding second grid area, and the deformation error information represents the deformation error between the first grid vertex and the second grid vertex in the current grid vertex pair.
[0016] Optionally, the position of the sampling pixel point in each first grid area is obtained by interpolating all first grid vertices in each first grid area.
[0017] Optionally, the displacement adjustment amount acquisition unit can be configured to: acquire the current displacement adjustment amount of the second mesh vertex in the current mesh vertex pair undergoing iterative optimization; calculate, based on the current displacement adjustment amount, an update amount of the current displacement adjustment amount of the second mesh vertex in the current mesh vertex pair undergoing iterative optimization through the optimization matrix; adjust the update amount on the current displacement adjustment amount to obtain an updated displacement adjustment amount; use the updated displacement adjustment amount as the current displacement adjustment amount, and repeat the steps of calculating the update amount and obtaining the updated displacement adjustment amount until the iteration end condition is met, thereby obtaining the displacement adjustment amount of the second mesh vertex in the current mesh vertex pair relative to the corresponding first mesh vertex.
[0018] Optionally, the displacement adjustment amount acquisition unit may be further configured to: determine that the iteration end condition is met when the value of the update amount is less than or equal to a preset threshold.
[0019] Optionally, the position adjustment unit can be configured to: add the position coordinates of the first mesh vertex in each mesh vertex pair to the corresponding displacement adjustment amount to obtain a pre-adjusted position of the second mesh vertex in each mesh vertex pair in the second image; and adjust the position of each second mesh vertex to the pre-adjusted position to obtain a target image aligned with the first image.
[0020] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory storing computer-executable instructions, wherein the computer-executable instructions, when executed by the at least one processor, prompt the at least one processor to execute the image processing method according to the present disclosure.
[0021] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium storing instructions is provided, characterized in that when the instructions are executed by at least one processor, the at least one processor is prompted to execute the image processing method according to the present disclosure.
[0022] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, wherein instructions in the computer program product can be executed by a processor of a computer device to complete the image processing method according to the present disclosure.
[0023] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects:
[0024] According to the image processing method, device, electronic device and storage medium disclosed in the present invention, in the process of video image alignment using a grid optical flow algorithm based on image photometric error, the derivation method in the optimization process of photometric error is improved, and there is no need to repeatedly calculate the image gradient derivative, thereby greatly reducing the computational complexity of the algorithm and accelerating the running speed of the algorithm. Thus, while taking into account the accuracy of video image alignment, the processing time of video image alignment can be reduced, the probability of mobile phone heating due to video image alignment is reduced, and the real-time performance is better.
[0025] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0027] Figure 1 is a flowchart illustrating an image processing method according to an exemplary embodiment of the present disclosure.
[0028] Figure 2 is a schematic diagram illustrating a first image and a second image in an image alignment process according to an exemplary embodiment of the present disclosure.
[0029] Figure 3 is a block diagram illustrating an image processing apparatus according to an exemplary embodiment of the present disclosure.
[0030] Figure 4 is a block diagram illustrating an electronic device 400 according to an exemplary embodiment of the present disclosure.
[0031] Figure 5 is a block diagram illustrating a computer-readable storage medium 500 according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0032] In order to enable ordinary people in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0033] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation methods described in the following examples do not represent all implementation methods consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.
[0034] It should be noted that the phrase "at least one of the items" in this disclosure includes three types of parallel situations: "any one of the items", "a combination of any multiple items of the items", and "all of the items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. For another example, "performing at least one of step 1 and step 2" includes the following three parallel situations: (1) performing step 1; (2) performing step 2; and (3) performing steps 1 and 2.
[0035] When shooting mobile videos, image alignment is often required to obtain clear, stable, and realistic images. Grid optical flow, as a video alignment algorithm, is widely used in scenarios such as video stabilization, video completion, and video tracking during mobile video shooting. This requires the algorithm to complete alignment within a specified timeframe while ensuring accurate alignment. It must also run for extended periods without overheating the phone and requiring too much memory.
[0036] Currently, there are three main approaches to grid optical flow algorithms: one based on optical flow point tracking and clustering, the other based on image photometric error tracking, and the third based on deep learning. The first approach is computationally inefficient but requires rich video image textures. It struggles to achieve good results in areas with weak textures, and alignment accuracy is low. The second approach, based on image photometric error tracking, addresses weak texture tracking while ensuring alignment accuracy. For example, in an existing grid optical flow calculation method based on image photometric error, the image is divided into M*N grids. The pixel photometric errors within the grids in adjacent image frames are used to calculate a one-to-one mapping between grids. This is then achieved by using a single mapping matrix within each grid to achieve inter-frame image alignment. However, this approach is computationally expensive, and the iterative least squares optimization process requires repeated calculations of image gradient derivatives, resulting in poor real-time image alignment. The third approach, based on deep learning, is computationally expensive, making real-time alignment difficult on mobile phones and easily causing the phone to overheat.
[0037] In order to speed up the alignment of video images and reduce the probability of mobile phones heating up during the alignment process, the present disclosure proposes an image processing method, device, electronic device and storage medium. Specifically, in the process of video image alignment using a grid optical flow algorithm based on image photometric error, the derivation method in the optimization process of photometric error is improved, and there is no need to repeatedly calculate the image gradient derivative, thereby greatly reducing the amount of calculation of the algorithm and speeding up the algorithm. Thus, while taking into account the accuracy of video image alignment, the processing time of video image alignment can be reduced, the probability of mobile phones heating up due to video image alignment is reduced, and the real-time performance is better. Below, reference will be made to Figures 1 to 4 An image processing method, apparatus, electronic device, and storage medium according to exemplary embodiments of the present disclosure are described in detail.
[0038] Figure 1 is a flowchart illustrating an image processing method according to an exemplary embodiment of the present disclosure.
[0039] Reference Figure 1In step 101, a first image and a second image in a video are obtained, wherein the first image is a previous frame image of the second image, and the first image and the second image are respectively divided into a plurality of first grid areas and a plurality of second grid areas corresponding to each other. The corresponding first grid areas and second grid areas constitute a grid area pair, and the corresponding first grid vertices in the first grid areas and the second grid vertices in the second grid areas constitute a grid vertex pair.
[0040] During video capture, real-time image alignment of video frames is often required to obtain clearer and more stable video images. During image alignment, the first and second images in the video can be first acquired. Here, the first image is the previous frame of the second image. The first image is used as the reference frame for image alignment (i.e., the alignment target), while the second image is the frame to be optimized for alignment with the first image. By constructing an objective function between the second and first images and optimizing the objective function using a least squares optimization algorithm, the second image can be aligned with the first image.
[0041] According to an exemplary embodiment of the present disclosure, in order to reduce the amount of calculation for image alignment, the first dynamic image area in the first image and the second dynamic image area in the second image can be first removed respectively, where the first dynamic image area corresponds to the second dynamic image area. Then, the area in the first image other than the first dynamic image area is divided into a plurality of first grid areas, and the area in the second image other than the second dynamic image area is divided into a plurality of second grid areas corresponding one to one to the plurality of first grid areas. Therefore, the corresponding first grid areas and second grid areas constitute a grid area pair, and the corresponding first grid vertex in the first grid area and the second grid vertex in the second grid area constitute a grid vertex pair. Here, the size and division method of the first grid area and the second grid area can be determined according to the actual image processing scenario. In some embodiments, the first dynamic image area and the second dynamic area are moving portrait areas. In this case, a portrait segmentation algorithm can be used to remove the portrait area.
[0042] According to an exemplary embodiment of the present disclosure, image alignment between frames can be performed by bringing the position of the second grid vertex in each grid vertex pair closer to the position of the corresponding first grid vertex. Specifically, pixel sampling can first be performed every predetermined number of pixels in the first image to obtain a plurality of uniformly distributed sampling pixels. For example, but not limited to, sampling can be performed every two pixels, so that each first grid area has a sampling pixel. That is, the sampling pixel points in the first grid area are obtained by performing pixel sampling every predetermined number of pixels in the first image. The position of the sampling pixel point in each first grid area can be obtained by interpolating all grid vertices of the first grid area where the sampling pixel point is located. Specifically, any interpolation method in the prior art can be used, for example, but not limited to bilinear interpolation to obtain the position of the sampling pixel point in each first grid area.
[0043] In order to clearly illustrate the solution of the present disclosure, the following description is combined with Figure 2 to proceed. Figure 2 is a schematic diagram illustrating a first image and a second image in an image alignment process according to an exemplary embodiment of the present disclosure. Figure 2 , image T is the target frame in image alignment (i.e., the first image), image I is the frame to be optimized (i.e., the second image), the portrait areas in image T and image I have been removed, leaving the image background area for image alignment. Image T and image I are divided into M*N first grid areas and second grid areas corresponding to each other, and the corresponding first grid areas and second grid areas form a grid area pair, and the corresponding first grid vertex in the first grid area and the second grid vertex in the second grid area form a grid vertex pair. Each first grid vertex in image T can be denoted as V i , each corresponding second mesh vertex in image I can be recorded as The position coordinates of each sampling pixel in the image T can be obtained by interpolating the four grid vertices of the first grid area where it is located. For each sampling pixel in the image T, the interpolation coefficient can be recorded as a = (a1, a2, a3, a4). Therefore, the coordinates of each sampling pixel can be expressed as:
[0044]
[0045] Among them, p represents the coordinates of the sampling pixel point; a i Represents the interpolation coefficient of the sampling pixel point; V i Indicates the coordinates of the first grid vertex in the first grid area where the sampling pixel point is located.
[0046] The coordinates of the sampling pixel points in image I corresponding to the sampling pixel points in image T can be expressed as:
[0047]
[0048] Where p represents the coordinates of the sampling pixel point corresponding to the sampling pixel point in the image T; a i V represents the interpolation coefficient of the sampling pixel point corresponding to the sampling pixel point in the image T; i x represents the coordinates of the first grid vertex of the first grid area where the sampling pixel point in the image T is located; i Represents the displacement adjustment of each second mesh vertex in image I relative to the corresponding first mesh vertex. By optimizing the displacement adjustment x of each second mesh vertex coordinate in image I i , image alignment between image T and image I can be achieved.
[0049] Return to reference Figure 1 In step 102, an objective function may be constructed for each mesh vertex pair, wherein the objective function represents a displacement error between the second mesh vertex and the first mesh vertex in the current mesh vertex pair.
[0050] According to an exemplary embodiment of the present disclosure, the objective function is a displacement error function between the first image and the second image, and at the algorithm execution level, it is the displacement error between the second mesh vertex and the first mesh vertex in each mesh vertex pair. In some embodiments, the objective function can be constructed based on the photometric error information and the deformation error information, so that the objective function can more comprehensively express the displacement error between the second mesh vertex and the first mesh vertex, so that the image alignment can be performed using the objective function to obtain a better alignment effect. Here, the photometric error information represents the photometric error between the sampling pixel points in at least one first mesh area near the current mesh vertex pair and the sampling pixel points in the corresponding second mesh area, and the deformation error information represents the deformation error between the first mesh vertex and the second mesh vertex in the current mesh vertex pair. Specifically, the movement between the corresponding pixel points between the first image and the second image can be transmitted to the mesh vertices, so the movement between the first mesh vertex and the corresponding second mesh vertex in each mesh vertex pair can reflect the movement direction and distance (i.e., photometric error) between the sampling pixel points in the first image and the sampling pixel points in the corresponding second image near the mesh vertex pair. Figure 2, the positions of the mesh vertex pairs include the four vertex positions of the image, the four side length positions, and the internal positions of the image excluding the four vertex positions and the four side length positions. Therefore, for the mesh vertex pairs at the four vertex positions, there is only one mesh area pair nearby; for the mesh vertices at the four side length positions, there are two mesh area pairs nearby; and for the mesh vertex pairs inside the image, there are four mesh area pairs nearby. In some scenarios, for the mesh vertex pairs inside the image, three of the four nearby mesh area pairs can be selected to construct the objective function. In each mesh area pair, due to differences in actual shooting conditions (for example, lens structure, ambient light, and shooting techniques, etc.), the corresponding first mesh and second mesh will also produce mesh deformation. The mesh deformation amount (i.e., deformation error) can also be determined by the positions of the first mesh vertex and the second mesh vertex in each mesh vertex pair. The constructed objective function can be expressed, for example, but not limited to, as follows:
[0051]
[0052] in, Represents the displacement error between the second mesh vertex and the first mesh vertex in the current mesh vertex pair; Represents the photometric error between the second mesh vertex and the first mesh vertex in the current mesh vertex pair; Indicates the deformation error between the second mesh vertex and the first mesh vertex in the current mesh vertex pair.
[0053] Next, the objective function is optimized using a least squares optimization algorithm. For example, but not limited to, the Gauss-Newton method, gradient descent method, or Levenberg-Marquardt method can be used to optimize the objective function. During the iterative optimization process, the objective function is first subjected to a first-order Taylor expansion. The first-order Taylor expansion of the photometric error can be expressed as:
[0054]
[0055] in, represents the photometric error between the second mesh vertex and the first mesh vertex in the current mesh vertex pair; T(W(V,0)) represents the coordinates of the first mesh vertex in the current mesh vertex pair; represents the derivative of the first mesh region in at least one mesh region pair near the current mesh vertex pair with respect to the aforementioned displacement adjustment amount (i.e., the displacement adjustment amount of the second mesh vertex relative to the first mesh vertex); Δx represents the update amount of the current displacement adjustment amount; and I(W(V,x)) represents the coordinates of the second mesh vertex in the current mesh vertex pair.
[0056] The deformation error describes the deformation error between the second mesh vertex and the first mesh vertex in the current mesh vertex pair caused by mesh deformation. For example, but not limited to, it can be expressed as:
[0057]
[0058] in, represents the deformation error caused by mesh deformation between the second mesh vertex and the first mesh vertex in the current mesh vertex pair; V1 represents the true position of the second mesh vertex in the current mesh vertex pair in the second image; V2 and V3 respectively represent the positions of any two mesh vertices of the other three second mesh vertices in the second mesh area where the second mesh vertex in the current mesh vertex pair is located in the corresponding first mesh area; u and v respectively represent the coordinate values of the first mesh vertex in the current mesh vertex pair when represented by any two mesh vertices of the other three first mesh vertices in the first mesh area where the first mesh vertex is located; R 90 Represents the rotation matrix, whose value is
[0059] In step 103 , an optimization matrix of the current mesh vertex pair is determined by using at least one first mesh region near the current mesh vertex pair.
[0060] Here, the optimization matrix is a matrix used when iteratively optimizing the objective function of the current mesh vertex pair. In some embodiments, the optimization matrix is, for example, a Hessian matrix. The inverse matrix of the Hessian matrix can be expressed as, for example, but not limited to:
[0061]
[0062] Among them, H -1 represents the inverse matrix of the Hessian matrix; A Jacobian matrix representing a first mesh region in at least one mesh region pair near a current mesh vertex pair undergoing iterative optimization with respect to the aforementioned displacement adjustment amount (i.e., the displacement adjustment amount of the second mesh vertex relative to the first mesh vertex); A derivative matrix representing a first mesh region in at least one mesh region pair near a current mesh vertex pair undergoing iterative optimization with respect to the aforementioned displacement adjustment amount; A Jacobian matrix representing the deformation of any three second mesh vertices in a second mesh region of at least one mesh region pair near the current mesh vertex pair with respect to the displacement adjustment amount; represents the derivative of the shape variable.
[0063] In the current iterative optimization process of the objective function, image alignment is performed by adding the updated displacement adjustment amount obtained in each iteration to image I. Therefore, the positions of the grid points on image I are updated after each iterative optimization, resulting in the need to recalculate the inverse matrix of the optimization matrix for each optimization iteration, resulting in a large amount of computation, thus affecting the real-time performance of image alignment. However, the present disclosure calculates the inverse matrix of the optimization matrix using at least one first grid region near the current grid vertex pair. Since the first grid region is a fixed region, the value of this matrix can be calculated before iterative optimization of the current grid vertex pair. The calculated value of the Hessian matrix can be directly reused each time the iterative optimization is used. This significantly reduces the computational effort during the iterative optimization process, improves the computational speed, and thus achieves better real-time performance.
[0064] In step 104 , the objective function is iteratively optimized using the optimization matrix to obtain a displacement adjustment amount of the second mesh vertex in the current mesh vertex pair relative to the corresponding first mesh vertex.
[0065] According to an exemplary embodiment of the present disclosure, a current displacement adjustment amount of the second mesh vertex in the current mesh vertex pair undergoing iterative optimization may be obtained, and based on the current displacement adjustment amount, an update amount of the current displacement adjustment amount of the second mesh vertex in the current mesh vertex pair undergoing iterative optimization may be calculated using the optimization matrix determined in step 103. The update amount may be adjusted on the current displacement adjustment amount to obtain an updated displacement adjustment amount. The updated displacement adjustment amount may be used as the current displacement adjustment amount, and the steps of calculating the update amount and obtaining the updated displacement adjustment amount may be repeated until an iteration end condition is met, thereby obtaining the displacement adjustment amount of the second mesh vertex in the current mesh vertex pair relative to the corresponding first mesh vertex.
[0066] According to an exemplary embodiment of the present disclosure, when the value of the update amount obtained by a certain iterative optimization is less than or equal to a preset threshold, it is determined that the iterative end condition has been met, and the iterative optimization of the current mesh vertex pair can be terminated. Here, the update amount of the current displacement adjustment amount can be expressed as, for example, but not limited to:
[0067]
[0068] Where Δx represents the update amount; H -1 represents the inverse matrix of the Hessian matrix; represents the Jacobian matrix of the first mesh region in at least one mesh region pair near the current mesh vertex pair undergoing iterative optimization with respect to the displacement adjustment amount; I(W(V,x)) represents the position of the sampling pixel points in the second mesh region in at least one mesh region pair near the current mesh vertex pair undergoing iterative optimization; T(V) represents the position of the first mesh vertex in the current mesh vertex pair undergoing iterative optimization; E d Describes the deformation amounts of any three second mesh vertices in a second mesh region in at least one mesh region pair near the current mesh vertex pair; represents the derivative of the shape variable.
[0069] After obtaining the updated amount of the current displacement adjustment amount in each iterative optimization, the current displacement adjustment amount may be updated once. The updating method may be, for example, but not limited to, expressed as:
[0070] x i+1 =x i -Δx (8)
[0071] Among them, x i Indicates the current displacement adjustment amount; Δx indicates the update amount of the current displacement adjustment amount; x i+1 represents the updated displacement adjustment for this iterative optimization. Subtraction is used to update the displacement adjustment here because a reverse derivative (i.e., derivative with respect to image I) is used. Therefore, the direction of the iterative optimization is opposite to the direction of motion of the second mesh vertex in the current mesh vertex pair relative to the corresponding first mesh vertex.
[0072] In step 105 , the position of the second mesh vertex in each mesh vertex pair relative to the corresponding first mesh vertex is adjusted by a corresponding displacement adjustment amount to obtain a target image aligned with the first image.
[0073] According to an exemplary embodiment of the present disclosure, the position coordinates of the first mesh vertex in each mesh vertex pair can be added to the corresponding displacement adjustment amount to obtain a pre-adjusted position of the second mesh vertex in each mesh vertex pair in the second image. The position of each second mesh vertex is then adjusted to the pre-adjusted position to obtain a target image aligned with the first image. Here, the position of the second mesh vertex in each mesh vertex pair relative to the corresponding first mesh vertex can be adjusted each time an updated displacement adjustment amount of the second mesh vertex relative to the corresponding first mesh vertex is obtained during iterative optimization of the current mesh vertex pair, and the position adjustment can be performed until the iterative optimization of the current mesh vertex pair is completed. Alternatively, after the iterative optimization of the current mesh vertex pair is completed, the updated displacement adjustment amount obtained from the last iterative optimization is used as the final displacement adjustment amount to be performed on the second mesh vertex in the current mesh vertex pair relative to the corresponding first mesh vertex, for position adjustment. For all mesh vertex pairs, the position of the second mesh vertex can be adjusted to the corresponding pre-adjusted position when the iterative optimization for the current mesh vertex pair is completed. Alternatively, after the iterative optimization for all mesh vertex pairs is completed, the position of the second mesh vertex in all mesh vertex pairs is adjusted to the corresponding pre-adjusted position, thereby obtaining a target image aligned with the first image.
[0074] Figure 3 is a block diagram illustrating an image processing apparatus according to an exemplary embodiment of the present disclosure.
[0075] Reference Figure 3 According to an exemplary embodiment of the present disclosure, the image processing apparatus 300 may include an image acquiring unit 301 , an objective function constructing unit 302 , an optimization matrix determining unit 303 , a displacement adjustment amount acquiring unit 304 , and a position adjusting unit 305 .
[0076] The image acquisition unit 301 can acquire a first image and a second image in a video, wherein the first image is a previous frame image of the second image, and the first image and the second image are respectively divided into a plurality of first grid areas and a plurality of second grid areas corresponding to each other. The corresponding first grid areas and second grid areas constitute a grid area pair, and the corresponding first grid vertices in the first grid areas and the second grid vertices in the second grid areas constitute a grid vertex pair.
[0077] During video capture, real-time image alignment of video frames is often required to obtain clearer and more stable video images. During image alignment, the first and second images in the video are first acquired. Here, the first image is the previous frame of the second image. The first image is used as the reference frame for image alignment (i.e., the alignment target), while the second image is the frame to be optimized for alignment with the first image. By constructing an objective function between the second and first images and optimizing the objective function using a least-squares optimization algorithm, the second image can be aligned with the first image.
[0078] According to an exemplary embodiment of the present disclosure, in order to reduce the computational complexity of image alignment, the image processing apparatus 300 further includes an image removal unit 306 ( Figure 3 (not shown), the image removal unit 306 may first remove the first dynamic image area in the first image and the second dynamic image area in the second image, respectively, where the first dynamic image area corresponds to the second dynamic image area. The image acquisition unit 301 then divides the area in the first image other than the first dynamic image area into a plurality of first grid areas, and divides the area in the second image other than the second dynamic image area into a plurality of second grid areas, where the plurality of first grid areas correspond one to one with the plurality of second grid areas. Therefore, the corresponding first grid area and the second grid area constitute a grid area pair, and the corresponding first grid vertex in the first grid area and the second grid vertex in the second grid area constitute a grid vertex pair. Here, the size and division method of the first grid area and the second grid area may be determined according to the actual image processing scenario. In some embodiments, the first dynamic image area and the second dynamic area are moving portrait areas. In this case, a portrait segmentation algorithm may be used to remove the portrait area.
[0079] The objective function constructing unit 302 may construct an objective function for each mesh vertex pair. Here, the objective function represents the displacement error between the second mesh vertex and the first mesh vertex in the current mesh vertex pair.
[0080] According to an exemplary embodiment of the present disclosure, the objective function construction unit 302 can construct an objective function based on photometric error information and deformation error information. Here, the photometric error information represents the photometric error between the sampled pixels in at least one first grid area near the current grid vertex pair and the sampled pixels in the corresponding second grid area, and the deformation error information represents the deformation error between the first grid vertex and the second grid vertex in the current grid vertex pair. Here, the sampled pixels in the first grid area are obtained by sampling pixels every predetermined number of pixels in the first image. The position of the sampled pixel in each first grid area can be obtained by interpolating all grid vertices in the first grid area where the sampled pixel is located. Specifically, any interpolation method in the prior art can be used, such as, but not limited to, bilinear interpolation, to obtain the position of the sampled pixel in each first grid area.
[0081] The optimization matrix determining unit 303 may determine the optimization matrix of the current mesh vertex pair through at least one first mesh region near the current mesh vertex pair.
[0082] The displacement adjustment amount acquisition unit 304 may iteratively optimize the objective function using the optimization matrix to obtain the displacement adjustment amount of the second mesh vertex in the current mesh vertex pair relative to the corresponding first mesh vertex.
[0083] According to an exemplary embodiment of the present disclosure, the displacement adjustment amount acquisition unit 304 may acquire a current displacement adjustment amount for the second mesh vertex in the current mesh vertex pair undergoing iterative optimization; based on the current displacement adjustment amount, calculate an update amount for the current displacement adjustment amount of the second mesh vertex in the current mesh vertex pair undergoing iterative optimization using the optimization matrix determined in step 103; adjust the update amount based on the current displacement adjustment amount to obtain an updated displacement adjustment amount, use the updated displacement adjustment amount as the current displacement adjustment amount, and repeat the steps of calculating the update amount and obtaining the updated displacement adjustment amount until an iterative termination condition is met, thereby obtaining the displacement adjustment amount of the second mesh vertex in the current mesh vertex pair relative to the corresponding first mesh vertex. According to an exemplary embodiment of the present disclosure, when the value of the update amount obtained from a certain iterative optimization is less than or equal to a preset threshold, it is determined that the iterative termination condition has been met, and the iterative optimization of the current mesh vertex pair may be terminated.
[0084] Here, the update amount of the current displacement adjustment amount in each iterative optimization process can be obtained according to formula (7) in the previous image processing method, and the optimization matrix in each iterative optimization process can be obtained according to formula (6) in the previous image processing method.
[0085] The position adjustment unit 305 may adjust the position of the second mesh vertex in each mesh vertex pair relative to the corresponding first mesh vertex by a corresponding displacement adjustment amount to obtain a target image aligned with the first image.
[0086] According to an exemplary embodiment of the present disclosure, the position adjustment unit 305 may add the position coordinates of the first mesh vertex in each mesh vertex pair to the corresponding displacement adjustment amount to obtain a pre-adjusted position of the second mesh vertex in each mesh vertex pair in the second image, and adjust the position of each second mesh vertex to the pre-adjusted position to obtain a target image aligned with the first image. Here, the position adjustment of the second mesh vertex in each mesh vertex pair relative to the corresponding first mesh vertex may be performed each time an updated displacement adjustment amount of the second mesh vertex relative to the corresponding first mesh vertex is obtained during the iterative optimization of the current mesh vertex pair, until the iterative optimization of the current mesh vertex pair is completed. Alternatively, after the iterative optimization of the current mesh vertex pair is completed, the updated displacement adjustment amount obtained from the last iterative optimization is used as the final displacement adjustment amount to be performed on the second mesh vertex in the current mesh vertex pair relative to the corresponding first mesh vertex, for position adjustment. For all mesh vertex pairs, the position of the second mesh vertex can be adjusted to the corresponding pre-adjusted position when the iterative optimization for the current mesh vertex pair is completed. Alternatively, after the iterative optimization for all mesh vertex pairs is completed, the position of the second mesh vertex in all mesh vertex pairs is adjusted to the corresponding pre-adjusted position, thereby obtaining a target image aligned with the first image.
[0087] Figure 4 is a block diagram of an electronic device 400 according to an exemplary embodiment of the present disclosure.
[0088] Reference Figure 4 The electronic device 400 includes at least one memory 401 and at least one processor 402, wherein the at least one memory 401 stores a set of computer-executable instructions. When the computer-executable instruction set is executed by the at least one processor 402, the image processing method according to the exemplary embodiment of the present disclosure is executed.
[0089] As an example, the electronic device 400 may be a PC, a tablet device, a personal digital assistant, a smart phone, or other device capable of executing the above-mentioned instruction set. Here, the electronic device 400 is not necessarily a single electronic device, but may also be any device or circuit that can execute the above-mentioned instructions (or instruction set) individually or in combination. The electronic device 400 may also be part of an integrated control system or system manager, or may be configured as a portable electronic device that is interconnected with a local or remote (e.g., via wireless transmission) interface.
[0090] In electronic device 400, processor 402 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0091] The processor 402 can execute instructions or codes stored in the memory 401, wherein the memory 401 can also store data. Instructions and data can also be sent and received over the network via the network interface device, wherein the network interface device can use any known transmission protocol.
[0092] Memory 401 may be integrated with processor 402, for example, by placing RAM or flash memory within an integrated circuit microprocessor or the like. Furthermore, memory 401 may comprise a separate device, such as an external disk drive, a storage array, or any other storage device usable by a database system. Memory 401 and processor 402 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, or the like, such that processor 402 can access files stored in memory.
[0093] In addition, the electronic device 400 may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the electronic device 400 may be connected to each other via a bus and / or a network.
[0094] Figure 5 is a block diagram illustrating a computer-readable storage medium 500 according to an exemplary embodiment of the present disclosure.
[0095] Reference Figure 5The computer-readable storage medium 500 stores instructions 501, wherein when the instructions 501 are executed by at least one processor, the at least one processor is prompted to perform the image processing method according to the present disclosure. Examples of the computer-readable storage medium 500 include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card storage (such as, multimedia card, secure digital (SD) card or ultra-fast digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, any other device configured to store the computer program and any associated data, data files and data structures in a non-transitory manner and provide the computer program and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the above-mentioned computer-readable storage medium 500 can be run in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system so that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.
[0096] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided. Instructions in the computer program product may be executed by a processor of a computer device to implement the image processing method according to an exemplary embodiment of the present disclosure.
[0097] According to the image processing method, device, electronic device and storage medium disclosed in the present invention, in the process of video image alignment using a grid optical flow algorithm based on image photometric error, the derivation method in the optimization process of photometric error is improved, and there is no need to repeatedly calculate the image gradient derivative, thereby greatly reducing the computational complexity of the algorithm and accelerating the running speed of the algorithm. Thus, while taking into account the accuracy of video image alignment, the processing time of video image alignment can be reduced, the probability of mobile phone heating due to video image alignment is reduced, and the real-time performance is better.
[0098] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0099] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An image processing method, characterized in that: include: Acquire a first image and a second image in a video, wherein the first image is a previous frame of the second image, and the first image and the second image are respectively divided into a plurality of first grid areas and a plurality of second grid areas in one-to-one correspondence, wherein the corresponding first grid areas and second grid areas constitute a grid area pair, and the corresponding first grid vertex in the first grid area and the second grid vertex in the second grid area constitute a grid vertex pair; For each mesh vertex pair, constructing an objective function, wherein the objective function represents a displacement error between a second mesh vertex and a first mesh vertex in the current mesh vertex pair; Before iteratively optimizing the objective function, determining a value of an optimization matrix of the current mesh vertex pair through at least one first mesh region near the current mesh vertex pair, wherein the optimization matrix is a matrix used when iteratively optimizing the objective function; Iteratively optimizing the objective function using the value of the optimization matrix to obtain a displacement adjustment amount of the second mesh vertex in the current mesh vertex pair relative to the corresponding first mesh vertex; adjusting the position of the second mesh vertex in each mesh vertex pair relative to the corresponding first mesh vertex by a corresponding displacement adjustment amount to obtain a target image aligned with the first image; The iteratively optimizing the objective function using the value of the optimization matrix to obtain a displacement adjustment amount of the second mesh vertex in the current mesh vertex pair relative to the corresponding first mesh vertex includes: Obtaining a current displacement adjustment value of a second mesh vertex in the current mesh vertex pair undergoing iterative optimization; Calculating, based on the current displacement adjustment amount, an update amount of the current displacement adjustment amount of the second mesh vertex in the current mesh vertex pair undergoing iteration by using a value of the optimization matrix; Adjusting the update amount based on the current displacement adjustment amount to obtain an updated displacement adjustment amount; The updated displacement adjustment amount is used as the current displacement adjustment amount, and the steps of calculating the update amount and obtaining the updated displacement adjustment amount are repeatedly performed until the iteration end condition is met, thereby obtaining the displacement adjustment amount of the second mesh vertex in the current mesh vertex pair relative to the corresponding first mesh vertex.
2. The image processing method according to claim 1, wherein: Before the first image and the second image are divided into a plurality of first grid areas and a plurality of second grid areas corresponding to each other, the method includes: Removing a first dynamic image region in the first image and a second dynamic image region in the second image respectively, the first dynamic image region corresponding to the second dynamic image region; The first image and the second image are respectively divided into a plurality of first grid areas and a plurality of second grid areas corresponding to each other, including: The area of the first image except the first dynamic image area is divided into a plurality of first grid areas; The area of the second image except the second dynamic image area is divided into a plurality of second grid areas, and the plurality of second grid areas correspond one-to-one to the plurality of first grid areas.
3. The image processing method according to claim 1, wherein: The constructing of the objective function comprises: The objective function is constructed based on photometric error information and deformation error information, wherein the photometric error information represents a photometric error between a sampling pixel point in at least one first grid area near the current grid vertex pair and a sampling pixel point in a corresponding second grid area, and the deformation error information represents a deformation error between a first grid vertex and a second grid vertex in the current grid vertex pair.
4. The image processing method according to claim 3, wherein: The position of the sampling pixel point in each first grid area is obtained by interpolating all first grid vertices in each first grid area.
5. The image processing method according to claim 1, wherein: The step of repeatedly executing the step of calculating the update amount and the step of obtaining the updated displacement adjustment amount until an iteration end condition is reached includes: When the value of the update amount is less than or equal to a preset threshold, it is determined that the iteration end condition is met.
6. The image processing method according to claim 1, wherein: The step of adjusting the position of the second mesh vertex in each mesh vertex pair relative to the corresponding first mesh vertex by a corresponding displacement adjustment amount to obtain a target image aligned with the first image includes: adding the position coordinates of the first mesh vertex in each mesh vertex pair to the corresponding displacement adjustment amount to obtain a pre-adjusted position of the second mesh vertex in each mesh vertex pair in the second image; The position of each second mesh vertex is adjusted to the pre-adjusted position to obtain a target image aligned with the first image.
7. An image processing device, characterized in that: include: an image acquisition unit configured to: acquire a first image and a second image in a video, wherein the first image is a previous frame image of the second image, the first image and the second image are respectively divided into a plurality of first grid areas and a plurality of second grid areas corresponding to each other, the corresponding first grid areas and the corresponding second grid areas constitute a grid area pair, and the corresponding first grid vertex in the first grid area and the second grid vertex in the second grid area constitute a grid vertex pair; An objective function construction unit is configured to: construct an objective function for each mesh vertex pair, wherein the objective function represents a displacement error between the second mesh vertex and the first mesh vertex in the current mesh vertex pair; an optimization matrix determining unit, configured to: before iteratively optimizing the objective function, determine a value of an optimization matrix of the current mesh vertex pair using at least one first mesh region near the current mesh vertex pair, wherein the optimization matrix is a matrix used when iteratively optimizing the objective function; a displacement adjustment amount obtaining unit configured to: iteratively optimize the objective function using the value of the optimization matrix to obtain a displacement adjustment amount of the second mesh vertex in the current mesh vertex pair relative to the corresponding first mesh vertex; a position adjustment unit configured to: adjust the position of the second mesh vertex in each mesh vertex pair relative to the corresponding first mesh vertex by a corresponding displacement adjustment amount to obtain a target image aligned with the first image; Wherein, the displacement adjustment amount acquisition unit is configured to: acquire the current displacement adjustment amount of the second mesh vertex in the current mesh vertex pair undergoing iterative optimization; calculate, based on the current displacement adjustment amount, an update amount of the current displacement adjustment amount of the second mesh vertex in the current mesh vertex pair undergoing iteration through the value of the optimization matrix; adjust the update amount on the current displacement adjustment amount to obtain an updated displacement adjustment amount; use the updated displacement adjustment amount as the current displacement adjustment amount, and repeat the steps of calculating the update amount and obtaining the updated displacement adjustment amount until the iteration end condition is met, thereby obtaining the displacement adjustment amount of the second mesh vertex in the current mesh vertex pair relative to the corresponding first mesh vertex.
8. The image processing device according to claim 7, wherein Also includes: an image removal unit configured to, before the first image and the second image are divided into a plurality of first grid areas and a plurality of second grid areas corresponding to each other, remove a first dynamic image area in the first image and a second dynamic image area in the second image, respectively, the first dynamic image area corresponding to the second dynamic image area; The image acquisition unit is configured to: divide the area of the first image except the first dynamic image area into a plurality of first grid areas; The area of the second image except the second dynamic image area is divided into a plurality of second grid areas, and the plurality of second grid areas correspond one-to-one to the plurality of first grid areas.
9. The image processing device according to claim 7, wherein: The objective function building unit is configured as follows: The objective function is constructed based on photometric error information and deformation error information, wherein the photometric error information represents a photometric error between a sampling pixel point in at least one first grid area near the current grid vertex pair and a sampling pixel point in a corresponding second grid area, and the deformation error information represents a deformation error between a first grid vertex and a second grid vertex in the current grid vertex pair.
10. The image processing device according to claim 9, wherein The position of the sampling pixel point in each first grid area is obtained by interpolating all first grid vertices in each first grid area.
11. The image processing device according to claim 7, wherein The displacement adjustment amount acquiring unit is further configured to: When the value of the update amount is less than or equal to a preset threshold, it is determined that the iteration end condition is met.
12. The image processing device according to claim 7, wherein The position adjustment unit is configured to: adding the position coordinates of the first mesh vertex in each mesh vertex pair to the corresponding displacement adjustment amount to obtain a pre-adjusted position of the second mesh vertex in each mesh vertex pair in the second image; The position of each second mesh vertex is adjusted to the pre-adjusted position to obtain a target image aligned with the first image.
13. An electronic device, characterized in that: include: at least one processor; at least one memory storing computer-executable instructions, Wherein, when the computer executable instructions are executed by the at least one processor, the at least one processor is prompted to perform the image processing method according to any one of claims 1 to 6.
14. A computer-readable storage medium storing instructions, characterized in that: When the instructions are executed by at least one processor, the at least one processor is prompted to perform the image processing method according to any one of claims 1 to 6.
15. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by at least one processor, the image processing method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Visual positioning method and device, electronic equipment and readable storage medium
CN111862206A
Medical image registration method, electronic device and storage medium
CN112116642A
Image processing method and device, electronic equipment and storage medium
CN112884664A