Image alignment method, image alignment device and terminal equipment

By extracting feature points and calculating motion vectors in image alignment technology, the problem of large calculation and low efficiency in the prior art is solved, and efficient image alignment is achieved.

CN114187333BActive Publication Date: 2025-08-15TCL TECHNOLOGY GROUP CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010963340.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-14
Publication Date
2025-08-15
Estimated Expiration
2040-09-14

AI Technical Summary

Technical Problem

The existing image alignment technology has high computational capacity and low efficiency, making it difficult to meet the needs of fast image processing.

Method used

By extracting feature points in the image frame to be aligned and the reference image frame, the motion vector is calculated, and the image frame to be aligned to the reference image frame according to the motion vector, the calculation amount is reduced and efficiency is improved.

Benefits of technology

It greatly reduces the calculation amount of image alignment, improves the efficiency of image alignment, and is suitable for fast image processing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114187333B_ABST
    Figure CN114187333B_ABST
Patent Text Reader

Abstract

This application applies to the field of image processing technology and provides an image alignment method, an image alignment device, a terminal device, and a computer-readable storage medium. The method includes: obtaining an image frame to be aligned and a reference image frame; extracting a first feature point from the image frame to be aligned and a second feature point from the reference image frame; calculating a motion vector for the image frame to be aligned based on the first and second feature points; and aligning the image frame to be aligned with the reference image frame based on the motion vector. This method can significantly reduce the computational complexity of image alignment and improve the efficiency of image alignment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image processing technology, and in particular relates to an image alignment method, an image alignment device, a terminal device, and a computer-readable storage medium. Background Art

[0002] Image alignment technology has many applications in daily life. For example, when shooting with a mobile phone, hand tremors can cause slight differences between multiple frames of images taken consecutively. To achieve better image quality, image alignment technology is needed to align multiple frames to the same coordinate system.

[0003] Existing image alignment technologies mainly include optical flow-based image alignment technology, feature extraction-based image alignment technology, pyramid-based image alignment technology, etc. However, these image alignment technologies are computationally intensive and inefficient. Summary of the Invention

[0004] In view of this, the present application provides an image alignment method, an image alignment apparatus, a terminal device, and a computer-readable storage medium, which can greatly reduce the computational complexity of image alignment and improve the efficiency of image alignment.

[0005] In a first aspect, the present application provides an image alignment method, comprising:

[0006] Obtaining the image frame to be aligned and the reference image frame;

[0007] Extracting a first feature point from the image frame to be aligned and a second feature point from the reference image frame;

[0008] Calculating the motion vector of the image frame to be aligned according to the first feature point and the second feature point;

[0009] The image frame to be aligned is aligned with the reference image frame according to the motion vector.

[0010] In a second aspect, the present application provides an image alignment device, comprising:

[0011] An image acquisition unit, configured to acquire an image frame to be aligned and a reference image frame;

[0012] A feature point extraction unit, configured to extract first feature points from the image frame to be aligned and second feature points from the reference image frame;

[0013] a motion vector calculation unit, configured to calculate the motion vector of the image frame to be aligned based on the first feature point and the second feature point;

[0014] The image alignment unit is configured to align the image frame to be aligned with the reference image frame according to the motion vector.

[0015] In a third aspect, the present application provides a terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method provided in the first aspect is implemented.

[0016] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method provided in the first aspect is implemented.

[0017] In a fifth aspect, the present application provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute the method provided in the first aspect above.

[0018] As can be seen from the above, in the present application scheme, the image frame to be aligned and the reference image frame are first obtained, the first feature point in the image frame to be aligned and the second feature point in the reference image frame are extracted, and then the motion vector of the image frame to be aligned is calculated based on the first feature point and the second feature point. Finally, the image frame to be aligned is aligned to the reference image frame based on the motion vector. The present application scheme can greatly reduce the amount of computation for image alignment and improve the efficiency of image alignment by extracting feature points representing high-frequency information in the image frame to be aligned and the reference image frame, and calculating the motion vector of the image frame to be aligned based on the extracted feature points. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1 Schematic diagram of the image alignment method provided in the embodiment of the present application;

[0021] Figure 2 This is an example diagram of characteristic points provided in an embodiment of the present application;

[0022] Figure 3 This is a schematic diagram of the feature descriptor generation principle provided in an embodiment of the present application;

[0023] Figure 4 is an example diagram of an image block provided in an embodiment of the present application;

[0024] Figure 5 This is an example diagram of a pixel block provided in an embodiment of the present application;

[0025] Figure 6 This is a structural block diagram of an image alignment device provided in an embodiment of the present application;

[0026] Figure 7 It is a structural diagram of the terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0028] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0029] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0030] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0031] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0032] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0033] Figure 1 A flowchart of an image alignment method provided in an embodiment of the present application is shown, and is described in detail as follows:

[0034] Step 101: obtaining an image frame to be aligned and a reference image frame;

[0035] In an embodiment of the present application, the terminal device needs to acquire at least two frames of images. For example, the at least two frames of images can be at least two frames of images continuously captured during the process of the terminal device taking pictures. After the terminal device acquires at least two frames of images, it selects one frame of image from the at least two frames of images as a reference image frame, and the remaining images are all used as image frames to be aligned, and then each frame of image frames to be aligned is aligned to the reference image frame. Optionally, the method for selecting the reference image frame can be to calculate the clarity of each frame of the at least two frames of images respectively, and select the frame of image with the highest clarity as the reference image frame to ensure that the image frame to be aligned with the reference image frame is clear. Alternatively, the resolution of each frame of image in the at least two frames of images can be calculated respectively, and the frame of image with the highest resolution can be used as the reference image frame. The method for selecting the reference image frame is not limited here.

[0036] For example, the terminal device can use the continuous shooting mode of its built-in camera to continuously shoot multiple frames of images of the same scene in a very short time, such as shooting 30 frames of facial images of the same face in 1 second, namely facial image 1, facial image 2, facial image 3, ..., facial image 30. Then, the clarity of the 30 frames of facial images can be calculated separately, and the frame with the highest clarity among the 30 frames of facial images can be used as the reference image frame, and the remaining 29 frames of images can be used as image frames to be aligned. For example, assuming that among the 30 frames of facial images, facial image 1 has the highest clarity, therefore, facial image 1 can be used as the reference image frame, and facial image 2, facial image 3, facial image 4, ..., facial image 30 can all be used as image frames to be aligned.

[0037] Step 102: extracting first feature points in the image frame to be aligned and second feature points in the reference image frame;

[0038] In the embodiment of the present application, since each frame of the image frame to be aligned needs to be aligned to the reference image frame respectively, the operations performed on each frame of the image frame to be aligned are the same or similar. Therefore, for the convenience of explanation, the embodiment of the present application will take one of the several frames of the image frame to be aligned as an example for introduction. Among them, the first feature point represents the high-frequency information in the image frame to be aligned, and the second feature point represents the high-frequency information in the reference image frame. By performing feature point detection on the image frame to be aligned and the reference image frame, the first feature point in the image frame to be aligned and the second feature point in the reference image frame can be obtained. For details, please refer to Figure 2 , Figure 2 is an example of the first feature point in the image frame to be aligned, Figure 2 Each small circle in represents a first feature point. Optionally, in order to achieve a balance between the effect and efficiency of feature point detection, feature point detection can be performed on the image frame to be aligned and the reference image frame using the Features From Accelerated Segment Test (FAST) algorithm.

[0039] Take the image frame to be aligned as an example, please refer to Figure 3 , Figure 3Each square in represents a pixel. For a certain pixel point p in the image frame to be aligned, a circle with a radius of 3 is determined with the pixel point p as the center, and there are 16 pixels on the circle. The absolute value d of the difference between the pixel value of each pixel point and the pixel point p in the 16 pixels on the circle is calculated. If the number of pixels whose corresponding absolute value d is greater than the threshold t among the 16 pixels on the circle is greater than or equal to 12, the pixel point p is determined as a corner point. At the same time, the confidence of the pixel point p is calculated based on the sum of the absolute values d corresponding to the 16 pixels on the circle. The greater the sum of the absolute values d corresponding to the 16 pixels on the circle, the greater the confidence. Based on this, all corner points in the image frame to be aligned are detected. If the number of detected corner points is less than the preset number of corner points (such as 500), the threshold t is reduced, and all corner points in the image frame to be aligned are detected again in the above manner. If the number of detected corner points is still less than the preset number of corner points, the threshold t is reduced again, and all corner points in the image frame to be aligned are detected again in the above manner, and so on, until the number of corner points in the image frame to be aligned is greater than or equal to the preset number of corner points. After detecting that the number of corner points in the image frame to be aligned is greater than or equal to the threshold number of corner points, the corner points are sorted in descending order according to their confidence, and the preset number of corner points in the front are taken as the first feature points, thereby improving the accuracy of the FAST algorithm. Since the method of detecting feature points for the reference image frame is the same as that for the image frame to be aligned, it will not be repeated here.

[0040] For example, assuming the preset number of corner points is 500, and the number of corner points detected in the image frames to be aligned with a threshold t of 10 is 300, since 300 is less than 500, the threshold t is reduced to 6. The number of corner points detected in the image frames to be aligned with a threshold t of 6 is 800, and since 800 is greater than 500, the corner points are sorted in descending order according to their confidence level, and the first 500 corner points are taken as the first feature points.

[0041] Step 103, calculating the motion vector of the image frame to be aligned based on the first feature point and the second feature point;

[0042] In an embodiment of the present application, the first feature point can represent high-frequency information in the image frame to be aligned, that is, the characteristic contour of the object in the image frame to be aligned, and the second feature point can represent high-frequency information in the reference image frame, that is, the characteristic contour of the object in the reference image frame. Therefore, by analyzing the first feature point and the second feature point, the offset between the object in the image frame to be aligned and the object in the reference image frame can be obtained, and then the motion vector of the image frame to be aligned can be obtained.

[0043] Optionally, the above step 103 may specifically include:

[0044] A1. Calculating a global motion vector of the image frame to be aligned based on the first feature point and the second feature point;

[0045] A2. Divide the image frame to be aligned into a preset number of image blocks;

[0046] A3. Calculating the local motion vector of each image block in the image frame to be aligned based on the global motion vector and the first feature point;

[0047] A4. Calculate a motion vector based on the global motion vector and the local motion vector.

[0048] In the embodiment of the present application, by analyzing the first feature point and the second feature point, the offset between the image frame to be aligned and the reference image frame can be obtained, and then the global motion vector of the image frame to be aligned can be obtained. Next, the image frame to be aligned needs to be divided into a preset number of image blocks, and the area of each image block is equal. Please refer to Figure 4 , Figure 4 This is an example of an image block in an image frame to be aligned, where each square represents an image block. For example, the image frame to be aligned can be divided into 32×32 image blocks. For an image frame to be aligned with a size of 4000×3000 pixels, the size of each image block is 125×94 pixels. Among them, the global motion vector represents the offset between the image frame to be aligned and the reference image frame, and the first feature point represents the high-frequency information in the image frame to be aligned, that is, the outline (such as text and icons, etc.). Therefore, based on the global motion vector and the first feature point in the image block, an image block similar to the image block in the image frame to be aligned can be matched in the reference image frame, and the two similar image blocks have similar high-frequency information. Based on the offset between the two similar image blocks, the local motion vector of the image block in the image frame to be aligned can be obtained. Finally, based on the global motion vector and the local motion vector, the motion vector can be calculated. It should be noted that in the embodiment of the present application, a motion vector is calculated for each image block respectively, that is, the number of motion vectors is the same as the number of image blocks in the image frame to be aligned.

[0049] For example, the motion vector of each image block can be the sum of the local motion vector of the image block and the global motion vector of the image frame to be aligned. Taking image block A as an example, assuming that the local motion vector of image block A is (0, -1) and the global motion vector of the image frame to be aligned is (2, 2), the calculated motion vector of image block A is (2, 1).

[0050] Optionally, the above step A1 may specifically include:

[0051] A11. Generate a feature descriptor for each first feature point and a feature descriptor for each second feature point;

[0052] A12. Determine a target first feature point that matches a target second feature point;

[0053] A13. Calculate the offset between the first target feature point and the second target feature point, and add the offset to the offset list;

[0054] A14. Determine the global motion vector of the image frame to be aligned according to the offset list.

[0055] In an embodiment of the present application, the feature descriptor of the first feature point and the feature descriptor of the second feature point are generated in the same way. The following describes the generation method of the feature descriptor of the first feature point by taking a first feature point q in the image frame to be aligned as an example. First, a neighborhood window is determined with the first feature point q as the center, and 256 pairs of pixels are randomly selected within the neighborhood window. Then, for each pair of pixels selected, the size relationship of the pixel values of the two pixels in the pair of pixels (such as pixel 1 and pixel 2) is determined. If the pixel value of pixel 1 is less than pixel 2, the corresponding value of the pair of pixels is recorded as 1. If the pixel value of pixel 1 is not less than pixel 2, the corresponding value of the pair of pixels is recorded as 0. Finally, the corresponding values of the 256 pairs of pixels are spliced into a 256-bit binary code, which is the feature descriptor of the first feature point q.

[0056] After generating the feature descriptors of each first feature point and the feature descriptors of each second feature point, it is necessary to match the first feature point with the second feature point based on the feature descriptors. For ease of explanation, a second feature point in the reference image frame will be taken as an example, and the second feature point is the target second feature point. It can be understood that in practice, the operation of the target second feature point will be performed on each second feature point in the reference image frame. Specifically, the target first feature point is determined as follows: the Hamming distance between the feature descriptor of the target second feature point and the feature descriptors of each first feature point in the image frame to be aligned is calculated, and the first feature point with the smallest corresponding Hamming distance is determined as the target first feature point, that is, the feature descriptor of the target second feature point is matched with the feature descriptor of the target first feature point.

[0057] After determining the target first feature point that matches the target second feature point, it is also necessary to calculate the offset between the target second feature point and the target first feature point. For example, assuming that the coordinates of the target second feature point in the reference image frame are (50, 50) and the coordinates of the target first feature point in the image frame to be aligned are (48, 49), then the offset between the target second feature point and the target first feature point is (2, 1). The calculated offset between the target second feature point and the target first feature point will be added to the preset offset list. After performing the same operation as the target second feature point on each second feature point in the reference image frame, the offset list will contain several offsets. By analyzing the offsets contained in the offset list, the global motion vector of the image frame to be aligned can be determined.

[0058] For example, the offset that appears the most times in the offset list can be determined as the global motion vector. For example, assuming that the offset list contains 500 offsets, of which 400 are (1, 2), 80 are (0, 1), and 20 are (3, 2), then the offset (1, 2) that appears the most times is determined as the global motion vector.

[0059] Optionally, the above step A3 may specifically include:

[0060] B1, determining a first pixel block in the target image block with the first feature point in the target image block as the center;

[0061] B2. determining, in the reference image frame according to the global motion vector, a second pixel block that matches the first pixel block;

[0062] B3. Determine the offset between the first pixel block and the second pixel block as the local motion vector of the target image block.

[0063] In the embodiment of the present application, for ease of explanation, an image block in the image frame to be aligned will be taken as an example for introduction, and the image block is the target image block. It can be understood that, in practice, the operations of the target image block will be performed on each image block in the image frame to be aligned. Specifically, first, with the first feature point in the target image block as the center, a first pixel block is determined in the target image block according to a preset pixel block size, wherein the pixel block size can be set according to actual conditions, such as a pixel block size of 32×32 pixels. Then, based on the global motion vector, a second pixel block matching the first pixel block can be determined in the reference image frame, and the size of the second pixel block is the same as that of the first pixel block. Finally, the offset between the first pixel block and the second pixel block can be calculated through the coordinates of the first pixel block in the image frame to be aligned and the coordinates of the second pixel block in the reference image frame, and the offset is determined as the local motion vector of the target image block. Because the first pixel block is determined in the target image block with the first feature point in the target image block as the center, and block matching is performed using the first pixel block, rather than directly using the image blocks, the image alignment method in the embodiment of the present application has a very small computational load and is very efficient. Furthermore, because the first feature point in the target image block represents the high-frequency information in the target image block, the first pixel block determined with the first feature point as the center will contain very little low-frequency information. Therefore, using the first pixel block for block matching can also reduce interference from low-frequency information in the target image block, thereby improving the accuracy of the calculated local motion vector of the target image block.

[0064] For example, assuming that the coordinates of the center point of the second pixel block in the reference image frame are (50, 50), and the coordinates of the center point of the first pixel block in the image frame to be aligned are (48, 49), then the offset between the first pixel block and the second pixel block is (2, 1), that is, the local motion vector of the target image block is (2, 1). As an example, Figure 4 Each image block (ie, square) in is marked with a local motion vector of each image block.

[0065] It should be noted that since the first feature points are not necessarily evenly distributed in the image frames to be aligned, some image blocks may not contain a first feature point, some image blocks may contain one first feature point, and some image blocks may contain two or more first feature points. Based on this, before executing the above step B1, the following steps are also included:

[0066] If the target image block does not contain the first feature point, the local motion vectors of the image blocks surrounding the target image block are directly used as the local motion vector of the target image block. Specifically, the average value of the local motion vectors of the four image blocks adjacent to the upper, lower, left and right of the target image block can be used as the local motion vector of the target image block. If the four image blocks adjacent to the upper, lower, left and right of the target image block do not contain the first feature point, the average value of the local motion vectors of the four image blocks adjacent to the upper left corner, lower left corner, upper right corner and lower right corner of the target image block are used as the local motion vector of the target image block. If the four image blocks adjacent to the upper left corner, lower left corner, upper right corner and lower right corner of the target image block still do not contain the first feature point, the global motion vector of the image frame to be aligned is used as the local motion vector of the target image block. This avoids the block effect.

[0067] If the target image block contains the first feature point, it is determined whether the number of the first feature points in the target image block is greater than a preset feature point number threshold. Taking the feature point number threshold as 3 as an example, if the number of the first feature points in the target image block is less than or equal to 3, all the first feature points in the target image block are retained.

[0068] If the number of first feature points contained in the target image block is greater than 3, the first feature points at the edge of the target image block are removed, wherein the edge of the target image block can be the outermost circle of the target image block or the outermost two circles of the target image block. The definition of the edge can be set according to needs and is not specifically limited here. Then, the first feature points are sorted in descending order according to the confidence of each first feature point. In the target image block, only the first three first feature points are retained, and the first feature points other than the first three first feature points in the target image block are removed. It should be noted that if the first feature points at the edge of the target image block are removed, the number of first feature points remaining in the target image block is less than 3, then all the remaining first feature points in the target image block are retained. In the above manner, the number of first feature points retained in each image block does not exceed the feature point number threshold, thereby reducing the amount of calculation in subsequent steps B1, B2 and B3.

[0069] It should be understood that since the target image block may contain more than one first feature point, more than one first pixel block may be determined in the target image block. If more than two first pixel blocks are determined in the target image block, then for each first pixel block, a second pixel block that matches the first pixel block is determined in the reference image frame, and the offset between the first pixel block and the second pixel block is calculated. After obtaining the offsets corresponding to the first pixel blocks in the target image block, the average of the offsets can be determined as the local motion vector of the target image block.

[0070] Optionally, the above step B2 may specifically include:

[0071] B21. Determine a search range in the reference image frame according to the global motion vector and the first feature point in the target image block;

[0072] B22. Search for a second pixel block matching the first pixel block within the search range.

[0073] In an embodiment of the present application, a search range can be determined in a reference image frame based on a global motion vector and a first feature point in a target image block. Then, a candidate second pixel block can be determined with each pixel point in the search range as the center, wherein the size of each candidate second pixel block is the same as the size of the first pixel block. Finally, the similarity between each candidate second pixel block and the first pixel block is calculated, and the candidate second pixel block with the highest similarity to the first pixel block is determined as the second pixel block that matches the first pixel block. Optionally, the Euclidean space distance between the candidate second pixel block and the first pixel block can be used as the similarity between the candidate second pixel block and the first pixel block. Specifically, assuming that the size of the first pixel block is 32×32 pixels, the similarity between the candidate second pixel block and the first pixel block can be calculated according to the following formula:

[0074]

[0075] Among them, L is the similarity, P r is the pixel value of the pixel point in the candidate second pixel block, P t is the pixel value of the pixel in the first pixel block.

[0076] For example, the search range in the reference image frame may be determined by selecting a center point in the reference image frame, where the offset between the center point and the first feature point in the target image block is equal to the global motion vector. Next, a search range is determined in the reference image frame based on the center point and a preset search range size, where the search range size may be 16×16 pixels.

[0077] For example, please refer to Figure 5 , Figure 5The target image block in the image contains a first feature point A, the coordinates of which are (7, 7). Assuming the global motion vector is (2, 3), and the preset search range size is 3×3, a search range centered on the center point a is determined in the reference image frame based on the first feature point A, where the coordinates of the center point a are (7-2, 7-3) = (5, 4). The search range contains 9 pixels, and a candidate second pixel block is determined with each of the 9 pixels as the center. The similarity between the 9 candidate second pixel blocks and the first pixel block is then calculated. Among them, the candidate second pixel block centered on pixel point m has the highest similarity to the first pixel block. Therefore, the candidate second pixel block centered on pixel point m is used as the second pixel block that matches the first pixel block.

[0078] Step 104, aligning the image frame to be aligned with the reference image frame according to the motion vector;

[0079] In an embodiment of the present application, the image frame to be aligned includes several image blocks. After calculating the motion vector of each image block, the image frame to be aligned can be aligned with the reference image frame based on the motion vector of each image block. Specifically, the image block can be aligned with the reference image frame based on the motion vector of each image block. For example, assuming that the motion vector of an image block A in the image frame to be aligned is (m, n), after aligning image block A with the reference image frame, image block A' is obtained. Then, the pixel value of the pixel point with coordinates (x, y) in image block A' is equal to the pixel value of the pixel point with coordinates (x+m, y+n) in image block A.

[0080] As can be seen from the above, in the present application scheme, the image frame to be aligned and the reference image frame are first obtained, the first feature point in the image frame to be aligned and the second feature point in the reference image frame are extracted, and then the motion vector of the image frame to be aligned is calculated based on the first feature point and the second feature point. Finally, the image frame to be aligned is aligned to the reference image frame based on the motion vector. The present application scheme can greatly reduce the amount of computation for image alignment and improve the efficiency of image alignment by extracting feature points representing high-frequency information in the image frame to be aligned and the reference image frame, and calculating the motion vector of the image frame to be aligned based on the extracted feature points.

[0081] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0082] Figure 6 A structural block diagram of an image alignment device provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.

[0083] The image alignment device 600 includes:

[0084] An image acquisition unit 601 is configured to acquire an image frame to be aligned and a reference image frame;

[0085] A feature point extraction unit 602 is configured to extract first feature points from the image frame to be aligned and second feature points from the reference image frame;

[0086] A motion vector calculation unit 603 is configured to calculate a motion vector of the image frame to be aligned based on the first feature point and the second feature point;

[0087] The image alignment unit 604 is configured to align the image frame to be aligned with the reference image frame according to the motion vector.

[0088] Optionally, the motion vector calculation unit 603 further includes:

[0089] A global motion vector calculation subunit, configured to calculate the global motion vector of the image frame to be aligned based on the first feature point and the second feature point;

[0090] An image segmentation subunit, configured to segment the image frame to be aligned into a preset number of image blocks;

[0091] A local motion vector calculation subunit, configured to calculate a local motion vector of each image block in the image frame to be aligned based on the global motion vector and the first feature point;

[0092] The motion vector calculation subunit is used to calculate the motion vector according to the global motion vector and the local motion vector.

[0093] Optionally, the global motion vector calculation subunit includes:

[0094] a descriptor generating subunit, configured to generate a feature descriptor for each first feature point and a feature descriptor for each second feature point;

[0095] a feature point matching subunit, configured to determine a target first feature point that matches a target second feature point, wherein the target second feature point is any second feature point in the reference image frame, and a feature descriptor of the target second feature point matches a feature descriptor of the target first feature point;

[0096] an offset calculation subunit, configured to calculate an offset between the first target feature point and the second target feature point, and add the offset to an offset list;

[0097] The global motion vector determining subunit is configured to determine the global motion vector of the image frame to be aligned according to the offset list.

[0098] Optionally, the global motion vector determination subunit is specifically configured to determine the offset that appears the most times in the offset list as the global motion vector.

[0099] Optionally, the local motion vector calculation subunit includes:

[0100] A first pixel block determining subunit is configured to determine a first pixel block in a target image block with a first feature point in the target image block as the center, where the target image block is any image block in the image frame to be aligned;

[0101] a second pixel block determining subunit, configured to determine, in the reference image frame according to the global motion vector, a second pixel block that matches the first pixel block, wherein the second pixel block has the same size as the first pixel block;

[0102] The local motion vector determining subunit is configured to determine the offset between the first pixel block and the second pixel block as the local motion vector of the target image block.

[0103] Optionally, the second pixel block determining subunit includes:

[0104] A search range determining subunit, configured to determine a search range in the reference image frame according to the global motion vector and the first feature point in the target image block;

[0105] The second pixel block search subunit is configured to search within the search range for a second pixel block that matches the first pixel block.

[0106] Optionally, the above-mentioned search range determination subunit is specifically used to select a center point in the above-mentioned reference image frame, and the offset between the above-mentioned center point and the first feature point in the above-mentioned target image block is equal to the above-mentioned global motion vector; the above-mentioned search range is determined in the above-mentioned reference image frame according to a preset size, and the above-mentioned search range is centered on the above-mentioned center point.

[0107] Optionally, the image acquisition unit 601 includes:

[0108] An image acquisition subunit, configured to acquire at least two frames of images;

[0109] The image frame determination subunit is configured to determine one of the at least two image frames as the reference image frame, and determine the remaining image frames as the image frames to be aligned.

[0110] As can be seen from the above, in the present application scheme, the image frame to be aligned and the reference image frame are first obtained, the first feature point in the image frame to be aligned and the second feature point in the reference image frame are extracted, and then the motion vector of the image frame to be aligned is calculated based on the first feature point and the second feature point. Finally, the image frame to be aligned is aligned to the reference image frame based on the motion vector. The present application scheme can greatly reduce the amount of computation for image alignment and improve the efficiency of image alignment by extracting feature points representing high-frequency information in the image frame to be aligned and the reference image frame, and calculating the motion vector of the image frame to be aligned based on the extracted feature points.

[0111] Figure 7 This is a schematic diagram of the structure of a terminal device provided in one embodiment of the present application. Figure 7 As shown, the terminal device 7 of this embodiment includes: at least one processor 70 ( Figure 7 Only one is shown), a memory 71, and a computer program 72 stored in the memory 71 and executable on the at least one processor 70. When the processor 70 executes the computer program 72, the following steps are implemented:

[0112] Obtaining the image frame to be aligned and the reference image frame;

[0113] Extracting a first feature point from the image frame to be aligned and a second feature point from the reference image frame;

[0114] Calculating the motion vector of the image frame to be aligned according to the first feature point and the second feature point;

[0115] The image frame to be aligned is aligned with the reference image frame according to the motion vector.

[0116] Assuming that the above is the first possible implementation, in a second possible implementation provided on the basis of the first possible implementation, calculating the motion vector of the image frame to be aligned based on the first feature point and the second feature point includes:

[0117] Calculating the global motion vector of the image frame to be aligned based on the first feature point and the second feature point;

[0118] Dividing the image frame to be aligned into a preset number of image blocks;

[0119] Calculating the local motion vector of each image block in the image frame to be aligned according to the global motion vector and the first feature point;

[0120] The motion vector is calculated based on the global motion vector and the local motion vector.

[0121] In a third possible implementation provided on the basis of the second possible implementation, the step of calculating the global motion vector of the image frame to be aligned based on the first feature point and the second feature point includes:

[0122] Generate a feature descriptor for each first feature point and a feature descriptor for each second feature point;

[0123] Determining a target first feature point that matches a target second feature point, where the target second feature point is any second feature point in the reference image frame, and a feature descriptor of the target second feature point matches a feature descriptor of the target first feature point;

[0124] Calculate the offset between the first target feature point and the second target feature point, and add the offset to the offset list;

[0125] The global motion vector of the image frame to be aligned is determined according to the offset list.

[0126] In a fourth possible implementation provided on the basis of the third possible implementation, determining the global motion vector of the image frame to be aligned according to the offset list includes:

[0127] The offset that appears the most times in the offset list is determined as the global motion vector.

[0128] In a fifth possible implementation provided on the basis of the second possible implementation, calculating the local motion vector of each image block in the image frame to be aligned according to the global motion vector and the first feature point includes:

[0129] Determine a first pixel block in the target image block with a first feature point in the target image block as the center, where the target image block is any image block in the image frame to be aligned;

[0130] determining, in the reference image frame according to the global motion vector, a second pixel block that matches the first pixel block, wherein the second pixel block has the same size as the first pixel block;

[0131] The offset between the first pixel block and the second pixel block is determined as the local motion vector of the target image block.

[0132] In a sixth possible implementation provided on the basis of the fifth possible implementation, determining, in the reference image frame according to the global motion vector, a second pixel block matching the first pixel block includes:

[0133] Determining a search range in the reference image frame according to the global motion vector and the first feature point in the target image block;

[0134] A second pixel block matching the first pixel block is searched within the search range.

[0135] In a seventh possible implementation provided on the basis of the sixth possible implementation, determining the search range in the reference image frame according to the global motion vector and the first feature point in the target image block includes:

[0136] Selecting a center point in the reference image frame, where the offset between the center point and the first feature point in the target image block is equal to the global motion vector;

[0137] The search range is determined in the reference image frame according to a preset size, and the search range is centered on the center point.

[0138] In an eighth possible implementation provided on the basis of the first possible implementation, or the second possible implementation, or the third possible implementation, or the fourth possible implementation, or the fifth possible implementation, or the sixth possible implementation, or the seventh possible implementation, the acquiring of the image frame to be aligned and the reference image frame includes:

[0139] Acquire at least two frames of images;

[0140] One of the at least two frames of images is determined as the reference image frame, and the remaining images are determined as the image frames to be aligned.

[0141] The terminal device 7 can be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The terminal device may include, but is not limited to, a processor 70 and a memory 71. Those skilled in the art will understand that Figure 7 It is only an example of the terminal device 7 and does not constitute a limitation on the terminal device 7. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.

[0142] The processor 70 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.

[0143] In some embodiments, the memory 71 may be an internal storage unit of the terminal device 7, such as a hard disk or memory of the terminal device 7. In other embodiments, the memory 71 may also be an external storage device of the terminal device 7, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 7. Furthermore, the memory 71 may include both an internal storage unit of the terminal device 7 and an external storage device. The memory 71 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 71 may also be used to temporarily store data that has been output or is about to be output.

[0144] As can be seen from the above, in the present application scheme, the image frame to be aligned and the reference image frame are first obtained, the first feature point in the image frame to be aligned and the second feature point in the reference image frame are extracted, and then the motion vector of the image frame to be aligned is calculated based on the first feature point and the second feature point. Finally, the image frame to be aligned is aligned to the reference image frame based on the motion vector. The present application scheme can greatly reduce the amount of computation for image alignment and improve the efficiency of image alignment by extracting feature points representing high-frequency information in the image frame to be aligned and the reference image frame, and calculating the motion vector of the image frame to be aligned based on the extracted feature points.

[0145] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0146] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the above-mentioned device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0147] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0148] An embodiment of the present application provides a computer program product. When the computer program product is run on a terminal device, the terminal device executes the steps in the above-mentioned various method embodiments.

[0149] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The above-mentioned computer program can be stored in a computer-readable storage medium, and the computer program, when executed by the processor, can implement the steps of the above-mentioned various method embodiments. Among them, the above-mentioned computer program includes computer program code, and the above-mentioned computer program code can be in source code form, object code form, executable file or some intermediate form. The above-mentioned computer-readable medium may at least include: any entity or device capable of carrying the computer program code to the terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, a computer-readable medium cannot be an electric carrier signal and a telecommunication signal.

[0150] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0151] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0152] In the embodiments provided in this application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely illustrative. For example, the division of the above modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0153] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0154] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. An image alignment method, characterized in that: include: Obtaining the image frame to be aligned and the reference image frame; Extracting first feature points in the image frame to be aligned and second feature points in the reference image frame; Calculating a global motion vector of the image frame to be aligned according to the first feature point and the second feature point; Dividing the image frame to be aligned into a preset number of image blocks; Calculating a local motion vector of each image block in the image frame to be aligned according to the global motion vector and the first feature point; Calculating the motion vector of the image frame to be aligned according to the global motion vector and the local motion vector, wherein the motion vector of each image block is a sum of the local motion vector of the image block and the global motion vector of the image frame to be aligned; The image frame to be aligned is aligned with the reference image frame according to the motion vector.

2. The image alignment method according to claim 1, wherein: The calculating the global motion vector of the image frame to be aligned according to the first feature point and the second feature point includes: Generate a feature descriptor for each first feature point and a feature descriptor for each second feature point; Determining a target first feature point that matches a target second feature point, where the target second feature point is any second feature point in the reference image frame, and a feature descriptor of the target second feature point matches a feature descriptor of the target first feature point; Calculating an offset between the first target feature point and the second target feature point, and adding the offset to an offset list; The global motion vector of the image frame to be aligned is determined according to the offset list.

3. The image alignment method according to claim 2, wherein: The determining the global motion vector of the image frame to be aligned according to the offset list includes: The offset that appears the most times in the offset list is determined as the global motion vector.

4. The image alignment method according to claim 1, wherein: The calculating, based on the global motion vector and the first feature point, the local motion vector of each image block in the image frame to be aligned includes: Determine a first pixel block in the target image block with a first feature point in the target image block as the center, where the target image block is any image block in the image frame to be aligned; determining, in the reference image frame according to the global motion vector, a second pixel block that matches the first pixel block, wherein the second pixel block has the same size as the first pixel block; An offset between the first pixel block and the second pixel block is determined as a local motion vector of the target image block.

5. The image alignment method according to claim 4, characterized in that: The determining, in the reference image frame according to the global motion vector, a second pixel block that matches the first pixel block comprises: determining a search range in the reference image frame according to the global motion vector and a first feature point in the target image block; A second pixel block matching the first pixel block is searched within the search range.

6. The image alignment method according to claim 5, characterized in that: The step of determining a search range in the reference image frame according to the global motion vector and the first feature point in the target image block includes: Selecting a center point in the reference image frame, where an offset between the center point and a first feature point in the target image block is equal to the global motion vector; The search range is determined in the reference image frame according to a preset size, and the search range is centered on the center point.

7. The image alignment method according to any one of claims 1 to 6, characterized in that: The obtaining of the image frame to be aligned and the reference image frame includes: Acquire at least two frames of images; One of the at least two frames of images is determined as the reference image frame, and the remaining images are determined as the image frames to be aligned.

8. An image alignment device, characterized in that: include: An image acquisition unit, configured to acquire an image frame to be aligned and a reference image frame; a feature point extraction unit, configured to extract first feature points from the image frame to be aligned and second feature points from the reference image frame; a global motion vector calculation subunit, configured to calculate a global motion vector of the image frame to be aligned based on the first feature point and the second feature point; An image segmentation subunit, configured to segment the image frame to be aligned into a preset number of image blocks; a local motion vector calculation subunit, configured to calculate a local motion vector of each image block in the image frame to be aligned according to the global motion vector and the first feature point; a motion vector calculation subunit, configured to calculate the motion vector of the image frame to be aligned based on the global motion vector and the local motion vector, wherein the motion vector of each image block is the sum of the local motion vector of the image block and the global motion vector of the image frame to be aligned; An image alignment unit is configured to align the image frame to be aligned with the reference image frame according to the motion vector.

9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method and apparatus for processing global motion vector of screen image

    CN107197278A

  • Image alignment method and mobile terminal

    CN107527360A