An image stitching method and device, a computer device, and a storage medium

By applying self-learning iterative regression and teacher-student self-learning model in image stitching, the problems of stitching errors and RANSAC threshold adjustment caused by structural differences between images are solved, and a more efficient and robust image stitching effect is achieved.

CN114066793BActive Publication Date: 2025-06-24SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111358575.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-16
Publication Date
2025-06-24
Estimated Expiration
2041-11-16

AI Technical Summary

Technical Problem

Existing image stitching methods can easily lead to ghosting or distortion in the stitching results when processing structural differences between images, and require manual adjustment of the RANSAC algorithm threshold to obtain better performance.

Method used

The image stitching model based on the self-learning iterative regression idea is adopted, and the feature point pair set is updated through iterative regression of the teacher-student self-learning model, an image transformation model is generated, and the matching feature point pair is initialized using the local feature point symmetric constraint idea.

Benefits of technology

The problem of manually adjusting the threshold of the RANSAC algorithm is effectively avoided by learning different models, improving the accuracy of the initial matching point pairs, and optimizing the image stitching model to reduce parameter sensitivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114066793B_ABST
    Figure CN114066793B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses an image stitching method, device, computer device, and storage medium, including: regarding the model solving problem in image stitching as a regression problem through an image stitching model learning and solving method based on the idea of self-learning iterative regression, and gradually approaching and solving the optimal image transformation model by iteratively updating the teacher-student model multiple times, effectively avoiding the problem of manually adjusting the RANSAC algorithm threshold when learning different models, and being able to optimize and perform regression solving for different image transformation models. In addition, a teacher model initialization method using the idea of local feature point symmetry constraint is used to solve the initial matching feature point pairs, improving the accuracy rate of the initial matching point pairs compared to the RANSAC algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of computer technology, and in particular, to an image stitching method, apparatus, computer device, and storage medium. Background Art

[0002] In certain specific scenarios in real life, such as monitoring, due to the limitations of camera hardware conditions, when multiple cameras capture videos or images, the obtained image data will have brightness differences due to changes in environmental light sources or camera exposure parameters. In addition, the image data will also have structural differences due to different angles and positions of the cameras during shooting, that is, the same object will have different shapes or positions in different images. Therefore, directly combining and stitching the image data will result in serious stitching errors. To solve the above problems, a variety of effective image stitching methods have been proposed, which can be mainly divided into the following three categories:

[0003] 1. Global stitching algorithm. Such algorithms first extract the feature points of the images, then match the feature point pairs, use the correct feature point pairs to solve the projection transformation matrix between the images, transform all the images to a reference image plane according to the projection transformation matrix, and finally perform a smooth transition process on the overlapping regions of the images to finally obtain a high-quality stitched image. This type of method uses a globally consistent projection transformation matrix and related image fusion algorithms, and has a good stitching effect for images without parallax changes in the overlapping regions. However, this method cannot handle the structural difference phenomenon between images, which will cause ghosting or distortion in the stitching results of images with parallax changes in the overlapping regions.

[0004] 2. Optimal seam algorithm. Such algorithms hope to find an optimal dividing line within the overlapping regions of the input image data, and separate the areas with structural differences through this line to obtain divided regions with different structures. If this method can successfully calculate the optimal dividing line, it can successfully avoid the problem of structural differences and obtain excellent stitched images. However, often many image data do not have such a dividing line, so the optimal seam algorithm will still cause tearing and bending of object edges in the stitching results.

[0005] III. Local structure stitching algorithm. The core idea of this type of algorithm is to calculate multiple local transformation projection matrices to meet different image structures, thereby eliminating the structural differences in the overlapping regions of the input image data and finally obtaining a stitched image with good quality. However, the local structure stitching algorithm uses the RANSAC algorithm to obtain matching feature point pairs and thus learn different mathematical models. When there is no depth change in the overlapping region scene of the input data, that is, when the input data is taken from a single viewpoint of the camera, the threshold of the RANSAC algorithm needs to be set very small to ensure that all the matching point pairs are correct; while when there is depth change in the overlapping region scene of the input data and the input data is taken from multiple viewpoints of the camera, the RANSAC algorithm needs to set a relatively large threshold to ensure that the correct matching point pairs are not filtered out. At this time, the threshold is usually 10 to 100 times that of the single-viewpoint data. If there are incorrect matches among the matching point pairs at this time, the stitching result will have serious errors. Therefore, this type of method requires manual adjustment of the threshold to have good performance on each dataset.

[0006] In summary, the existing image stitching methods mainly have the following problems: abnormal phenomena such as blurring and ghosting appear in the stitched panoramic image, and manual adjustment of the threshold is required to have good performance on each dataset. Summary of the Invention

[0007] Embodiments of the present invention provide an image stitching method, device, computer device, and storage medium.

[0008] To solve the above technical problems, an embodiment of the present invention adopts a technical solution: providing an image stitching method, including the following steps:

[0009] Obtain a reference image and an image to be stitched, and respectively detect the feature points of the reference image and the image to be stitched according to a preset scale-invariant feature transform algorithm to obtain a reference image feature point set and an image to be stitched feature point set;

[0010] Detect the reference image feature point set and the image to be stitched feature point set according to a preset nearest neighbor saliency algorithm to obtain a rough detection matching point pair set;

[0011] Divide the elements in the rough detection matching point pair set into a pseudo-correct feature point pair set and a pseudo-incorrect feature point pair set according to a preset local feature point symmetry constraint teacher model;

[0012] Iteratively regress and update the pseudo-correct feature point pair set and the incorrect feature point pair set according to a preset teacher-student self-learning model to obtain a correct point pair set;

[0013] Generate an image transformation model according to the correct point pair set, and use the image transformation model to transform and stitch the image to be stitched to obtain a stitched image.

[0014] Further, detecting the feature points of the reference image and the image to be stitched respectively according to the preset Scale-Invariant Feature Transform (SIFT) algorithm to obtain a reference image feature point set and an image to be stitched feature point set includes:

[0015] Calculating the respective Gaussian operators of the reference image I(x, y) and the image to be stitched J(x, y) using the preset Scale-Invariant Feature Transform (SIFT) algorithm to establish a Difference of Gaussian (DoG) pyramid;

[0016] Obtaining the neighborhood gradient direction m and amplitude θ of the reference image feature points and the image to be stitched respectively through the Difference of Gaussian (DoG) pyramid;

[0017] Combining the abscissa x, ordinate y, neighborhood gradient direction m, and amplitude in the reference image I(x, y) into a 128-dimensional feature descriptor to generate the reference image feature point set F I , F I ={u i}={(x i , y i )}, u i is the reference image feature point, and x i and y i are the coordinates of the reference image feature point respectively;

[0018] Combining the abscissa x, ordinate y, neighborhood gradient direction m, and amplitude in the image to be stitched J(x, y) into a 128-dimensional feature descriptor to generate the image to be stitched feature point set F J , F J ={v j}={(x j , y j )}, v j is the reference image feature point, and x j and y j are the coordinates of the reference image feature point respectively.

[0019] Further, detecting the reference image feature point set and the image to be stitched feature point set according to the preset Nearest Neighbor with Ratio Test (NNDR) algorithm to obtain a rough detection matching point pair set includes:

[0020] The rough detection matching point pair set F i =(u i , v′ i ),

[0021] where v′ i ∈F J is the feature point roughly matched with the reference image feature point, and satisfies di1 = ||u i - v i '||², d i2 = ||u i - v j2 ||², v i ' and v j2 respectively represent the feature point with the smallest Euclidean distance to u and the feature point with the second smallest Euclidean distance in F J in the Euclidean distance. i

[0022] Furthermore, the coarse detection and matching point pair set is divided into a pseudo-correct feature point pair set and a pseudo-incorrect feature point pair set according to the preset local feature point symmetry constraint teacher model, including:

[0023] Using the preset local feature point symmetry constraint teacher model H t to transform the reference image feature point u i to obtain the matching detected feature point v i ;

[0024] Search for the truly corresponding matching feature point v' i in the preset circular area adjacent to v i ;

[0025] Perform local area constraint and symmetry constraint on the truly corresponding matching feature point v' i to obtain the pseudo-correct feature point pair set F1, and use the set composed of the feature points in the reference image feature point set F i except the pseudo-correct feature points as the pseudo-incorrect feature point pair set.

[0026] Furthermore, the image transformation model includes multiple local transformation models. Using the image transformation model to transform and splice the to-be-spliced image to obtain a spliced image, including:

[0027] Divide the to-be-spliced image into multiple local areas according to the processing area of the local transformation model:

[0028] Use the local transformation model to perform transformations on the corresponding local areas respectively;

[0029] Align the transformed multiple local areas on the reference image to obtain a spliced image.

[0030] To solve the above technical problems, an embodiment of the present invention also provides an image splicing device, including:

[0031] An acquisition module, configured to acquire a reference image and an image to be stitched, and respectively detect feature points of the reference image and the image to be stitched according to a preset scale-invariant feature transform algorithm to obtain a reference image feature point set and an image to be stitched feature point set;

[0032] A processing module, configured to detect the reference image feature point set and the image to be stitched feature point set according to a preset nearest neighbor saliency algorithm to obtain a rough detection matching point pair set;

[0033] The processing module is configured to divide elements in the rough detection matching point pair set into a pseudo-correct feature point pair set and a pseudo-incorrect feature point pair set according to a preset local feature point symmetry constraint teacher model;

[0034] The processing module is configured to iteratively regress and update the pseudo-correct feature point pair set and the incorrect feature point pair set according to a preset teacher-student self-learning model to obtain a correct point pair set;

[0035] An execution module, configured to generate an image transformation model according to the correct point pair set, and use the image transformation model to transform and stitch the image to be stitched to obtain a stitched image.

[0036] Further, the processing module includes:

[0037] A first acquisition sub-module, configured to respectively calculate the reference image I(x, y) and the image to be stitched J(x, y) by using a preset scale-invariant feature transform algorithm to obtain their respective Gaussian operators and establish a Gaussian difference pyramid;

[0038] A first processing sub-module, configured to respectively obtain the neighborhood gradient direction m and amplitude θ of the reference image feature points and the image to be stitched through the Gaussian difference pyramid;

[0039] A first execution sub-module, configured to combine the abscissa x, ordinate y, neighborhood gradient direction m and amplitude in the reference image I(x, y) into a 128-dimensional feature descriptor to generate the reference image feature point set F I , F I ={u i}={(x i , y i )}, u i is the reference image feature point, x i and y i are respectively the coordinates of the reference image feature point;

[0040] A second execution sub-module, configured to combine the abscissa x, ordinate y, neighborhood gradient direction m and amplitude θ in the image to be stitched J(x, y) into a 128-dimensional feature descriptor to generate the image to be stitched feature point set F J , FJ = {v j} = {(x j , y j )}, where v j is the reference image feature point, and x j and y j are the coordinates of the reference image feature point respectively.

[0041] Furthermore, the rough detection matching point pair set F i = (u i , v' i ),

[0042] where v' i ∈ F J is the feature point roughly matched with the reference image feature point, satisfying d i1 = ‖u i - v i '‖2, d i2 = ||u i - v j2 ||2, v i ' and v j2 respectively represent the feature point with the smallest Euclidean distance and the feature point with the second smallest Euclidean distance to u J in F i .

[0043] Furthermore, the processing module includes:

[0044] The second acquisition sub-module is used to transform the reference image feature point u t by using the preset local feature point symmetry constraint teacher model H i to obtain the matching detected feature point v i ;

[0045] The second processing sub-module is used to search for the truly corresponding matching feature point v' i in the preset circular area adjacent to v i ;

[0046] The third execution sub-module is used to perform local area constraint and symmetry constraint on the truly corresponding matching feature point v' i to obtain the pseudo-correct feature point pair set F1, and use the set composed of the feature points in the reference image feature point set F i except the pseudo-correct feature points as the pseudo-incorrect feature point pair set.

[0047] Furthermore, the execution module includes:

[0048] A third acquisition sub-module, configured to divide the to-be-stitched image into multiple local regions according to the processing region of the local transformation model:

[0049] A third processing sub-module, configured to use the local transformation model to perform transformations on corresponding local regions respectively;

[0050] A fourth execution sub-module, configured to align the multiple transformed local regions on the reference image to obtain a stitched image.

[0051] To solve the above technical problems, an embodiment of the present invention further provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the above-described image stitching method.

[0052] The beneficial effects of the embodiments of the present invention are as follows: Through the image stitching model learning and solving method based on the self-learning iterative regression idea, the model solving problem in image stitching is regarded as a regression problem, and the teacher-student model is gradually updated through multiple iterations to gradually approximate and solve the optimal image transformation model, effectively avoiding the problem of manually adjusting the RANSAC algorithm threshold when learning different models, and being able to optimize and perform regression solving for different image transformation models. In addition, the teacher model initialization method using the local feature point symmetry constraint idea is used to solve the initial matching feature point pairs, which improves the accuracy rate of the initial matching point pairs compared with the RANSAC algorithm. Description of the Drawings

[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0054] Figure 1 It is a schematic flowchart of the basic process of an image stitching method provided by an embodiment of the present invention;

[0055] Figure 2 It is a basic structural block diagram of an image stitching device provided by an embodiment of the present invention;

[0056] Figure 3 It is a basic structural block diagram of a computer device provided by an embodiment of the present invention. Detailed Embodiments

[0057] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention.

[0058] In some of the processes described in the specification, claims, and the above-mentioned drawings of the present invention, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" herein are used to distinguish different messages, devices, modules, etc., do not represent a sequence, and do not limit that "first" and "second" are of different types.

[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0060] Embodiment

[0061] Those skilled in the art can understand that the "terminal" and "terminal device" used herein include both devices with a wireless signal receiver, which only have a wireless signal receiver without transmission capabilities, and devices with receiving and transmitting hardware, which have receiving and transmitting hardware capable of performing two-way communication on a two-way communication link. Such devices may include: cellular or other communication devices, which have a single-line display or a multi-line display or a cellular or other communication device without a multi-line display; PCS (Personal Communications Service), which can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which may include a radio frequency receiver, a pager, Internet / intranet access, a web browser, a notepad, a calendar, and / or a GPS (Global Positioning System) receiver; conventional laptop and / or palm computers or other devices, which are conventional laptop and / or palm computers or other devices with and / or including a radio frequency receiver. The "terminal" and "terminal device" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to operate locally and / or in a distributed manner at any other location on the earth and / or in space. The "terminal" and "terminal device" used herein can also be a communication terminal, an Internet access terminal, a music / video playback terminal, such as a PDA, a MID (Mobile Internet Device), and / or a mobile phone with music / video playback functions, or can also be a smart TV, a set-top box, and other devices.

[0062] As Figure 1 shown, Figure 1 is a schematic diagram of the basic process of an image stitching method provided by an embodiment of the present invention, which is characterized by including the following steps:

[0063] S1. Obtain a reference image and an image to be stitched, and respectively detect the feature points of the reference image and the image to be stitched according to a preset scale-invariant feature transform algorithm to obtain a reference image feature point set and an image to be stitched feature point set;

[0064] In an embodiment of the present invention, a reference image I and an image J to be stitched are input. The two input images, the reference image I and the image J to be stitched, must have a considerable proportion of overlapping areas, and the resolution sizes of the images must be the same.

[0065] S2. Detect the reference image feature point set and the image to be stitched feature point set according to a preset nearest neighbor saliency algorithm to obtain a rough detection matching point pair set;

[0066] Use the existing Scale-Invariant Feature Transform (SIFT) algorithm to detect feature points in I and J respectively, and obtain two sets of feature point sets F I and F J . For F I and F J Use the existing Nearest Neighbor Significance (2-NN) algorithm to obtain a set of rough detection matching point pairs F i .

[0067] In an embodiment of the present invention, the set of rough detection matching point pairs F i =(u i , v′ i ),

[0068] where v′ i ∈F J is the feature point that is roughly matched with the feature point of the reference image, and satisfies d i1 =‖u i -v i ′‖2, d i2 =||u i -v j2 ||2, v i ′ and v j2 respectively represent the feature point with the smallest Euclidean distance and the feature point with the second smallest Euclidean distance from u J in F i . Preferably, δ = 0.8. It should be noted that the feature point v j2 with the second smallest Euclidean distance is the feature point selected as the second one in the ascending order of Euclidean distance.

[0069] S3. Divide the elements in the set of rough detection matching point pairs into a set of pseudo-correct feature point pairs and a set of pseudo-incorrect feature point pairs according to a preset local feature point symmetry constraint teacher model;

[0070] S4. Iteratively regress and update the set of pseudo-correct feature point pairs and the set of incorrect feature point pairs according to a preset teacher-student self-learning model to obtain a set of correct point pairs;

[0071] S5. Generate an image transformation model according to the set of correct point pairs, and use the image transformation model to transform and splice the to-be-spliced image to obtain a spliced image.

[0072] An image stitching method provided by an embodiment of the present invention uses an image stitching model learning and solving method based on the idea of self-learning iterative regression. It regards the model solving problem in image stitching as a regression problem, and gradually approaches and solves the optimal image transformation model by iteratively updating the teacher-student model multiple times. This effectively avoids the problem of manually adjusting the RANSAC algorithm threshold when learning different models, and can optimize and perform regression solving for different image transformation models. In addition, the initialization method of the teacher model using the idea of local feature point symmetry constraint is used to solve the initial matching feature point pairs, which improves the accuracy rate of the initial matching point pairs compared with the RANSAC algorithm.

[0073] In one embodiment of the present invention, the steps of respectively detecting the feature points of the reference image and the image to be stitched according to the preset scale-invariant feature transform algorithm to obtain the reference image feature point set and the image to be stitched feature point set specifically include:

[0074] Step 1: Use the preset scale-invariant feature transform algorithm to calculate the respective Gaussian operators of the reference image I(x, y) and the image to be stitched J(x, y) to establish a Gaussian difference pyramid;

[0075] Step 2: Obtain the neighborhood gradient direction m and amplitude θ of the reference image feature points and the image to be stitched through the Gaussian difference pyramid;

[0076] Step 3: Combine the abscissa x, ordinate y, neighborhood gradient direction m, and amplitude in the reference image I(x, y) into a 128-dimensional feature descriptor to generate the reference image feature point set F I , F I ={u i}={(x i ,y i )}, u i is the reference image feature point, x i and y i are the coordinates of the reference image feature point respectively;

[0077] Step 4: Combine the abscissa x, ordinate y, neighborhood gradient direction m, and amplitude θ in the image to be stitched J(x, y) into a 128-dimensional feature descriptor to generate the image to be stitched feature point set F J , F J ={v j}={(x j ,y j )}, v j is the reference image feature point, x j and y j are the coordinates of the reference image feature point respectively.

[0078] In an embodiment of the present invention, an original input image, i.e., a reference image I(x, y), is provided, where x and y respectively represent the horizontal and vertical coordinates of pixel points in the image. At this time, the existing Scale-Invariant Feature Transform (SIFT) algorithm is used to calculate the reference image I(x, y) and the image to be stitched J(x, y) to obtain a Gaussian operator and establish a Gaussian difference pyramid, thereby obtaining the gradient direction m and amplitude θ of the feature point neighborhood. The combination of (x, y, m, θ) forms a 128-dimensional feature descriptor, and a set of feature points F is generated. I and F J , and each feature point has a corresponding 128-dimensional feature descriptor.

[0079] F I ={u i}={(x i ,y i )},

[0080] F J ={v j}={(x j ,y j )},

[0081] where u i is the feature point of the reference image, x i and y i are the coordinates of the feature point of the reference image respectively, v j is the feature point of the reference image, x j and y j are the coordinates of the feature point of the reference image respectively.

[0082] In an embodiment of the present invention, according to a preset local feature point symmetry constraint teacher model, the elements in the set of rough detection and matching point pairs are divided into a set of pseudo-correct feature point pairs and a set of pseudo-incorrect feature point pairs, including the following steps:

[0083] Step 1: Use a preset local feature point symmetry constraint teacher model H t to transform the feature point u i of the reference image to obtain a matching feature point v i to be detected;

[0084] In practical applications, when taking images, the camera coordinate system 0 with the Z-axis coinciding with the camera optical axis is used as the reference. The coordinates of a point P in the real space in the camera coordinate system 0 are U=(X0, Y0, Z0) T , and the homogeneous coordinates on the imaging plane are u i =(x0, y0, 1) T , and the coordinates in the camera coordinate system 1 are V=(X1, Y1, Z1) T , and the homogeneous coordinates on the imaging plane are v j=(x1,y1,1) T Then, according to the camera projection model, the camera coordinates and image coordinates in different coordinate systems can be expressed as follows:

[0085]

[0086]

[0087] i and j are expressed as the serial numbers of the feature points in the sets F I and F J Suppose the rotation angle of the camera coordinate system 1 relative to the camera coordinate system 0 in the horizontal direction is θ x , and the rotation angle in the vertical direction is θ y , then the coordinates of point P in different camera coordinate systems satisfy the following formula:

[0088]

[0089] Based on the above formula, the following teacher model H can be obtained t transforms the reference image feature point u i to obtain the matching feature point v to be detected i ,

[0090]

[0091] where K is the internal parameter matrix of the camera, which is an inherent parameter of the camera and can be obtained using the Zhang Zhengyou calibration method. H t represents the transformation model of the feature point pair and is also the initial teacher model

[0092] Step 2: Search for the truly corresponding matching feature point v' within a preset circular area adjacent to v i ; i

[0093] Step 3: Apply local area constraint and symmetry constraint to the truly corresponding matching feature point v' i to obtain the set F1 of pseudo-correct feature point pairs, and use the set composed of the feature points in the reference image feature point set F i except for the pseudo-correct feature points as the set of pseudo-incorrect feature point pairs

[0094] The correct point set in the local constraint area can be described as:

[0095] F LR ={(u i ,v′ i ), where ‖v i -v′ i ‖2≤l int , generally l int =15}​

[0096] Meanwhile, symmetric constraints are also required. Therefore, we can obtain the set F1 of pseudo-correct feature point pairs and the set F2 of pseudo-incorrect feature point pairs:

[0097] F1 = {(u i , v′ i ), where ‖u i - u′ i ‖2 ≤ l int , ‖v i - v′ i ‖2 ≤ l int , and generally l int = 15};

[0098] F2 = F i - F1.

[0099] In an embodiment of the present invention, the set of pseudo-correct feature point pairs and the set of incorrect feature point pairs are iteratively regressed and updated according to a preset teacher-student self-learning model to obtain a set of correct point pairs.

[0100] Among them, the obtained set F1 of pseudo-correct feature point pairs generates a student model H s , and based on H s , F1 and F2 are updated to obtain the updated feature point sets F 1_update and F 2_update .

[0101] First, generate the student model H s using the existing model solution formula as follows:

[0102]

[0103]

[0104] Among them, ([x1, y1, 1], [x2, y2, 1]) = ((u i , v′ i )) = F1.

[0105] For the i-th point pair (u Wi , v′ Wi ) in the set F2 of incorrect point pairs, the point u Wi on the reference image plane is transformed by H s to obtain the transformed point coordinates u WHi ,

[0106] u WHi = H s · u wi

[0107] Therefore, the updated correct point pair set F can be obtained. 1_update and the updated incorrect point pair set F 2_update are:

[0108] F 1_update = {(u Wi , v′ Wi ), ‖u WHi - v′ Wi ‖2 ≤ l s , generally l s = 30}

[0109] F 2_update = F i - F 1_update

[0110] Input the obtained F 1_update and F 2_update into the student model H s and continuously update it. When F 1_update no longer changes, the F 1_update obtained from the last update is regarded as the final correct feature point pair set F 1_final .

[0111] In an embodiment of the present invention, the image transformation model includes multiple local transformation models. Using the image transformation model to perform transformation and splicing on the to-be-spliced image to obtain a spliced image includes the following steps:

[0112] Step 1: Divide the to-be-spliced image into multiple local regions according to the processing area of the local transformation model:

[0113] Step 2: Use the local transformation model to perform transformations on the corresponding local regions respectively;

[0114] Step 3: Align the transformed multiple local regions on the reference image to obtain a spliced image.

[0115] In the embodiment of the present invention, according to the finally updated correct point pair set F 1_final , generate the final image transformation model H final . Align the image J final obtained after being transformed by H warp to I to obtain the final spliced image (I, J warp ).

[0116] The formula for generating the final image transformation model H final is the existing model solving formula as shown below:

[0117]

[0118]

[0119] where ([x1, y1, 1], [x2, y2, 1]) = ((u i , v′ i )) = F I .

[0120] Let the image transformation model H final be composed of N local transformation models H n . The image J to be stitched can also be divided into N corresponding local regions J n . Then, there is the image to be stitched after transformation:

[0121]

[0122] Then, the image to be stitched at this time can be aligned to the reference image to obtain the finally stitched output image, which is denoted as (I, J warp ).

[0123] In order to improve the stitching accuracy of the local structure image stitching algorithm and reduce the parameter sensitivity of the existing algorithm, the present invention proposes a teacher-student self-learning model regression image stitching method. First, the image stitching transformation model can be regarded as a regression problem, and the optimal image stitching transformation model can be gradually approximated and solved through multiple iterative regressions. Therefore, this approach can avoid the problem of manually adjusting the RANSAC threshold when learning different models. Based on the initial set of matching feature point pairs, the method makes a preliminary division based on the local symmetry constraint teacher model of the present invention to obtain a set of pseudo-correct point pairs and a set of pseudo-incorrect point pairs, and then generates an image stitching transformation model using the set of pseudo-correct point pairs. On this basis, the set of pseudo-incorrect point pairs is refined again to update the set of pseudo-correct point pairs and the set of pseudo-incorrect point pairs. Finally, after multiple iterative refinements, the final set of correct point pairs and the image stitching transformation model are obtained, thereby generating the final stitched transformation image.

[0124] To solve the above technical problems, an embodiment of the present invention further provides an image stitching device, including:

[0125] An embodiment of the present invention further provides an image stitching device. Specifically, please refer to Figure 2 , Figure 2 which is the basic structural block diagram of the image stitching device in this embodiment.

[0126] As shown in Figure 2As shown in the figure, an image stitching device includes: an acquisition module 2100, a processing module 2200, and an execution module 2300. Among them, the acquisition module 2100 is used to acquire a reference image and an image to be stitched, and respectively detect feature points of the reference image and the image to be stitched according to a preset scale-invariant feature transform algorithm to obtain a reference image feature point set and an image to be stitched feature point set; the processing module 2200 is used to detect the reference image feature point set and the image to be stitched feature point set according to a preset nearest neighbor saliency algorithm to obtain a rough detection matching point pair set; the processing module 2200 is used to divide elements in the rough detection matching point pair set into a pseudo-correct feature point pair set and a pseudo-wrong feature point pair set according to a preset local feature point symmetry constraint teacher model; the processing module 2200 is used to iteratively regress and update the pseudo-correct feature point pair set and the wrong feature point pair set according to a preset teacher-student self-learning model to obtain a correct point pair set; the execution module 2300 is used to generate an image transformation model according to the correct point pair set, and use the image transformation model to transform and stitch the image to be stitched to obtain a stitched image.

[0127] An image stitching device provided by an embodiment of the present invention, through an image stitching model learning and solving method based on the idea of self-learning iterative regression, regards the model solving problem in image stitching as a regression problem, and gradually approaches and solves the optimal image transformation model by iteratively updating the teacher-student model multiple times, effectively avoiding the problem of manually adjusting the RANSAC algorithm threshold when learning different models, and being able to optimize and regressively solve for different image transformation models. In addition, the teacher model initialization method using the idea of local feature point symmetry constraint is used to solve the initial matching feature point pair, which improves the correct rate of the initial matching point pair compared with the RANSAC algorithm.

[0128] In some embodiments, the processing module includes: a first acquisition sub-module, which is used to calculate the respective Gaussian operators of the reference image I(x, y) and the image to be stitched J(x, y) according to a preset scale-invariant feature transform algorithm to establish a Gaussian difference pyramid; a first processing sub-module, which is used to respectively obtain the neighborhood gradient direction m and amplitude θ of the reference image feature points and the image to be stitched through the Gaussian difference pyramid; a first execution sub-module, which is used to combine the abscissa x, ordinate y, neighborhood gradient direction m and amplitude in the reference image I(x, y) into a 128-dimensional feature descriptor to generate the reference image feature point set F I , F I ={u i}={(x i , y i )}, u i is the reference image feature point, x i and y iThey are the coordinates of the reference image feature points respectively; the second execution sub-module is used to combine the abscissa x, ordinate y, neighborhood gradient direction m, and amplitude θ in the image J(x, y) to be stitched into a 128-dimensional feature descriptor, and generate the set F of feature points of the image to be stitched J , F J ={v j}={(x j , y j )}, v j is the reference image feature point, and x j and y j are the coordinates of the reference image feature point respectively.

[0129] In some embodiments, the set F of roughly detected matching point pairs i =(u i , v' i ), where v' i ∈F J is the feature point roughly matched with the reference image feature point, and satisfies d i1 =‖u i -v i '‖2, d i2 =||u i -v j2 ||2, v i ' and v j2 respectively represent the feature point with the smallest Euclidean distance and the feature point with the second smallest Euclidean distance from u J in F i .

[0130] In some embodiments, the processing module includes: a second acquisition sub-module, which is used to use a preset local feature point symmetry constraint teacher model H t to transform the reference image feature point u i to obtain the matching feature point v i to be detected; a second processing sub-module, which is used to search for the truly corresponding matching feature point v' i in a preset circular area adjacent to v i ; a third execution sub-module, which is used to perform local area constraint and symmetry constraint on the truly corresponding matching feature point v' i to obtain a set F1 of pseudo-correct feature point pairs, and use the set composed of the feature points in the reference image feature point set F i except the pseudo-correct feature points as the set of pseudo-incorrect feature point pairs.

[0131] In some embodiments, the execution module includes: a third acquisition sub-module, configured to divide the image to be stitched into a plurality of local regions according to the processing region of the local transformation model; a third processing sub-module, configured to use the local transformation model to perform transformations on corresponding local regions respectively; and a fourth execution sub-module, configured to align the transformed plurality of local regions on the reference image to obtain a stitched image.

[0132] To solve the above technical problems, an embodiment of the present invention further provides a computer device. For details, please refer to Figure 3 , Figure 3 which is a basic structural block diagram of the computer device in this embodiment.

[0133] As Figure 3 shown, it is a schematic internal structure diagram of the computer device. As Figure 3 shown, the computer device includes a processor, a non-volatile storage medium, a memory, and a network interface connected through a system bus. Among them, the non-volatile storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The control information sequence can be stored in the database. When the computer-readable instructions are executed by the processor, the processor can implement an image stitching method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute an image stitching method. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art can understand that Figure 3 the structure shown in

[0134] In this embodiment, the processor is used to execute Figure 2 the specific content of the acquisition module 2100, the processing module 2200, and the execution module 2300 in

[0135] The computer device provided by the embodiment of the present invention regards the model solving problem in image stitching as a regression problem through an image stitching model learning and solving method based on the idea of self-learning iterative regression. By iteratively updating the teacher-student model multiple times to gradually approximate and solve the optimal image transformation model, it effectively avoids the problem of manually adjusting the RANSAC algorithm threshold when learning different models, and can optimize and perform regression solving for different image transformation models. In addition, the initialization method of the teacher model using the idea of local feature point symmetry constraint is used to solve the initial matching feature point pairs, which improves the accuracy rate of the initial matching point pairs compared with the RANSAC algorithm.

[0136] The present invention also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the image stitching method described in any one of the above embodiments.

[0137] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.

[0138] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limitation, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0139] The above are only some embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An image stitching method, characterized in that, Including the following steps: Obtain a reference image and an image to be stitched, and respectively detect the feature points of the reference image and the image to be stitched according to a preset Scale-Invariant Feature Transform (SIFT) algorithm to obtain a reference image feature point set and an image to be stitched feature point set; Detect the reference image feature point set and the image to be stitched feature point set according to a preset Nearest Neighbor Salience (NNS) algorithm to obtain a roughly detected matching point pair set; Classify the elements in the roughly detected matching point pair set into a pseudo-correct feature point pair set and a pseudo-incorrect feature point pair set according to a preset local feature point symmetry constraint teacher model; Iteratively regress and update the pseudo-correct feature point pair set and the incorrect feature point pair set according to a preset teacher-student self-learning model to obtain a correct point pair set; Generate an image transformation model according to the correct point pair set, and use the image transformation model to transform and stitch the image to be stitched to obtain a stitched image; The step of respectively detecting the feature points of the reference image and the image to be stitched according to a preset Scale-Invariant Feature Transform (SIFT) algorithm to obtain a reference image feature point set and an image to be stitched feature point set includes: Respectively calculate the Gaussian operators of the reference image I(x, y) and the image to be stitched J(x, y) using a preset Scale-Invariant Feature Transform (SIFT) algorithm to establish a Difference of Gaussian (DoG) pyramid; Respectively obtain the neighborhood gradient direction m and amplitude θ of the reference image feature points and the image to be stitched through the Difference of Gaussian (DoG) pyramid; Combine the abscissa \(x\), ordinate \(y\), neighborhood gradient direction \(m\) and amplitude in the reference image \(I(x,y)\) into a 128 - dimensional feature descriptor to generate the reference image feature point set \(F\). I , \(F\) I =\(\{u\) i \}=\{(x i ,y i )\}, \(u\) i is the reference image feature point, \(x\) i and \(y\) i are the coordinates of the reference image feature point respectively; Combine the abscissa \(x\), ordinate \(y\), neighborhood gradient direction \(m\), and amplitude \(\theta\) in the image \(J(x,y)\) to be stitched into a 128 - dimensional feature descriptor, and generate the set \(F\) of feature points of the image to be stitched J , \(F\) J =\(\{v\) j \}\)=\(\{(x\) j ,y\) j )\}, \(v\) j is the feature point of the reference image, \(x\) j and \(y\) j are the coordinates of the feature point of the reference image respectively.

2. The image stitching method according to claim 1, wherein The step of detecting the reference image feature point set and the image to be stitched feature point set according to a preset Nearest Neighbor Salience (NNS) algorithm to obtain a roughly detected matching point pair set includes: The rough detection matching point pair set F i =(u i , v i ′), where v i ′ ∈ F J is the feature point that is roughly matched with the reference image feature point and satisfies d i1 = ‖u i - v i ′‖², d i2 = ||u i - v j2 ||², v′ i and v j2 respectively represent the feature point with the smallest Euclidean distance and the feature point with the second smallest Euclidean distance from u J in F i to u.

3. The image stitching method according to claim 1, wherein The step of classifying the elements in the roughly detected matching point pair set into a pseudo-correct feature point pair set and a pseudo-incorrect feature point pair set according to a preset local feature point symmetry constraint teacher model includes: Using the preset local feature point symmetry constraint teacher model H t Perform transformation on the reference image feature point u i To obtain the matching feature point v to be detected i ; Search for the truly corresponding matching feature points within the adjacent preset circular area at v i Search for the truly corresponding matching feature points within the adjacent preset circular area For the matching feature points corresponding to the truth Perform local area constraint and symmetry constraint to obtain a set F1 of pseudo-correct feature point pairs, and use the set composed of the feature points in the reference image feature point set F i Except for the pseudo-correct feature points as the set of pseudo-incorrect feature point pairs.

4. The image stitching method according to claim 1, characterized in that, The image transformation model includes multiple local transformation models. Using the image transformation model to transform and stitch the image to be stitched to obtain a stitched image includes: Divide the image to be stitched into multiple local regions according to the processing regions of the local transformation models; Respectively transform the corresponding local regions using the local transformation models; Align the transformed multiple local regions on the reference image to obtain a stitched image.

5. An image stitching device, characterized in that, Including: An acquisition module, configured to obtain a reference image and an image to be stitched, and respectively detect the feature points of the reference image and the image to be stitched according to a preset Scale-Invariant Feature Transform (SIFT) algorithm to obtain a reference image feature point set and an image to be stitched feature point set; A processing module, configured to detect the reference image feature point set and the image to be stitched feature point set according to a preset Nearest Neighbor Salience (NNS) algorithm to obtain a roughly detected matching point pair set; The processing module is configured to classify the elements in the roughly detected matching point pair set into a pseudo-correct feature point pair set and a pseudo-incorrect feature point pair set according to a preset local feature point symmetry constraint teacher model; The processing module is configured to iteratively regress and update the pseudo-correct feature point pair set and the incorrect feature point pair set according to a preset teacher-student self-learning model to obtain a correct point pair set; An execution module, configured to generate an image transformation model according to the correct point pair set, and use the image transformation model to transform and splice the to-be-spliced image to obtain a spliced image; The processing module includes: A first acquisition sub-module, configured to use a preset scale-invariant feature transform algorithm to calculate the reference image I(x, y) and the to-be-spliced image J(x, y) respectively to obtain their respective Gaussian operators and establish a Gaussian difference pyramid; A first processing sub-module, configured to respectively obtain the neighborhood gradient direction m and amplitude θ of the reference image feature points and the to-be-spliced image through the Gaussian difference pyramid; The first execution sub-module is used to combine the abscissa x, ordinate y, neighborhood gradient direction m, and amplitude in the reference image I(x, y) into a 128-dimensional feature descriptor, and generate the reference image feature point set F I , F I ={u i}={(x i , y i )}, u i is the reference image feature point, and x i and y i are the coordinates of the reference image feature point respectively; The second execution sub-module is configured to combine the abscissa x, ordinate y, neighborhood gradient direction m, and amplitude θ in the to-be-stitched image J(x, y) into a 128-dimensional feature descriptor, and generate the to-be-stitched image feature point set F J , F J ={v j}={(x j , y j )}, where v j is the reference image feature point, and x j and y j are the coordinates of the reference image feature point respectively.

6. The image splicing device according to claim 5, wherein The set of roughly detected matching point pairs F i =(u i , v' i ), where \(v'\) i ∈F J is the feature point roughly matched with the reference image feature point, satisfying d i1 =‖u i -v i '‖², d i2 =||u i -v j2 ||², v i ' and v j2 respectively represent the feature point with the smallest Euclidean distance and the feature point with the second smallest Euclidean distance to u J in F i .

7. A computer device, comprising a memory and a processor, wherein computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the image splicing method according to any one of claims 1 to 4.

8. A storage medium storing computer-readable instructions, wherein when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the image splicing method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Image registration method and system based on Gaussian field constraint and manifold regularization

    CN109448031A

  • Gimbal system and image processing method therefor, and unmanned aerial vehicle

    WO2020037615A1