An Image Stitching Method and Device Based on Stable Linear Structures and Suture Line Estimation
Through the image stitching method based on stable linear structure and suture estimation, the characteristics are detected using SIFT, RANSAC and LSD algorithms, and combined with the APAP algorithm to optimize the transformation matrix, the parallax problem caused by depth of field difference is solved and the image stitching quality is improved.
Patent Information
- Application Number
- CN202111199009.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-14
- Filing Date
- 2021-10-14
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-10-14
AI Technical Summary
In the prior art, parallax problems caused by the difference in depth of field make the images captured by ordinary cameras poorly stitched, especially in the case of large parallaxes, it is difficult to meet the expected requirements.
The image stitching method based on stable linear structure and suture estimation is adopted, point features are obtained through SIFT and RANSAC algorithms, line features are detected by LSD algorithm, and homography matrix H is estimated in combination with APAP algorithm, sutures are obtained, and stable sutures are judged and spliced.
It effectively alleviates the difference in depth of field, reduces parallax, and improves the effect of image stitching, especially in the case of large parallax, which enhances the stability and quality of stitching.
Smart Images

Figure CN114202458B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image stitching, and in particular to an image stitching method and device based on stable linear structures and suture line estimation. Background Art
[0002] Image stitching technology is to stitch two or more images with overlapping regions into an image with a wider view angle and higher resolution. It is a key research content in the field of computer vision and also has extensive applications in life, such as virtual reality, panoramic video, and aerial photography image processing.
[0003] There are mainly two methods for obtaining large-view-angle images. One is to use a wide-angle lens (such as a fish-eye lens) to capture an image with an ultra-large field of view. The other is to use an ordinary camera to take multiple photos with overlapping regions, and then stitch them pairwise through stitching technology to finally synthesize an image with a wider view angle. However, wide-angle lenses are costly and can only be used in specific scenarios. Ordinary cameras are convenient to use and are more suitable for ordinary people's daily life. However, compared with the first method, the view angle of ordinary cameras is relatively narrow. Although the camera view angle can be increased by adjusting the camera focal length, the resolution of the images captured in this way is relatively low. Therefore, it is more difficult to stitch the images obtained by ordinary cameras. Currently, the research on image stitching technology mainly focuses on improving the algorithm performance, and the implementation of the present invention is also based on the images obtained by ordinary cameras.
[0004] Image stitching technology has been widely applied at an early stage, with both civilian and military applications, such as remote sensing image stitching and tank panoramic systems. However, the users of image stitching technology are mainly national institutions and research units. In recent years, with the progress of computer vision and image processing related technologies, as well as the development of supporting hardware, image stitching technology has gradually become a "popular black technology" in people's lives. People have gradually become the direct users of this technology. For example, virtual reality (VR), which used to require expensive and bulky equipment to achieve, can now be experienced by ordinary people only through lightweight VR glasses. Applications in daily life have made image stitching technology become a hot spot in computer vision again, and have thus promoted the development and evolution of more applications.
[0005] Since people take photos at relatively random angles, the scenes captured are likely to be located on different planes, which will cause differences in depth of field, resulting in large parallax, creating a huge obstacle to image registration and causing the final stitching effect to fail to meet the expected requirements. Therefore, image stitching under large parallax is still a very challenging topic to date. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide an image stitching method and device based on a stable linear structure and suture estimation, so as to alleviate the depth-of-field difference in the prior art and reduce parallax.
[0007] The present invention provides an image stitching method based on a stable linear structure and suture estimation, including:
[0008] S1: Obtain two images with an overlapping area, use SIFT to match the feature points of the two images, then use the RANSAC algorithm to eliminate false features to obtain point features, and use the LSD algorithm to detect and match line features;
[0009] S2: Segment the target image in the two images to obtain a grid, and combine the point features and line features to use the APAP algorithm to estimate the homography transformation matrix H between the source image and the target image;
[0010] S3: Based on the homography transformation matrix H, obtain the slope relationship between the source image and the target image and obtain cross features;
[0011] S4: Combine the point features, line features and cross features to obtain the expression of the grid, and at the same time obtain the suture between the two images;
[0012] S5: Determine whether the suture is a stable suture,
[0013] If so, stitch the source image and the target image.
[0014] Preferably, the step of using the LSD algorithm to detect and match line features includes:
[0015] Reduce the source image and the target icon image according to a preset ratio, and obtain the gradient value of each point;
[0016] Perform pseudo-sorting on the target image based on the gradient value of each point, set a gradient threshold, and divide it into N bins according to the gradient value. Obtain the maximum gradient value in each bin, and sort all the bins according to the maximum gradient value in each bin;
[0017] Obtain the bin with the largest gradient value, select a seed point from the bin with the largest gradient value, and determine the θ of the seed point l and the difference between the regional angle is less than a preset value;
[0018] If so, add the seed point to the regional update set, and re-execute the step of selecting a seed point from the bin with the largest gradient value and determining whether the difference between the θ of the seed point l and the regional angle is less than a preset value;
[0019] Fit the region update set to a rectangle and determine whether the number of points within the rectangle is greater than a point threshold;
[0020] If so, fit the rectangle to a straight line.
[0021] Preferably, the steps of obtaining two images with an overlapping region, using SIFT to match the feature points of the two images, using the RANSAC algorithm to remove false features to obtain point features, and using the LSD algorithm to detect and match line features include:
[0022] S101: Obtain the scale space of two images with an overlapping region and generate a difference-of-Gaussians pyramid;
[0023] S102: Calculate spatial extreme points based on the difference-of-Gaussians pyramid to obtain feature points;
[0024] S103: Obtain the amplitude, position, and direction of the feature points, describe and match the feature points using the KNN algorithm.
[0025] Preferably, the steps of segmenting the target image in the two images to obtain a grid, and estimating the homography transformation matrix H between the source image and the target image using the APAP algorithm by combining point features and line features include:
[0026] Establish an inlier set library I';
[0027] Randomly select 4 sample feature points to obtain a transformation matrix H, calculate the projection error of the remaining feature points with respect to the sample feature points, and if the projection error is less than a projection threshold, mark the feature points as the inlier set I;
[0028] Obtain the number of the inlier set I, and if the number of the inlier set I is greater than the number of feature points in the inlier set library I', update the inlier set library I' to the inlier set I, record the transformation matrix H at this time, and execute the step of randomly selecting 4 sample feature points to obtain a transformation matrix H, calculating the projection error of the remaining feature points with respect to the sample feature points, and if the projection error is less than a projection threshold, marking the feature points as the inlier set I.
[0029] Preferably, the steps of segmenting the target image in the two images to obtain a grid, and estimating the homography transformation matrix H between the source image and the target image using the APAP algorithm by combining point features and line features further include
[0030] Optimize the transformation matrix H using the APAP matrix:
[0031] The steps of optimizing the transformation matrix H using the APAP matrix include:
[0032] Homogenize the feature points of the target image and the source image, and standardize the coordinates of the target image and the source image;
[0033] Divide the target image with standardized coordinates into grids, calculate the Gaussian weights of all feature points in each grid, and perform singular value decomposition to obtain the optimized transformation matrix H.
[0034] Preferably, the steps of obtaining the slope relationship between the source image and the target image and obtaining the cross features based on the homography transformation matrix H include:
[0035] Obtain the slope of the matrix before and after the homography transformation of the target image;
[0036] Based on the slope of the matrix before and after the homography transformation of the target image, obtain the dividing lines of the overlapping area and the non-overlapping area between the target image and the source image;
[0037] Preferably, using the grids in step S3, the optimization of point features and line features is respectively carried out by using the following expressions:
[0038]
[0039] Among them, p′ is the matching point of the target image corresponding to the reference image, indicating the gap between the feature point of the homography transformation matrix H after the homography transformation and the actual matching feature point
[0040]
[0041] Among them, assume {p k} represents all the endpoints of the line segment cut by the grid in the source image, represented by The corresponding matching line segment in the target image is represented by a j x + b j y + c j = 0;
[0042] Perform perspective optimization of the cross feature by using the following formula:
[0043]
[0044] Among them, and are the normal vectors of the two line segments of the cross feature. A pair of cross features is composed of and The two line segments are represented by the endpoints as and
[0045] Reduce the projection distortion by the following formula:
[0046]
[0047] The steps of simultaneously obtaining the stitching line between two images include:
[0048] Calculating the stitching line, and the steps of calculating the stitching line include defining labels and allocating the labels;
[0049] Evaluating the stitching line, and using the following formula to evaluate the cost of the pixel blocks along the stitching line through structural similarity consistency:
[0050]
[0051] where SSIM represents the structural similarity, and its value range is [-1, 1], and the value is 1 if and only if the two pixel blocks are exactly the same;
[0052] Using the following formula to evaluate the cost of the stitching line evaluation pixel points through the difference of adjacent pixel values, and the specific expression is as follows:
[0053]
[0054] Using the following formula to construct a hybrid evaluation model, and the expression is:
[0055]
[0056] where represents the denoising cost.
[0057] On the other hand, the present invention provides an image stitching device based on a stable linear structure and a stitching line estimation, including:
[0058] Line feature matching module: Obtaining two images with an overlapping area, using SIFT to match the feature points of the two images, and then using the RANSAC algorithm to remove the false features to obtain point features, and using the LSD algorithm to detect and match the line features;
[0059] Homography transformation matrix acquisition module: Used to segment the target image in the two images to obtain a grid, and combining the point features and the line features to estimate the homography transformation matrix H between the source image and the target image by using the APAP algorithm;
[0060] Cross feature acquisition module: Used to obtain the slope relationship between the source image and the target image and obtain the cross feature based on the homography transformation matrix H;
[0061] Stitching line acquisition module: Combining the point features, the line features and the cross features to obtain the expression of the grid, and simultaneously obtaining the stitching line between the two images;
[0062] Stitching line determination module: Judging whether the stitching line is a stable stitching line,
[0063] If so, the source image of the quantity source is stitched with the target image.
[0064] The embodiments of the present invention bring the following beneficial effects: The present invention provides an image stitching method and device based on stable linear structures and suture estimation, which relates to the technical field of image stitching, and includes: obtaining two images with an overlapping area, using SIFT to match the feature points of the two images, and then using the RANSAC algorithm to remove false features to obtain point features, and detecting and matching line features; segmenting the target image in the two images to obtain a grid, and combining the point features and line features to estimate the homography transformation matrix H between the source image and the target image: based on the homography transformation matrix H, obtaining the slope relationship between the source image and the target image and obtaining cross features: combining the point features, line features and cross features to obtain the expression of the grid, and obtaining the suture between the two images; determining whether the suture is a stable suture, and if so, obtaining the stitching of the source image and the target image. Through the present invention, the technical problems of depth of field difference and parallax reduction in the prior art can be alleviated.
[0065] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are realized and obtained by the structures specifically pointed out in the specification, claims and drawings.
[0066] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0067] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0068] Figure 1 It is a flowchart of an image stitching method based on stable linear structures and suture estimation provided by an embodiment of the present invention; Detailed Embodiments
[0069] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0070] At present, since people take pictures at random angles, the scenery captured is likely to be located in different planes, which will cause differences in depth of field, resulting in a large parallax, creating a huge obstacle to image registration, and causing the final stitching effect to fail to meet the expected requirements. Based on this, an image stitching method and device based on stable linear structures and suture estimation provided by an embodiment of the present invention are used to alleviate the depth of field differences in the prior art and reduce parallax.
[0071] To facilitate the understanding of this embodiment, first, a detailed introduction to an image stitching method based on stable linear structures and suture estimation disclosed in an embodiment of the present invention will be given.
[0072] Embodiment 1:
[0073] The present invention provides an image stitching method based on stable linear structures and suture estimation, including:
[0074] S1: Obtain two images with overlapping regions, use SIFT to match the feature points of the two images, then use the RANSAC algorithm to eliminate false features to obtain point features, and use the LSD algorithm to detect and match line features;
[0075] Furthermore, the SIFT algorithm searches for stable points whose positions do not change during scale transformation in different scale spaces. Assuming I(x, y) is any input image, the scale space of an image can be defined as:
[0076] L(x,y,σ) = G(x,y,σ) * I(x,y)
[0077] where * represents the convolution operation, σ represents the size of the scale space. The larger σ is, the more blurred the image is and the easier it is to obtain overall information; the smaller σ is, the clearer the image is and the easier it is to obtain detailed information. G(x, y, σ) is the Gaussian function, defined as:
[0078]
[0079] After a series of scale space transformations (i.e., changing σ) and multiple downsamplings (the number is uncertain, generally 2 times, and downsampling means halving the image size), a Gaussian pyramid is obtained.
[0080] To find stable and invariant extreme points in the scale space, the Difference of Gaussian (DOG) function D(x, y, σ) needs to be used, defined as:
[0081] D(x,y,σ) = [G(x,y,kσ) - G(x,y,σ)] * I(x,y)
[0082] = L(x, y, kσ) - L(x, y, σ)
[0083] Among them, kσ and σ represent two consecutive smoothing scales of the image, and the obtained results are stored in the Difference of Gaussian (DoG) pyramid. In summary, the general structure of the DoG pyramid includes four layers with scales increasing from small to large and usually downsampled by 4 (which can be determined according to the image size). Each layer includes s groups of images processed by kσ several times;
[0084] Compare each pixel with a total of 26 pixels, including its 8 adjacent pixels and 9 pixels corresponding to its upper and lower layers respectively, to determine whether it is an extreme value;
[0085] To achieve the rotation invariance of feature points, define the local region range as:
[0086] r = 3 * 1.5σ
[0087] where r is the radius of the local circular region and σ is the scale.
[0088] It is necessary to obtain the local features in this local region, namely the gradient magnitude and the gradient direction. The expression for the gradient magnitude is:
[0089]
[0090] where x and y are the horizontal and vertical coordinates respectively;
[0091] Among them, the expression of L was mentioned before and it is a convolution operation with a Gaussian convolution kernel.
[0092] The expression for the gradient direction is:
[0093]
[0094] Then specify a direction for the key point according to the gradient direction histogram (the histogram represents the values within a certain range). The main direction is the direction corresponding to the value with the largest ordinate in the gradient histogram. If there is a value in the gradient histogram and this value is 80% of the maximum value, then the direction corresponding to this value is considered as the secondary direction.
[0095] After completing this step, each feature point has three pieces of information: position, scale, and direction;
[0096] In the embodiments provided by the present invention, it is necessary to correct the rotation main direction to ensure rotational invariance. Let the 8×8 pixel area around any feature point be the total calculation area. One candidate point can be obtained for each 4×4 pixel area, so there are 4 candidate points in an 8×8 area. The gradient histogram of each candidate point is counted and set to 8 bins (i.e., 8 main directions). SIFT generally uses a 4×4 candidate point to describe a key point, and each candidate point has 8 main directions, so a total of 128-dimensional feature vector of 4×4×8 is formed. Finally, normalization processing is performed to normalize the length of the feature vector to further remove the influence of illumination.
[0097] S2: Segment the target images in the two images to obtain grids, and estimate the homography transformation matrix H between the source image and the target image by combining point features and line features using the APAP algorithm;
[0098] S3: Based on the homography transformation matrix H, obtain the slope relationship between the source image and the target image and obtain the cross feature;
[0099] S4: Combine point features, line features and cross features to obtain the expression of the grid, and at the same time obtain the stitching line between the two images;
[0100] S5: Judge whether the stitching line is a stable stitching line,
[0101] If so, splice the source image and the target image.
[0102] Preferably, the steps of detecting and matching line features using the LSD algorithm include:
[0103] Reduce the source image and the target icon image in accordance with a preset ratio, and obtain the gradient value of each point;
[0104] Furthermore, reduce both horizontally and vertically by 0.8, and calculate the gradient value through the following formula:
[0105] Calculate the gradient value and gradient direction of this point through the four points in the lower right corner of each point. For the image I(x, y), there is:
[0106]
[0107] where g x and g y are the gradient values in the horizontal and vertical directions of the pixel point (x, y) respectively, and the gradient direction is expressed as:
[0108]
[0109] The gradient value is expressed as:
[0110]
[0111] Pseudo - sort the target image based on the gradient value of each point, set a gradient threshold, divide it into N bins equally according to the gradient value, obtain the maximum gradient value in each bin, and pseudo - sort all the bins according to the maximum gradient value in each bin;
[0112] Furthermore, because the number of pixel points in an image is large, pseudo - sorting can reduce the time complexity to O(n). First, evenly divide the range from 0 to the previously obtained maximum gradient value into 1024 equal parts. Each of these 1024 parts is denoted as a bin (it is easy to understand that each bin contains several gradient values). Find the maximum value in a bin, and sort the 1024 bins according to the maximum value (the gradient values within each bin are not sorted), then the result of pseudo - sorting can be obtained;
[0113] Obtain the bin with the largest gradient value, select a seed point from the bin with the largest gradient value, and determine whether the difference between the gradient direction and the regional angle is less than a preset value;
[0114] A smaller gradient value means a flat gradient region, which will produce a larger error when calculating the gradient. Therefore, in the LSD process, points with a gradient value less than ρ do not participate in the final construction process. Assume \(I_p\) is an ideal image without noise and other contaminations, and \(n_c\) is the noise, then we have:
[0115]
[0116] Among them, represents the image after being contaminated by noise (generally, in the process of image processing, it will be more or less affected by noise), represents the gradient calculation. Pixel points will have an angular error under the influence of noise. Therefore, we believe that under the quantification of the influence of noise, points with an angular error less than the threshold are added to the line segment detection process. Thus, qualified points need to satisfy:
[0117]
[0118] Among them, q is the boundary of (usually taken as 2 in practical situations). The right - hand part of the above formula is called the tolerance, denoted as τ, that is, |error_angle| ≤ τ. Then the calculation method of the threshold is:
[0119]
[0120] Therefore, it can be equivalently screened by comparing the preset value with the threshold ρ.
[0121] If so, add the seed point to the region update set, and re - execute the step of selecting the seed point from the bin with the largest gradient value and determining whether the difference between the gradient direction and the region angle is less than the preset value;
[0122] In the embodiment provided by the present invention, check whether, among the neighborhood of the seed point (the 8 pixel points around the seed point), the difference between the θ of these points l and the region angle is less than the aforementioned τ (generally set to 22.5, which can be set according to the actual situation). If it is satisfied, add it to the current region update set, and each time it is added, update the region angle to:
[0123]
[0124] where j is the subscript of the pixel point in the current region update set.
[0125] Furthermore, through the above steps, points with the difference between the obtained gradient threshold and the region angle less than the preset value can be obtained, and through iteration, execute the step of obtaining the bin with the largest gradient value, selecting the seed point from the bin with the largest gradient value, and determining whether the difference between the gradient direction and the region angle is less than the preset value;
[0126] In the embodiment provided by the present invention, since the result output by the aforementioned steps is subjected to rectangular fitting, the following scheme is required
[0127] 1) Calculate the center of the rectangle:
[0128]
[0129] where G is the gradient value obtained in step 2, and x and y are the abscissa and ordinate of the corresponding pixel respectively.
[0130] (5) It is necessary to obtain the direction of the rectangle and then perform the rectangle fitting operation. The direction of the matrix can be expressed as the angle of the corresponding eigenvector, and the corresponding eigenvector is the eigenvalue of the following matrix:
[0131]
[0132] where, m xx , m xy , m yy The expressions of are respectively:
[0133]
[0134]
[0135] After obtaining the expression of M, it is easy to obtain the eigenvector and the angle. The angle is the corresponding angle, and the solution process will not be introduced in detail.
[0136] Therefore, the minimum bounding rectangle corresponding to a region is obtained, and these two processes are also called rectangle fitting;
[0137] After rectangle fitting, it is necessary to judge whether the number of points in the fitted rectangle region is greater than the threshold R (which can be set differently according to the actual situation). If it is satisfied, the next NFA calculation is performed. If it is not satisfied, the region is truncated (that is, the rectangle region that meets the conditions within the region is intercepted as the output); specifically as follows:
[0138]
[0139] Among them, m and n are the sizes of the input image I, p is the probability that a certain point in the rectangle conforms to the value range of [0, 2π], and the expression is p = τ / π. k is the number of feature points in the fitted rectangle, and n is the number of pixel points in the fitted rectangle.
[0140] If the fitted rectangle region satisfies NFA ≤ ε, it is considered that the fitting is successful (generally, ε is taken as 1). Otherwise, p = p / 2, one row is reduced from the long side of the fitted rectangle, one row is reduced from the short side, and one row is reduced from the other long side, and NFA is calculated until the condition is met.
[0141] The line features can be fitted into a rectangle through the foregoing solution.
[0142] Fit the region update set into a rectangle, and determine whether the number of points in the rectangle is greater than the point threshold;
[0143] If it is satisfied, the rectangle is fitted into a straight line.
[0144] Preferably, the steps of obtaining two images with an overlapping region, using SIFT to match the feature points of the two images, using the RANSAC algorithm to remove false features to obtain point features, and using the LSD algorithm to detect and match line features include:
[0145] S101: Obtain the scale space of two images with an overlapping region and generate a Gaussian difference pyramid;
[0146] S102: Calculate the spatial extreme points based on the Gaussian difference pyramid to obtain feature points;
[0147] S103: Obtain the amplitude, position, and direction of the feature points, describe and match the feature points using the KNN algorithm.
[0148] Preferably, the steps of segmenting the target image in the two images to obtain a grid, and using the APAP algorithm to estimate the homography transformation matrix H between the source image and the target image by combining point features and line features include:
[0149] Build an inlier set library I';
[0150] Randomly select 4 sample feature points to obtain the transformation matrix H, calculate the projection error between the remaining feature points and the sample feature points. If the projection error is less than the projection threshold, mark the feature points as the inlier set I;
[0151] Furthermore, calculate the transformation matrix H through the following formula:
[0152]
[0153] where, is the rotation matrix, representing the amount of rotation of the image, represents the amount of horizontal and vertical displacement of the image, [h 31 h 32 represents the amount of deformation of the image in the horizontal and vertical directions. Generally, h 33 is set to 1.
[0154] The estimation method of H is as follows:
[0155]
[0156] where, [x, y, 1] T is the coordinate point of the selected target image, [x', y', 1] T is the corresponding coordinate point in the reference image that matches the target image;
[0157] The calculation formula of the projection error is as follows:
[0158]
[0159] where, n is the number of feature points in the feature point set, (x, y) is the coordinate of the feature point in the target image, and (x', y') is the coordinate of the corresponding feature point in the reference image
[0160] Obtain the number of the inlier set I. If the number of the inlier set I is greater than the number of feature points in the inlier set library I', update the inlier set library I' to the inlier set I, record the transformation matrix H at this time, and execute the step of randomly selecting 4 sample feature points to obtain the transformation matrix H, calculating the projection error between the remaining feature points and the sample feature points. If the projection error is less than the projection threshold, mark the feature points as the inlier set I.
[0161] Preferably, the step of segmenting the target image in the two images to obtain a grid, and estimating the homography transformation matrix H between the source image and the target image by combining point features and line features using the APAP algorithm further includes
[0162] Optimize the transformation matrix H using the APAP matrix:
[0163] The steps of optimizing the transformation matrix H using the APAP matrix include:
[0164] Homogenize the feature points of the target image and the source image, and normalize the coordinates of the target image and the source image;
[0165] Divide the target image with normalized coordinates into grids, calculate the Gaussian weights of all feature points in each grid, and perform singular value decomposition to obtain the optimized transformation matrix H.
[0166] Furthermore, the following method is adopted:
[0167] (1) For the input target image I(x, y) and reference image I′(x′, y′), homogenize the feature points and normalize the coordinates of the feature points of the two images, and the method is:
[0168]
[0169] where μ is the mean of the abscissas of the feature points, σ is the standard deviation of the abscissas of the feature points, and the processing method for the ordinates is the same as that for the abscissas.
[0170] (2): For the target image, divide the image into grids of size 40×40 (the size can be set arbitrarily), calculate the Gaussian weights of each grid and all the feature points therein, and the expression is as follows:
[0171]
[0172] where, x * is the central coordinate of the grid, and x i is the coordinate of the feature point in the target image after being screened by the RANSAC algorithm.
[0173] (3): Perform wsvd for singular value decomposition, and the corresponding expression is:
[0174] h * = argmin h ||W * Ah|| 2
[0175] The solution is the right singular eigenvector corresponding to the minimum eigenvalue of W * A, where the weight matrix h represents the form of H obtained in the RANSAC algorithm sorted by columns. Here, the form of A is:
[0176]
[0177] where, 0 1×3 represents a 1×3 zero matrix, It represents the homogeneous form of the coordinates of a certain feature point (denoted as p) in the target image, and (x′, y′) is the feature point in the reference image corresponding to p.
[0178] (4): Solving step 4 can obtain the column-wise representation form of Hmdlt.
[0179] Preferably, the steps of obtaining the slope relationship between the source image and the target image and obtaining the cross feature based on the homography transformation matrix H include:
[0180] Obtain the slopes before and after the homography transformation of the target image.
[0181] Based on the verified conclusion that there is only one cluster of parallel lines in an image, and the parallel relationship is still guaranteed after the rigid transformation of the image. The slopes before and after the transformation can be calculated by the following formula:
[0182]
[0183] Obtain the dividing lines of the overlapping area and the non-overlapping area between the target image and the source image based on the slopes before and after the homography transformation of the target image.
[0184] Specifically, set the slope k1, and set l v as the dividing line closest to the overlapping and non-overlapping areas, and l u as the line perpendicular to l v , which can be calculated by the following formula:
[0185] k1·k2 = -1, s(x, y, k1)·s(x, y, k2) = -1
[0186] Preferably, using the grid in step S3, optimize the point feature and the line feature respectively using the following expressions:
[0187]
[0188] Among them, p′ is the matching point in the reference image corresponding to the target image, representing the gap between the feature point of the homography transformation matrix H after the homography transformation and the actual matching feature point.
[0189]
[0190] Among them, assume that {p k} represents all the endpoints of the line segments cut by the grid in the source image, which is represented by , and the corresponding matching line segment in the target image is represented by a j x + b j y + c j = 0;
[0191] Cross feature perspective optimization is performed using the following formula:
[0192]
[0193] Wherein, and are the normal vectors of the two line segments of the cross feature. A pair of cross features consists of and The two line segments are represented by endpoints as and
[0194] Projection distortion is reduced using the following formula:
[0195]
[0196] Preferably, the step of simultaneously obtaining the stitching line between two images includes:
[0197] Calculating the stitching line, and the step of calculating the stitching line includes defining labels and assigning the labels;
[0198] Evaluating the stitching line, and the cost of the pixel block along the stitching line is evaluated using the following formula through structural similarity consistency:
[0199]
[0200] Wherein, SSIM represents structural similarity, and its value range is [-1, 1]. The value is 1 if and only if the two pixel blocks are exactly the same;
[0201] The cost of the stitching line evaluation pixel point is evaluated using the following formula through the difference of adjacent pixel values, and the specific expression is as follows:
[0202]
[0203] A hybrid evaluation model is constructed using the following formula, and the expression is:
[0204]
[0205] Wherein represents the denoising cost.
[0206] Embodiment 2:
[0207] Embodiment 2 of the present invention provides an image stitching device based on stable linear structures and stitching line estimation, including:
[0208] Line feature matching module: Obtain two images with an overlapping area, and use SIFT to match the two
[0209] Test data Overlap rate (%) Resolution Number of matching points Number of matching lines
[0210] After the feature points of the image are obtained, the RANSAC algorithm is used to eliminate the false features to obtain point features, and the LSD algorithm is used to detect and match line features;
[0211] Homography transformation matrix acquisition module: used to segment the target images in two images to obtain a grid, and combined with point features and line features, the APAP algorithm is used to estimate the homography transformation matrix H between the source image and the target image;
[0212] Cross feature acquisition module: used to obtain the slope relationship between the source image and the target image and obtain cross features based on the homography transformation matrix H;
[0213] Suture acquisition module: combines point features, line features and cross features to obtain the expression of the grid, and at the same time obtains the suture between the two images;
[0214] Suture determination module: judge whether the suture is a stable suture,
[0215] If so, splice the source image and the target image.
[0216] In order to verify the technical effects provided by the present invention, Table 1 and Table 2 are combined;
[0217] Table 1 Details of test data
[0218] Temple 52.49 730×487 323 11 Parking lot 24.71 653×490 293 49 Guardrail 66.75 1280×960 637 12 Ceiling 62.93 800×600 363 51 Graffiti building 59.32 600×400 106 4 Station 51.38 1440×1080 1191 37 Community 22.60 1080×1440 545 21 Building 20.37 1080×1440 1071 34 Bookcase 60.74 1440×1080 414 17 Office building 12.45 1440×1080 545 21
[0219] Table 2 Peak signal-to-noise ratio index (PSNR) in special cases
[0220] Method (Year proposed) PSNR (Parking lot data) PSNR (Bookcase data) AANAP (2015) 13.679 8.746 APAP (2013) 15.811 14.316 ELA (2017) 25.731 15.023 TFT (2019) 14.786 17.382 PS (2019) 30.703 21.948 Method of this paper 61.037 Inf
[0221] The higher the score, the more similar the object before splicing is to the object after splicing (the Inf value is infinity, indicating exactly the same), that is, the smaller the deformation.
[0222] Table 3 Comparison of PSNR and structural similarity consistency index (SSIM)
[0223]
[0224]
[0225] This table shows the PSNR and SSIM indexes of all test data in the overlapping area. The value range of SSIM is [0,1], and 1 means that the two images are exactly the same. Our method has obtained the highest score in most cases (the bold value indicates the highest). Among them, although the best result was not achieved in the ceiling data test, the visual result is better than that of ELA.
[0226] Unless otherwise specifically stated, the relative steps, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0227] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.
[0228] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0229] In addition, in the description of the embodiments of the present invention, unless otherwise clearly defined and limited, the terms "install", "connect", and "couple" shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be a direct connection or an indirect connection through an intermediate medium, and it may be the internal communication of two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0230] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance.
[0231] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any technician familiar with the technical field of the present invention can still modify the technical solutions recorded in the foregoing embodiments or easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An image stitching method based on a stable linear structure and suture line estimation, characterized in that, Including: S1: Obtain two images with an overlapping area. After using SIFT to match the feature points of the two images, use the RANSAC algorithm to eliminate false features to obtain point features, and use the LSD algorithm to detect and match line features; S2: Segment the target image in the two images to obtain a grid, and combine the point features and line features to estimate the homography transformation matrix H between the source image and the target image using the APAP algorithm; S3: Based on the homography transformation matrix H, obtain the slope relationship between the source image and the target image and obtain cross features; S4: Combine the point features, line features, and cross features to obtain the expression of the grid, and at the same time obtain the stitching line between the two images; S5: Determine whether the stitching line is a stable stitching line, If so, splice the source image and the target image.
2. The method according to claim 1, wherein Including: The step of using the LSD algorithm to detect and match line features includes: Shrink the source image and the target icon image according to a preset ratio, and obtain the gradient value of each point; Perform pseudo-sorting on the target image based on the gradient value of each point, set a gradient threshold, and divide it into N bins according to the gradient value. Obtain the maximum gradient value in each bin, and sort all the bins according to the maximum gradient value in each bin; Obtain the bin with the largest gradient value, select a seed point from the bin with the largest gradient value, and determine the θ of the seed point l and whether the difference between the regional angle is less than a preset value; If so, add the seed point to the region update set, and re-execute the step of selecting a seed point from the bin with the largest gradient value and determining whether the difference between the θ of the seed point l and the regional angle is less than a preset value; Fit the region update set into a rectangle, and determine whether the number of points in the rectangle is greater than the point number threshold; If satisfied, fit the rectangle into a straight line.
3. The method according to claim 1, characterized in that, The step of obtaining two images with an overlapping area, using SIFT to match the feature points of the two images, using the RANSAC algorithm to eliminate false features to obtain point features, and using the LSD algorithm to detect and match line features includes: S101: Obtain the scale space of two images with an overlapping area, and generate a Gaussian difference pyramid; S102: Based on the Gaussian difference pyramid, calculate the spatial extreme points to obtain feature points; S103: Obtain the amplitude, position, and direction of the feature points, describe and match the feature points using the KNN algorithm.
4. The method according to claim 1, wherein The step of segmenting the target image in the two images to obtain a grid, and combining the point features and line features to estimate the homography transformation matrix H between the source image and the target image using the APAP algorithm includes: Establish an inlier set library I'; Randomly select 4 sample feature points to obtain the transformation matrix H, calculate the projection error between the remaining feature points and the sample feature points. If the projection error is less than the projection threshold, mark the feature points as the inlier set I; Obtain the number of the inlier set I. If the number of the inlier set I is greater than the number of feature points in the inlier set library I', update the inlier set library I' to the inlier set I, record the transformation matrix H at this time, and execute the step of randomly selecting 4 sample feature points to obtain the transformation matrix H, calculate the projection error between the remaining feature points and the sample feature points. If the projection error is less than the projection threshold, mark the feature points as the inlier set I.
5. The method according to claim 3, wherein The step of segmenting the target image in the two images to obtain a grid, and combining the point features and line features to estimate the homography transformation matrix H between the source image and the target image using the APAP algorithm further includes, Optimize the transformation matrix H using the APAP matrix: The steps of optimizing the transformation matrix H using the APAP matrix include: Homogenize the feature points of the target image and the source image, and standardize the coordinates of the target image and the source image; Divide the target image with standardized coordinates into grids, calculate the Gaussian weights of all feature points in each grid, and perform singular value decomposition to obtain the optimized transformation matrix H.
6. The method according to claim 4, characterized in that The steps of obtaining the slope relationship and the cross feature between the source image and the target image based on the homography transformation matrix H include: Obtain the slopes of the matrices before and after the homography transformation of the target image; Based on the slopes of the matrices before and after the homography transformation of the target image, obtain the dividing lines of the overlapping area and the non-overlapping area between the target image and the source image.
7. The method according to claim 2, wherein Include: Using the grids in step S3, optimize the point features and line features respectively using the following expressions: Where p′ is the matching point of the target image corresponding to the reference image, representing the gap between the feature points after homography transformation and the actual matching feature points; Among them, assume that {p k} represents all the endpoints of the line segments cut by the grid in the source image, which is represented by . The corresponding matching line segments in the target image are represented by a j x + b j y + c j = 0; Perform perspective optimization of the cross feature using the following formula: Among them, and are the normal vectors of the two line segments of the cross feature. A pair of cross features is composed of and The two line segments are represented by endpoints as and Reduce the projection distortion by the following formula:
8. The method according to claim 1, wherein The steps of simultaneously obtaining the stitching line between two images include: Calculate the stitching line, and the steps of calculating the stitching line include defining labels and assigning the labels; Evaluate the stitching line, and use the following formula to evaluate the cost of the pixel block along the stitching line through structural similarity consistency: Where SSIM represents the structural similarity, and the value range is [-1, 1]. The value is 1 if and only if the two pixel blocks are exactly the same; Use the following formula to evaluate the cost of the pixel points of the stitching line through the difference of adjacent pixel values. The specific expression is as follows: Construct a hybrid evaluation model using the following formula. The expression is: Among them represents the denoising cost.
9. An image stitching device based on a stable linear structure and suture estimation, characterized in that, Include: Line feature matching module: Obtain two images with an overlapping area, use SIFT to match the feature points of the two images, then use the RANSAC algorithm to remove the false features to obtain point features, and use the LSD algorithm to detect and match line features; Homography transformation matrix acquisition module: Used to segment the target image in the two images to obtain grids, and estimate the homography transformation matrix H between the source image and the target image by combining point features and line features using the APAP algorithm; Cross feature acquisition module: Used to obtain the slope relationship and the cross feature between the source image and the target image based on the homography transformation matrix H; Stitching line acquisition module: Combine point features, line features and cross features to obtain the expression of the grid, and simultaneously obtain the stitching line between the two images; Stitching line determination module: Judge whether the stitching line is a stable stitching line, If so, splice the source image and the target image.
Citation Information
Patent Citations
Method and system for stitching text images
CN102074001A
image splicing method based on hybrid transformation
CN109658370A