Scene Image Stitching Method, System and Electronic Device Based on Feature Point Matching
Through the scene image stitching method based on feature point matching, iterative updates are performed using the pose matrix and matching point transformation loss, and the problems of high computational complexity and inaccurate description of geometric relationships in the prior art are solved, and efficient and high-quality image stitching is achieved.
Patent Information
- Application Number
- CN202411240509.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-05
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-09-05
AI Technical Summary
The prior art has high computational complexity when processing large-scale image data, which is difficult to meet the needs of real-time applications. In scenarios with complex geometric relationships, it is difficult for the holographic matrix transformation model to accurately describe the geometric relationships between images.
The scene image stitching method based on feature point matching is adopted, and the feature points in the stitching image are transformed through the stitching matrix, the matching point transformation loss is calculated, and the stitching matrix is iteratively updated to optimize the calculation efficiency of the stitching algorithm and improve the quality of the image stitching.
The calculation efficiency of the stitching algorithm is optimized and the quality of image stitching is improved. Especially in complex geometric relationship scenarios, the geometric relationship between images can be described more accurately.
Smart Images

Figure CN119359536B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of binocular vision. Specifically, it relates to a method for stitching scene images based on feature point matching, a scene image stitching based on feature point matching, and an electronic device. Background Art
[0002] With the rapid development of autonomous driving technology, the intelligence level of vehicles has been continuously improved. Autonomous driving technology has become one of the important research directions in the future transportation field. To achieve efficient, safe, and reliable autonomous driving, vehicles need to have accurate environmental perception capabilities to comprehensively and accurately understand and judge the surrounding environment. Among many perception technologies, the scene perception technology based on Bird's Eye View (BEV) has become an important research direction in the field of autonomous driving because it can provide environmental information with a global perspective.
[0003] The acquisition of BEV images is based on fusing data collected by multiple sensors (especially cameras) at a certain moment, and its range is limited. Therefore, to obtain BEV images in the global scene, scene stitching technology is needed to stitch and fuse BEV images obtained from different positions to obtain a global BEV image.
[0004] Existing scene image stitching methods are mainly divided into two categories: feature point-based methods and direct stitching methods. Feature point-based methods mainly calculate the transformation relationship between images by detecting and matching feature points in the images to achieve stitching; while direct stitching methods achieve stitching by directly optimizing the overlapping area of image pixels. In addition, deep learning technology has also started to emerge in the field of image stitching. By using convolutional neural networks (CNNs) for feature extraction and image registration, the accuracy and speed of stitching have been further improved.
[0005] For example, the automatic panoramic image generation system proposed by Brown and Lowe in 2007. This system first uses the SIFT (Scale-Invariant Feature Transform) algorithm to extract feature points in the image, then matches the feature points through nearest neighbor search, then uses the RANSAC (Random Sample Consensus) algorithm to calculate the homography matrix between images, and finally processes the overlapping area of the image through a multi-band fusion algorithm to generate a seamless panoramic image. This method has high robustness and accuracy in feature point detection, matching, and image fusion, and is widely used in actual panoramic image generation systems.
[0006] In the existing technology, although significant progress has been made in image stitching technology, it still faces some challenges in practical applications. First, traditional stitching methods have a high computational complexity when dealing with large-scale image data, making it difficult to meet the requirements of real-time applications. Second, for scenes with complex geometric relationships, such as scenes with significant depth changes, the existing homography matrix transformation model is difficult to accurately describe the geometric relationships between images (feature points).
[0007] Therefore, how to optimize the computational efficiency of the stitching algorithm and improve the quality of image stitching remains an important topic in the research of image stitching technology. Summary of the Invention
[0008] The objective of this application is: how to optimize the computational efficiency of the stitching algorithm and improve the quality of image stitching.
[0009] The technical solution of the first aspect of this application is: to provide a method for stitching scene images based on feature point matching, the method comprising: Step 1: Obtain a set of images in the BEV view captured by each camera at the current moment; Step 2: Adopt a traversal stitching method, sequentially select the next image in the set of images as the image to be stitched, and calculate the feature matching point pairs between the image to be stitched and the stitched image, where the stitched image is the first image in the set of images during the first traversal stitching; Step 3: Based on the pose matrix and the feature matching point pairs, perform pose transformation on the feature points in the image to be stitched, and update the feature matching point pairs; Step 4: Calculate the matching point transformation loss according to the updated feature matching point pairs, and iteratively update the pose matrix based on the matching point transformation loss; Step 5: Perform pose transformation on the image to be stitched based on the updated pose matrix, and stitch the transformed image to be stitched to the stitched image, and re-execute Step 2 until all the images in the set of images are stitched.
[0010] In any of the above technical solutions, further, in step 2, calculating the feature matching point pairs between the image to be stitched and the stitched image specifically includes: Step 21: Using the feature map extraction network, respectively obtain the feature maps of the image to be stitched and the stitched image, generate a plurality of feature map groups, and sequentially denote the plurality of feature map groups in ascending order of resolution as the first feature map group, the second feature map group, and the third feature map group; Step 22: Using the query vector, calculate the first similarity map on the first feature map group; Step 23: Calculate the two feature points with the highest similarity to any query vector in the first similarity map respectively, and form the first similar point pair; Step 24: Based on the positions of the first similar point pair, intercept the second feature map on the second feature map group, and based on the second feature map, calculate the second similarity map to generate the second similar point pair; Step 25: Based on the positions of the second similar point pair, intercept the third feature map on the third feature map group, and based on the third feature map, calculate the third similarity map to generate the third similar point pair, and denote the third similar point pair as the feature matching point pair.
[0011] In any of the above technical solutions, further, step 24 further includes: Based on the first threshold, sequentially determine whether the similarity of the point pairs in the second similarity map is greater than the first threshold; if not, discard the point pairs, if so, denote the point pairs as the second similar point pair.
[0012] In any of the above technical solutions, further, step 25 further includes: Using one layer of convolution and the sigmoid activation function, calculate the edge degree map corresponding to the third feature map group; Weight the third similarity map based on the edge degree map; Based on the second threshold, sequentially determine whether the similarity of the point pairs in the weighted third similarity map is greater than the second threshold; if not, discard the point pairs, if so, denote the point pairs as the third similar point pair.
[0013] In any of the above technical solutions, further, step 25 further includes: Using one layer of convolution and the sigmoid activation function, calculate the background degree map corresponding to the third feature map group; Weight the third similarity map based on the background degree map; Based on the third threshold, sequentially determine whether the similarity of the point pairs in the weighted third similarity map is greater than the third threshold; if not, discard the point pairs, if so, denote the point pairs as the third similar point pair.
[0014] In any of the above technical solutions, further, the calculation formula of the matching point transformation loss is:
[0015]
[0016] In the formula, is the feature point located in the stitched image in the feature matching point pair, is the feature point located in the image to be stitched after pose transformation, p = 1, 2,..., P, and P is the total number of feature matching point pairs.
[0017] The technical solution of the second aspect of the present application is: to provide a scene image stitching system based on feature point matching, the system includes: an acquisition unit, which is configured to acquire an image set in the BEV view captured by each camera at the current moment; a matching unit, which is configured to adopt a traversal stitching method, and sequentially select the next image in the image set as the image to be stitched, and calculate the feature matching point pairs between the image to be stitched and the stitched image. Among them, when traversing and stitching for the first time, the stitched image is the first image in the image set; a transformation unit, which is configured to perform pose transformation on the feature points in the image to be stitched based on the pose matrix and the feature matching point pairs, and update the feature matching point pairs; an iteration unit, which is configured to calculate the matching point transformation loss according to the updated feature matching point pairs, and iteratively update the pose matrix based on the matching point transformation loss; a stitching unit, which is configured to perform pose transformation on the image to be stitched based on the updated pose matrix, and stitch the transformed image to be stitched to the stitched image until all the images in the image set are stitched.
[0018] In any of the above technical solutions, further, the matching unit is further configured to: use a feature map extraction network to respectively obtain the feature maps of the image to be stitched and the stitched image, generate a plurality of feature map groups, and sequentially record the plurality of feature map groups in ascending order of resolution as the first feature map group, the second feature map group, and the third feature map group; use a query vector to calculate a first similarity map on the first feature map group; respectively calculate the two feature points with the highest similarity to any query vector in the first similarity map to form a first similar point pair; based on the positions of the first similar point pairs, intercept a second feature map on the second feature map group, and calculate a second similarity map based on the second feature map to generate a second similar point pair; based on the positions of the second similar point pairs, intercept a third feature map on the third feature map group, and calculate a third similarity map based on the third feature map to generate a third similar point pair, and record the third similar point pair as the feature matching point pair.
[0019] In any of the above technical solutions, further, the matching unit is further configured to: sequentially determine whether the similarity of the point pairs in the second similarity map is greater than a first threshold based on the first threshold; if not, discard the point pairs, and if so, record the point pairs as the second similar point pairs.
[0020] The technical solution of the third aspect of the present application is: to provide an electronic device, the electronic device includes: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to implement any of the scene image stitching methods based on feature point matching in the first aspect technical solution when executing the executable instructions.
[0021] The beneficial effects of the present application are:
[0022] In the technical solution of this application, the pose matrix is used to perform pose transformation on the feature points in the images to be stitched, obtain the matching results from the BEV perspective, and construct a matching point transformation loss. The pose matrix is iteratively updated by maximizing the coincidence degree of feature points, which optimizes the calculation efficiency of the stitching algorithm and improves the quality of image stitching.
[0023] This application also optimizes the calculation process of feature matching point pairs. During the calculation of feature matching point pairs between images, a query vector is used as a medium, and the two feature points with the highest similarity to the query vector are respectively calculated in the image group composed of the left and right feature maps to form the first similar point pair. Then, by using the "coarse-grained - fine-grained" iterative update method, the more accurately positioned and final feature matching point pairs are gradually obtained, rather than directly performing dense matching on the two images. This not only helps to reduce the calculation amount, but also enables the matching results to have high precision while having a certain degree of robustness, thereby improving the quality of image stitching. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The above and / or additional aspects of the present application will become obvious and easy to understand when combined with the description of the embodiments in conjunction with the following drawings, where:
[0025] Figure 1 is a schematic flowchart of a method for stitching scene images based on feature point matching according to an embodiment of the present application;
[0026] Figure 2 is a schematic diagram of the calculation process of the second similar point pair according to an embodiment of the present application;
[0027] Figure 3 is a schematic block diagram of a system for stitching scene images based on feature point matching according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] In order to more clearly understand the above objects, features, and advantages of the present application, the present application will be further described in detail below in conjunction with the drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.
[0029] In the following description, many specific details are set forth in order to fully understand the present application. However, the present application may be implemented in other ways different from those described herein. Therefore, the protection scope of the present application is not limited by the specific embodiments disclosed below.
[0030] Embodiment 1:
[0031] As Figure 1As shown in the figure, this embodiment provides a method for stitching scene images based on feature point matching. The method includes:
[0032] Step 1: Obtain the set of images in the BEV view captured by each camera at the current moment;
[0033] Specifically, this embodiment takes the in-vehicle surround view camera as an example for illustration. It is assumed that the number of in-vehicle surround view cameras is 6, and they are numbered in a clockwise direction in sequence. After the in-vehicle surround view cameras are started, 6 images in the BEV view centered on the vehicle will be obtained 1 ≤ i ≤ C, C = 6, thus forming the set of images in the BEV view
[0034] Step 2: Adopt the method of traversing and stitching. Successively select the next image in the set of images as the image to be stitched, and calculate the feature matching point pairs between the image to be stitched and the stitched image. Among them, when traversing and stitching for the first time, the stitched image is the first image in the set of images;
[0035] Specifically, the 6 images obtained in the BEV view are stitched in sequence according to the camera numbers. During the stitching process, the first image is used as the initial stitched image, and then the remaining images are stitched in sequence. Now, take any image other than the first image as an example for illustration. The label sub BEV , is introduced to represent the current image to be stitched, and I BEV is introduced to represent the current stitched image.
[0036] Based on the feature point matching algorithm, perform feature point matching on the current image to be stitched the current stitched image I BEV for the two images to obtain the matched feature matching point pairs Among them, P is the total number of successfully matched point pairs, is the p-th matching point in the current image to be stitched , is the matching point in the current stitched image I BEV that matches .
[0037] It should be noted that in the feature point matching algorithm, only some feature points in the image can be matched.
[0038] Step 3: Based on the pose matrix and the feature matching point pairs, perform pose transformation on the feature points in the image to be stitched and update the feature matching point pairs;
[0039] Specifically, the pose matrix includes a rotation matrix R and a translation matrix T. Among them, the rotation matrix R is initialized as the identity matrix, and the translation matrix T is initialized as the unit vector. This pose matrix represents the current image to be stitched The current stitched image I BEV The relative pose between two images.
[0040] Therefore, based on the pose matrix, the p-th feature point in the current image to be stitched can be transformed to the current stitched image I BEV to obtain the updated feature matching point pairs
[0041] Step 4: Calculate the matching point transformation loss based on the updated feature matching point pairs, and iteratively update the pose matrix based on the matching point transformation loss;
[0042] Specifically, since the transformation of the above feature points is based on the initial pose matrix, there is a deviation between the feature points and the transformed feature points Based on this deviation, the matching point transformation loss can be calculated, and the corresponding calculation formula can be:
[0043]
[0044] In the formula, is the feature point located in the stitched image in the feature matching point pair, is the feature point located in the image to be stitched after pose transformation, p = 1, 2,..., P, and P is the total number of feature matching point pairs.
[0045] After that, an iterative algorithm, such as the Newton iterative method, is introduced for iterative calculation to update the pose matrix, so that the deviation between the feature points and the transformed feature points is minimized to ensure that the current image to be stitched after pose transformation can be transformed to the current stitched image I BEV .
[0046] Step 5: Perform pose transformation on the image to be stitched based on the updated pose matrix, stitch the transformed image to be stitched to the stitched image, and select the next unstitched image with the label sub BEV +1 as the new image to be stitched, and re-execute Step 2 until all images in the image set are stitched.
[0047] Considering that factors such as perspective differences, lighting changes, and dynamic objects that may exist during the image shooting process will lead to a decrease in the accuracy of feature point matching, thus affecting the overall effect of image stitching. Therefore, this embodiment also shows a method for image feature point matching. In this embodiment, the resolution of the image to be stitched and the stitched image is the same, both (W, H). In the above step 2, calculating the feature matching point pairs between the image to be stitched and the stitched image specifically includes:
[0048] Step 21: Using the feature map extraction network, respectively obtain the feature maps of the image to be stitched and the stitched image, generate multiple groups of feature maps, and sequentially record the multiple groups of feature maps in ascending order of resolution as the first feature map group, the second feature map group, and the third feature map group. Among them, the resolution of the third feature map group is the same as the resolution (W, H) of the image to be stitched;
[0049] Specifically, based on the existing U-Net network framework, the image to be stitched Stitched image I BEV Input the two images into the network, and three feature maps with different resolutions can be obtained, which are sequentially recorded as the first feature map group (F A1 , F B1 ); the second feature map group (F A2 , F B2 ); the third feature map group (F A3 , F B3 ), that is, {F Ai , F Bi |1 ≤ i ≤ 3}, and the resolutions are sequentially The feature dimensions are d 1 , d 2 , d 3 .
[0050] Step 22: Using the query vector, calculate the first similarity map on the first feature map group;
[0051] Specifically, initialize M query vectors Calculate the similarity with each point on two feature maps F A1 , F B1 in the first feature map group respectively. During the calculation process, first stack all the query vectors in row-major order to obtain the query matrix Then multiply the query matrix Q with two feature maps F A1 , F B1 respectively to obtain the first similarity maps Sim A1 , Sim B1 , and the corresponding calculation formula is:
[0052]
[0053] Among them, the Softmax() function can be used to normalize the two similarity maps.
[0054] Step 23: Calculate the two feature points with the highest similarity to any query vector in the first similarity map respectively, and form a first similarity point pair;
[0055] Specifically, the first similarity maps Sim A1 and Sim B1 are 2M similarity maps. Among them, the first similarity map corresponding to the i 1 th query vector is respectively For these two first similarity maps, calculate the two feature points with the largest similarity value respectively. The corresponding calculation formula is:
[0056]
[0057] In the formula, represents the horizontal and vertical coordinates corresponding to the calculation of when it is the largest, is the horizontal and vertical coordinates corresponding to the first similarity map , is the coordinate corresponding to the calculation of when it is the largest; the calculation in the similarity map is the same by analogy.
[0058] respectively represent the positions of the points with the highest similarity to the i A1 th query vector B1 in the first feature maps (F 1 and F ). Since the similarities of these two points to the same query vector are both the highest, it can be inferred that these two feature points also have a relatively high similarity. Denote these two points as a first similarity point pair, that is
[0059] Therefore, in this embodiment, these two feature points are combined into a first similarity point pair Through this indirect calculation method, using the query vector as an intermediate variable, search for the feature points with the highest similarity to the same query vector in the two feature maps respectively, so as to indirectly obtain a pair of similarity point pairs. Compared with the dense matching method in the traditional feature point matching process, such a method can greatly reduce the calculation amount and optimize the calculation efficiency.
[0060] Step 24: Based on the positions of the first similarity point pairs, intercept the second feature maps on the second group of feature maps, and calculate the second similarity map based on the second feature maps to generate the second similarity point pairs;
[0061] Specifically, as Figure 2 shown, since the downsampling ratio (resolution) of the second group of feature maps (F A2 , F B2 ) is 2 times higher than that of the first group of feature maps (F A1 , F B1 ), that is, one feature point in (F A1 , F B1 ) corresponds to a 2x2 region in (F A2 , F B2 ).
[0062] Considering that the first similarity point pairs pair A1 , F B1 have been calculated in the above process, but because its downsampling ratio is not high enough, pixel-level matching cannot be obtained. Therefore, based on the positions of the first similarity point pairs 1 , intercept two 2x2-sized second feature maps from the corresponding 2x2 regions in the second group of feature maps (\ , F A2 , F B2 ) Perform dense matching calculation on these two second feature maps to obtain a 4x4-sized second similarity map The corresponding calculation formula is:
[0063]
[0064] In the formula, the Flatten() operation is to flatten the original -sized feature map into a -sized feature map, and the two calculate the similarity map of size R 4×4 Among them, 1 ≤ i 1 ≤ M.
[0065] To obtain the second similarity point pairs, it is necessary to find the position with the maximum similarity from the second similarity map, and the corresponding calculation formula is:
[0066] x = x A * 2 + y A
[0067] y = x B * 2 + y B
[0068]
[0069] Among them, x A , y A , x B , y B ∈ [0, 2) is the coordinate corresponding to the second feature map in it, x, y ∈ [0, 4) is the second similarity map corresponding horizontal and vertical coordinates. By calculating the position corresponding to the maximum similarity in the second similarity map the position of this point corresponding to the second feature map can be deduced backwards:
[0070]
[0071] Among them, / / represents the integer division operation, and % represents the remainder operation. The position in the feature map is obtained, and the coordinates corresponding to it in the second feature map (F A2 , F B2 ) can be deduced backwards:
[0072]
[0073] So far, the second similar point pair is obtained:
[0074] Step 25: Repeat the above process. Based on the positions of the second similar point pair, intercept the third feature map on the third feature map group, and calculate the third similarity map based on the third feature map to generate the third similar point pair, and record the third similar point pair as the feature matching point pair.
[0075] Specifically, based on the positions of the two similar points in the second similar point pair pair 2 in it intercept two 2x2-sized third feature maps from the corresponding regions in the third feature map group (F A3 , F B3 ) Calculate the dense matching of the two third feature maps to obtain a 4x4-sized third similarity map The corresponding calculation formula is:
[0076]
[0077] Through the above method, we obtain the similarity of the similar point pair in the second similar point pair pair 2 in the third feature map group (F A3 , F B3 ). Each point in this third similarity map corresponds to a group of 16 combinations of similar points in the two 2x2-sized third feature maps intercepted above, that is:
[0078]
[0079] To obtain the third similarity point pair, it is necessary to find the position with the maximum similarity in the third similarity graph, and the corresponding calculation formula is:
[0080] x = x A * 2 + y A
[0081] y = x B * 2 + y B
[0082]
[0083] where x A , y A , x B , y B ∈[0, 2) are the corresponding coordinates in the third feature map and x, y ∈[0, 4) are the horizontal and vertical coordinates corresponding to the third similarity graph. By calculating the position corresponding to the maximum similarity in the third similarity graph the position of this point corresponding to the third feature map can be deduced inversely:
[0084]
[0085] where, / / represents the integer division operation, and % represents the remainder operation. After obtaining its position in the feature map, the coordinates corresponding to it in the second feature maps (F A3 , F B3 ) can be deduced inversely:
[0086]
[0087]
[0088] Thus far, the third similarity point pair is obtained:
[0089] In this embodiment, by introducing an intermediate variable query vector feature points with the highest similarity to any query vector are respectively found in two low-resolution feature maps, and the first similarity point pair is formed. Then, through the upsampling method, dense matching is performed in the high-resolution and small-region (second and third) feature maps to obtain the final feature matching point pair, reducing the computational complexity of the feature point matching calculation process and optimizing the computational efficiency of the stitching algorithm. In any of the above embodiments, further, after obtaining the second similarity graph
[0090] After that, in order to obtain more matching point pairs and enable the model to match as many similar points in the two images as possible, rather than being limited to obtaining only one pair of matching points on each similarity map, step 24 further includes:
[0091] Based on the first threshold θ 2 , successively determine whether the similarity of the point pairs in the second similarity map is greater than the first threshold θ 2 ;
[0092] If not, it means that this point pair is very likely not a group of similar point pairs, and this point pair can be discarded; if so, it means that this point pair is very likely a group of similar point pairs, and this point pair is retained and denoted as the second similar point pair.
[0093] In any of the above embodiments, further, considering the edge points in the graph, such as the corners of the wall, the success rate of matching is relatively high. Step 25 further includes:
[0094] Use a layer of convolution Conv1() and sigmoid activation function to calculate the edge degree map corresponding to the third feature map group, and obtain the probability that each point in the third feature map group is located on the edge, where the resolution of the third feature map group is the same as the resolution of the spliced image;
[0095] Specifically, use a layer of convolution Conv1() to transform the original third feature map group (F A3 , F B3 ) into two corresponding edge degree maps edgeMap A , edgeMap B ∈ R H×W , the input channel is d 3 during the calculation, the output channel is 1, and then a per-pixel Sigmoid() activation function is connected to convert the output into a probability within the range of [0,1]. Its calculation formula is:
[0096]
[0097] During the training process of the above layer of convolution Conv1(), the true value of the edge degree map calculated by the traditional algorithm can be used for supervision, so that the parameters in this convolution can focus on learning the information related to the edge in the image features.
[0098] Weight the third similarity map based on the edge degree map, and the corresponding calculation formula is:
[0099]
[0100] In the formula, is the weighted third similarity map.
[0101] Specifically, appropriate weights are added to the point pairs on the third similarity map based on the edge degree map. The pixel points of non-edge points are more difficult to match successfully. Relatively speaking, even if the matching is successful, the credibility is not high. Therefore, a lower weight is added to them; while the credibility of edge points is higher, and their weights can be appropriately increased.
[0102] Multiply the similarity at each position by the corresponding edge degrees in images A and B, and the weighted third similarity map is obtained.
[0103] Based on the second threshold θ 3 , sequentially determine whether the similarity of the point pairs in the weighted third similarity map is greater than the second threshold θ 3 ; if not, it means that this point pair is very likely not a set of similar point pairs, and the point pair can be discarded; if so, it means that this point pair is very likely a set of similar point pairs, and the point pair is retained and denoted as the third similar point pair, that is
[0104] In any of the above embodiments, further, considering that there may be dynamic objects in the image, even if such objects are successfully matched in the two images, since the dynamic objects in the two images may not be in the same position, such point pairs are of no value for pose estimation. Step 25 further includes:
[0105] Using a layer of convolution Conv1() and a sigmoid activation function, calculate the background degree map corresponding to the third feature map group, and obtain the probability that each point in the third feature map group is a background point;
[0106] Specifically, using a layer of convolution Conv1(), the original third feature maps (F A3 , F B3 ) are transformed into two corresponding background degree maps backMap A , backMap B . In the calculation process, the input channel is d 3 , the output channel is 1, and then a per-pixel Sigmoid() function is connected to transform the output into a probability in the range of [0,1]. The calculation formula is:
[0107]
[0108] During the training process of the above layer of convolution Conv1(), the ground truth of the background degree map calculated by the existing instance segmentation algorithm can be used for supervision to guide the learning of this layer of convolution, so that the parameters in the convolution can focus on learning the information related to the background in the image features.
[0109] Weight the third similarity map based on the background degree map; it should be noted that the third similarity map can be weighted solely using the background degree map, or the third similarity map can be weighted simultaneously using the edge degree map and the background degree map. The corresponding calculation formulas are as follows:
[0110]
[0111] Or
[0112]
[0113] Based on the third threshold, sequentially determine whether the similarity of the point pairs in the weighted third similarity map is greater than the third threshold; if not, it indicates that this point pair is very likely not a set of similar point pairs, and the point pair can be discarded; if so, it indicates that this point pair is very likely a set of similar point pairs, and the point pair is retained and denoted as the third similar point pair.
[0114] Embodiment 2:
[0115] As Figure 3 shown, this embodiment provides a scene image stitching system 100 based on feature point matching. The system 100 includes:
[0116] An acquisition unit 101, which is configured to acquire a set of images in the BEV view captured by each camera at the current moment;
[0117] A matching unit 102, which is configured to adopt a traversal stitching method to sequentially select the next image in the set of images as the image to be stitched, and calculate the feature matching point pairs between the image to be stitched and the stitched image. Among them, when traversing and stitching for the first time, the stitched image is the first image in the set of images;
[0118] A transformation unit 103, which is configured to perform pose transformation on the feature points in the image to be stitched based on the pose matrix and the feature matching point pairs, and update the feature matching point pairs;
[0119] An iteration unit 104, which is configured to calculate the matching point transformation loss based on the updated feature matching point pairs, and iteratively update the pose matrix based on the matching point transformation loss;
[0120] A stitching unit 105, which is configured to perform pose transformation on the image to be stitched based on the updated pose matrix, and stitch the transformed image to be stitched to the stitched image until the stitching of all images in the set of images is completed.
[0121] In any of the above embodiments, further, the matching unit 102 is further configured to: use a feature map extraction network to respectively obtain the feature maps of the images to be stitched and the stitched image, generate a plurality of feature map groups, and sequentially denote the plurality of feature map groups in ascending order of resolution as the first feature map group, the second feature map group, and the third feature map group; use a query vector to calculate a first similarity map on the first feature map group; respectively calculate the two feature points with the highest similarity to any query vector in the first similarity map to form a first similar point pair; based on the positions of the first similar point pair, intercept a second feature map on the second feature map group, and based on the second feature map, calculate a second similarity map to generate a second similar point pair; based on the positions of the second similar point pair, intercept a third feature map on the third feature map group, and based on the third feature map, calculate a third similarity map to generate a third similar point pair, and denote the third similar point pair as a feature matching point pair.
[0122] In any of the above embodiments, further, the matching unit 102 is further configured to: sequentially determine whether the similarity of the point pairs in the second similarity map is greater than a first threshold based on the first threshold; if not, discard the point pairs, and if so, denote the point pairs as second similar point pairs.
[0123] In any of the above embodiments, further, the matching unit 102 is further configured to: use a single layer of convolution and a sigmoid activation function to calculate an edge degree map corresponding to the third feature map group; weight the third similarity map based on the edge degree map; sequentially determine whether the similarity of the point pairs in the weighted third similarity map is greater than a second threshold based on the second threshold; if not, discard the point pairs, and if so, denote the point pairs as third similar point pairs.
[0124] In any of the above embodiments, further, the matching unit 102 is further configured to: use a single layer of convolution and a sigmoid activation function to calculate a background degree map corresponding to the third feature map group; weight the third similarity map based on the background degree map; sequentially determine whether the similarity of the point pairs in the weighted third similarity map is greater than a third threshold based on the third threshold; if not, discard the point pairs, and if so, denote the point pairs as third similar point pairs.
[0125] In any of the above embodiments, further, the calculation formula of the matching point transformation loss is:
[0126]
[0127] In the formula, is the feature point located in the stitched image in the feature matching point pair, is the feature point located in the image to be stitched after pose transformation, p = 1, 2,..., P, and P is the total number of feature matching point pairs.
[0128] Embodiment 3:
[0129] This embodiment provides an electronic device, which includes: a processor; a memory for storing executable instructions that can be executed by the processor; wherein, when the processor is configured to execute the executable instructions, the image stitching method according to any one of the above-mentioned Embodiment 1 is implemented.
[0130] It should be noted that the number of processors can be one or more. At the same time, in the electronic device of the embodiment of the present application, an input device and an output device may also be included. Among them, the processor, the memory, the input device and the output device can be connected through a bus or in other ways, which is not specifically limited here.
[0131] So far, the embodiments of the present application have been described in detail. In order to avoid obscuring the concept of the present application, some details known in the art are not described. Those skilled in the art can fully understand how to implement the technical solutions disclosed here based on the above description.
[0132] Although some specific embodiments of the present application have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for the purpose of illustration and not for the purpose of limiting the scope of the present application.
[0133] The steps in the present application can be adjusted, combined and deleted according to actual needs.
[0134] Although the present application has been disclosed in detail with reference to the accompanying drawings, it should be understood that these descriptions are merely exemplary and not used to limit the application of the present application. The protection scope of the present application is defined by the appended claims and may include various variations, modifications and equivalent solutions made to the invention without departing from the protection scope and spirit of the present application.
Claims
1. A scene image stitching method based on feature point matching, characterized in that: The method comprises: Step 1: Obtain the image set from the BEV perspective captured by each camera at the current moment; Step 2: adopting a traversal stitching method, sequentially selecting the next image in the image set as the image to be stitched, and calculating the feature matching point pairs between the image to be stitched and the stitched image, wherein the stitched image is the first image in the image set during the first traversal stitching; Step 3: Based on the pose matrix and the feature matching point pairs, perform pose transformation on the feature points in the image to be stitched, and update the feature matching point pairs; Step 4: Calculate the matching point transformation loss according to the updated feature matching point pair, and iteratively update the pose matrix based on the matching point transformation loss; Step 5: performing posture transformation on the images to be stitched based on the updated posture matrix, and stitching the transformed images to be stitched to the stitched images, and re-performing step 2 until all images in the image set are stitched; Wherein, in step 2, calculating the feature matching point pairs between the image to be stitched and the stitched image specifically includes: Step 21: using a feature map extraction network, respectively obtaining the feature maps of the image to be stitched and the stitched image, generating a plurality of feature map groups, and recording the plurality of feature map groups in order of resolution from small to large as a first feature map group, a second feature map group, and a third feature map group; Step 22: using the query vector, calculating a first similarity graph on the first feature graph group; Step 23: respectively calculating two feature points in the first similarity graph that are most similar to any of the query vectors to form a first similarity point pair; Step 24: based on the position of the first similar point pair, intercepting a second feature map on the second feature map group, and based on the second feature map, calculating a second similarity map to generate a second similar point pair; Step 25: Based on the position of the second similar point pair, intercept the third feature map on the third feature map group, and based on the third feature map, calculate the third similarity map, generate a third similar point pair, and record the third similar point pair as the feature matching point pair.
2. The scene image stitching method based on feature point matching according to claim 1, characterized in that: The step 24 further includes: Based on a first threshold, determining in sequence whether the similarities of the point pairs in the second similarity graph are greater than the first threshold; If not, discard the point pair. If so, record the point pair as the second similar point pair.
3. The scene image stitching method based on feature point matching according to claim 1, characterized in that: The step 25 further includes: Calculate an edge degree map corresponding to the third feature map group using a layer of convolution and a sigmoid activation function; weighting the third similarity graph based on the edge degree graph; Based on the second threshold, determining in sequence whether the weighted similarities of the point pairs in the third similarity graph are greater than the second threshold; If not, discard the point pair. If so, record the point pair as the third similar point pair.
4. The scene image stitching method based on feature point matching according to claim 1 or 3, characterized in that: The step 25 further includes: Calculate the background degree map corresponding to the third feature map group using a layer of convolution and a sigmoid activation function; weighting the third similarity map based on the background degree map; Based on a third threshold, determining in sequence whether the weighted similarities of the point pairs in the third similarity graph are greater than the third threshold; If not, discard the point pair. If so, record the point pair as the third similar point pair.
5. The scene image stitching method based on feature point matching according to claim 1, characterized in that: The calculation formula of the matching point transformation loss is: In the formula, is a feature point in the feature matching point pair located in the stitched image, is the feature point located in the image to be stitched after the posture transformation, p=1,2,…,P, and P is the total number of the feature matching point pairs.
6. A scene image stitching system based on feature point matching, characterized in that: The system comprises: An acquisition unit, the acquisition unit is configured to acquire a set of images captured by each camera at a current moment under the BEV perspective; A matching unit, wherein the matching unit is configured to sequentially select the next image in the image set as the image to be stitched, and calculate the feature matching point pairs between the image to be stitched and the stitched image, wherein the stitched image is the first image in the image set during the first traversal stitching; A transformation unit, wherein the transformation unit is configured to perform a posture transformation on the feature points in the to-be-joined images based on a posture matrix and the feature matching point pairs, and update the feature matching point pairs; An iteration unit, the iteration unit being configured to calculate a matching point transformation loss according to the updated feature matching point pairs, and iteratively update the pose matrix based on the matching point transformation loss; a stitching unit, wherein the stitching unit is configured to perform a posture transformation on the images to be stitched based on the updated posture matrix, and stitch the transformed images to be stitched to the stitching image until the stitching of all images in the image set is completed, Wherein, the matching unit is further configured as: Using a feature map extraction network, respectively obtaining feature maps of the image to be stitched and the stitched image, generating a plurality of feature map groups, and recording the plurality of feature map groups in order of resolution from small to large as a first feature map group, a second feature map group, and a third feature map group; Using the query vector, calculating a first similarity graph on the first feature graph group; Calculate two feature points in the first similarity graph that are most similar to any of the query vectors to form a first similarity point pair; Based on the position of the first similar point pair, intercepting a second feature map on the second feature map group, and based on the second feature map, calculating a second similarity map to generate a second similar point pair; Based on the position of the second similar point pair, a third feature map is intercepted on the third feature map group, and based on the third feature map, a third similarity map is calculated to generate a third similar point pair, and the third similar point pair is recorded as the feature matching point pair.
7. The scene image stitching system based on feature point matching according to claim 6, characterized in that: The matching unit is further configured to: Based on a first threshold, determining in sequence whether the similarities of the point pairs in the second similarity graph are greater than the first threshold; If not, discard the point pair. If so, record the point pair as the second similar point pair.
8. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; Wherein, the processor is configured to implement the scene image stitching method based on feature point matching as described in any one of claims 1 to 5 above when executing the executable instructions.
Citation Information
Patent Citations
Image detection method used for improving feature matching accuracy
CN107451610A
Image alignment method and device based on multi-granularity processing fusion and medium
CN115205516A