A panoramic spliced image registration method based on an attention mechanism
Through the feature matching network based on the attention mechanism, the problems of high computational complexity and low precision of the traditional image registration algorithm in the ship deck environment are solved, high-precision image registration and fusion are achieved, and the visual effect of ship monitoring is improved.
Patent Information
- Application Number
- CN202510234669.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-02-28
AI Technical Summary
Traditional image registration algorithms have high computational complexity in ship deck environments and are limited by computing resources. In addition, they are prone to splicing misalignment and image blurring under conditions of lighting changes and occlusions. They cannot meet high-precision registration requirements and affect the image fusion effect.
A feature matching network based on the attention mechanism is adopted to perform image registration through the feature extraction network and the attention graph neural network. Self-attention and cross-attention are combined to enhance the feature matching performance, solve the affine transformation relationship between adjacent images, reduce noise and perform pixel fusion.
It improves the accuracy and precision of image registration, reduces the degree of distortion, enhances the user's visual experience, and ensures safe navigation of ships and efficient deck operations.
Smart Images

Figure CN120163703B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent image processing, and in particular to a panoramic stitching image registration method based on an attention mechanism. Background Art
[0002] In the field of video surveillance, panoramic cameras are deployed at commanding heights within the surveillance area and use multiple imaging sensors to generate a panoramic stitched image, enabling large-scale scene monitoring and providing users with a global perspective. Panoramic stitching primarily involves image acquisition, image preprocessing, image registration, and image fusion. Image registration, the core of panoramic stitching, calculates the corresponding image transformation relationship based on the overlapping portions of two adjacent images. Its efficiency and quality determine the final image application effect and play a key role in the overall consistency and degree of distortion of the final stitched image.
[0003] Image registration calculates the pixels with obvious features in the images to be stitched and matches the pixels between the images to obtain the transformation relationship between the images. Traditional image registration can be divided into five steps: (1) extracting feature points. Common feature point extraction algorithms include Scale-Invariant Feature Transform (SIFT), Oriented FAST and Rotated BRIEF (ORB), and Speeded-Up Robust Features (SURF); (2) calculating descriptors; (3) nearest neighbor matching; (4) filtering outliers; and (5) solving geometric constraints.
[0004] The panoramic stitching camera's placement on a ship is limited in height. To meet the requirements of large-scale deck surveillance, a large number of imaging sensors are required for image stitching. The deck's numerous objects and the unpredictable weather conditions at sea (including frequent rain and fog) create varying degrees of interference with the consistency, accuracy, and stability of image feature extraction and matching. Traditional image matching algorithms perform poorly in this environment and are unable to support the high-precision registration requirements. Furthermore, they are prone to problems such as misaligned stitching and blurred images due to lighting changes and occlusions.
[0005] Furthermore, ship deck operations are frequent, and panoramic stitching registration results require high instantaneous accuracy. Traditional registration algorithms are very sensitive to small-scale perspective errors caused by sensor errors, and the spatial relationship mapping between the two is prone to jitter, affecting the image fusion effect.
[0006] In short, traditional methods have high computational complexity and usually require a large amount of computing resources to process data, which puts a lot of pressure on the limited computing resources of ships. Summary of the Invention
[0007] The embodiment of the present invention provides a panoramic stitching image registration method based on the attention mechanism, which ensures the high consistency of the picture while achieving the purpose of panoramic surveillance, reduces the degree of distortion, improves the user's visual experience, and plays a role in ensuring the safe navigation of ships and efficient deck operations.
[0008] In a first aspect, the present invention provides a panoramic stitching image registration method based on an attention mechanism, comprising:
[0009] Acquire the multi-channel target images required for stitching, pre-process the multi-channel target images, and perform projection coordinate transformation on the pre-processed images;
[0010] Through the feature matching network based on the attention mechanism, the transformation relationship between the images after coordinate transformation is obtained for image stitching;
[0011] The stitched images are pixel-fused to eliminate stitching traces and generate a panoramic image that conforms to human vision.
[0012] In some examples, the step of obtaining the transformation relationship between the coordinate-transformed images through a feature matching network based on an attention mechanism to perform image stitching includes:
[0013] The image after coordinate transformation is subjected to denoising and redundant feature filtering, and key point features are extracted from the image after denoising and redundant feature filtering through a feature extraction network to obtain feature descriptors;
[0014] The feature matching network is used to perform similarity matching and filtering based on the results of feature extraction to obtain the matching relationship of feature points;
[0015] The matching relationship of the feature points is further filtered to solve the affine transformation relationship between adjacent images and determine the spatial position relationship mapping.
[0016] In some instances, the feature extraction network includes a basic encoding network and a feature extraction network, wherein the basic encoding network is used to detect corner points of the input image as candidate feature points; the feature extraction network is used to output feature points and descriptors, wherein the feature point extraction network and the descriptor extraction network share a single forward encoder, and adopt different structures in the decoder part.
[0017] In some instances, the feature matching network includes an attention graph neural network and an optimal matching layer, wherein the attention graph neural network is used to encode feature points and descriptors into a vector, and use self-attention and cross-attention to enhance the feature matching performance of the vector; the optimal matching layer obtains a matching score matrix by calculating the inner product of the feature matching vector, and finally solves the optimal feature allocation matrix.
[0018] In some instances, in the attention graph neural network, it is assumed that each feature point in an input image is a node, the input image is abstracted into a complete graph, self-attention operation is used to connect all feature points inside the input image, and cross-attention calculation is used to connect the input image feature points with all feature points of another complete graph, and self-attention operation and cross-attention operation are performed alternately.
[0019] In some instances, Get the output of the attention graph neural network Depend on Get the output of the attention graph neural network in, is the calculation result of the i-th initial feature on image A at layer l, is the calculation result of the i-th initial feature on image B at layer l, W is the weight, and b is the bias.
[0020] In some examples, at the best matching layer, a soft assignment matrix is constructed by calculating a score matrix and maximizing the overall score.
[0021] In some instances, the aggregate and Calculate the score matrix and introduce an underflow mechanism in the last row or column of the score matrix to filter out incorrect matching relationships.
[0022] In some examples, performing pixel fusion on the stitched images to eliminate stitching traces and generate a panoramic image that conforms to human vision includes:
[0023] Mapping each pixel of the image to be stitched to each coordinate point corresponding to the reference coordinate system according to the spatial transformation relationship;
[0024] The weighted average method is used to fuse the stitched images, and the overlapping parts of the stitched images are smoothed.
[0025] In some examples, the fusing of the images to be stitched by using a weighted average method includes:
[0026] The overlapping pixels in the image to be spliced and the reference image are assigned their own weights, where the weight of the image to be spliced w is m and the reference image weight w n satisfy
[0027] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:
[0028] To address the severe distortion of camera images caused by the height restrictions of panoramic cameras deployed on ships, and to address the need for large-scale scene monitoring, a panoramic image registration method based on an attention mechanism was first employed. Compared to other mainstream algorithms, this method significantly improves the accuracy of the corresponding transformation relationships between the stitched images and the high-precision registration capability of multiple images under these unfavorable conditions. This method achieves visual consistency in large-scale scene monitoring and reduces distortion, playing a positive role in improving the user visual experience, ensuring safe navigation, and efficient deck operations. This invention is suitable for application on various types of surface ships. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0030] Figure 1 is a flow chart of a panoramic image stitching method provided by an embodiment of the present invention;
[0031] Figure 2 This is a schematic diagram of the equidistant cylindrical projection principle provided by an embodiment of the present invention;
[0032] Figure 3 is a flow chart of an image registration algorithm provided by an embodiment of the present invention;
[0033] Figure 4 This is a flow chart of the feature extraction network and feature matching network algorithm provided by an embodiment of the present invention. The feature matching network includes an attention graph neural network and an optimal matching layer;
[0034] Figure 5 This is a flow chart of the feature point and descriptor network algorithm provided by an embodiment of the present invention;
[0035] Figure 6 is a flow chart of a feature matching network algorithm provided by an embodiment of the present invention;
[0036] Figure 7 Schematic diagram of feature matching based on the attention mechanism provided by an embodiment of the present invention;
[0037] Figure 8 This is a schematic diagram of feature matching using ORB provided by an embodiment of the present invention;
[0038] Figure 9 This is an image schematic diagram of a panoramic stitching image registration method based on an attention mechanism provided by an embodiment of the present invention;
[0039] Figure 101 is an image schematic diagram of an ORB panoramic stitching image registration method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0040] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0041] In the following description, specific embodiments of the present invention will be described with reference to steps and symbols performed by one or more computers, unless otherwise specified. Therefore, these steps and operations will be mentioned several times as being performed by a computer, and computer execution as referred to herein includes operations by a computer processing unit that represents electronic signals of data in a structured form. This operation converts the data or maintains it at a location in the computer's memory system, which can be reconfigured or otherwise change the operation of the computer in a manner familiar to testers in the field. The data structure in which the data is maintained is a physical location in the memory, which has specific characteristics defined by the data format. However, the principles of the present invention are described in the above text, which does not represent a limitation, and testers in the field will understand that the various steps and operations below can also be implemented in hardware.
[0042] As used herein, the terms "module" or "unit" may be considered software objects executed on the computing system. The various components, modules, engines, and services herein may be considered implementation objects on the computing system. While the devices and methods herein are preferably implemented in software, they may also be implemented in hardware and remain within the scope of protection of the present invention.
[0043] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.
[0044] For large-scale ship deck scenes with multiple complex objects, a feasible method is to achieve comprehensive monitoring of a large area by stitching images from multiple cameras under the condition of limited camera height. However, the disadvantage is that the cameras need to take into account the entire scene, resulting in severe image distortion. Moreover, due to the need to stitch multiple images together, image consistency is poor. Targeting such scenes, the present invention employs a panoramic stitching image registration method based on an attention mechanism for the first time. This method can complete multi-camera image registration operations in scenes with severe camera image distortion and large scale. While achieving the purpose of panoramic surveillance, it ensures image height consistency, reduces distortion, and improves the user's visual experience, thus ensuring safe navigation of ships and efficient deck operations.
[0045] The key technical point of the image registration process of the present invention is the design of a feature matching network based on an attention mechanism, a neural network technology based on deep learning. The feature matching network design is combined with a feature extraction network. It can address adverse conditions such as high distortion and the need to stitch multiple images together. It can also extract features from the multiple images to be stitched together, and then perform feature matching between the extracted feature points and descriptors to obtain a matching relationship, thereby solving the spatial position relationship mapping. Compared with other mainstream algorithms, the present invention can significantly improve the accuracy of the corresponding transformation relationship between the images to be stitched under the above-mentioned adverse conditions and the high-precision registration capability of multiple images.
[0046] like Figure 1 As shown, the panoramic image stitching provided by the embodiments of the present invention mainly includes four steps: image acquisition, image preprocessing, image registration, and image fusion. The image acquisition step refers to acquiring the multiple high-definition images required for stitching. The image preprocessing step performs distortion correction, exposure correction, and other operations on the images from the multiple imaging sensors built into the panoramic camera, as well as image projection coordinate conversion. The image registration step uses a matching algorithm to obtain the transformation relationship between the images. The image fusion step uses an image fusion algorithm to perform pixel fusion on the stitched images, eliminating stitching artifacts and ultimately generating a panoramic image that conforms to human vision.
[0047] In the embodiment of the present invention, in the pre-processing step, equidistant cylindrical projection is used as the projection transformation method to achieve coordinate unification and prepare for the subsequent image stitching. This method has the characteristics of prominent image features and simple calculation.
[0048] like Figure 2 The equidistant cylindrical projection principle diagram shows that surface CD is the original image plane, and its projected cylindrical surface is surface EF. The radius of the projected cylinder is r, the ratio of OX′ to OX is k, and the coordinates of the target point are x and y. Based on the principle of triangle similarity, the projected x′ and y′ coordinates can be calculated. The formula for equidistant cylindrical projection coordinate conversion is as follows: W is the image width, and H is the image height.
[0049]
[0050] In an embodiment of the present invention, in the image registration link, in order to address the problem of difficulty in matching caused by environmental features such as multiple stitching images and large degree of distortion, as well as the different imaging modes, image features, angles and lighting of images collected by each image sensor in the panoramic camera, an image registration algorithm based on deep learning and attention mechanism is designed. The algorithm can simultaneously extract features from multiple images, and perform feature matching between the extracted feature points and descriptors to obtain the matching relationship between feature points of adjacent images, and then solve the affine transformation relationship between adjacent images to determine the spatial position relationship mapping.
[0051] The image registration algorithm is implemented by the registration network. It is generally divided into five parts, including the registration network image preprocessing module, feature extraction network, feature matching network, post-processing module and spatial transformation relationship, such as Figure 3 As shown. The image preprocessing module of the registration network is used to reduce noise and filter redundant features of the image after image preprocessing, retain and enhance the key point features in the image; the feature extraction network is used to extract key point features of the image data and obtain feature descriptors; the feature matching network performs similarity matching and filtering based on the results of feature extraction to obtain the matching relationship of feature points; the post-processing module further filters the matching information output by the feature matching module; the spatial transformation relationship solves the corresponding spatial position relationship mapping. Among them, the feature extraction network and feature matching network constructed using the neural network receptive field and feature expression ability are as follows. Figure 4 As shown in Figure 2, it is an important component of the registration network, and the feature matching network is the core content of this paper.
[0052] The feature point,descriptor extraction network consists of two parts, such as Figure 5 As shown in the figure, one part is the basic encoding network (encoder), which is used to detect corner points in the input image as candidate feature points; the other part is the feature extraction network (decoder), which is used to output feature points and descriptors. Note that the feature point detection network and the descriptor extraction network share a single forward encoder, while the decoder adopts a different structure.
[0053] The feature matching network includes the Attentional Graph Neural Network (AGNN) and the Optimal Matching Layer, such as Figure 6 As shown in the figure, the attention graph neural network encodes feature points and descriptors into a vector and uses self-attention and cross-attention to enhance the feature matching performance of the vector. The optimal matching layer calculates the inner product of the feature matching vector to obtain the matching score matrix, and finally solves the optimal feature allocation matrix.
[0054] The specific technical implementation of the feature extraction network and feature matching network is as follows:
[0055] In the image feature coding, given two images A and B, the feature point position on the image is recorded as p, and the corresponding descriptor is d, so the image feature can be represented by (p, d). For the i-th image feature point, it can be represented as p i =(x, y, c), where c represents the confidence of the feature point and (x, y) represents the coordinates of the feature point; descriptor d∈R D , D represents the dimension of the feature. According to the above description, the initial feature corresponding to each feature point is:
[0056] x i =d i +MLP enc (p i )
[0057] Among them, MLP stands for Multi-Layer Perceptron. For each feature point, the multi-layer perceptron is used to encode the location information of the feature point and upgrade the dimension to a size that matches the descriptor vector.
[0058] In the attention graph neural network for feature matching, each feature point in an input image is assumed to be a node, and the input image can be abstracted into a complete graph. The self-attention operation can connect all feature points within the image, while the cross-attention operation can connect the feature points of this image with all feature points of another complete graph. During the attention calculation process, the query information key query (query) is input, and the value corresponding to the target attribute key (key) is retrieved, and the final weighted average is obtained as the output. The specific calculation method is:
[0059]
[0060] Among them, ε represents the type of attention calculation; attention weight α ij Represents the similarity between key and query:
[0061]
[0062] The query, key, and value are all obtained through linear projection of the input features. The specific calculation method is as follows:
[0063]
[0064] When performing self-attention calculation, the query, key, and value are all calculated from the input corresponding to the same image; when performing cross-attention calculation, the query needs to be calculated for the input corresponding to the current image, while the key and value are calculated from the input corresponding to the other image.
[0065] In the design, the self-attention operation and the cross-attention operation are performed alternately for a total of L times. Assume is the calculation result of the i-th initial feature on image A at the l-th layer, and its corresponding update calculation method for the l+1-th layer is:
[0066]
[0067] Where || represents the concatenation of self-attention and cross-attention. For image A, the output of the attention graph neural network is:
[0068]
[0069] in, To match the description vector for feature matching, the features in image B also have a similar output form.
[0070] In feature matching, we can calculate the score matrix S and maximize the overall score ∑ i,j S i,j P i,j Construct the soft assignment matrix P. Based on the above feature encoding method and attention graph neural network structure, aggregation and Calculate the score matrix:
[0071]
[0072] An underflow mechanism is introduced in the last row / column of the score matrix S. When the score is not "passing", it is removed from the feature set to filter out incorrect matching relationships and improve the matching accuracy. The matching relationships are further filtered by the set threshold to obtain the feature point pairs with successful matching:
[0073]
[0074] In the embodiment of the present invention, a Random Sample Consensus (RANSAC) algorithm is used in the post-processing phase to eliminate false matches introduced by noise and other factors. The optimal model parameters are determined by randomly sampling observation data and comparing them. The comparison conditions are designed to comply with two conditions: first, for a single model, noisy feature points are not always included in the comparison; second, there are enough feature points to determine the correct model parameters.
[0075] In the spatial transformation relationship link, based on the above filtering results, the affine transformation parameters and the precise spatial position relationship mapping between the images to be stitched are obtained to determine the spatial transformation relationship.
[0076] In the image fusion process, each pixel of the image to be stitched is mapped to each coordinate point corresponding to the reference coordinate system based on the spatial transformation relationship. At the same time, in order to eliminate obvious gaps, the design uses a weighted average method to fuse the images. The overlapping pixels in the image to be stitched m(x, y) and the reference image n(x, y) are assigned their own weights, and the overlapping parts of the image to be stitched are smoothed. The fused pixel is represented as:
[0077]
[0078] Among them, the weight w m and w n Make the following constraints:
[0079]
[0080] The present invention will be further described below with reference to the accompanying drawings and embodiments:
[0081] like Figure 1 ,The flow diagram of panoramic stitching image ,processing includes image acquisition, image preprocessing, image registration and image fusion ,processing. Image registration is the core.
[0082] like Figure 2 ,The image registration algorithm is divided into five parts, registration network image ,preprocessing, feature extraction, feature matching, post-processing and spatial ,transformation relationship, and each part is processed serially.
[0083] like Figure 5 In the feature extraction stage, the feature points and descriptors are calculated. It is necessary to extract images from multiple images that need to be spliced at the same time and calculate the feature points and descriptors of each image.
[0084] like Figure 6 The feature matching process is divided into two parts: the attention graph neural network and the optimal matching layer. In the attention graph neural network, self-attention and cross-attention are used to enhance the registration performance of the features to be spliced; in the optimal matching layer, the feature allocation matrix is calculated.
[0085] In the post-processing stage, the RANSAC algorithm is used to randomly extract observation data to estimate the various parameters of the model and suppress the external point noise.
[0086] The spatial transformation relationship link determines the affine transformation parameters and precise spatial position relationship mapping between the images to be stitched, and determines the spatial transformation relationship.
[0087] Examples:
[0088] Input the images to be stitched into the image registration module and set the parameters required by the algorithm. The main parameters of the image registration module are set as follows:
[0089] (1) The key point detection threshold keypoint_threshold is 0.005
[0090] (2) Set the NMS suppression area radius nms_radius to 4
[0091] (3) The maximum number of detection points max_keypoints is 512
[0092] (4) Feature matching threshold match_threshold is 0.1
[0093] Note: The above simulation takes 69ms (14.5fps) for 512 points on an NVIDIA GeForce GTX 1080 GPU.
[0094] Verification results are as follows Figure 7 、 Figure 8 The two figures show the results of image registration using an attention-based mechanism and traditional ORB. The two results show that the attention-based image registration method can achieve high-precision image registration on extremely wide outdoor images. When using traditional ORB for image registration in the same scene, more outliers and mismatches are observed.
[0095] Combined with the comparative analysis of similar work mentioned above, a quantitative description of the registration using different methods is obtained, see Table 1 below:
[0096] Table 1
[0097]
[0098] It can be seen that the panoramic stitching image registration method based on the attention mechanism is superior to the traditional image registration method based on ORB and SIFT in terms of pose error, matching accuracy and matching score.
[0099] Figure 9 , Figure 10The figure shows a schematic diagram of the panoramic stitching results after using the attention mechanism-based registration method and the ORB method respectively. It can be seen from the figure that when using ORB for image registration, the registration accuracy is not high at the image fusion point, resulting in image noise and edge jagged phenomena in the final presentation effect. After using the panoramic stitching image registration method based on the attention mechanism, a relatively accurate registration result can be obtained, thereby generating a panoramic stitching image with no obvious seams in the stitching area, no obvious image distortion, and clear area.
[0100] The above is a detailed introduction to a panoramic stitching image registration method based on the attention mechanism provided by an embodiment of the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A panoramic stitching image registration method based on attention mechanism, characterized in that: include: Acquire the multi-channel target images required for stitching, pre-process the multi-channel target images, and perform projection coordinate transformation on the pre-processed images; Through the feature matching network based on the attention mechanism, the transformation relationship between the images after coordinate transformation is obtained for image stitching; The weighted average method is used to perform pixel fusion on the stitched images to eliminate stitching traces and generate a panoramic image that conforms to human vision; The feature extraction network includes a basic encoding network and a feature extraction network. The basic encoding network is used to detect corner points of the input image as candidate feature points; the feature extraction network is used to output feature points and descriptors. The feature point extraction network and the descriptor extraction network share a single forward encoder, but use different structures in the decoder part. For each feature point, a multi-layer perceptron is used to encode the location information of the feature point and upgrade the dimension to a size that matches the descriptor vector. The feature matching network includes an attention graph neural network and an optimal matching layer. The attention graph neural network is used to encode feature points and descriptors into a vector and use self-attention and cross-attention to enhance the feature matching performance of the vector. The optimal matching layer calculates the inner product of the feature matching vector to obtain a matching score matrix, and finally solves the optimal feature allocation matrix. In the attention graph neural network, each feature point in an input image is assumed to be a node. The input image is abstracted into a complete graph. Self-attention operation is used to connect all feature points within the input image. Cross-attention operation is used to connect the feature points of the input image with all feature points of another complete graph. Self-attention operation and cross-attention operation are performed alternately. Depend on Get the output of the attention graph neural network ,Depend on Get the output of the attention graph neural network ,in, is the calculation result of the i-th initial feature on image A at layer l, is the calculation result of the i-th initial feature on image B at layer l, is the weight, is bias; In the optimal matching layer, the soft assignment matrix is constructed by calculating the score matrix and maximizing the overall score; polymerization and Calculate the score matrix and introduce an underflow mechanism in the last row or column of the score matrix to filter out incorrect matching relationships.
2. The method according to claim 1, characterized in that The feature matching network based on the attention mechanism is used to obtain the transformation relationship between the images after coordinate transformation to perform image stitching, including: The image after coordinate transformation is subjected to denoising and redundant feature filtering, and key point features are extracted from the image after denoising and redundant feature filtering through a feature extraction network to obtain feature descriptors; The feature matching network is used to perform similarity matching and filtering based on the results of feature extraction to obtain the matching relationship of feature points; The matching relationship of the feature points is further filtered to solve the affine transformation relationship between adjacent images and determine the spatial position relationship mapping.
3. The method according to claim 1 or 2, characterized in that The pixel fusion of the stitched images is performed using a weighted average method to eliminate stitching traces and generate a panoramic image that conforms to human vision, including: Mapping each pixel of the image to be stitched to each coordinate point corresponding to the reference coordinate system according to the spatial transformation relationship; The weighted average method is used to fuse the stitched images, and the overlapping parts of the stitched images are smoothed.
4. The method according to claim 3, characterized in that The weighted average method is used to fuse the stitched images, including: The overlapping pixels in the image to be spliced and the reference image are assigned their own weights, where the weight of the image to be spliced is and reference image weights satisfy .
Citation Information
Patent Citations
Feature point matching method of cross-spectrum image
CN116051872A
Multi-camera image splicing method based on traffic road
CN117974437A