Panoramic spliced image registration method based on attention mechanism
By introducing a feature matching network based on attention mechanism in the traditional image registration algorithm, the problem of traditional methods performing poorly in the ship deck environment is solved, and high-precision image registration and spatial position relationship mapping is achieved, which improves user visual experience and ship safety.
Patent Information
- Application Number
- CN202510234669.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-28
AI Technical Summary
Traditional image registration algorithms perform poorly in ship deck environments, and are greatly affected by light changes, occlusion and sensor errors, resulting in problems such as misalignment of splicing and blurred image. The calculation complexity is high, making it difficult to meet the needs of ships' limited computing resources.
The panoramic stitching image registration method based on attention mechanism is adopted to obtain the transformation relationship between images through the feature matching network, and combine the feature extraction network and the optimal matching layer to achieve high-precision image registration and spatial position relationship mapping.
It improves image registration accuracy and multi-picture high-precision registration capabilities under adverse conditions, reduces the degree of distortion, ensures picture consistency, improves user visual experience, and is suitable for safe navigation of ships and efficient deck operations.
Smart Images

Figure CN120163703A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent image processing, and particularly relates to a panoramic stitching image registration method based on an attention mechanism. Background Art
[0002] In the field of video surveillance, a panoramic camera is deployed at the high point of a surveillance area, and a plurality of imaging sensors are used to generate a panoramic stitching image to achieve large-scene surveillance and provide a global view for users. Panoramic stitching mainly includes links such as image acquisition, image preprocessing, image registration, and image fusion. Among them, the graphic registration link is a process of calculating the corresponding image transformation relationship according to the overlapping part between two adjacent images. As the core of panoramic stitching, the registration efficiency and quality determine the final image application effect and play a key role in aspects such as the overall consistency and distortion degree of the final stitched image.
[0003] Image registration calculates the pixel points with obvious features in the images to be stitched and matches the pixel points between the images, so as to obtain the transformation relationship between the images. Traditional image registration can be divided into 5 steps: (1) Extract feature points. Common feature point extraction algorithms include Scale-Invariant Feature Transform (SIFT), Oriented FAST and Rotated BRIEF (ORB), and Speeded-Up Robust Features (SURF), etc.; (2) Calculate descriptors; (3) Nearest neighbor matching; (4) Filter out outliers; (5) Solve geometric constraints.
[0004] The height of the position where the panoramic stitching camera is installed on the ship is limited. To meet the large-scale surveillance of the deck surface, more imaging sensors need to be configured for image stitching; there are many miscellaneous objects on the deck surface, and the marine weather is changeable (more rain and fog), which causes various degrees of interference to the consistency, accuracy, and stability of image feature extraction and feature matching. In this environment, the performance of traditional image matching algorithms is poor and cannot support the high-precision registration requirements. Even in the case of light changes and occlusions, problems such as stitching misalignment and image blurring are likely to occur.
[0005] Moreover, the operations on the ship's deck are frequent, and the instantaneous requirement for the panoramic stitching registration result is strong. Traditional registration algorithms are very sensitive to small-scale perspective errors caused by sensor errors, and the spatial relationship mapping between the two is prone to jitter, affecting the image fusion effect.
[0006] In short, the traditional method has a high computational complexity and usually requires a large amount of computing resources to process data, which causes great pressure on the limited computing resources of the ship. Summary of the Invention
[0007] An embodiment of the present invention provides a panoramic stitching image registration method based on an attention mechanism, which, while achieving the purpose of panoramic surveillance, ensures a highly consistent picture, reduces the degree of distortion, improves the user's visual experience, and plays a guarantee role for the safe navigation of ships, efficient operation on the deck, etc.
[0008] In a first aspect, the present invention provides a panoramic stitching image registration method based on an attention mechanism, including:
[0009] Obtain multiple target images required for stitching, preprocess the multiple target images, and perform projection coordinate transformation on the preprocessed images;
[0010] Through a feature matching network based on an attention mechanism, obtain the transformation relationship between the images after coordinate transformation for image stitching;
[0011] Perform pixel fusion on the stitched image to eliminate stitching traces and generate a panoramic image that conforms to human vision.
[0012] In some examples, the step of obtaining the transformation relationship between the images after coordinate transformation through a feature matching network based on an attention mechanism for image stitching includes:
[0013] Perform noise reduction and redundant feature filtering on the images after coordinate transformation, and perform key point feature extraction on the images after noise reduction and redundant feature filtering through a feature extraction network to obtain feature descriptors;
[0014] Perform similarity matching and filtering based on the results of feature extraction through a feature matching network to obtain the matching relationship of feature points;
[0015] Further filter the matching relationship of feature points, and then solve the affine transformation relationship between adjacent images to determine the mapping of spatial position relationships.
[0016] In some examples, the feature extraction network includes a basic encoding network and a feature extraction network. The basic encoding network is used to detect the corner points of the input image as candidate feature points; the feature extraction network is used to output feature points and descriptors. Among them, the feature point extraction network and the descriptor extraction network share a single forward encoder, and different structures are adopted in the decoder part.
[0017] In some examples, the feature matching network includes an attention graph neural network and an optimal matching layer. Among them, the attention graph neural network is used to encode feature points and descriptors into a vector, and uses self-attention and cross-attention to enhance the feature matching performance of the vector; the optimal matching layer calculates the inner product of the feature matching vectors to obtain a matching degree score matrix, and finally solves the optimal feature assignment matrix.
[0018] In some instances, in the attention graph neural network, assuming that each feature point in an input image is a node, the input image is abstracted as a complete graph. Self-attention operations are used to connect all feature points within the input image, and cross-attention calculations are used to connect the feature points of the input image with all feature points of another complete graph. The self-attention operations and cross-attention operations are alternated.
[0019] In some instances, from obtain the output of the attention graph neural network From obtain the output of the attention graph neural network where is the calculation result of the i-th initial feature on image A at the l-th layer, is the calculation result of the i-th initial feature on image B at the l-th layer, W is the weight, and b is the bias.
[0020] In some instances, in the optimal matching layer, a soft assignment matrix is constructed by calculating a score matrix and maximizing the overall score.
[0021] In some instances, aggregate and calculate the score matrix, and introduce an underflow mechanism in the last row or column of the score matrix to filter out incorrect matching relationships.
[0022] In some instances, performing pixel fusion on the spliced image to eliminate the splicing trace and generate a panoramic image that conforms to human eye vision includes:
[0023] Mapping each pixel of the image to be spliced to each coordinate point of the reference coordinate system according to the spatial transformation relationship;
[0024] Fusing the image to be spliced using the weighted average method, and simultaneously smoothing the overlapping part of the image to be spliced.
[0025] In some instances, the step of fusing the image to be spliced using the weighted average method includes:
[0026] Respectively assign respective weight values to the overlapping pixel points in the image to be spliced and the reference image, where the weight value w m of the image to be spliced and the weight value w n of the reference image satisfy
[0027] Generally speaking, compared with the prior art by the above technical solution conceived by the present invention, the following beneficial effects can be achieved:
[0028] To solve the problem of severe distortion of the camera image caused by the height limitation of the panoramic camera installed on the ship and the demand for large-scale scene monitoring, a panoramic stitching image registration method based on the attention mechanism is adopted for the first time. Compared with other mainstream algorithms, this method can greatly improve the accuracy of the corresponding transformation relationship between the images to be stitched and the high-precision registration ability of multiple images under the above-mentioned adverse conditions, realize the visual consistency of large-scale scene monitoring, reduce the degree of distortion, and play a positive role in improving the user's visual experience, ensuring the safe navigation of the ship, and efficient operation on the deck. The invention can be popularized and applied to various types of surface ships. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0030] Figure 1 It is a flowchart of the panoramic stitching image method provided by the embodiment of the present invention;
[0031] Figure 2 It is a schematic diagram of the equidistant cylindrical projection principle provided by the embodiment of the present invention;
[0032] Figure 3 It is a flowchart of the image registration algorithm provided by the embodiment of the present invention;
[0033] Figure 4 It is a flowchart of the feature extraction network and the feature matching network algorithms provided by the embodiment of the present invention. The feature matching network includes an attention graph neural network and an optimal matching layer;
[0034] Figure 5 It is a flowchart of the feature point and descriptor network algorithm provided by the embodiment of the present invention;
[0035] Figure 6 It is a flowchart of the feature matching network algorithm provided by the embodiment of the present invention;
[0036] Figure 7 It is a schematic diagram of the feature matching based on the attention mechanism provided by the embodiment of the present invention;
[0037] Figure 8 It is a schematic diagram of the feature matching using ORB provided by the embodiment of the present invention;
[0038] Figure 9 It is a schematic diagram of the image using the panoramic stitching image registration method based on the attention mechanism provided by the embodiment of the present invention;
[0039] Figure 10It is a schematic diagram of an image using the ORB panoramic stitching image registration method provided by an embodiment of the present invention. Detailed implementation manners
[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts belong to the protection scope of the present invention.
[0041] In the following description, specific embodiments of the present invention will be described with reference to steps and symbols executed by one or more computers, unless otherwise specified. Therefore, these steps and operations will be mentioned several times as being executed by a computer. What is meant by a computer execution herein includes operations of a computer processing unit representing electronic signals in a structured form of data. This operation transforms the data or maintains it at a position in the computer's memory system, which can be reconfigured or otherwise changed in a manner well known to those skilled in the art to change the operation of the computer. The data structure maintained by the data is a physical location in the memory, which has specific characteristics defined by the data format. However, the principles of the present invention are described in the above text, which does not represent a limitation. Those skilled in the art will understand that the following various steps and operations can also be implemented in hardware.
[0042] The term "module" or "unit" used herein can be regarded as a software object executed on the computing system. Different components, modules, engines, and services herein can be regarded as implementation objects on the computing system. The devices and methods herein are preferably implemented in software, and of course can also be implemented in hardware, all within the protection scope of the present invention.
[0043] Those skilled in the art of this technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0044] For large-scale warship deck scenes with multiple types of complex objects, under the condition of limited camera installation height, a feasible method is to achieve overall monitoring of a large area through multi-camera image stitching. However, the disadvantage is that the camera needs to cover the entire scene, resulting in serious image distortion, and due to the involvement of multi-image stitching, the image consistency is poor. For this type of scene, the present invention first adopts a panoramic stitching image registration method based on an attention mechanism, which can complete the multi-camera image registration operation under the conditions of serious camera image distortion and large-scale scenes. While achieving the purpose of panoramic monitoring, it ensures high image consistency, reduces the distortion degree, improves the user's visual experience, and plays a guarantee role for the safe navigation of warships and efficient deck operations.
[0045] In the image registration link of the present invention, the design content of a feature matching network based on an attention mechanism is the technical key point, and the attention mechanism is a neural network technology based on deep learning. The feature matching network design combines with the feature extraction network, which can extract features from multiple images to be stitched simultaneously for adverse conditions such as high distortion and the need for multi-image stitching, and perform feature matching on the extracted feature points and descriptors to obtain the matching relationship, and then solve the mapping of the spatial position relationship. Compared with other mainstream algorithms, the present invention can greatly improve the accuracy of the corresponding transformation relationship between the images to be stitched and the high-precision multi-image registration ability under the above adverse conditions.
[0046] As Figure 1 shown, the panoramic image stitching provided by the embodiment of the present invention mainly includes four links: image acquisition, image preprocessing, image registration, and image fusion. Among them, the image acquisition link refers to acquiring multiple high-definition images required for stitching; the image preprocessing link corrects the distortion, exposure, etc. of the images of multiple imaging sensors built in the panoramic camera, and performs image projection coordinate conversion; the image registration link obtains the transformation relationship between the images through a matching algorithm; image fusion is to perform pixel fusion on the stitched images through an image fusion algorithm to eliminate the stitching marks and finally generate a panoramic image that conforms to the human eye vision.
[0047] In the embodiment of the present invention, in the preprocessing link, equidistant cylindrical projection is adopted as the projection transformation method to achieve coordinate unification and make preparations for subsequent image stitching. This method has characteristics such as prominent image features and simple calculation.
[0048] As Figure 2 shown in the schematic diagram of equidistant cylindrical projection, the CD plane is the original image plane, and its projection cylinder surface is the EF curved surface. Among them, the radius of the projection cylinder is r, the ratio of OX′ to OX is k, and the coordinates of the target point are x and y. According to the principle of similar triangles, the coordinates of x′ and y′ after projection can be calculated, and the formula for coordinate conversion of equidistant cylindrical projection is as follows. Among them, W is the image width and H is the image height.
[0049]
[0050] In the embodiments of the present invention, in the image registration link, aiming at the environmental characteristics of a large number of splicing pictures and a large degree of distortion, as well as the problems of difficult matching caused by different imaging modes, image features, angle illumination, etc. of the images collected by each image sensor in the panoramic camera, an image registration algorithm based on the attention mechanism of deep learning is designed, which can simultaneously extract features from multiple images, perform feature matching on the extracted feature points and descriptors, obtain the matching relationship of adjacent image feature points, and then solve the affine transformation relationship between adjacent images to determine the mapping of the spatial position relationship.
[0051] The image registration algorithm is implemented by a registration network. Generally, it is divided into five parts, including a registration network image preprocessing module, a feature extraction network, a feature matching network, a post-processing module, and a spatial transformation relationship, as Figure 3 shown. The registration network image preprocessing module is used to denoise and filter redundant features of the preprocessed image, retain and enhance the key point features in the image; the feature extraction network is used to extract key point features from the image data and obtain feature descriptors; the feature matching network performs similarity matching and filtering based on the results of feature extraction to obtain the matching relationship of feature points; the post-processing module further filters the matching information output by the feature matching module; the spatial transformation relationship solves the corresponding mapping of the spatial position relationship. Among them, the feature extraction network and the feature matching network constructed by using the receptive field and feature expression ability of the neural network, as Figure 4 shown, are important components of the registration network, and the feature matching network among them is the core content of the present invention.
[0052] The feature point and descriptor extraction network includes two parts, as Figure 5 shown. One part is the basic coding network (encoder), which is used to detect the corner points of the input image as candidate feature points; the other part is the feature extraction network (decoder), which is used to output feature points and descriptors. Note that the feature point detection network and the descriptor extraction network share a single forward encoder, and different structures are adopted in the decoder part.
[0053] The feature matching network includes an Attentional Graph Neural Network (AGNN) and an Optimal Matching Layer, as Figure 6 shown. Among them, the Attentional Graph Neural Network encodes the feature points and descriptors into a vector and enhances the feature matching performance of the vector by using self-attention and cross-attention; the Optimal Matching Layer calculates the inner product of the feature matching vectors to obtain the matching degree score matrix, and finally solves the optimal feature assignment matrix.
[0054] The specific technical implementation of the feature extraction network and the feature matching network is as follows:
[0055] In the feature encoding of an image, given two pictures A and B, the positions of the feature points on the pictures are denoted as p, and the corresponding descriptors are d. Therefore, the image features can be represented by (p, d). For the i-th image feature point, it can be expressed as p i =(x, y, c), where c represents the confidence of the feature point, and (x, y) represents the coordinates of the feature point; the descriptor d ∈ R D , and D represents the dimension of the feature. According to the above description, the initial feature corresponding to each feature point is:
[0056] x i =d i +MLP enc (p i )
[0057] where MLP represents the Multi-Layer Perceptron. For each feature point, the multi-layer perceptron is used to encode the position information of the feature point and dimension it up to a size that matches the descriptor vector.
[0058] In the attention graph neural network for feature matching, assuming that each feature point in an input image is a node, the input image can be abstracted as a complete graph. Using self-attention operations can connect all the feature points inside the image, while cross-attention calculations can connect the feature points of this graph with all the feature points of another complete graph. During the attention calculation process, the input query information key query retrieves the value (value) corresponding to the attribute based on the target attribute key, and finally, the weighted average is obtained as the output. The specific calculation method is as follows:
[0059]
[0060] where ε represents the type of attention calculation; the attention weight α ij represents the similarity between the key and the query:
[0061]
[0062] The query, key, and value are all obtained through the linear projection operation of the input features. The specific calculation method is as follows:
[0063]
[0064] When performing self-attention calculation, the query, key, and value are all calculated from the input corresponding to the same image; when performing cross-attention calculation, the query needs to be calculated from the input corresponding to the current image, while the key and value are calculated from the input corresponding to another image.
[0065] In the design, the self-attention operation and the cross-attention operation are alternated and executed L times. Suppose is the calculation result of the i-th initial feature on image A at the l-th layer, and its update calculation method corresponding to the l+1-th layer is:
[0066]
[0067] where, || represents the concatenation of self-attention and cross-attention. For image A, the output of the attention graph neural network is:
[0068]
[0069] where, is the matching description vector for feature matching, and the features in image B also have a similar output form.
[0070] In feature matching, the score matrix S can be calculated, and the soft assignment matrix P can be constructed by maximizing the overall score ∑ i,j S i,j P i,j Based on the above feature encoding method and the attention graph neural network structure, aggregate and to calculate the score matrix:
[0071]
[0072] An underflow mechanism is introduced in the last row / column of the score matrix S. When the score is "failed", it is removed from the feature set to filter out incorrect matching relationships and improve the accuracy of matching. The matching relationships are further filtered through a set threshold to obtain the successfully matched feature point pairs:
[0073]
[0074] In the embodiment of the present invention, in the post-processing link, the Random Sample Consensus (RANSAC) algorithm is used to eliminate incorrect matches introduced due to noise and other reasons. By randomly extracting observation data and selecting the optimal model parameters, for the selection conditions, two conditions are designed. One is that for a single model, the noise feature points will not always participate in the selection; the other is that there are enough feature points to determine the correct model parameters.
[0075] In the spatial transformation relationship step, based on the above filtering results, the affine transformation parameters and the mapping of the precise spatial position relationship between the images to be stitched are obtained to determine the spatial transformation relationship.
[0076] In the image fusion step, each pixel of the images to be stitched is mapped to the corresponding coordinate points in the reference coordinate system according to the spatial transformation relationship. At the same time, in order to eliminate obvious gaps, the weighted average method is designed to fuse the images. For the overlapping pixel points in the images to be stitched m(x, y) and the reference image n(x, y), their respective weights are assigned, and the overlapping part of the images to be stitched is smoothed. The fused pixel is expressed as:
[0077]
[0078] Among them, for the weights w m and w n the following constraints are imposed:
[0079]
[0080] The present invention will be further described below in conjunction with the accompanying drawings and embodiments:
[0081] As Figure 1 , the panoramic stitching image flow chart includes steps such as image acquisition, image preprocessing, image registration, and image fusion. Image registration is the core.
[0082] As Figure 2 , the image registration algorithm is divided into five parts: preprocessing of the registration network image, feature extraction, feature matching, post-processing, and spatial transformation relationship, and each part is processed serially.
[0083] As Figure 5 , in the feature extraction step, the feature points and descriptors are calculated. It is necessary to extract images from multiple images to be stitched at the same time, and calculate the feature points and descriptors of each image.
[0084] As Figure 6 , in the feature matching step, it is divided into two parts: the attention map neural network and the optimal matching layer. In the attention map neural network, the feature registration performance of the images to be stitched is enhanced through self-attention and cross-attention; in the optimal matching layer, the feature assignment matrix is calculated.
[0085] In the post-processing step, the RANSAC algorithm is used to randomly extract observation data to estimate the parameters of the model, thereby suppressing the outlier noise.
[0086] In the spatial transformation relationship step, the affine transformation parameters and the mapping of the precise spatial position relationship between the images to be stitched are determined to determine the spatial transformation relationship.
[0087] Example:
[0088] Input the image to be stitched into the image registration module and set the parameters required by the algorithm. The main parameters of the image registration module are set as follows:
[0089] (1) The keypoint detection threshold keypoint_threshold is 0.005
[0090] (2) Set the NMS suppression region radius nms_radius to 4
[0091] (3) The maximum number of detected points max_keypoints is 512
[0092] (4) The feature matching threshold match_threshold is 0.1
[0093] Note: The above simulation is carried out in the NVIDIA GeForce GTX 1080 GPU environment, and the time taken for 512 points is 69 ms (14.5 fps).
[0094] The verification results are as Figure 7 、 Figure 8 , and the two figures are the schematic diagrams of the results of registration using the attention mechanism and the results of image registration using the traditional ORB respectively. It can be seen from the two results that the image registration method based on the attention mechanism can perform high-precision image registration on extremely wide outdoor images; when using the traditional ORB for image registration in the same scenario, there are more outliers and mismatching phenomena compared.
[0095] Combined with the comparative analysis of similar work above, the quantitative description of registration using different methods is obtained, as shown in Table 1 below:
[0096] Table 1
[0097]
[0098] It can be seen that the panoramic stitching image registration method based on the attention mechanism is superior to the traditional image registration methods mainly based on ORB and SIFT in terms of pose error, matching accuracy, and matching score.
[0099] Figure 9 , Figure 10The following is a schematic diagram of the results of panoramic stitching after using the registration method based on the attention mechanism and the ORB method respectively. It can be seen from the figure that when using ORB for image registration, due to its low registration accuracy at the image fusion area, there are phenomena such as image noise and edge sawtooth in the final presentation effect. However, after using the panoramic stitching image registration method based on the attention mechanism, a relatively accurate registration result can be obtained, thereby generating a panoramic stitching image with no obvious seams in the stitching area, no obvious distortion in the image, and clear regions.
[0100] The above has introduced in detail a panoramic stitching image registration method based on the attention mechanism provided by the embodiments of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A panoramic stitching image registration method based on attention mechanism, characterized in that: include: Acquire multiple target images required for stitching, preprocess the multiple target images, and perform projection coordinate transformation on the preprocessed images; Through the feature matching network based on the attention mechanism, the transformation relationship between the images after coordinate transformation is obtained to perform image stitching; The stitched images are pixel-fused to eliminate stitching marks and generate a panoramic image that conforms to human vision.
2. The method according to claim 1, characterized in that The feature matching network based on the attention mechanism is used to obtain the transformation relationship between the images after coordinate transformation to perform image stitching, including: Denoise and filter redundant features on the image after coordinate transformation, extract key point features from the image after denoising and filtering redundant features through a feature extraction network and obtain feature descriptors; The feature matching network is used to perform similarity matching and filtering based on the results of feature extraction to obtain the matching relationship of feature points; The matching relationship of the feature points is further filtered to solve the affine transformation relationship between adjacent images and determine the spatial position relationship mapping.
3. The method according to claim 2, characterized in that The feature extraction network includes a basic encoding network and a feature extraction network. The basic encoding network is used to detect corner points of an input image as candidate feature points; the feature extraction network is used to output feature points and descriptors, wherein the feature point extraction network and the descriptor extraction network share a single forward encoder, and different structures are used in the decoder part.
4. The method according to claim 3, characterized in that: The feature matching network includes an attention graph neural network and an optimal matching layer, wherein the attention graph neural network is used to encode feature points and descriptors into a vector, and use self-attention and cross-attention to enhance the feature matching performance of the vector; the optimal matching layer obtains a matching score matrix by calculating the inner product of the feature matching vector, and finally solves the optimal feature allocation matrix.
5. The method according to claim 4, characterized in that In the attention graph neural network, it is assumed that each feature point in an input image is a node, the input image is abstracted as a complete graph, self-attention operation is used to connect all feature points inside the input image, and cross-attention calculation is used to connect the feature points of the input image with all feature points of another complete graph, and self-attention operation and cross-attention operation are performed alternately.
6. The method according to claim 5, characterized in that Depend on Get the output of the attention graph neural network Depend on Get the output of the attention graph neural network in, is the calculation result of the i-th initial feature on image A at layer l, is the calculation result of the i-th initial feature on image B at layer l, W is the weight, and b is the bias.
7. The method according to claim 6, characterized in that In the best matching layer, the soft assignment matrix is constructed by calculating the score matrix and maximizing the overall score.
8. The method according to claim 7, characterized in that polymerization and The score matrix is calculated, and an underflow mechanism is introduced in the last row or column of the score matrix to filter out erroneous matching relationships.
9. The method according to any one of claims 1 to 8, characterized in that: The pixel fusion of the stitched images is performed to eliminate stitching traces and generate a panoramic image that conforms to human vision, including: Mapping each pixel of the image to be stitched to each coordinate point corresponding to the reference coordinate system according to the spatial transformation relationship; The weighted average method is used to fuse the stitched images, and the overlapping parts of the stitched images are smoothed.
10. The method according to claim 9, characterized in that The method of fusing the stitched images by using a weighted average method includes: The overlapping pixels in the image to be stitched and the reference image are assigned their own weights, where the weight of the image to be stitched w m and the reference image weight w n satisfy
Citation Information
Patent Citations
Feature point matching method of cross-spectrum image
CN116051872A
Multi-camera image splicing method based on traffic road
CN117974437A
Deep learning method for multiple object tracking from video
US20240144489A1