Unmanned aerial vehicle multi-modal image registration method and system based on dual-channel attention network

The infrared and visible image features are extracted through a dual-channel attention network, combined with feature pyramids and adaptive matching strategies, the accuracy and anti-interference problems of multimodal image registration in drone aerial photography are solved, and high-precision image registration is achieved.

CN120355759APending Publication Date: 2025-07-22JILIN ELECTRIC POWER RES INST LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510425001.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently carry out multimodal registration of infrared and visible light images, especially in drone aerial photography, which has problems such as small number of feature points, low registration accuracy and complex background interference. Traditional methods are prone to introduce artificial errors and nonlinearization of the grayscale mapping relationship of multiple sources of images, resulting in registration failure.

Method used

Using a dual-channel attention network method, feature points are extracted through position attention and channel attention models, combined with feature pyramid module and adaptive matching strategy, feature representation and registration accuracy are enhanced, and feature point detection and descriptor matching are optimized using loss function.

Benefits of technology

It improves the accuracy and anti-interference ability of multimodal image registration, enhances the accuracy of feature point extraction and matching in complex backgrounds, and improves the registration effect of aerial images of drones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355759A_ABST
    Figure CN120355759A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-modal image registration, in particular to an unmanned aerial vehicle multi-modal image registration method and system based on a double-channel attention network, and the method comprises the steps: inputting an infrared image and a visible light image into a feature point detection module, and extracting the feature points of the two images; the two-channel attention network can extract features and descriptors of infrared and visible light images at the same time, and the global features after decoding are enhanced by adopting a mode of combining position attention with channel attention, so that the model can resist interference of a complex background, and the registration precision is improved; a feature pyramid module is adopted to carry out multi-scale feature registration on the image so as to extract target features of different scales, especially small targets, the number of extracted feature points is increased, and the mismatching rate is reduced; by adopting a homography adaptive matching strategy, the recheck rate of feature points and the cross-domain practicability are enhanced, and the registration precision of a multi-modal image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal image registration, and particularly relates to a method and system for unmanned aerial vehicle multimodal image registration based on a dual-channel attention network. Background Art

[0002] Renewable energy sources such as wind power, photovoltaic power generation, and hydropower are natural green energy sources. Wind turbine blades are key components for wind turbines to capture wind energy. Due to long-term exposure to harsh outdoor working environments, the surfaces of wind turbine blades are prone to damage such as corrosion and wear, gel coat peeling, and cracks. According to investigations, the cost of wind turbine blades accounts for about 20% of the overall cost, and the accident rate accounts for 30% of the total accidents. Wind turbine collapse incidents caused by blades occur from time to time, which may cause significant economic and property losses, and even casualties. Therefore, timely detection and maintenance of wind turbine blade damage are of great significance for ensuring the safe operation of wind turbines.

[0003] With the popularization of unmanned aerial vehicles and the development of optical technologies, the use of unmanned aerial vehicles to carry multiple sensors to aerial photograph the same scene has been widely applied. Image registration technology geometrically aligns two or more images of the same scene obtained at the same time from different perspectives or different sensors, which is the key to ensuring subsequent tasks such as image stitching, image fusion, and defect recognition. However, due to different sensor imaging modes and differences in the perception environment, multimodal image registration of infrared and visible light images has always been one of the difficult problems in the field of computer vision. In recent years, deep learning technologies have been widely applied in the field of image registration, and a series of algorithm models based on deep learning have emerged.

[0004] To address the problem of difficult extraction of effective features in infrared and visible light images, there are currently five solutions as follows: an infrared and visible light image registration method based on a residual dense network, which uses a feature extraction network designed based on a residual dense network to extract hierarchical features of the image pair; a convolutional neural network structure for geometric matching, which is based on three main components, simulates the feature extraction and matching processes, and at the same time includes standard steps for internal detection and model parameter estimation, and additionally supports end-to-end trainability; an unsupervised deep homography estimation network, which, on the basis of completing feature extraction, adds a mask structure to complete the function of the random sample consensus algorithm, and then realizes the outlier filtering task during registration; a method for enhancing local feature descriptors, which includes cross-modal context, regional information, and geometric structure information to surpass traditional local detail representations; the D2-Net method, which aims to perform feature detection and dense feature description simultaneously. This method delays feature detection to a later stage and obtains more stable and reliable key points through this model, solving the problem of difficult finding of reliable pixel points under complex imaging conditions.

[0005] For the registration of UAV aerial images, most of the above methods adopt a single traditional method or a combination of traditional methods and deep learning. For the problems faced by multi-modal image registration, such as few extracted feature points, low registration accuracy, and interference from complex backgrounds, in addition, due to the differences in the imaging principles of multi-source images such as infrared and visible light, traditional registration algorithms based on image gray level and features are difficult to achieve accurate registration. There are the following problems in the registration of multi-source images: on the one hand, the registration system of multi-source images often relies on manual registration. By manually marking the coordinates of key matching points to calculate the image coordinate transformation matrix, this method is prone to introducing human errors and is time-consuming and laborious in the registration process. On the other hand, the gray mapping relationship of such multi-source images for the same object is non-linear, making it difficult to describe features based on the gray level of the domain, which in turn leads to registration failure or low registration accuracy. Summary of the Invention

[0006] The purpose of the present invention is to overcome the defects of the prior art and propose a UAV multi-modal image registration method based on a dual-channel attention network, which can improve the registration accuracy of multi-modal images.

[0007] To achieve the above purpose, the present invention adopts the following specific technical solutions:

[0008] The UAV multi-modal image registration method based on a dual-channel attention network provided by the present invention includes the following steps:

[0009] S1. Extract feature points of visible light images and infrared images;

[0010] S2. Model the dependence relationships of the spatial dimension and the channel dimension of the extracted feature points through a position attention model and a channel attention model, and perform rough matching on the outputs of the position attention model and the channel attention model to enhance the feature representation;

[0011] The position attention model is used to determine the spatial dependence relationship between any two positions, and the channel attention model is used to capture the channel dependence relationship between any two channel mappings and update it using the weights of all channel mappings;

[0012] S3. Perform optimal value screening and homography matrix solution on the roughly matched feature points to select the features with the best matching results, and register the images according to the homography matrix; the optimal value screening scores the confidence of the feature points obtained by rough matching to evaluate the accuracy of feature point matching, and performs more refined matching to obtain an accurate homography matrix;

[0013] S4. Construct a loss function, and use a spatial transformation network to map the image to be registered into the result image through data sets, evaluation indicators, and ablation experiment analysis results.

[0014] Further, in step S1, the SuperPoint model is used to extract the feature points of the visible light image and the infrared image. The SuperPoint model includes an encoder, a feature pyramid, a feature point decoder, and a feature description decoder;

[0015] The encoder includes a convolutional layer, a max pooling layer, and a non-linear activation function. The encoder uses three max pooling layers to reduce the size of the input image to 1 / 8 of the original image size. Each pooling layer is a 2×2 non-overlapping max pooling layer;

[0016] The feature pyramid first reduces the number of channels of the feature map to 64 and obtains the feature map through bilinear interpolation Then, multi-scale feature registration is performed to obtain the feature map Finally, the feature map After convolution, the obtained feature map Is used for decoding;

[0017] The feature point decoder includes two 3×3 convolutional layers. The second layer is a dimensionality reduction layer with a stride of 1. First, the feature map After Conv1 operation, it is converted into a 65-channel feature map Through The coordinates X pos , Y pos Of the feature points are obtained, and then through the reconstruction operation for upsampling, the resolution is adjusted back to the original image size to obtain Finally, feature point calculation is performed on the full-resolution map;

[0018] In the feature description decoder, the input feature map Is output through the Conv2 operation; Conv2 contains two convolutional layers. The parameters of the first layer are the same as those of Conv1, and the second layer is still a dimensionality reduction layer. The number of output channels is set to 256 for subsequent transplantation operations; The feature map output by Conv2 First, it is restored to the full resolution through bilinear interpolation, then normalized to the unit length through L2-Norm, and finally a feature map with the same size as the original image and 256 channels is output

[0019] Further, specifically in step S2, the fused feature map Performs a dimensionality reduction operation in the attention module to obtain Processed through a convolutional layer on To obtain three mappings B, C, and D. B and C are reshaped from a 3-order tensor to a 2-order tensor. B is transposed and multiplied by C to obtain R (H×W)×(H×W)For the matrix, the spatial attention map S is calculated through the softmax layer. The transpose of S is used for matrix multiplication with D, and element-wise addition is performed to obtain the output feature E. The calculation processes of the spatial attention map and the output feature are as follows:

[0020]

[0021] Then, the value of the position attention is given the weight α and multiplied by the feature map after dimensionality reduction to obtain the output feature E;

[0022]

[0023] where S mn represents the correlation between the nth position and the mth position. The larger the value, the greater the similarity. The weight α represents the scaling factor;

[0024] Different from the position attention, the channel attention module uses the dimensionality-reduced feature map multiplied by its own transpose matrix; then, the channel attention map X ∈ R C×C is obtained through the softmax layer; finally, the transpose of A is used for matrix multiplication with X, and then element-wise addition with A is performed to obtain E. The calculation process is as follows:

[0025]

[0026] where X mn represents the correlation between the nth channel and the mth channel, and β represents the scaling factor.

[0027] Furthermore, in step S3, the optimal value screening is specifically as follows:

[0028] The feature points and descriptions are combined into the initial representation x i of each feature point:

[0029]

[0030] where MLP is a multi-layer perceptron used to increase the dimension of low-dimensional features by coupling the positions and visual representations of feature points;

[0031] By repeatedly enhancing the feature matching mechanism of vectors, self-attention and cross-attention are used to describe the matching relationship between feature points; by aggregating information and cross-information, the message m ε→i can be obtained as the weighted average of v j in the attention mechanism:

[0032] m ε→i = ∑ j:(i,j)∈ε α ij v j ;

[0033] Among them, the attention weight α ij is the softmax of the similarity between the query q i and the retrieved object key value k j , that is, α ij = softmax j (q i T k j );

[0034] Assume that the feature point i to be queried is located on the target image Q to be registered, and all source feature points are located on the source image S. W is the network parameter in matrix format, W1 is the set of W2 and W3, b is the bias, the key k j , the query q i and the value v j can be written as:

[0035]

[0036] Each loop l has its corresponding projection parameter, which is shared by all feature points; q i corresponds to the feature representation of the feature point i on the image to be registered, k i k j and v i v j are all a kind of mapping from the recalled image feature point j; α ij represents the similarity of these two features, which is calculated from q i and k j . The larger the value, the more similar the two feature points are. Then, the similarity is used to perform weighted summation on v j to obtain m ε→i ; The confidence score of the feature point pair can be expressed as:

[0037]

[0038] Among them, <·,·> is the inner product, f i S and are the final matching descriptors; According to the confidence S i,j , the Sinkhorn algorithm is used to generate the optimal feature assignment matrix, where the sum of each row and column of the final matrix is 1; By continuously scaling and updating S i,j until it converges completely;

[0039] The specific solution of the homography matrix is as follows:

[0040] The homography matrix is solved by constructing a soft matching matrix H, and the soft matching matrix is obtained by calculating the score matrix Implementation, where M is the number of rows of the matrix and N is the number of columns; by maximizing the overall score ∑ i,j S i,j H i,j the homography matrix H can be obtained.

[0041] Furthermore, in step S4, the overall loss L of the loss function all is calculated to include the positional decoder loss L d and the descriptor decoder loss L f , and the specific formula is as follows:

[0042] L all = α d L d + α f L f ;

[0043] To balance the two parts of the loss, two additional weight parameters α d and α f are also set to ensure the correct convergence of the loss function;

[0044] The calculation of the positional decoder loss L d includes calculating the scores of two frames of images after homography transformation and calculating the error and confidence between the predicted feature point values and the true values. The formula is as follows:

[0045] L d = L ds (x oi , y oi ) + L ds (x tri , y tri ) + L dp (x of , y of ) + L dp (x trf , y trf );

[0046] Among them, L ds represents the positional score loss, L dp represents the positional error loss of the feature points, (x oi , y oi ) represents the source image, (x tri , y tri ) represents the transformed image, (x of , y of ) represents the predicted value of the source image, and (x trf , y trf ) represents the predicted value of the transformed image;

[0047] The focal loss function is used to improve the convergence speed of feature point detection:

[0048]

[0049] F ft = -α θ (1 - PF) λ log(PF);

[0050] In the formula, U ω = H / 8, V ω = W / 8, y is the label of the true value of the feature point grounded, and α θ is used to suppress the imbalance number of positive and negative samples, and λ is used to control the imbalance number of difficult and easy samples;

[0051] By adding a soft - argmax function in the 5×5 patch near each feature point, the estimation accuracy of the feature point is further improved, the coordinate position of each feature point is refined, and the coordinates are updated with sub - pixel accuracy. The calculation process is as follows:

[0052] T pi = T o +(△x, △y);

[0053]

[0054] Among them, T pi is the coordinate of the updated predicted value, T o is the central pixel coordinate value, T g is the pixel value at the position of the heat map g, and T t is the true value;

[0055] L soarg = L s-arg1 + L s-arg2 + L s-arg3 ;

[0056]

[0057] L s-arg3 = φ∑(L s-arg - L s-arg2 );

[0058] Among them, L s-arg1 is the cumulative error, ξ is the weight coefficient, L s-arg2 is the error mean, L s-arg3 represents the confidence level, and φ is the weight coefficient;

[0059] Descriptor decoder loss L fThe calculation is as follows: The two image frames for calculating the descriptor decoder loss are the image pair after homography transformation; by calculating the homography corresponding point pairs between the two images, the final descriptor loss is obtained. To improve the training accuracy and convergence speed of the network, by setting the error loss term of the descriptor, the descriptor in the original image is The descriptor in the transformed image is During training, the matching threshold T is set to 4 pixel values, specifically as follows:

[0060]

[0061] M = ‖N - N h ‖;

[0062] L = ‖N - N h ‖²;

[0063] where M represents the interval between two pixels, N represents the central coordinate of a pixel in a unit after homography transformation, N h represents the central pixel coordinate of the corresponding unit in the image after homography transformation, and L represents the 2-norm of the pixel spacing;

[0064] Since the number of corresponding points between frames is less than the number of non-corresponding points, by setting a modulation factor χ to reduce the influence of the high loss of non-corresponding points on the entire descriptor loss function, a hinge loss function is used to add upper and lower bounds for prediction, r t and r b are the upper and lower bounds respectively;

[0065]

[0066] F m = χ·Z·max(0, r t - d T d h ) + (1 - Z)·max(0, d T d h - r b );

[0067] In the formula, K and K h represent the number of corresponding points and non-corresponding points in the original image, respectively, and the two frames of images after homography transformation; ψ is the weight coefficient of the descriptor cumulative error loss term.

[0068] Further, in step S4, the dataset is experimented on the UAV-MM-FB dataset to verify the effectiveness of the dual-channel attention registration model; the evaluation metrics use the root mean square error, peak signal-to-noise ratio, mutual information, correct matching point count, and target registration error methods to conduct quantitative index comparisons to compare the effects of different methods and evaluate the performance of the dual-channel attention registration model; the ablation experiment verifies the effectiveness of the dual-channel attention registration model by conducting ablation experiments on the feature pyramid, dual-channel attention, and homography matrix.

[0069] The calculation method of the root mean square error RMSE is as follows:

[0070]

[0071] The smaller the root mean square error value, the better the registration effect and the higher the registration accuracy.

[0072] The calculation method of the peak signal-to-noise ratio PSNR is as follows:

[0073]

[0074] Among them, MAX is the maximum possible pixel value in the image, MAX is set to 255, and MSE is the mean square error between images; the higher the peak signal-to-noise ratio value, the lower the distortion generated during the registration process and the more similar the registered images.

[0075] The calculation method of the mutual information MI is as follows:

[0076] MI(R,F) = H(R) + H(F) - H(R,F);

[0077] Among them, H(R) represents the entropy of the reference image, H(F) represents the entropy of the image to be registered, and H(R,F) represents their joint entropy; the smaller the joint entropy of the images, the larger the mutual information value and the higher the similarity degree between the images.

[0078] The correct matching point count NOCC is used to evaluate the number of feature points that can be correctly matched extracted by the model.

[0079] The calculation method of the target registration error TRE is as follows:

[0080]

[0081] The target registration error represents the difference of the same registration point in two images. The calculation method of the target registration error is to calculate the Euclidean distance between the original point and the registered point, and then calculate its mean and variance.

[0082] The present invention also provides a UAV multi-modal image registration system based on a dual-channel attention network, which applies the above-mentioned UAV multi-modal image registration method based on a dual-channel attention network, and includes:

[0083] Feature point detection module: used to extract feature points of visible light images and infrared images;

[0084] Dual-channel attention registration network module: used to model the dependence relationships of the spatial dimension and channel dimension of the feature points of the visible light image and the infrared image extracted by the feature point detection module through a position attention model and a channel attention model, and perform rough matching on the outputs of the position attention model and the channel attention model to enhance the feature representation; the position attention model is used to determine the spatial dependence relationship between any two positions, the channel attention model is used to capture the channel dependence relationship between any two channel mappings, and is updated using the weights of all channel mappings;

[0085] Adaptive matching module: performs optimal value screening and homography matrix solution on the roughly matched feature points to select the features with the optimal matching results, and registers the images according to the homography matrix; the optimal value screening performs confidence scoring on the feature points obtained by rough matching to evaluate the accuracy of feature point matching, and performs more refined matching to obtain an accurate homography matrix;

[0086] Spatial transformation network module, which uses a spatial transformation network to map the image to be registered into the result image.

[0087] The present invention can achieve the following technical effects:

[0088] The UAV multi-modal image registration method and system based on a dual-channel attention network provided by the present invention can simultaneously extract the features and descriptors of infrared and visible light images, and adopt the method of combining position attention and channel attention to enhance the global features after decoding, so that the model can resist the interference of complex backgrounds and improve the registration accuracy; adopt a feature pyramid module to perform multi-scale feature registration on the images to extract target features of different scales, especially small targets, increase the number of extracted feature points, and reduce the false matching rate; by adopting a homography adaptive matching strategy, the re-inspection rate of feature points and the cross-domain practicability are enhanced to improve the registration accuracy of multi-modal images. Description of the Drawings

[0089] Figure 1 is a schematic flowchart of the UAV multi-modal image registration method based on a dual-channel attention network provided by an embodiment of the present invention;

[0090] Figure 2 is a schematic structural diagram of the feature point detection module provided by an embodiment of the present invention;

[0091] Figure 3 is a schematic structural diagram of an encoder provided according to an embodiment of the present invention;

[0092] Figure 4 is a schematic structural diagram of a feature pyramid provided according to an embodiment of the present invention;

[0093] Figure 5 is a schematic structural diagram of a position decoder provided according to an embodiment of the present invention;

[0094] Figure 6 is a schematic structural diagram of a feature description decoder provided according to an embodiment of the present invention;

[0095] Figure 7 is a schematic structural diagram of a dual-channel attention registration network provided according to an embodiment of the present invention;

[0096] Figure 8 is a schematic diagram of a position attention model provided according to an embodiment of the present invention;

[0097] Figure 9 is a schematic diagram of a channel attention model provided according to an embodiment of the present invention;

[0098] Figure 10 is a comparative diagram of visual effects of ablation experiments conducted on the UAV-MM-FB dataset according to an embodiment of the present invention;

[0099] Figure 11 is a comparative diagram of visual effects on the UAV-MM-FB dataset according to an embodiment of the present invention;

[0100] Figure 12 is a diagram of quantitative comparison results of ablation experiments according to an embodiment of the present invention;

[0101] Figure 13 is a diagram of comparative results of quantitative metrics on the UAV-MM-FB dataset according to an embodiment of the present invention. Detailed implementation manners

[0102] In the following, embodiments of the present invention will be described with reference to the accompanying drawings. In the following description, the same modules are denoted by the same reference numerals. In the case of the same reference numerals, their names and functions are also the same. Therefore, their detailed descriptions will not be repeated.

[0103] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention.

[0104] An embodiment of the present invention provides a method and system for multi-modal image registration of unmanned aerial vehicles based on a dual-channel attention network. The basic structure of the system and the method flow are as follows Figure 1 As shown, first, the infrared and visible light images are input into the feature point detection module to extract the feature points of the two images; then, the extracted feature points are input into the dual-channel attention registration network for rough matching; immediately afterwards, the adaptive matching module is used to screen the optimal values of the paired feature points and solve the homography matrix, effectively selecting the features with the best matching results, and registering the images according to the homography matrix; finally, the spatial transformation network is used to map the image to be registered into the result image.

[0105] The specific steps of the multi-modal image registration method of unmanned aerial vehicles based on the dual-channel attention network are as follows:

[0106] S1. Extract the feature points of the visible light image and the infrared image.

[0107] The feature point detection module is mainly composed of two inputs, infrared and visible light. Each input has to go through the SuperPoint model for feature point extraction. The structure of the feature point detection module is as follows Figure 2 As shown. The SuperPoint model consists of an encoder, a feature pyramid, a feature point decoder, and a descriptor decoder to achieve the feature point extraction task of visible light and infrared images.

[0108] The structure of the encoder is as follows Figure 3 As shown, the encoder consists of a convolutional layer, a max-pooling layer, and a non-linear activation function. The encoder uses three max-pooling layers to reduce the size of the input image to 1 / 8 of the original image size. Each pooling layer is a 2×2 non-overlapping max-pooling layer.

[0109] Construct a hierarchical feature extraction ability through the combination of the convolutional layer, the max-pooling layer, and the non-linear activation function. Use three non-overlapping max-pooling layers to gradually compress the input size to 1 / 8 of the original image, significantly reducing the computational complexity while maintaining the key feature information, enhancing the model's abstract expression ability of the local features of the image, and effectively suppressing noise interference through gradual dimensionality reduction, improving the robustness of feature encoding.

[0110] The structure of the feature pyramid is as follows Figure 4 As shown, the feature pyramid first reduces the dimension of the feature map F3 to 64 channels, and obtains the feature map through bilinear interpolation Then, perform multi-scale feature registration on the image to obtain the feature map Finally, perform convolution on the feature map again to obtain the feature map Use to perform subsequent decoding work.

[0111] By reducing the feature map to 64 channels, computational redundancy is reduced, effectively improving the recognition accuracy of multi-scale targets in object detection or segmentation tasks.

[0112] The structure of the position decoder is as Figure 5 shown. The feature point decoder includes two 3×3 convolutional layers; the second layer is a dimensionality reduction layer with a stride of 1. First, the feature map is converted into a 65-channel feature map after Conv1 operation through to obtain the coordinates X pos and Y pos of the feature points. Then, through upsampling by reconstruction operation (reshape), the resolution is adjusted back to the original image size to obtain Finally, feature point calculations are performed on the full-resolution map. By setting two convolutional layers, efficient decoding of feature points can be achieved.

[0113] The structure of the feature description decoder is as Figure 6 shown. The feature description decoder is similar to the processing process of the detector. The input feature map needs to be output through Conv2 operation; Conv2 contains two convolutional layers. The parameters of the first layer are the same as those of Conv1, and the second layer is still a dimensionality reduction layer, but the number of output channels is set to 256, which is also for the convenience of subsequent transplantation operations; the feature map output by Conv2 is first restored to full resolution through bilinear interpolation; then, it is normalized to unit length through L2-Norm, and finally a feature map with the same size as the original image and 256 channels is output

[0114] Through the feature description decoder, the discriminability and matching robustness of the feature descriptor are enhanced, taking into account multi-task compatibility and cross-scale geometric consistency, providing a high-discriminative dense feature expression basis for image registration or retrieval tasks.

[0115] S2. Model the dependence relationships of the spatial dimension and channel dimension of the extracted feature points through the position attention model and the channel attention model, and perform rough matching on the outputs of the position attention model and the channel attention model to enhance the feature representation;

[0116] The position attention model is used to determine the spatial dependence relationship between any two positions, and the channel attention model is used to capture the channel dependence relationship between any two channel maps and update using the weights of all channel maps.

[0117] To enhance the global features of the decoded feature map, the dependence relationships of the spatial dimension and the channel dimension are respectively modeled through the position attention model and the channel attention model; the position attention module determines the spatial dependence relationship between any two positions; the channel attention model captures the channel dependence relationship between any two channel mappings and uses the weights of all channel mappings to update them, and the outputs of the two attention modules are roughly matched to enhance the feature representation. The structure of the dual-channel attention registration network is as Figure 7 shown, the position attention model is as Figure 8 shown, and the channel attention model is as Figure 9 shown.

[0118] The weight calculation of the dual-channel attention network is as follows. The fused feature map needs to be dimension-reduced in the attention module to obtain Process through a convolutional layer to obtain three mappings B, C, and D. B and C are reshaped from a 3-order tensor to a 2-order tensor. Transpose B and multiply it by C to become an R (H×W)×(H×W) matrix. Then, calculate the spatial attention map S through the softmax layer, perform matrix multiplication on D with the transpose of S, and perform pixel-wise addition to obtain the output feature E. The calculation processes of the spatial attention map and the output feature are:

[0119]

[0120] Then assign the value of the position attention to the weight α and multiply it by the dimension-reduced feature map to obtain the output feature E;

[0121]

[0122] where S mn represents the correlation between the nth position and the mth position. The larger the value, the greater the similarity. The weight α represents the scaling factor;

[0123] Different from the position attention, the channel attention module multiplies the dimension-reduced feature map by its own transpose matrix; then, obtain the channel attention map X ∈ R C×C through the softmax layer; finally, perform matrix multiplication on X with the transpose of A, and then perform element-wise addition with A to obtain E. The calculation process is:

[0124]

[0125] where X mn represents the correlation between the nth channel and the mth channel, and β represents the scaling factor.

[0126] The position attention module adds context information to the local features to enhance their representation, as Figure 8 shown; the channel attention module improves the feature representation by establishing interdependencies between channels, as Figure 9 shown.

[0127] S3. Perform optimal value screening and homography matrix solving on the coarsely matched feature points to select the features with the optimal matching results, and register the images according to the homography matrix; the optimal value screening scores the confidence of the feature points obtained by the coarse matching to evaluate the accuracy of the feature point matching, and performs a more refined matching to obtain an accurate homography matrix.

[0128] Optimal value screening: After the coarse matching by the dual-channel attention registration network to obtain partially matched feature points, it is also necessary to score the confidence of these feature points to evaluate the accuracy of the feature point matching, and perform a more refined matching to obtain a more accurate homography matrix (HomographyNet). Since the combination of the feature point position and description will obtain stronger feature matching specificity, the feature points and descriptions are merged into the initial representation x of each feature point i ;

[0129]

[0130] where MLP is a multi-layer perceptron used to increase the dimension of the low-dimensional features by coupling the position and visual representation of the feature points. Subsequently, by repeatedly enhancing the feature matching mechanism of the vector, self-attention and cross-attention are used to describe the matching relationship between the feature points; by aggregating information and cross-information, the message m ε→i is obtained as the weighted average of v j in the attention mechanism:

[0131] m ε→i = ∑ j:(i,j)∈ε α ij v j ; (6)

[0132] where the attention weight α ij is the softmax of the similarity between the query q i and the retrieved object key value k j , that is, α ij = softmax j (q i T k j ). Suppose the feature point i to be queried is located on the target image Q to be registered, and all source feature points are located on the source image S, W is the network parameter in matrix format, W1 is the set of W2 and W3, b is the bias, so the key kj Query q i Sum value v j can be written as:

[0133]

[0134] Each loop l has its corresponding projection parameters, which are shared by all feature points; q i The feature representation corresponding to the feature point i on the image to be registered, k i k j and v i v j are both a mapping from the recalled image feature point j; α ij represents the similarity of these two features, which is calculated from q i and k j The larger the value, the more similar the two feature points are. Then, the similarity is used to perform weighted summation on v j to obtain m ε→i ;

[0135] The confidence score of the feature point pair can be expressed as:

[0136]

[0137] where <·,·> is the inner product, f i S and are the final matching descriptors; with the confidence S i,j , then the Sinkhorn algorithm is used to generate the optimal feature assignment matrix, where the sum of each row and column of the final matrix is 1; this process is achieved by continuously scaling and updating S i,j until it converges completely.

[0138] Solving the homography matrix: The homography matrix is solved by constructing a soft matching matrix H; for the registration task of the present invention, this soft matching matrix can be realized by calculating the score matrix where M is the number of rows of the matrix and N is the number of columns; specifically, it is to maximize the overall score ∑ i,j S i,j H i,j to obtain the homography matrix H.

[0139] S4. Construct the loss function. Through the dataset, evaluation metrics, and ablation experiment analysis results, the image to be registered is mapped to the result image using the spatial transformation network.

[0140] The overall loss calculation of the loss function consists of the position decoder loss L d and the descriptor decoder loss L fComposed as follows: The specific formula is as follows:

[0141] L all = α d L d + α f L f ; (10)

[0142] By adopting a training form that simultaneously optimizes these two parts of the loss, in order to balance the two parts of the loss, two additional weight parameters α d and α f are also set to ensure the correct convergence of the loss function.

[0143] The loss calculation of the position decoder is divided into two parts. The first part is to calculate the scores of the two frames of images after the homography transformation, and the second part is to calculate the error and confidence between the predicted values and the true values of the feature points. The loss calculation of the position decoder is as follows:

[0144] L d = L ds (x oi , y oi ) + L ds (x tri , y tri ) + L dp (x of , y of ) + L dp (x trf , y trf ); (11)

[0145] Among them, L ds represents the position score loss, L dp represents the position error loss of the feature points, (x oi , y oi ) represents the source image, (x tri , y tri ) represents the transformed image, (x of , y of ) represents the predicted value of the source image, (x trf , y trf ) represents the predicted value of the transformed image.

[0146] Feature point detection can be regarded as a binary classification problem. Most previous works are based on the cross-entropy loss function. However, when the sample distribution is uneven, it is easy to have biases. Therefore, the focal loss function is used in model training to improve the convergence speed of feature point detection. The loss calculation is as follows:

[0147]

[0148] Fft = -α θ (1 - PF) λ log(PF); (14)

[0149] Where U ω = H / 8, V ω = W / 8, y is the label of the true value of the feature point grounded, α θ is used to suppress the imbalance number of positive and negative samples, and λ is used to control the imbalance number of difficult and easy samples;

[0150] To further improve the estimation accuracy of the feature points, by adding a soft-argmax function in the 5×5 patch near each feature point, further refine the coordinate position of each feature point, and update the coordinates with sub-pixel accuracy. The calculation process is as follows:

[0151] T pi = T o + (△x, △y); (15)

[0152]

[0153] Where T pi is the updated predicted value coordinate, T o is the central pixel coordinate value, T g is the pixel value at the position of the heat map g, T t is the true value.

[0154] The error loss of the feature point coordinate value includes error accumulation, error mean and confidence. The role of error accumulation is to ensure the overall accuracy of feature point prediction, and the error mean and confidence loss terms are to ensure the stability of the accuracy of the feature points extracted by the training network;

[0155] L soarg = L s-arg1 + L s-arg2 + L s-arg3 ; (18)

[0156]

[0157]

[0158] L s-arg3 = φ∑(L s-arg - L s-arg2 ); (21)

[0159] Where L s-arg1 is the cumulative error, ξ is the weight coefficient, L s-arg2 is the error mean, L s-arg3 represents the confidence, and φ is the weight coefficient.

[0160] Descriptor decoder loss: The two image frames for calculating the descriptor decoder loss are the image pair after homography transformation; the final descriptor loss is obtained by calculating the homography corresponding point pairs between the two images. Meanwhile, to improve the training accuracy and convergence speed of the network, by setting the error loss term of the descriptor, the descriptor in the original image is The descriptor in the transformed image is During training, the matching threshold T is set to 4 pixel values, and the specific calculation process is as follows:

[0161]

[0162] M = ‖N - N h ‖; (23)

[0163] L = ‖N - N h ‖2; (24)

[0164] where, M represents the interval between two pixels, N represents the central coordinate of a pixel in a unit after homography transformation, N h represents the central pixel coordinate of the corresponding unit in the image after homography transformation, and L represents the 2-norm of the pixel spacing.

[0165] Since the number of corresponding points between frames is significantly less than the number of non-corresponding points, a modulation factor χ is set to reduce the influence of the high loss of non-corresponding points on the entire descriptor loss function. Meanwhile, the hinge loss function is used to add upper and lower bounds for prediction, r t and r b are the upper and lower bounds respectively;

[0166]

[0167] F m = χ·Z·max(0, r t - d T d h ) + (1 - Z)·max(0, d T d h - r b );(26)

[0168] In the formula, K and K h represent the number of corresponding points and non-corresponding points in the original image, and the two frames of images after homography transformation respectively; ψ is the weight coefficient of the descriptor cumulative error loss term.

[0169] The design of the descriptor decoder loss and the position decoder loss forms a synergistic effect through weighted joint optimization. The position loss precisely constrains the spatial distribution of feature points, and the descriptor loss strengthens the modality invariance of feature representations. This dual-supervision mechanism effectively improves the robustness of the model in scenarios where there are large gray-scale differences and unequal texture information between infrared and visible-light images, laying a solid foundation for the high-precision solution of the subsequent homography matrix.

[0170] Dataset: Experiments were conducted on the UAV-MM-FB dataset to verify the effectiveness of the dual-channel attention registration model and the superiority of the state-of-the-art methods.

[0171] Evaluation metrics: The root mean square error (RMSE), peak signal-to-noise ratio (PSNR), mutual information (MI), number of correct correspondences (NOCC), and target registration error (TRE) were used to quantitatively compare different methods, enabling a fair comparison of the effects of different methods and thus evaluating the performance of the dual-channel attention registration model.

[0172] Ablation experiment: Ablation experiments were carried out on the feature pyramid (FP), dual-channel attention (Tw-A), and homography matrix (HN) to verify the effectiveness of the dual-channel attention registration model. The visual effect comparison diagrams of the ablation experiments on the UAV-MM-FB dataset are as Figure 10 shown, and the visual effect comparison diagrams on the UAV-MM-FB dataset are as Figure 11 shown.

[0173] By combining the position attention and channel attention mechanisms to enhance the global feature representation after feature decoding, using the feature pyramid module to extract multi-scale target features to increase the number of feature points, and introducing a homography adaptive matching strategy to optimize the feature point recheck rate and cross-domain usability, the problems of few feature points, low accuracy, and complex background interference in existing multi-modal registration algorithms are solved.

[0174] The calculation formulas for the quantitative metrics of the evaluation metrics are as follows:

[0175] The root mean square error is used to measure the accuracy of image registration, and the calculation method of the root mean square error is as follows:

[0176]

[0177] The smaller the root mean square error value, the better the registration effect and the higher the registration accuracy;

[0178] The peak signal-to-noise ratio is an objective standard for evaluating images. The higher the peak signal-to-noise ratio value, the lower the distortion generated during the registration process, and the more similar the registered images. The calculation method of the peak signal-to-noise ratio is as follows:

[0179]

[0180] Among them, MAX is the maximum possible pixel value in the image, MAX is set to 255, and MSE is the mean square error between images;

[0181] Mutual information is an evaluation method for measuring the quality of image registration. It evaluates the quality of image registration by calculating the mutual information between two images. The mutual information calculation method is as follows:

[0182] MI(R,F) = H(r) + H(F) - H(R,F); (29)

[0183] Among them, H(R) represents the entropy of the reference image, H(F) represents the entropy of the image to be registered, and H(R,F) represents the joint entropy of the two; the smaller the joint entropy of the image, the larger the mutual information value, indicating a higher degree of similarity between the images;

[0184] The number of correct matching points is used to evaluate the number of feature points that can be correctly matched extracted by the model;

[0185] The target registration error formula is as follows:

[0186]

[0187] The target registration error represents the difference of the same registration point in two images. The calculation method of the target registration error is to calculate the Euclidean distance between the original point and the registered point, and then calculate its mean and variance.

[0188] The evaluation in terms of visual effect cannot objectively compare the advantages and disadvantages of different registration methods. Therefore, a comparison of quantitative indicators is needed to fairly compare the effects of different methods.

[0189] The results of the ablation experiment show that when the FP module is removed from the network, the feature expression ability of the network becomes poor, resulting in a significant increase in the RMSE value and a decrease in the registration accuracy; after removing Tw-A, in addition to the increase in RMSE, the signal-to-noise ratio PSNR decreases significantly, and the mutual correlation MI also decreases, while TRE increases, indicating a decrease in the registration accuracy; after removing HN, the performance of the network is affected by the mismatched points, resulting in a significant decline in the indicators.

[0190] The quantitative comparison of the ablation experiment is as Figure 12 shown, where RMER is 110.9, PSNR is 36.5, MI is 0.49, NOCC is 141, and TRE is 32.6 for the best result; for intuitive comparison, two groups of images are selected from UAV-MM-FB for visual effect analysis. The analysis results are as Figure 10As shown, where (a) is the visible light image, (b) is the infrared image, (c) is the removal of FP, (d) is the removal of Tw-A, (e) is the removal of HN, and (f) is the method of the present invention. When FP is removed, the decline in feature expression ability leads to the loss of information, resulting in fewer registration points. When Tw-A is removed, the scale features cannot be fully registered, which also brings the problem of background interference. However, its negative impact on the network is less than the removal of FP. After removing HN, the network discards important parameters in the second-stage training, resulting in a large parameter error in the homography matrix solution and a low re-inspection rate. Especially in the case of cross-domain, by using the method of the present invention, it can be seen that the registration effect has been effectively enhanced.

[0191] The comparison results of quantitative indicators on the UAV-MM-FB dataset are as Figure 13 shown. The feature points proposed by the present invention are around 1100, which is significantly higher than other methods. Therefore, the RMSE of the dual-channel attention registration model is 79.7, which is significantly lower than traditional registration methods such as the ORB method and SURF, and also lower than deep learning registration methods such as ALIKE and L2Net, indicating that the dual-channel attention registration model has higher accuracy, which is 42.8% higher than the ORB method, 38.5% higher than the SURF method, 11% higher than the ALIKE method, and 18.3% higher than the L2Net method. In terms of PSNR and MI, the PSNR and MI indicators of the dual-channel attention registration model are also significantly better than other methods.

[0192] By selecting two groups of typical image pairs from the UAV-MM-FB dataset for comparison, the comparison effect is as Figure 11 shown. The registration points in the figure have been marked with straight lines. In the figure, (a) and (b) are the original visible light and infrared images. The result of registering using the ORB method is shown in (c) of the figure. It can be seen that the extracted feature points are very few and the registration effect is not good. The result of registering using the SURF method is shown in (d) of the figure. The extracted feature points are slightly more than the ORB method, but there are still many mis-matching points, resulting in a distorted registration result. The result of registering using the ALIKE method is shown in (e) of the figure. It can be seen that due to the interference of the complex background, the matching distortion is obvious. The result of registering using the L2Net method is shown in (f) of the figure. Although this method can also extract more feature points, it is seriously interfered by the complex background, resulting in low registration accuracy. The registration result of the dual-channel attention registration model is shown in (g) of the figure. It can be seen that the dual-channel attention registration model is superior to other methods in terms of the number of feature extraction points, registration accuracy, and anti-complex background interference.

[0193] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0194] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

[0195] The above specific implementation manners of the present invention do not constitute a limitation on the protection scope of the present invention. Any other corresponding changes and deformations made according to the technical concept of the present invention should be included in the protection scope of the claims of the present invention.

Claims

1. A method for multi-modal image registration of unmanned aerial vehicles based on a dual-channel attention network, characterized in that It includes the following steps: S1. Extract the feature points of the visible light image and the infrared image; S2. Model the dependence relationships of the spatial dimension and the channel dimension for the extracted feature points through the position attention model and the channel attention model, and perform rough matching on the outputs of the position attention model and the channel attention model to enhance the feature representation; The position attention model is used to determine the spatial dependence relationship between any two positions, and the channel attention model is used to capture the channel dependence relationship between any two channel mappings and update using the weights of all channel mappings; S3. Perform optimal value screening and homography matrix solution on the roughly matched feature points to select the features with the optimal matching results, and register the images according to the homography matrix; the optimal value screening performs confidence scoring on the feature points obtained by rough matching to evaluate the accuracy of feature point matching, and performs more refined matching to obtain an accurate homography matrix; S4. Construct a loss function, analyze the results through the dataset, evaluation metrics and ablation experiments, and use the spatial transformation network to map the image to be registered into the result image.

2. The method for multi-modal image registration of an unmanned aerial vehicle based on a dual-channel attention network according to claim 1, wherein In step S1, the SuperPoint model is used to extract the feature points of the visible light image and the infrared image. The SuperPoint model includes an encoder, a feature pyramid, a feature point decoder and a feature description decoder; The encoder includes a convolutional layer, a max pooling layer and a non-linear activation function. The encoder uses three max pooling layers to reduce the size of the input image to 1 / 8 of the original image size. Each pooling layer is a 2×2 non-overlapping max pooling layer; The feature pyramid first reduces the dimensionality of the feature map to 64 channels and obtains the feature map through bilinear interpolation Then, multi-scale feature registration is performed to obtain the feature map Finally, for the feature map The feature map obtained after convolution is used for decoding; The feature point decoder includes two 3×3 convolutional layers. The second layer is a dimensionality reduction layer with a stride of 1. First, the feature map is converted into a 65-channel feature map after the Conv1 operation Through the coordinates X pos , Y pos of the feature points are obtained, and then through the reconstruction operation for upsampling, the resolution is adjusted back to the original image size to obtain Finally, feature point calculation is performed on the full-resolution map; The feature map input in the feature description decoder is output through the Conv2 operation; Conv2 contains two convolutional layers. The parameters of the first layer are the same as those of Conv1, and the second layer is still a dimensionality reduction layer with the number of output channels set to 256 for subsequent transplantation operations; the feature map output by Conv2 is first restored to the full resolution through bilinear interpolation, then normalized to the unit length through L2-Norm, and finally a feature map with the same size as the original image and 256 channels is output 3. The method for multi-modal image registration of an unmanned aerial vehicle based on a dual-channel attention network according to claim 2, wherein Specifically in step S2, the fused feature map is dimensionally reduced in the attention module to obtain Through the convolutional layer, is processed to obtain three mappings B, C, and D. B and C are reshaped from 3-order tensors into 2-order tensors. B is transposed and multiplied by C to obtain the matrix of (H×W)×(H×W) R. Then, the spatial attention map S is calculated through the softmax layer. The transpose of S is used to perform matrix multiplication on D, and a per-pixel addition operation is performed to obtain the output feature E. The calculation processes of the spatial attention map and the output feature are as follows: Then, the value of the position attention is given a weight α and multiplied by the feature map after dimensionality reduction to obtain the output feature E; Among them, S mn represents the correlation between the nth position and the mth position. The larger the value, the greater the similarity. The weight α represents the proportionality factor; Different from the position attention, the channel attention module utilizes the dimensionality-reduced feature map multiplied by its own transposed matrix; and then obtains the channel attention map X ∈ R C×C through the softmax layer; finally, performs matrix multiplication on X with the transpose of A, and then performs element-wise addition with A to obtain E, and the calculation process is as follows: Among them, X mn represents the correlation between the nth channel and the mth channel, and β represents the scaling factor.

4. The method for multi-modal image registration of an unmanned aerial vehicle based on a dual-channel attention network according to claim 1, wherein In step S3, the optimal value screening is specifically as follows: Combine the feature points and descriptions into an initial representation x of each feature point i : Among them, the MLP is a multi-layer perceptron, which is used to increase the dimension of the low-dimensional features by coupling the position and visual representation of the feature points; By repeatedly enhancing the feature matching mechanism of the vector, self-attention and cross-attention are used to describe the matching relationship between feature points; by aggregating information and cross-information, the message m can be obtained ε→i As v in the attention mechanism j The weighted average of: m ε→i = ∑ j:(i,j)∈ε α ij v j ; Among them, the attention weight α ij is the softmax of the similarity between the query q i and the retrieved object key value k j , that is, α ij = softmax j (q i T k j ); Suppose the feature point i to be queried is located on the target image Q to be registered, and all source feature points are located on the source image S. W is the network parameter in matrix format, W1 is the set of W2 and W3, b is the bias, the key k j , the query q i , and the value v j can be written as: For each cycle l, there is a corresponding projection parameter, which is shared by all feature points; q i The feature representation corresponding to the feature point i on the image to be registered, k i k j And v i v j Are all mappings from the recalled image feature point j; α ij Indicates the similarity of these two features, which is determined by q i And k j Calculated. The larger the value, the more similar the two feature points. Then, the similarity is used to perform weighted summation on v j To obtain m ε→i ; The confidence score of the feature point pair can be expressed as: where <·,·> is the inner product, and f i s and f j Q are the final matching descriptors; according to the confidence S i,j , the optimal feature assignment matrix is generated using the Sinkhorn algorithm, where the sum of each row and column of the final matrix is 1; by continuously scaling and updating S i,j until it converges completely; The homography matrix solution is specifically as follows: The homography matrix is solved by constructing a soft matching matrix H, and the soft matching matrix is realized by calculating the score matrix , where M is the number of rows of the matrix and N is the number of columns; by maximizing the total score ∑ i,j S i,j H i,j the homography matrix H can be obtained.

5. The method for multi-modal image registration of an unmanned aerial vehicle based on a dual-channel attention network according to claim 1, wherein In step S4, the overall loss L of the loss function all Calculate including the position decoder loss L d and the descriptor decoder loss L f , and the specific formula is as follows: L all = α d L d + α f L f ; To balance the losses of the two parts, two additional weight parameters α d and α f are also set to ensure the correct convergence of the loss function; Pose decoder loss L d is calculated by computing the score of two frames of images after homography transformation and calculating the error and confidence between the predicted value and the true value of the feature points. The formula is as follows: L d = L ds (x oi , y oi ) + L ds (x tri , y tri ) + L dp (x of , y of ) + L dp (x trf , y trf ); Among them, L ds represents the location score loss, and L dp represents the location error loss of the feature points. (x oi , y oi ) represents the source image, (x tri , y tri ) represents the transformed image, (x of , y of ) represents the predicted value of the source image, and (x trf , y trf ) represents the predicted value of the transformed image; Use the loss function focalloss to improve the convergence speed of feature point detection: F ft = -α θ (1 - PF) λ log(PF); Wherein, U ω = H / 8, V ω = W / 8, y is the label of the true value of the feature point grounded, α θ is used to suppress the imbalance number of positive and negative samples, and λ is used to control the imbalance number of difficult and easy samples; By adding a soft-argmax function in a 5×5 patch near each feature point, further improve the estimation accuracy of the feature points, refine the coordinate positions of each feature point, and update the coordinates with sub-pixel accuracy. The calculation process is as follows: T pi = T o + (△x, △y); Among them, T pi is the updated predicted value coordinate, T o is the central pixel coordinate value, T g is the pixel value at the position of heat map g, T t is the true value; L soarg = L s-arg1 + L s-arg2 + L s-arg3 ; L s-arg3 = φ∑(L s-arg - L s-arg2 ); Among them, L s-arg1 is the cumulative error, ξ is the weight coefficient, L s-arg2 is the mean error, L s-arg3 represents the confidence level, and φ is the weight coefficient; Descriptor decoder loss L f is calculated as follows: The two image frames for calculating the descriptor decoder loss are the image pair after homography transformation; by calculating the homography corresponding point pairs between the two images, the final descriptor loss is obtained. To improve the training accuracy and convergence speed of the network, by setting the error loss term of the descriptor, the descriptor in the original image is and the descriptor in the transformed image is During training, the matching threshold T is set to 4 pixel values, specifically as follows: M = ‖N - N h ‖; L = ‖N - N h ‖2; where M represents the interval between two pixels, N represents the central coordinate of a pixel in a unit after the homography transformation, and N h represents the central pixel coordinate of the corresponding unit in the image after the homography transformation, and L represents the 2-norm of the pixel pitch; Since the number of corresponding points between frames is less than the number of non-corresponding points, a modulation factor χ is set to reduce the impact of the high loss of non-corresponding points on the entire descriptor loss function. The hinge loss function is used to add upper and lower bounds for prediction, where r t and r b are the upper and lower bounds respectively; F m = χ·Z·max(0, r t - d T d h ) + (1 - Z)·max(0, d T d h - r b ); where K and K h represent the number of corresponding points and non-corresponding points in the original image, and two frames of images after the homography transformation, respectively; ψ is the weight coefficient of the descriptor cumulative error loss term.

6. The method for multi-modal image registration of an unmanned aerial vehicle based on a dual-channel attention network according to claim 5, wherein In step S4, the dataset is used to conduct experiments on the UAV-MM-FB dataset to verify the effectiveness of the dual-channel attention registration model; the evaluation metrics use the root mean square error, peak signal-to-noise ratio, mutual information, number of correct matches, and target registration error methods to conduct quantitative index comparisons to compare the effects of different methods to evaluate the performance of the dual-channel attention registration model; the ablation experiment conducts ablation experiments on the feature pyramid, dual-channel attention and homography matrix to verify the effectiveness of the dual-channel attention registration model; The calculation method of the root mean square error RMSE is as follows: The smaller the root mean square error value, the better the registration effect and the higher the registration accuracy; The calculation method of the peak signal-to-noise ratio PSNR is as follows: Among them, MAX is the maximum possible pixel value in the image, MAX is set to 255, and MSE is the mean square error between the images; the higher the peak signal-to-noise ratio value, the lower the distortion generated during the registration process and the more similar the registered images are; The calculation method of mutual information MI is as follows: MI(R,F) = H(R) + H(F) - H(R,F); Among them, H(R) represents the entropy of the reference image, H(F) represents the entropy of the image to be registered, and H(R,F) represents the joint entropy of the two; the smaller the joint entropy of the image, the larger the mutual information value, and the higher the similarity degree between the images; The number of correct matching points NOCC is used to evaluate the number of feature points that can be correctly matched extracted by the model; The calculation method of the target registration error TRE is as follows: The target registration error represents the difference of the same registration point in two images. The calculation method of the target registration error is to calculate the Euclidean distance between the original point and the registered point, and then calculate its mean and variance.

7. A multi-modal image registration system for unmanned aerial vehicles based on a dual-channel attention network, applying the method for multi-modal image registration of unmanned aerial vehicles based on a dual-channel attention network according to any one of claims 1 to 6, characterized in that, Including: Feature point detection module: used to extract feature points of visible light images and infrared images; Dual-channel attention registration network module: used to model the dependence relationship of the spatial dimension and the channel dimension of the feature points of the visible light image and the infrared image extracted by the feature point detection module through the position attention model and the channel attention model, and perform rough matching on the outputs of the position attention model and the channel attention model to enhance the feature representation; The position attention model is used to determine the spatial dependence relationship between any two positions, and the channel attention model is used to capture the channel dependence relationship between any two channel mappings and update it using the weights of all channel mappings; Adaptive matching module: perform optimal value screening and homography matrix solution on the roughly matched feature points to select the features with the best matching results, and register the images according to the homography matrix; the optimal value screening scores the confidence of the feature points obtained by rough matching to evaluate the accuracy of feature point matching, and performs more refined matching to obtain an accurate homography matrix; Spatial transformation network module, which uses the spatial transformation network to map the image to be registered into the result image.

Citation Information

Patent Citations

  • Feature point matching method of cross-spectrum image

    CN116051872A

  • Multi-source image registration method based on deep learning

    CN119251269A

  • Coarse-to-fine heterologous image matching method based on edge guidance

    WO2024148969A1