Visible light and infrared image registration method based on super-resolution
By using three-dimensional radiation transmission model and super-resolution technology in visible light and infrared image registration, and integrating Transformer and SuperGlue networks for feature matching, the problems of insufficient training data and low registration accuracy in the prior art are solved, and higher registration accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202510536506.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art has problems in the registration of visible light and infrared images, the inability to process feature matching of images with different resolutions, and the low registration accuracy.
Generate simulated data through a three-dimensional radiation transmission model, combine super-resolution and pyramid technology to expand the multi-resolution image dataset, and integrate Transformer and SuperGlue networks for rough and fine matching to improve registration accuracy.
The training data is enriched, the model's adaptability to different scenarios and conditions is enhanced, and the accuracy and robustness of visible and infrared image registration are improved.
Smart Images

Figure CN120047506A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image registration, and particularly relates to a visible light and infrared image registration method based on super-resolution. Background Art
[0002] Visible light and infrared images are of great value in various applications, including target monitoring, disaster management, and environmental monitoring. Due to the differences in imaging devices and the influence of the natural environment, directly registering these two types of images is challenging. Efficient and accurate registration of visible light and infrared images can provide a good foundation for image applications.
[0003] Visible light and infrared image registration methods at home and abroad include feature-based registration methods, region-based registration methods, and deep learning-based methods. Feature-based registration methods mainly use feature points (such as corner points, edges) of images for registration. They have high computational efficiency and are suitable for real-time processing. However, for low-resolution infrared images, feature point extraction may be unstable, and it is difficult to handle the spectral differences between images.
[0004] Region-based registration methods achieve registration by matching the region information of images (such as template matching). They are suitable for cases where the differences between images are large. They are insensitive to the spectral differences of images, pay more attention to the overall regional features, and can also work effectively in images without significant feature points. However, their computational complexity is high, the matching takes time, and complex preprocessing is required.
[0005] Deep learning-based methods mainly use deep learning technologies such as convolutional neural networks (CNNs) for feature extraction and registration. They can automatically learn cross-modal feature representations, improve the robustness and accuracy of registration, can automatically learn complex cross-modal features, and improve the accuracy of registration. They can handle images under different scenarios and conditions. However, they rely on a large amount of training data, and the training process is time-consuming. Summary of the Invention
[0006] In view of this, the present invention aims to provide a visible light and infrared image registration method based on super-resolution. On the one hand, a simulation data set under various conditions is constructed with the help of a three-dimensional radiation transfer model, and at the same time, super-resolution and pyramid techniques are used to expand the multi-resolution image data set, effectively enhancing the expanded data set to make up for the lack of actual shooting data. On the other hand, a coarse matching feature point network and a false matching point detection network based on a deep learning network structure are integrated to improve the registration accuracy of visible light and infrared images with resolution differences and spectral differences.
[0007] To achieve the above object, the technical solution of the present invention is realized as follows: The present invention provides a visible light and infrared image registration method based on super-resolution, including: Obtain the measured visible light image and infrared image, as well as the simulated visible light image and infrared image, and form a set of initial samples with the visible light image and infrared image containing the same area; Use the resolution adjustment method to adjust the resolution of the visible light image and infrared image in each set of initial samples, generate multiple sets of visible light images and infrared images with different resolutions, form multiple sets of extended samples with different resolutions, and combine the initial samples and extended samples to form a training sample set; Use the training sample set to train the image registration model. The image registration model includes: a coarse matching feature extraction network constructed based on the Transformer network, which is used to extract and match features of the input visible light image and infrared image, and output a pair of coarse matching features; use the SuperGlue network to evaluate the pair of coarse matching features, filter out the pairs of coarse matching features with evaluation scores lower than the set threshold, and output the remaining pairs of coarse matching features as pairs of fine matching features; Input the visible light image and infrared image to be matched into the trained image registration model, generate an affine transformation matrix according to the pairs of fine matching features output by the trained image registration model, and perform registration transformation on the visible light image and infrared image to be matched using the affine transformation matrix.
[0008] Preferably, the measured visible light image and infrared image are obtained by: aerial photography by a drone.
[0009] Preferably, the simulated visible light image and infrared image are obtained by: Simulate and construct visible light images and infrared images through a three-dimensional radiative transfer model.
[0010] Preferably, use the three-dimensional radiative transfer model to generate multiple sets of visible light images and infrared images with different shooting angles, different resolutions, and / or different spectral attributes.
[0011] Preferably, use a super-resolution convolutional neural network to perform super-resolution reconstruction on the visible light image and infrared image in each set of initial samples to generate multiple sets of high-resolution extended samples; Use the pyramid technology in the GDAL library to reduce the resolution of the visible light image and infrared image in each set of initial samples to generate multiple sets of low-resolution extended samples.
[0012] Preferably, the super-resolution convolutional neural network includes a feature extraction layer, a non-linear mapping layer, and a reconstruction layer. The feature extraction layer extracts preliminary features from the low-resolution initial sample image using convolutional operations; the non-linear mapping layer is used to map the preliminary features to a high-dimensional feature space; the reconstruction layer uses convolutional operations to reconstruct the features in the high-dimensional feature space to generate a high-resolution extended sample.
[0013] Preferably, the pyramid technique uses the bilinear interpolation method to reduce the resolution. The calculation formula of bilinear interpolation is: ; where, represents the pixel value before interpolation, represents the pixel value after interpolation x and y represent the pixel coordinates.
[0014] Preferably, the coarse matching feature extraction network includes a feature extractor, a Transformer module, and a feature matching layer. The feature extractor extracts features from the visible light image and the infrared image based on the ResNET network; the Transformer module performs interactive matching on the features output by the feature extractor based on the self-attention mechanism; the feature matching layer is used to match the features of the visible light image and the infrared image to form a pair of coarse matching features.
[0015] Preferably, the loss function of the coarse matching feature extraction network Loss is: ; where, N represents the total number of pairs of coarse matching features output by the coarse matching feature extraction network, represents the distance between the i-th pair of features, represents the actual matching label of the i-th pair of features, 1 represents that the i-th pair of features are actually matched, 0 represents that the i-th pair of features are actually not matched, and m is a set threshold.
[0016] Preferably, the SuperGlue network generates a feature map of the visible light image and a feature map of the infrared image according to the pair of coarse matching features output by the coarse matching feature extraction network, and updates the feature map of the visible light image and the feature map of the infrared image through the self-attention mechanism and the cross-attention mechanism. After the update, the evaluation score of each feature in the feature map of the visible light image and the corresponding feature in the feature map of the infrared image is calculated.
[0017] Compared with the prior art, the present invention can achieve the following beneficial effects: In the process of obtaining training data for the present invention, based on the actual captured data, simulated and emulated data for training is supplemented through the three-dimensional radiative transfer model DART. In addition, multiple sets of extended samples with different resolutions are generated through a super-resolution convolutional neural network and pyramid technology, greatly enriching the data samples. This not only makes up for the deficiencies of the actual captured data but also enhances the adaptability of the model to different scenarios, different angles, and different meteorological conditions, ensuring that the model can maintain high robustness in complex scenarios.
[0018] The image registration model of the present invention integrates the Transformer deep learning network and the SuperGlue network. Based on the Transformer network, feature extraction and rough matching are performed on the input visible light image and infrared image, and then the SuperGlue network evaluates the rough-matched feature pairs, enabling fine-grained feature matching, effectively removing mis-matched feature pairs, improving the accuracy of the matching result, and enhancing the registration accuracy of visible light and infrared images with resolution differences and spectral differences. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 is a detailed working flowchart of the super-resolution-based visible light and infrared image registration method provided by an embodiment of the present invention; Figure 2 is a schematic diagram of the data processing process of the super-resolution-based visible light and infrared image registration method provided by an embodiment of the present invention; Figure 3 is a schematic diagram of the registration result of the actually measured visible light image and infrared image of a drone provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention. Similar elements in different embodiments are labeled with related similar element numbers. In the following embodiments, many details are described to enable a better understanding of the present invention. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other elements, materials, or methods. In some cases, some operations related to the present invention are not shown or described in the specification, which is to avoid the core part of the present invention being overwhelmed by excessive description. For those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations based on the description in the specification and the general technical knowledge in the field.
[0021] It should be noted that, without conflict, the embodiments and features in the embodiments of the present invention can be combined with each other to form various embodiments. At the same time, the steps or actions in the method description can also be reordered or adjusted in an obvious manner by those skilled in the art. Therefore, the various sequences in the specification and drawings are only for clearly describing a certain embodiment and do not mean that they are the necessary sequences, unless it is stated that a certain sequence must be followed.
[0022] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. In addition, terms such as "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise stated, the meaning of "a plurality" is two or more.
[0023] In the description of the present invention, it should be noted that, unless otherwise clearly specified and defined, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be a direct connection or an indirect connection through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0024] The present invention will be described in detail below with reference to the drawings and in conjunction with embodiments.
[0025] Please refer to Figure 1 and Figure 2 , in an embodiment of the present invention, a visible light and infrared image registration method based on super-resolution is provided, which is mainly used to solve the problems of insufficient training data volume in image matching by existing convolutional neural network models, inability to perform feature matching of images with different resolutions, and low image registration accuracy. The image registration method of the embodiment of the present invention includes: S1: First, construct a data set of visible light images and infrared images in different scenarios. Specifically, in the embodiment of the present invention, a part of measured visible light images and infrared images are obtained through drone aerial photography, and the aerial photography data of these measured visible light images and infrared images are composed into a measured data set. Since the aerial photography data is limited, and the aerial photography data is usually obtained by the same device at the same time and in the same area, its data has great limitations. Therefore, using it for training the image registration model will result in poor adaptability of the model to different scenarios, different angles, different meteorological conditions, etc. When registering visible light and infrared images with different resolutions and / or different spectra, the registration accuracy is low and it is difficult to meet the actual application requirements.
[0026] For the above reasons, the embodiments of the present invention design the following method to expand the training data volume. Specifically: on the basis of the measured data set, a simulated data set is introduced. The radiation transfer process in complex scenarios is simulated through a three-dimensional radiation transfer model to generate simulated visible light images and infrared images. In the embodiments of the present invention, the specific three-dimensional radiation transfer model adopts the DART model. First, a variety of different scenarios are constructed through the three-dimensional radiation transfer model DART, including but not limited to vegetation scenarios, forest scenarios, and urban scenarios, etc. The constructed scenarios should include different ground object types and structures as much as possible to represent various situations in actual applications. And different spectral attributes are constructed for different components (such as leaves, buildings, roads, animals, people, etc.) in each scenario, so that each component has different radiation characteristics in the visible light and thermal infrared bands to simulate the spectral differences in real situations. In addition, different meteorological conditions such as sunny days and cloudy days, as well as different shooting angles, are added to the scenarios to generate a large number of visible light images and infrared images at multiple resolutions, increasing the diversity of the data set as much as possible and improving the adaptability of the model to different environmental conditions.
[0027] In the above-mentioned measured data set and simulated data set, for any visible light image and infrared image, the visible light image and the infrared image containing the same area can form a group of initial samples.
[0028] To further expand the sample data volume, the embodiments of the present invention also adjust the resolution of the visible light images and infrared images in each group of initial samples through a resolution adjustment method to generate multiple groups of visible light images and infrared images with different resolutions, forming multiple groups of extended samples with different resolutions. Specifically, a super-resolution convolutional neural network (SCRNN) is used to perform super-resolution reconstruction on the visible light images and infrared images in each group of initial samples to generate multiple groups of high-resolution extended samples. The super-resolution convolutional neural network mainly includes a feature extraction layer, a non-linear mapping layer, and a reconstruction layer. Among them, the feature extraction layer uses convolutional operations to extract preliminary features from low-resolution images; the non-linear mapping layer maps the preliminarily extracted features to a high-dimensional feature space to enhance the expression ability of the features; the reconstruction layer maps the high-dimensional features in the high-dimensional feature space back to the image space through convolutional operations to reconstruct and generate multiple different high-resolution extended samples. At the same time, the pyramid technology in the GDAL library is also used to reduce the resolution of the visible light images and infrared images in each group of initial samples to generate multiple groups of low-resolution extended samples. The pyramid technology mainly generates lower-resolution images through a downsampling method. Specifically, the bilinear interpolation method is used for downsampling, and this method can smoothly reduce the image resolution and reduce image distortion. The calculation formula of bilinear interpolation is: ; Among them, represents the pixel value before interpolation, represents the pixel value after interpolationx and y represent pixel coordinates.
[0029] For visible light images and infrared images, the above bilinear interpolation calculation formula can be used to perform multiple iterative processes on the images to generate multiple low-resolution images.
[0030] In the embodiments of the present invention, through super-resolution and downsampling techniques, images with multiple resolutions such as 0.01m, 0.05m, 0.1m, 0.5m, 1m, 5m, 10m, and 50m are generated, thereby providing rich scale information for subsequent feature matching.
[0031] The initial samples and the extended samples obtained by the above expansion process are combined to form a training sample set.
[0032] S2: Construct an image registration model, and use the training sample set to train the image registration model. Input the samples in the training sample set into the image registration model for training. Each sample consists of a visible light image and an infrared image. The image registration model mainly includes two processes: rough matching and fine matching. Specifically, the rough matching process is mainly implemented by a rough matching feature extraction network constructed based on the Transformer network. The rough matching feature extraction network includes a feature extractor, a Transformer module, and a feature matching layer. Among them, the feature extractor mainly includes a ResNET network. The ResNET network is used to extract features from the visible light and infrared images respectively. After the feature extraction is completed, the feature extractor outputs the features. The Transformer module receives the features output by the feature extractor. The Transformer module uses the self-attention mechanism for interactive matching to capture the global relationship between the image features. The self-attention mechanism can dynamically adjust the feature weights according to the correlation between the feature points, enhance the expression of important features, and suppress the influence of unimportant features. Finally, the relationship between the features obtained by the Transformer module is estimated and paired through the feature matching layer to obtain the rough matching result of the image, forming a rough matching feature pair. The information of the rough matching feature pair includes the coordinate positions of the rough matching feature pair on the two images, and its corresponding feature descriptor. The feature descriptor is the feature vector of each feature, which is used to describe the local features of the point.
[0033] The rough matching feature extraction network constructs a loss function using contrastive loss. The loss function Loss is: ; where N represents the total number of rough matching feature pairs output by the rough matching feature extraction network, represents the distance between the i-th pair of features, Denote the actual matching label of the i-th pair of features. 1 indicates that the i-th pair of features actually match, 0 indicates that the i-th pair of features actually do not match, and m is a preset threshold.
[0034] To improve the accuracy of image registration, an embodiment of the present invention also designs a SuperGlue network at the backend of the coarse matching feature extraction network, and uses the SuperGlue network to detect and filter the mismatched features in the coarse matching feature pairs to obtain a more accurate fine-grained matching result. Specifically, the SuperGlue network receives the coarse matching feature pairs output by the coarse matching feature extraction network and identifies the coordinate positions and feature descriptors of the coarse matching feature pairs. The SuperGlue network constructs a graph G1 from the feature positions and feature descriptors of the visible light image, and constructs a graph G2 from the feature positions and feature descriptors of the infrared image. Each node in the graph represents a feature, and the edges between the nodes represent the relationships between the features. Graph G1 and graph G2 are respectively called the feature graph of the visible light image and the feature graph of the infrared image. Inside each graph (G1 and G2), through the self-attention mechanism, each node (feature) interacts with its neighboring nodes for information, updates its own feature representation, so that each feature not only contains its own feature information, but also fuses the surrounding feature information, enhancing the feature expression ability. Between graph G1 and graph G2, through the cross-attention mechanism, the features in the two graphs interact with each other, so that the features in G1 and G2 can refer to each other, further improving the accuracy of matching. After updating the feature graphs of the visible light image and the infrared image through the self-attention mechanism and the cross-attention mechanism, the evaluation score of each feature in the feature graph of the visible light image and the corresponding feature in the feature graph of the infrared image is calculated by the following formula: ; where is the matching score between the i-th feature in graph G1 and the j-th feature in graph G2, is the feature representation of the i-th feature in graph G1 after being updated by the self-attention and cross-attention mechanisms, represents the feature representation of the j-th feature in graph G2 after being updated by the self-attention and cross-attention mechanisms.
[0035] After obtaining the evaluation score, according to the set score threshold , filter out the low-confidence coarse matching feature pairs with evaluation scores lower than the set threshold, and output the remaining coarse matching features as fine matching feature pairs. The filtering discriminant is: ; where indicates whether the i-th feature in G1 matches the j-th feature in graph G2, is the exponential form of the matching score between the i-th feature and the j-th feature, It is the exponential sum of the matching scores between the $i$-th feature point in graph $G1$ and all feature points in graph $G2$.
[0036] S3: Input the visible light image and the infrared image to be matched into the trained image registration model. The image registration model registers the visible light image and the infrared image in the manner in S2, and outputs the accurately matched feature pairs. Use the least squares method to calculate the affine transformation matrix of the matched feature pairs, and apply the affine transformation matrix to the visible light image and the infrared image to be matched. The affine transformation matrix mainly includes operations such as image position, rotation, and scaling, and aligns the visible light image and the infrared image through registration transformation.
[0037] The visible light and infrared image registration method based on super-resolution given in the embodiments of the present invention has been actually tested. For details, please refer to Figure 3 , which is the registration result of the visible light image and the thermal infrared image of the UAV flight. The image to be registered is the visible light image, and the reference image is the infrared image. The registration result is obtained through the above method. It can be observed from the figure that the model has extremely strong detail retention and feature extraction capabilities, significantly improves the registration accuracy, performs excellently in the feature alignment of visible light and infrared images, can more accurately identify and match cross-modal key points, and shows strong adaptability when dealing with complex terrains and environmental changes.
[0038] In summary, the above are only the preferred embodiments of this specification, and are not intended to limit the protection scope of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this specification shall be included within the protection scope of this specification.
[0039] The system, device, module or unit illustrated in the above one or more embodiments can be specifically implemented by a computer chip or entity, or by a product with a certain function. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0040] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, the element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, commodity or device including the said element.
[0041] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content.
[0042] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A super-resolution-based visible light and infrared image registration method, characterized in that: include: Obtaining measured visible light images and infrared images, as well as simulated visible light images and infrared images, and forming a group of initial samples from visible light images and infrared images covering the same area; Using a resolution adjustment method to adjust the resolution of visible light images and infrared images in each group of initial samples, generating multiple groups of visible light images and infrared images with different resolutions, forming multiple groups of extended samples with different resolutions, and forming the initial samples and extended samples into a training sample set; The image registration model is trained using the training sample set, and the image registration model includes: a coarse matching feature extraction network constructed based on a Transformer network, the coarse matching feature extraction network is used to extract and match the input visible light image and infrared image, and output a coarse matching feature pair; the coarse matching feature pair is evaluated using a SuperGlue network, the coarse matching feature pairs with evaluation scores lower than a set threshold are filtered out, and the remaining coarse matching features are output as fine matching feature pairs; The visible light image and infrared image to be matched are input into the trained image registration model, and an affine transformation matrix is generated according to the precise matching feature pairs output by the trained image registration model, and the affine transformation matrix is used to perform registration transformation on the visible light image and infrared image to be matched.
2. The super-resolution-based visible light and infrared image registration method according to claim 1, characterized in that: The measured visible light images and infrared images are obtained by aerial photography using a drone.
3. The super-resolution-based visible light and infrared image registration method according to claim 1, characterized in that: The simulated visible light image and infrared image are obtained as follows: Visible light images and infrared images are constructed through three-dimensional radiation transfer model simulation.
4. The super-resolution-based visible light and infrared image registration method according to claim 3, characterized in that: A three-dimensional radiation transfer model is used to generate multiple sets of visible light images and infrared images with different shooting angles, different resolutions and / or different spectral properties.
5. The super-resolution-based visible light and infrared image registration method according to claim 1, characterized in that: The super-resolution convolutional neural network is used to perform super-resolution reconstruction on the visible light image and infrared image in each group of initial samples to generate multiple groups of high-resolution extended samples; The pyramid technology in the GDAL library is used to reduce the resolution of the visible light image and infrared image in each group of initial samples to generate multiple groups of low-resolution extended samples.
6. The super-resolution-based visible light and infrared image registration method according to claim 5, characterized in that: The super-resolution convolutional neural network includes a feature extraction layer, a nonlinear mapping layer and a reconstruction layer. The feature extraction layer uses a convolution operation to extract preliminary features from a low-resolution initial sample image; the nonlinear mapping layer is used to map the preliminary features to a high-dimensional feature space; The reconstruction layer utilizes convolution operations to reconstruct features in a high-dimensional feature space to generate high-resolution extended samples.
7. The super-resolution-based visible light and infrared image registration method according to claim 5, characterized in that: The pyramid technology uses a bilinear interpolation method to reduce the resolution. The calculation formula of the bilinear interpolation is: ; in, represents the pixel value before difference, Represents the pixel value after difference x and y Represents pixel coordinates.
8. The super-resolution-based visible light and infrared image registration method according to claim 1, characterized in that: The coarse matching feature extraction network includes a feature extractor, a Transformer module and a feature matching layer. The feature extractor extracts features of visible light images and infrared images based on the ResNET network; the Transformer module interactively matches the features output by the feature extractor based on the self-attention mechanism; and the feature matching layer is used to match the features of the visible light image and the infrared image to form a coarse matching feature pair.
9. The super-resolution-based visible light and infrared image registration method according to claim 8, characterized in that: The loss function of the coarse matching feature extraction network is Loss for: ; Wherein, N represents the total number of coarse matching feature pairs output by the coarse matching feature extraction network, represents the distance between the i-th pair of features, represents the actual matching label of the i-th pair of features, 1 represents the i-th pair of features actually matches, 0 represents the i-th pair of features actually does not match, and m is the set threshold.
10. The super-resolution-based visible light and infrared image registration method according to claim 1, characterized in that: The SuperGlue network generates a feature map of the visible light image and a feature map of the infrared image according to the coarse matching feature pairs output by the coarse matching feature extraction network, and updates the feature map of the visible light image and the feature map of the infrared image through a self-attention mechanism and a cross-attention mechanism, and after the update, calculates the evaluation score of each feature in the feature map of the visible light image and the corresponding feature in the feature map of the infrared image.
Citation Information
Patent Citations
SAR image super-resolution method based on neural network
CN110163802A
Multi-scale sample generation method for deep learning
CN113435487A
Infrared and visible light image registration method in electric power inspection scene
CN113628261A
Multi-temporal remote sensing image registration method based on cross attention and deformable convolution
CN116664892A
Self-supervised visual SLAM method based on graph neural network
CN118781464A