Infrared image and visible light image local feature matching method based on attention map neural network

Through the method based on attention map neural network, a matching model between infrared images and visible image features is established, which solves the problem of insufficient matching accuracy of infrared and visible image features in the prior art, and realizes high-precision feature matching and image fusion.

CN119992137APending Publication Date: 2025-05-13SUZHOU ZHIZHEN WEISHI OPTOELECTRONICS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411607092.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively match the feature point information in infrared images and visible light images, resulting in insufficient perception and robustness in complex environments.

Method used

Using an attention graph neural network method, a matching model between infrared image and visible image features is established by constructing a feature descriptor enhancement module, feature aggregation module and optimal matching module, a matching model between infrared images and visible image features is dynamically adjusted, the weight of connections between nodes is automatically learned, and the key feature points are performed.

Benefits of technology

The accuracy of matching infrared features and visible light features is improved, and high-precision matching between infrared images and visible light images is achieved, providing support for subsequent image alignment and image fusion tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992137A_ABST
    Figure CN119992137A_ABST
Patent Text Reader

Abstract

The invention relates to a local feature matching method for an infrared image and a visible light image based on an attention map neural network, and the method comprises the steps: firstly, carrying out the calculation of an image through a local feature extractor, so as to extract feature points and descriptor information of the feature points; secondly, enhancing the feature descriptors by using a feature descriptor enhancement module; thirdly, performing information aggregation on the local feature descriptors by using an attention module; and finally, performing matching by using an optimal matching layer based on a nearest neighbor matching strategy. According to the method, the feature matching method for local feature matching of the infrared image and the visible light image based on the attention mechanism is established, the structure is simple, precision is high, and the method is suitable for feature matching work of the infrared image and the visible light image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a local feature matching method of an infrared image and a visible light image based on an attention graph neural network, belongs to the field of image processing, relates to feature matching, and particularly relates to a feature matching method based on an attention graph neural network, which can be used to match feature point information in infrared images and visible light images. Background Art

[0002] With the development of sensor technology, the fusion of information from different sensors has become increasingly important in the field of image processing. Due to the different hardware structures and imaging principles of different types of sensors, different sensors can capture different information. At present, there is no sensor that can capture all the information in the same scene. For example, visible light sensors mainly capture the reflection information of light on the surface of an object, so they can capture the texture details of the surface of the object, but once affected by extreme weather, occlusion and lighting, a lot of information will be lost. Infrared sensors capture the thermal radiation information emitted by objects, so they can efficiently highlight targets under harsh conditions, but their ability to capture texture details is insufficient.

[0003] In view of the advantages and disadvantages of visible light and infrared imaging, image fusion technology has emerged as a key strategy for optimizing computer vision applications. By integrating the rich details of visible light images and the all-weather adaptability of infrared images, the fused image not only retains the key information of the original image, but also significantly improves the perception ability and robustness in complex environments, thus providing a more reliable data foundation for tasks such as target detection, classification, tracking and scene understanding.

[0004] The attention-based graph neural network (GNN) has shown great potential in processing the local feature matching problem between infrared images and visible light images. The core of this method is that it can identify and establish the correspondence between feature points in the two modal images, which is a key step in achieving image fusion and calibration.

[0005] In specific operations, the attention graph neural network first extracts local feature points from infrared images and visible light images respectively. These feature points usually carry rich local structural information, such as edges, corners or texture patterns. Next, GNN constructs a graph model to represent these feature points and their relationships with each other, where each node represents a feature point and the edge represents the similarity or correlation between feature points.

[0006] The attention mechanism is introduced into the graph neural network to dynamically adjust the weights of the connections between nodes, which means that the network can automatically learn which feature points are most critical to the matching process and give them higher weights. This adaptive weight distribution helps to suppress the influence of noise and irrelevant information, thereby improving the accuracy and robustness of matching.

[0007] Once the correspondence between the feature points in the infrared image and the visible light image is established, the next step is to estimate the homography matrix between the two images. The homography matrix describes the geometric transformation between the two images, including rotation, translation, scaling, and perspective distortion. The matrix can be solved by minimizing the reprojection error, that is, finding the optimal parameters to align the matching feature points as much as possible in the transformed image space.

[0008] With the homography matrix, the pixels in one image can be mapped to the corresponding positions in the other image, thereby achieving strict alignment of the two images. This process is crucial for subsequent image fusion, ensuring that information from different modalities can be accurately integrated to generate a consistent image output that contains both the details of the visible light image and the thermal characteristics of the infrared image, thereby improving the performance and applicability of the computer vision system. Summary of the invention

[0009] The present invention relates to a local feature matching method of infrared image and visible light image based on attention graph neural network.

[0010] A matching model between infrared image features and visible light image features was established.

[0011] The technical solution of the present invention is a local feature matching of infrared images and visible light images based on an attention graph neural network, and the implementation steps are as follows:

[0012] (1) constructing a neural network, wherein the neural network is composed of a feature descriptor enhancement module, a feature aggregation module, and an optimal matching module connected in sequence; the feature enhancement module is used to enhance key features, the feature aggregation module is used to update feature descriptor information, and the optimal matching module is used to match the updated descriptor;

[0013] (2) constructing a local feature matching dataset of infrared images and visible light images, wherein each pair of images in the feature matching dataset of infrared images and visible light images is annotated, and the annotation is to annotate the corresponding relationship between feature points in the two images;

[0014] (3) training the neural network;

[0015] (4) Perform feature matching on the images to be matched. Input the matching image pairs into the trained neural network and output the matching relationship between the feature points of the two images.

[0016] In step (1), the feature points of the infrared image A and the visible light image B are and its descriptor Obtained through existing local feature extraction algorithms, popular local feature extraction algorithms such as SIFT and SuperPoint can be used. Encode feature points and associate them with corresponding descriptors, and initialize the state for:

[0017]

[0018] Where I∈{A,B}, the descriptor and feature points The dimensions are all three-dimensional. The first dimension represents the number of batches, and the second dimension represents the number of feature points in the image. The third dimension represents the dimension of the local feature descriptor corresponding to each feature point, which is usually set to 256. The third dimension indicates that the coordinate dimension of the feature point is 2 and the shape is converted to the same as The final output is Dimensions and same.

[0019] The feature enhancement module in step (1) is composed of a focused linear self-attention module and a focused cross-attention module. The input data is three-dimensional, the first dimension represents the batch size, the second dimension represents the number of feature points in the image, and the third dimension represents the dimension of the local feature descriptor corresponding to each feature point, which is usually set to 256. The output size is the same as the input size.

[0020] The focused linear self-attention module is used to aggregate features within the image, and its calculation method is:

[0021] Q i =W Q x i ,

[0022] K i =W K x i ,

[0023] V i =W V x i ,

[0024]

[0025] x i =Sim(Q i ,k i )V i ,

[0026] After the focused linear self-attention module, there is a focused cross-attention module to associate information between images, which is calculated as:

[0027] QK i =W QK x i ,

[0028] V i =W V x i ,

[0029] QK j =W QK x j ,

[0030] V j =W V x j ,

[0031]

[0032] x i =Sim(QK i ,QK j )V i ,

[0033] x j =Sim(QK i ,QK j ) T V i ,

[0034] in

[0035] In step (1), the feature aggregation module consists of a self-attention module and a cross-attention module. The input data is three-dimensional. The first dimension represents the batch size, the second dimension represents the number of feature points in the image, and the third dimension represents the dimension of the local feature descriptor corresponding to each feature point. It is usually set to 256. The output size is the same as the input size. The formula is as follows:

[0036] x i =SelfAttention(x i ),

[0037] x i ,x j =CrossAttention(xi ,x j ).

[0038] In step (1), the optimal matching layer is composed of a matching matrix P. The calculation formula of P is as follows:

[0039] P ij =Linear(x i )Linear(x j ) T ,

[0040] Where Linear(·) is a linear transformation.

[0041] In step (2), the correspondence between feature points in each pair of infrared and visible light images to be matched needs to be clarified. The image pairs used are all collected by infrared and visible light binocular cameras to ensure the correspondence between feature points between the image pairs.

[0042] In the step (3), the feature points and their descriptor training set are input into the neural network, and the weights of the feature matching network are updated using the gradient descent method until the loss drops below 0.3, thereby obtaining a trained neural network.

[0043] The loss function defined in step (3) is:

[0044] Loss=Loss P +Loss pos

[0045] Among them, Loss P To evaluate the loss of matching accuracy, Loss pos is the loss of matching points.

[0046] Loss P The formula is as follows:

[0047]

[0048] Among them, P i,j is the predicted matching matrix, and M is the matching truth matrix.

[0049] Loss pos The formula is as follows:

[0050]

[0051] in, represents the set of unmatched points in image A, Represents the set of unmatched points in image B. σ i Indicates the matching score of the feature point.

[0052] The advantages of the present invention are: a local feature positive matching model between infrared images and visible light spectroscopic images is established, the accuracy of matching between infrared features and visible light features is improved, high-precision matching of infrared images and visible light images can be achieved, and support is provided for subsequent image alignment and image fusion tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0054] Figure 1 A computational flow chart of the local feature matching method for infrared images and visible light images based on the attention graph neural network provided in this application.

[0055] Figure 2 Model structure diagram of the local feature matching method of infrared images and visible light images based on attention graph neural network provided in this application.

[0056] Figure 3 This is a graph of matching results of the local feature matching method for infrared images and visible light images based on the attention graph neural network provided in this application. DETAILED DESCRIPTION

[0057] The specific implementation of the present invention is described below in conjunction with the accompanying drawings;

[0058] Figure 1 The calculation flow chart of the present invention is shown, and the specific implementation steps of the present invention are as follows:

[0059] (1) constructing a neural network, wherein the neural network is composed of a feature descriptor enhancement module, a feature aggregation module, and an optimal matching module connected in sequence; the feature enhancement module is used to enhance key features, the feature aggregation module is used to update feature descriptor information, and the optimal matching module is used to match the updated descriptor;

[0060] (2) constructing a local feature matching dataset of infrared images and visible light images, wherein each pair of images in the feature matching dataset of infrared images and visible light images is annotated, and the annotation is to annotate the corresponding relationship between feature points in the two images;

[0061] (3) training the neural network;

[0062] (4) Perform feature matching on the images to be matched. Input the matching image pairs into the trained neural network and output the matching relationship between the feature points of the two images.

[0063] Figure 2 The model structure diagram of the entire neural network is shown below:

[0064] Feature points of infrared image A and visible light image B and its descriptor Obtained through existing local feature extraction algorithms, popular local feature extraction algorithms such as SIFT and SuperPoint can be used. Encode feature points and associate them with corresponding descriptors, and initialize the state for:

[0065]

[0066] Where I∈{A,B}, the descriptor and feature points The dimensions are all three-dimensional. The first dimension represents the number of batches, and the second dimension represents the number of feature points in the image. The third dimension represents the dimension of the local feature descriptor corresponding to each feature point, which is usually set to 256. The third dimension indicates that the coordinate dimension of the feature point is 2 and the shape is converted to the same as The final output is Dimensions and same.

[0067] The feature enhancement module consists of a focused linear self-attention module and a focused cross-attention module. The input data is three-dimensional. The first dimension represents the batch size, the second dimension represents the number of feature points in the image, and the third dimension represents the dimension of the local feature descriptor corresponding to each feature point, which is usually set to 256. The output size is the same as the input size.

[0068] The focused linear self-attention module is used to aggregate features within the image, and its calculation method is:

[0069] Q i =W Q x i ,

[0070] K i =W K x i ,

[0071] V i =W V x i ,

[0072]

[0073] x i=Sim(Q i ,k i )V i ,

[0074] After the focused linear self-attention module, there is a focused cross-attention module to associate information between images, which is calculated as:

[0075] QK i =W QK x i ,

[0076] V i =W V x i ,

[0077] QK j =W QK x j ,

[0078] V j =W V x j ,

[0079]

[0080] x i =Sim(QK i ,QK j )V i ,

[0081] x j =Sim(QK i ,QK j ) T V j ,

[0082] in

[0083] The feature aggregation module consists of a self-attention module and a cross-attention module. The input data is three-dimensional. The first dimension represents the batch size, the second dimension represents the number of feature points in the image, and the third dimension represents the dimension of the local feature descriptor corresponding to each feature point. It is usually set to 256. The output size is the same as the input size. The formula is as follows:

[0084] x i =SelfAttention(x i ),

[0085] x i ,x j =CrossAttention(x i ,x j ).

[0086] In step (1), the optimal matching layer is composed of a matching matrix P. The calculation formula of P is as follows:

[0087] P ij =Linear(x i )Linear(x j ) T ,

[0088] Where Linear(·) is a linear transformation.

[0089] Then add the output of the feature aggregation module to the result of the feature enhancement module to get the updated state The whole module is repeated l times in total.

[0090] The feature points and their descriptor training set are input into the neural network, and the weights of the feature matching network are updated using the gradient descent method until the loss drops below 0.3, thereby obtaining a trained neural network.

[0091] The loss function defined in is:

[0092] Loss=Loss P +Loss pos

[0093] Among them, Loss P To evaluate the loss of matching accuracy, Loss pos is the loss of matching points.

[0094] Loss P The formula is as follows:

[0095]

[0096] Among them, P i,j is the predicted matching matrix, and M is the matching truth matrix.

[0097] Loss pos The formula is as follows:

[0098]

[0099] in, represents the set of unmatched points in image A, Represents the set of unmatched points in image B. σ i Indicates the matching score of the feature point.

[0100] The contents not described in detail in the specification of the present invention belong to the common knowledge of the professionals in this field.

Claims

1. A local feature matching method for infrared images and visible light images based on attention graph neural network, characterized in that: The implementation steps are as follows: (1) constructing a neural network, wherein the neural network is composed of a feature descriptor enhancement module, a feature aggregation module, and an optimal matching module connected in sequence; the feature enhancement module is used to enhance key features, the feature aggregation module is used to update feature descriptor information, and the optimal matching module is used to match the updated descriptor; (2) constructing a local feature matching dataset of infrared images and visible light images, wherein each pair of images in the infrared image and visible light image feature matching dataset is annotated, and the annotation is to annotate the corresponding relationship between feature points in the two images; (3) training the neural network; (4) Perform feature matching on the images to be matched. Input the matching image pairs into the trained neural network and output the matching relationship between the feature points of the two images.

2. The method for matching local features of infrared images and visible light images based on attention graph neural network according to claim 1, characterized in that: In step (1), the feature points of the infrared image A and the visible light image B are and its descriptor Obtained through existing local feature extraction algorithms, popular local feature extraction algorithms such as SIFT and SuperPoint can be used. Encode feature points and associate them with corresponding descriptors, and initialize the state for: Where I∈{A,B}, the descriptor and feature points The dimensions are all three-dimensional. The first dimension represents the number of batches, and the second dimension represents the number of feature points in the image. The third dimension represents the dimension of the local feature descriptor corresponding to each feature point, which is usually set to 256. The third dimension indicates that the coordinate dimension of the feature point is 2 and the shape is converted to the same as The final output is Dimensions and same.

3. The method for matching local features of infrared images and visible light images based on attention graph neural network according to claim 1, characterized in that: The feature enhancement module in step (1) is composed of a focused linear self-attention module and a focused cross-attention module. The input data is three-dimensional, the first dimension represents the batch size, the second dimension represents the number of feature points in the image, and the third dimension represents the dimension of the local feature descriptor corresponding to each feature point, which is usually set to 256. The output size is the same as the input size. The focused linear self-attention module is used to aggregate features within the image, and its calculation method is: Q i =W Q x i , K i =W K x i , V i =W V x i , x i =Sim(Q i ,K i )V i , After the focused linear self-attention module, there is a focused cross-attention module to associate information between images, which is calculated as: QK i =W QK x i , V i =W V x i , QK j =W QK x j , V j =W V x j , x i =Sim(QK i ,QK j )V i , x j =Sim(QK i ,QK j ) T V j , in 4. The method for matching local features of infrared images and visible light images based on attention graph neural network according to claim 1, characterized in that: The feature aggregation module in step (1) is composed of a self-attention module and a cross-attention module. The input data is three-dimensional. The first dimension represents the batch size, the second dimension represents the number of feature points in the image, and the third dimension represents the dimension of the local feature descriptor corresponding to each feature point. It is usually set to 256. The output size is the same as the input size. The formula is as follows: x i =SelfAttention(x i ), x i ,x j =CrossAttention(x i ,x j )。 5. The method for matching local features of infrared images and visible light images based on attention graph neural network according to claim 1, characterized in that: In step (1), the optimal matching layer is composed of a matching matrix P. The calculation formula of P is as follows: P ij =Linear(x i )Linear(x j ) T Where Linear(·) is a linear transformation.

6. The method for matching local features of infrared images and visible light images based on attention graph neural network according to claim 1, characterized in that: In step (2), each pair of infrared and visible light images to be matched needs to clarify the correspondence between the feature points in the image. The image pairs used are all collected by infrared and visible light binocular cameras to ensure the correspondence between the feature points between the image pairs.

7. The method for matching local features of infrared images and visible light images based on attention graph neural network according to claim 1, characterized in that: In the step (3), the feature points and their descriptor training set are input into the neural network, and the weights of the feature matching network are updated using the gradient descent method until the loss drops below 0.3, thereby obtaining a trained neural network.

8. The method for matching local features of infrared images and visible light images based on attention graph neural network according to claim 1, characterized in that: The loss function defined in step (3) is: Loss=Loss P +Loss pos Among them, Loss P To evaluate the loss of matching accuracy, Loss pos is the loss of matching points. Loss P The formula is as follows: Among them, P i,j is the predicted matching matrix, and M is the matching truth matrix. Loss pos The formula is as follows: in, represents the set of unmatched points in image A, Represents the set of unmatched points in image B. σ i Indicates the matching score of the feature point.