A method and device for fusion of unmanned aerial vehicle search and rescue images based on feature matching
By adopting a feature matching-based image fusion method in drone search and rescue scenarios, combined with Markov random airports, Bayesian networks and twin networks for feature matching and reconstruction, the problem that traditional algorithms are difficult to accurately detect rescue targets in complex environments is solved, and a more efficient and accurate search and rescue effect is achieved.
Patent Information
- Application Number
- CN202411622851.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-11-14
AI Technical Summary
In complex search and rescue scenarios, traditional deep learning algorithms find it difficult to accurately detect rescue targets under occlusion, action deformation, etc., resulting in insufficient search and rescue efficiency and accuracy.
The drone search and rescue image fusion method based on feature matching is adopted, and the RGB image and infrared image is preprocessed, and converted into feature maps and heat maps are combined with Markov random airports, Bayesian networks and twin networks to perform regional feature matching, and feature reconstruction is carried out through spatial attention and channel attention, and finally an image with precise positioning and significant rescue target features are generated.
It improves the accuracy and efficiency of personnel search and rescue in complex environments, can more significantly display the characteristics of rescue targets, provide intuitive references, and helps to search more efficiently and accurately.
Smart Images

Figure CN119131371B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a method and device for fusion of unmanned aerial vehicle search and rescue images based on feature matching. Background Art
[0002] In complex search and rescue scenarios, drones need to use algorithms to automatically identify people in images to improve the efficiency of rescue work. In the application of image recognition and neural network models, traditional target recognition algorithms based on deep learning occupy an important position. Among them, YOLOv5, as an advanced real-time target detection algorithm, is favored for its high efficiency and accuracy. It usually consists of three parts: feature extraction layer, feature matching layer and prediction layer. The feature extraction layer is responsible for extracting useful feature information from the input image; the feature matching layer uses this feature information to match and identify the target object; the prediction layer outputs information such as the location, category and confidence of the target object based on the matching results. However, although deep learning algorithms such as YOLOv5 have achieved remarkable results in target recognition, they still have differences in network structure and parameter settings, which leads to different algorithm performance and applicable scenarios. In addition, deep learning algorithms usually require a large amount of data sets and precise annotations to train the model so that the model can learn useful features and perform accurate classification. However, in actual rescue missions, due to complex situations and tight time, it is often difficult to obtain sufficient quantity and quality of training data, which limits the application of deep learning algorithms in the field of search and rescue.
[0003] Feature matching is one of the most popular directions in the field of computer vision in recent years. Its goal is to establish accurate correspondences between different images, which is crucial for applications such as visual positioning and target tracking. However, in practical applications, it is still very challenging to achieve accurate point matching due to the presence of matching various noises (such as scale changes, viewpoint and illumination changes, repetitive patterns, and poor textures). In search and rescue missions, the rescued persons may not be able to remain still or be in an obvious position due to injuries, entrapment, or panic. In addition, obstacles such as ruins and vegetation may also block the line of sight or body parts of the rescued persons, making it difficult for traditional image recognition algorithms to accurately detect their presence. In this case, even with advanced deep learning algorithms, accurate target detection may not be possible due to the lack of sufficient feature information.
[0004] Therefore, how to further improve the accuracy and efficiency of personnel search and rescue in complex situations such as occlusion and motion deformation is an urgent problem to be solved. Summary of the invention
[0005] In view of the above, the purpose of the present invention is to provide a method and device for UAV search and rescue image fusion based on feature matching. Through an improved feature matching fusion framework, feature matching is performed on feature maps and heat maps and the feature matching problem is converted into an energy minimization problem. The feature maps and heat maps are feature reconstructed in combination with spatial attention and channel attention. Finally, the feature matching results and feature reconstruction results are fused to obtain a generated image with precise positioning and significant features of the rescue target. The generated image can be applied to scenarios such as UAV search and rescue, provide an intuitive reference for units involved in the search and rescue, and help to search for target personnel more efficiently and accurately.
[0006] In order to achieve the above-mentioned invention object, the technical solution provided by the present invention is as follows:
[0007] An embodiment of the present invention provides a method for fusion of unmanned aerial vehicle search and rescue images based on feature matching, comprising the following steps:
[0008] The acquired RGB image and infrared image are preprocessed to obtain feature map and thermal map respectively;
[0009] The feature map and heat map are converted into a feature map graphical model and a heat map graphical model respectively. In each graphical model, the image area is represented by a node, and the image area is divided into multiple regional scale levels according to the size of different image areas. The inclusion relationship between nodes in different levels and the adjacency relationship between nodes in the same level are constructed by directed edges and undirected edges respectively.
[0010] The undirected edges of the feature map graphic model and the heat map graphic model are updated by a Markov random field based on the image region, and the directed edges are updated by a Bayesian network based on the image region to obtain an updated feature map graphic model and a heat map graphic model, and the similarity between all matching regions between the two updated graphic models is calculated by using a twin network, and the best region matching result is obtained by minimizing the global energy based on the similarity between the matching regions, and the feature map and the heat map are fused based on the best region matching result to obtain a first fused image;
[0011] After extracting image features and thermal features from the feature map and the thermal map respectively, a reconstructed feature map and a reconstructed thermal map are obtained through spatial attention operation, channel attention operation and feature reconstruction, and the reconstructed feature map and the reconstructed thermal map are fused to obtain a second fused image;
[0012] The first fused image and the second fused image are combined into a final generated image for search and rescue.
[0013] Preferably, the preprocessing of the acquired RGB image and infrared image to obtain a feature map and a thermal map respectively includes:
[0014] Key frame images are captured from the video taken by the drone's onboard camera. The key frame images include RGB images and infrared images taken synchronously. The RGB images are segmented using the U-Net model or the SAM model to obtain feature maps, and the infrared images are processed using the non-uniformity correction method to obtain thermal maps.
[0015] Preferably, the image regions are represented by nodes and divided into multiple regional scale levels according to the sizes of different image regions, including:
[0016] The feature map or heat map is divided into regions and the regions are generated by graph completion through the SAM model. According to the size of these regions, they are divided into L levels of regional scale levels. Each regional scale level corresponds to a different image scale. In each graphic model, nodes are used to represent the image region, so that the nodes in each graphic model are also divided into L levels.
[0017] Preferably, the updating of undirected edges of the feature map graphical model and the heat map graphical model through a Markov random field based on the image region comprises:
[0018] All edges in the feature graph model and the heat map model are considered as undirected edges and converted into undirected graphs, and random variables are introduced for all nodes in the undirected graphs of the two graph models. To indicate the matching status of these nodes with the source nodes, the region matching is performed to update the undirected edges by maximizing the joint probability distribution on the Markov random field based on the image region. The calculation formula is:
[0019] ,
[0020] According to the Hammersley-Clifford theorem, the probability distribution defined by the Markov random field based on the image region belongs to the Boltzmann distribution and is the exponential of the negative energy function, that is, , so the region matching is expressed as energy minimization, and the calculation formula is:
[0021] ,
[0022] According to the structure of the undirected graph, the energy function Node Energy and edge energy Two parts, the calculation formula is:
[0023] ,
[0024] in, is the parameter of the balance term, is the index of all neighbor node pairs in the undirected graph A collection of and Respectively i and j The random variables corresponding to the nodes;
[0025] Through this step, the undirected graph region matching problem is transformed into an energy minimization problem, and the nodes and undirected edges of the undirected graph are updated in reverse iteration according to the energy minimization result.
[0026] Preferably, updating the directed edges through the Bayesian network based on the image region comprises:
[0027] Only the directed edges in the feature map graphical model and the heat map graphical model are considered and converted into directed graphs. Region matching is performed by maximizing the joint probability distribution on the Bayesian network based on the image region to update the directed edges. The update method is the same as that of the undirected graph. The directed graph region matching problem is converted into an energy minimization problem, and the nodes and directed edges of the directed graph are updated in reverse iteration according to the energy minimization result.
[0028] Preferably, the step of calculating the similarity between all matching regions between the two updated graphic models using the Siamese network includes:
[0029] Input the updated feature graph model into the first branch of the twin network, the first branch includes a CNN module, a first L1 regularization module and a first activation module, and finally the first activation module outputs the activation value of each region in the updated feature graph model;
[0030] Input the updated heat map graphical model into the second branch of the twin network, the second branch includes a Transformer module, a second L1 regularization module and a second activation module, and finally the second activation module outputs the activation value of each region in the updated heat map graphical model;
[0031] The similarity between each matching area is obtained by calculating the similarity based on the activation values obtained from the two branches. , the calculation formula is:
[0032] ,
[0033] in, For expectations, and They are the regions in the feature graph model F m The activation value and heat map of the region in the graphical model H m The activation value of and They respectively represent the activation of region m in the feature map graphical model F to region m in the heat map graphical model H, and finally the similarity between all matching regions between the two graphical models is obtained.
[0034] Preferably, the twin network further includes: exchanging mutual information between the CNN module and the Transformer module through a feature coupling unit, and sharing weights between the first L1 regularization module and the second L1 regularization module through cross attention.
[0035] Preferably, obtaining the best region matching result by minimizing global energy based on the similarity between matching regions includes:
[0036] For each candidate node Calculate global matching energy , the calculation formula is:
[0037] ,
[0038] in, is the candidate node in the updated heat map graphical model H, is the weight of the balancing term, is the partition function;
[0039] is with and The energy associated with the matching probability between:
[0040] ,
[0041] in, is the source node in the updated feature graph model F, is the region corresponding to the source node in the feature graph model F The area corresponding to the candidate node in the heat map graphical model H The similarity between
[0042] is with and The energy associated with the matching probability between pairs of parent nodes is:
[0043] ,
[0044] in, For the heat map graphical model H The index set of the parent node of is the node index, for The area corresponding to the parent node in the heat map model H The similarity between
[0045] and same, yes and The energy related to the matching probability between the child node pairs, yes and The energy related to the matching probability between pairs of neighbor nodes;
[0046] By minimizing Find the node in the heat map graph model H corresponding to the best matching area , minimize the energy value The calculation formula is:
[0047] ,
[0048] like The corresponding global matching energy Greater than the set threshold parameter , then the source region node is considered No match, otherwise it is considered and The match was successful.
[0049] Preferably, the extracting of image features and thermal features from the feature map and the thermal map respectively and obtaining a reconstructed feature map and a reconstructed thermal map through spatial attention operation, channel attention operation and feature reconstruction include:
[0050] Image features and thermal features are extracted from feature maps and thermal maps respectively, and a common spatial attention map of image features and thermal features is obtained through spatial attention operation. The common spatial attention map is used as the weight of image features and thermal features to calculate the spatial attention map of image features and the spatial attention map of thermal features respectively.
[0051] For the spatial attention map of the image feature and the spatial attention map of the thermal feature, the channel attention map of the image feature and the channel attention map of the thermal feature are obtained respectively through channel attention operation;
[0052] The spatial attention map of the image feature is multiplied by the channel attention map of the image feature to obtain the reconstructed feature map, and the spatial attention map of the thermal feature is multiplied by the channel attention map of the thermal feature to obtain the reconstructed heat map.
[0053] To achieve the above-mentioned purpose of the invention, an embodiment of the present invention further provides a UAV search and rescue image fusion device based on feature matching, comprising: an image preprocessing module, a graphic model building module, a feature matching fusion module, a feature reconstruction optimization module and a search and rescue image generation module;
[0054] The image preprocessing module is used to preprocess the acquired RGB image and infrared image to obtain a feature map and a thermal map respectively;
[0055] The graphic model construction module is used to convert the feature map and the heat map into a feature map graphic model and a heat map graphic model respectively. In each graphic model, the image area is represented by a node, and the image area is divided into multiple regional scale levels according to the size of different image areas. The inclusion relationship between nodes in different levels and the adjacency relationship between nodes in the same level are respectively constructed by directed edges and undirected edges;
[0056] The feature matching fusion module is used to update the undirected edges of the feature map graphic model and the heat map graphic model through a Markov random field based on the image region, and update the directed edges through a Bayesian network based on the image region, so as to obtain an updated feature map graphic model and a heat map graphic model, calculate the similarity between all matching regions between the two updated graphic models by using a twin network, obtain the best region matching result based on the similarity between the matching regions by global energy minimization, and fuse the feature map and the heat map based on the best region matching result to obtain a first fused image;
[0057] The feature reconstruction optimization module is used to extract image features and thermal features from the feature map and the thermal map respectively, and then obtain a reconstructed feature map and a reconstructed thermal map through spatial attention operation, channel attention operation and feature reconstruction, and fuse the reconstructed feature map and the reconstructed thermal map to obtain a second fused image;
[0058] The search and rescue image generation module is used to merge the first fused image and the second fused image into a final generated image for search and rescue.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] 1) The present invention applies feature matching to the UAV search and rescue scenario. Due to the complex situations and terrain obstacles in actual rescue missions, the data processed by fusing the heat map and the feature map is of great help to the detection of the rescue target and can more significantly display the features of the rescue target; 2) The present invention simulates the relationship between different image regions through the node and edge relationships of the graph, and constructs a feature map graphic model and a heat map graphic model, which better handles the spatial position information between images, and transforms the graphic model to convert the matching task into an energy minimization task, thereby reducing the complexity of the task and improving the credibility of the results; 3) The present invention optimizes the intra-layer cross-modal features by reconstructing the features of the heat map and the feature map through spatial attention and channel attention, so as to pay more attention to the significant content in each modality and effectively enhance the significant features of the rescue target; 4) The present invention is also applicable to security work, game entertainment, film and television production and other fields through the improvement of feature matching. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0062] Figure 1 It is a flow chart of a method for fusion of unmanned aerial vehicle search and rescue images based on feature matching provided by an embodiment of the present invention;
[0063] Figure 2 Schematic diagram of the framework of the UAV search and rescue image fusion method based on feature matching provided by an embodiment of the present invention;
[0064] Figure 3 is a schematic diagram of a thermal map graphical model provided by an embodiment of the present invention;
[0065] Figure 4 is a schematic diagram of a feature graph model provided by an embodiment of the present invention;
[0066] Figure 5 It is a schematic diagram of a framework for feature matching between graphic models based on a twin network provided in an embodiment of the present invention;
[0067] Figure 6 It is a schematic diagram of using a feature coupling unit to communicate mutual information between a CNN module and a Transformer module according to an embodiment of the present invention;
[0068] Figure 7 It is a structural schematic diagram of a UAV search and rescue image fusion device based on feature matching provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0069] To make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific implementation methods described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0070] The inventive concept of the present invention is: in the post-disaster search and rescue process, drone search and rescue is one of the efficient and convenient means, but traditional drones are prone to misjudgment in the target recognition process, which causes rescuers to miss the rescue target. Based on this, in view of the problem of insufficient accuracy of personnel search and rescue in complex situations such as occlusion and motion deformation in the prior art, the embodiment of the present invention provides a drone search and rescue image fusion method and device based on feature matching, firstly using the SAM model to perform regional segmentation on the extracted feature map and heat map to improve the accuracy of the recognition result; secondly, the feature map and heat map are converted into a feature map graphic model and a heat map graphic model, and the regional feature matching problem is converted into an energy minimization problem by combining Markov random fields, Bayesian networks and similarity calculations; then the region-to-point matching framework is used to perform accurate feature map and heat map feature matching and image fusion, and at the same time, the feature map and heat map are reconstructed by combining spatial attention and channel attention, and finally the feature matching results and feature reconstruction results are fused to obtain a generated image to accurately identify the target person.
[0071] Figure 1 is a flow chart of a method for fusion of unmanned aerial vehicle search and rescue images based on feature matching provided by an embodiment of the present invention, Figure 2 FIG. 1 is a schematic diagram of a framework of a method for fusion of UAV search and rescue images based on feature matching provided by an embodiment of the present invention. Figure 1 and Figure 2 As shown, the embodiment provides a method for fusion of UAV search and rescue images based on feature matching, comprising the following steps:
[0072] S1, preprocess the acquired RGB image and infrared image to obtain feature map and thermal map respectively.
[0073] S1.1, extract the frame images from the video captured by the ordinary camera and thermal imaging camera carried by the drone, and intercept the video clips that may contain the target person, such as clips with obvious thermal image shapes or containing human body parts. Then, in each key clip, extract the frame image at a frame rate of 20fps until the end of the clip, and obtain multiple sets of corresponding RGB images and infrared images.
[0074] S1.2, using the U-Net model or the Segment Anything Model (SAM), the acquired RGB image and infrared image are taken as input, passed into the trained U-Net model or SAM model, and n feature maps rich in implicit semantic information are output. The incoming infrared image is preprocessed by the non-uniformity correction method to obtain n processed thermal maps. .
[0075] S2, the feature map and heat map are converted into a feature map graphical model and a heat map graphical model respectively. In each graphical model, the image area is represented by a node, and it is divided into multiple regional scale levels according to the size of different image areas. The inclusion relationship between nodes in different levels and the adjacency relationship between nodes in the same level are constructed by directed edges and undirected edges respectively.
[0076] S2.1, the feature map or heat map is divided into regions by the SAM model and the regions are generated by graph completion, and the image regions are used as nodes to simulate the relationship between regions by including edges and adjacent edges, so as to make it a multi-relationship graph. According to the size of these regions, they are divided into L levels of regional scale levels, and each regional scale level corresponds to a different image scale. In each graphical model, since the nodes represent the image regions, the nodes in each graphical model are also divided into L levels of levels. In the embodiment, Figure 3 and Figure 4 As shown, it is divided into 4 levels, namely L0, L1, L2, and L3.
[0077] The specific process of the graph completion algorithm is as follows:
[0078] (1) Starting from the lowest level, extract all isolated nodes. Assume that the set of these nodes is .
[0079] (2) Using the K-means algorithm, the nodes are clustered according to their center coordinates. The nodes in each cluster share similar center coordinates. The clusters after clustering are: , the specific process of clustering is expressed as:
[0080] ,
[0081] in, m is the number of cluster categories, Cluster center.
[0082] (3) In the same cluster, the node merges its corresponding region with its nearest neighbor, which is expressed as follows:
[0083] ,
[0084] in, For Node The area represented, is a parameter used to balance node attributes and distances. is the new region generated after fusion.
[0085] (4) Generate new nodes corresponding to the new areas.
[0086] (5) Use two types of edges, including edges and adjacent edges, to simulate the relationship between regions and connect new nodes.
[0087] S2.2, in the constructed feature graph model and heat map graph model, edges represent two types of relationships between regions, namely, inclusion and adjacency. Inclusion edges are directed, pointing from one region to one of its containing regions. They form hierarchical connections between graph nodes, and are particularly capable of robust and efficient region matching under scale changes. Adjacency edges are undirected, indicating that the regions it connects have common parts, but larger regions will not contain smaller regions. Possible edges are constructed based on the regions near the nodes.
[0088] S3, update the undirected edges of the feature map graphic model and the heat map graphic model through the Markov random field based on the image area, update the directed edges through the Bayesian network based on the image area, and obtain the updated feature map graphic model and the heat map graphic model, use the twin network to calculate the similarity between all matching areas between the two updated graphic models, obtain the best area matching result based on the similarity between the matching areas by global energy minimization, and fuse the feature map and the heat map based on the best area matching result to obtain the first fused image.
[0089] S3.1, in a given region, construct the edges between the initial nodes and the inferred possible corresponding regions. By considering the general adjacency relationship, all the edges in the feature graph model and the heat map graph model are regarded as undirected edges and converted into undirected graphs, and random variables are introduced for all nodes in the undirected graphs of the two graph models. To indicate the matching status of these nodes with the source nodes, the region matching is performed to update the undirected edges by maximizing the joint probability distribution on the Markov random field based on the image region. The calculation formula is:
[0090] ,
[0091] According to the Hammersley-Clifford theorem, the probability distribution defined by the Markov random field based on the image region belongs to the Boltzmann distribution and is the exponential of the negative energy function, that is, , so the region matching is expressed as energy minimization, and the calculation formula is:
[0092] ,
[0093] According to the structure of the undirected graph, the energy function Node Energy and edge energy Two parts, the calculation formula is:
[0094] ,
[0095] in, is the parameter of the balance term, is the index of all neighbor node pairs in the undirected graph A collection of and Respectively i and j The random variables corresponding to the nodes, for each graph node, when its matching probability is high, its energy is expected to be low.
[0096] Through this step, the undirected graph region matching problem is transformed into an energy minimization problem, and the nodes and undirected edges of the undirected graph are updated in reverse iteration according to the energy minimization result.
[0097] S3.2, only consider the directed edges in the feature graph model and the heat map model and convert them into directed graphs. If the region corresponding to the source node in the feature graph model The area corresponding to any node in the heat map graphical model If it does not match, Will not This phenomenon involves conditional independence. Based on this, the graphical model containing edges is converted into a Bayesian network based on image regions. The region matching is performed by maximizing the joint probability distribution on the Bayesian network based on image regions to update the directed edges. The calculation formula of the joint probability distribution is:
[0098] ,
[0099] Similar to the update method of undirected graphs, after obtaining the joint probability distribution, the directed graph region matching problem is transformed into an energy minimization problem, and the nodes and directed edges of the directed graph are updated in reverse iteration based on the energy minimization result.
[0100] S3.3, the similarity is calculated using the improved twin network, such as Figure 5 As shown, the updated feature map graphical model F and the heat map graphical model H are respectively input into the CNN module in the first branch of the twin network and the Transformer module in the second branch, and features are extracted respectively and the mutual information between them is communicated using the feature coupling unit (FCU), as shown in Figure 6 As shown in the figure, the CNN module passes the feature representation to the Transformer module through 1*1 convolution, downsampling, and layer normalization. The Transformer module passes the feature representation to the CNN module through upsampling, 1*1 convolution, and batch normalization to gradually fuse the image features in the CNN module with the thermal feature representation in the Transformer module.
[0101] The first L1 regularization module and the second L1 regularization module are connected after the CNN module and the Transformer module respectively. The base layer image with a stable structure and the detail layer image with a significant structure are retained through L1 regularization, and the weights are shared between the two regularization modules through cross attention. The calculation formula of L1 regularization is:
[0102] ,
[0103] ,
[0104] in, and are the L1 regularization loss calculated by the first L1 regularization module and the L1 regularization loss calculated by the second L1 regularization module, respectively. They are the loss of the feature map model F after the CNN module and the loss of the heat map model H after the Transformer module. is the regularization parameter, which is used to control the strength of the regularization term. and are the absolute values of the weights of each layer of the CNN module and the Transformer module, respectively. i and j Respectively represent the CNN modules The layer index in the layer and the Transformer module are The layer index in layers.
[0105] Finally, the first activation module and the second activation module are connected after the first L1 regularization module and the second L1 regularization module, respectively, to predict the probability of each block in one area appearing in another area, and output the activation value of each area in the updated feature map graphical model and the activation value of each area in the updated heat map graphical model. The similarity between each matching area is obtained by calculating the similarity based on the activation values obtained from the two branches. , the similarity is obtained by the expected product, which helps us to perform accurate region matching. The calculation formula is:
[0106] ,
[0107] in, For expectations, and They are the regions in the feature graph model F m The activation value and heat map of the region in the graphical model H m The activation value of and They respectively represent the activation of region m in the feature map graphical model F to region m in the heat map graphical model H, and finally the similarity between all matching regions between the two graphical models is obtained.
[0108] S3.4, the content obtained through the above steps is used to obtain the assumed region match, and finally, the best region match is determined by global energy minimization.
[0109] For each candidate node Calculate global matching energy , the calculation formula is:
[0110] ,
[0111] in, is the candidate node in the updated heat map graphical model H, is the weight of the balancing term, is the partition function.
[0112] is with and The energy associated with the matching probability between:
[0113] ,
[0114] in, is the source node in the updated feature graph model F, is the region corresponding to the source node in the feature graph model F The area corresponding to the candidate node in the heat map graphical model H The similarity between .
[0115] is with and The energy associated with the matching probability between pairs of parent nodes is:
[0116] ,
[0117] in, For the heat map graphical model H The index set of the parent node of is the node index, for The area corresponding to the parent node in the heat map model H The similarity between .
[0118] and same, yes and The energy related to the matching probability between the child node pairs, yes and The energy related to the matching probability between pairs of neighbor nodes.
[0119] By minimizing Find the node in the heat map graph model H corresponding to the best matching area , minimize the energy value The calculation formula is:
[0120] ,
[0121] like The corresponding global matching energy Greater than the set threshold parameter , then the source region node is considered No match, otherwise it is considered and The match was successful.
[0122] S3.5, matching the image area according to the matching nodes based on the best area matching result obtained in the above steps, and obtaining a first fused image by fusing the feature map and the heat map.
[0123] S4, after extracting image features and thermal features from the feature map and the heat map respectively, obtain a reconstructed feature map and a reconstructed heat map through spatial attention operation, channel attention operation and feature reconstruction, and fuse the reconstructed feature map and the reconstructed heat map to obtain a second fused image.
[0124] S4.1, extract image features from feature maps and heat maps respectively and thermal characteristics , the common spatial attention map of image features and thermal features is obtained through spatial attention operation ,in i Indicates i Layer characteristics, calculated as:
[0125] ,
[0126] in, is the element-wise multiplication operation, is the spatial attention operation, defined as:
[0127] ,
[0128] in, is the Sigmoid activation function, It is a 3*3 convolution operation. It is the maximum pooling operation along the channel direction.
[0129] Next, the public spatial attention map is used as the weight of the image feature and the thermal feature respectively, and updated in reverse to achieve feature alignment of the spatial part:
[0130] ,
[0131] ,
[0132] in, is the spatial attention map of image features, is the spatial attention map of thermal features.
[0133] S4.2, for the spatial attention map of image features and the spatial attention map of thermal features, the weight of the salient content is highlighted through channel attention operation:
[0134] ,
[0135] ,
[0136] in, is the channel attention map of image features, is the channel attention map of thermal features, is the channel attention operation, defined as:
[0137] ,
[0138] ,
[0139] in, is the Sigmoid activation function, It is a 1*1 convolution operation. It is the global maximum pooling operation.
[0140] S4.3, multiply the channel attention feature and the spatial attention feature to achieve feature reconstruction:
[0141] ,
[0142] ,
[0143] in, For the i The reconstructed feature map of the layer, For the i Reconstructed heatmap of layers.
[0144] After the above-mentioned spatial alignment, channel recalibration and feature reconstruction, the new reconstructed features show stronger representation capabilities, enabling further optimization of intra-layer cross-modal features to pay more attention to the salient content in each modality.
[0145] S4.4, fusing the reconstructed feature map and the reconstructed heat map to obtain a second fused image.
[0146] S5, merging the first fused image and the second fused image into a final generated image for search and rescue.
[0147] The first fused image and the second fused image are superimposed to obtain generated images that are accurately located and have significant features of the rescue target.
[0148] In summary, an embodiment of the present invention provides a method for fusion of UAV search and rescue images based on feature matching. Based on an improved feature matching and feature reconstruction framework, a generated image with precise positioning and significant features of the rescue target is obtained, and the most advanced basic model SAM is combined for region segmentation. This solves the problem that traditional models require a large amount of data sets and annotations. It can be better used in actual tasks of searching and rescuing personnel in complex situations such as occlusion and motion deformation, and provides an intuitive reference for units involved in the search and rescue, which helps to conduct more efficient and accurate searches for target personnel.
[0149] Based on the same inventive concept, Figure 7 As shown, an embodiment of the present invention also provides a UAV search and rescue image fusion device 700 based on feature matching, including: an image preprocessing module 710, a graphic model construction module 720, a feature matching fusion module 730, a feature reconstruction optimization module 740 and a search and rescue image generation module 750.
[0150] The image preprocessing module 710 is used to preprocess the acquired RGB image and infrared image to obtain a feature map and a thermal map respectively.
[0151] The graphic model construction module 720 is used to convert the feature map and the heat map into a feature map graphic model and a heat map graphic model respectively. In each graphic model, the image area is represented by a node, and is divided into multiple regional scale levels according to the size of different image areas. The inclusion relationship between nodes in different levels and the adjacency relationship between nodes in the same level are constructed by directed edges and undirected edges respectively.
[0152] The feature matching fusion module 730 is used to update the undirected edges of the feature map graphic model and the heat map graphic model through a Markov random field based on the image region, and to update the directed edges through a Bayesian network based on the image region, so as to obtain an updated feature map graphic model and a heat map graphic model, and to calculate the similarity between all matching regions between the two updated graphic models using a twin network, and to obtain the best region matching result based on the similarity between the matching regions by minimizing the global energy, and to fuse the feature map and the heat map based on the best region matching result to obtain a first fused image.
[0153] The feature reconstruction optimization module 740 is used to extract image features and thermal features from the feature map and the heat map respectively, and then obtain a reconstructed feature map and a reconstructed heat map through spatial attention operation, channel attention operation and feature reconstruction, and fuse the reconstructed feature map and the reconstructed heat map to obtain a second fused image.
[0154] The search and rescue image generation module 750 is used to merge the first fused image and the second fused image into a final generated image for search and rescue.
[0155] It should be noted that the UAV search and rescue image fusion device based on feature matching provided in the above embodiment belongs to the same inventive concept as the UAV search and rescue image fusion method based on feature matching. The specific implementation process is detailed in the embodiment of the UAV search and rescue image fusion method based on feature matching, which will not be repeated here.
[0156] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A UAV search and rescue image fusion method based on feature matching, characterized in that: The following steps are involved: The acquired RGB image and infrared image are preprocessed to obtain feature map and thermal map respectively; The feature map and heat map are converted into a feature map graphical model and a heat map graphical model respectively. In each graphical model, the image area is represented by a node, and the image area is divided into multiple regional scale levels according to the size of different image areas. The inclusion relationship between nodes in different levels and the adjacency relationship between nodes in the same level are constructed by directed edges and undirected edges respectively. The undirected edges of the feature map graphic model and the heat map graphic model are updated by a Markov random field based on the image region, and the directed edges are updated by a Bayesian network based on the image region to obtain an updated feature map graphic model and a heat map graphic model, and the similarity between all matching regions between the two updated graphic models is calculated by using a twin network, and the best region matching result is obtained by minimizing the global energy based on the similarity between the matching regions, and the feature map and the heat map are fused based on the best region matching result to obtain a first fused image; After extracting image features and thermal features from the feature map and the thermal map respectively, a reconstructed feature map and a reconstructed thermal map are obtained through spatial attention operation, channel attention operation and feature reconstruction, and the reconstructed feature map and the reconstructed thermal map are fused to obtain a second fused image; The first fused image and the second fused image are combined into a final generated image for search and rescue.
2. The method for fusion of unmanned aerial vehicle search and rescue images based on feature matching according to claim 1, characterized in that: The obtained RGB image and infrared image are preprocessed to obtain a feature map and a thermal map, respectively, including: Key frame images are captured from the video taken by the drone's onboard camera. The key frame images include RGB images and infrared images taken synchronously. The RGB images are segmented using the U-Net model or the SAM model to obtain feature maps, and the infrared images are processed using the non-uniformity correction method to obtain thermal maps.
3. The method for fusion of unmanned aerial vehicle search and rescue images based on feature matching according to claim 1, characterized in that: The image region is represented by a node, and is divided into multiple regional scale levels according to the size of different image regions, including: The feature map or heat map is divided into regions and the regions are generated by graph completion through the SAM model. According to the size of these regions, they are divided into L levels of regional scale levels. Each regional scale level corresponds to a different image scale. In each graphic model, nodes are used to represent the image region, so that the nodes in each graphic model are also divided into L levels.
4. The method for fusion of unmanned aerial vehicle search and rescue images based on feature matching according to claim 1, characterized in that: The method of updating the undirected edges of the feature map graphical model and the heat map graphical model through a Markov random field based on the image region includes: All edges in the feature graph model and the heat map model are considered as undirected edges and converted into undirected graphs, and random variables are introduced for all nodes in the undirected graphs of the two graph models. To indicate the matching status of these nodes with the source nodes, the region matching is performed to update the undirected edges by maximizing the joint probability distribution on the Markov random field based on the image region. The calculation formula is: , According to the Hammersley-Clifford theorem, the probability distribution defined by the Markov random field based on the image region belongs to the Boltzmann distribution and is the exponential of the negative energy function, that is, , so the region matching is expressed as energy minimization, and the calculation formula is: , According to the structure of the undirected graph, the energy function Node Energy and edge energy Two parts, the calculation formula is: , in, is the parameter of the balance term, is the index of all neighbor node pairs in the undirected graph A collection of and Respectively i and j The random variables corresponding to the nodes; Through this step, the undirected graph region matching problem is transformed into an energy minimization problem, and the nodes and undirected edges of the undirected graph are updated in reverse iteration according to the energy minimization result.
5. The method for fusion of unmanned aerial vehicle search and rescue images based on feature matching according to claim 4 is characterized in that: The method of updating the directed edge by using a Bayesian network based on the image region includes: Only the directed edges in the feature map graphical model and the heat map graphical model are considered and converted into directed graphs. Region matching is performed by maximizing the joint probability distribution on the Bayesian network based on the image region to update the directed edges. The update method is the same as that of the undirected graph. The directed graph region matching problem is converted into an energy minimization problem, and the nodes and directed edges of the directed graph are updated in reverse iteration according to the energy minimization result.
6. The method for fusion of unmanned aerial vehicle search and rescue images based on feature matching according to claim 1, characterized in that: The method of calculating the similarity between all matching regions between the two updated graphic models using the Siamese network includes: Input the updated feature graph model into the first branch of the twin network, the first branch includes a CNN module, a first L1 regularization module and a first activation module, and finally the first activation module outputs the activation value of each region in the updated feature graph model; Input the updated heat map graphical model into the second branch of the twin network, the second branch includes a Transformer module, a second L1 regularization module and a second activation module, and finally the second activation module outputs the activation value of each region in the updated heat map graphical model; The similarity between each matching area is obtained by calculating the similarity based on the activation values obtained from the two branches. , the calculation formula is: , in, For expectations, and They are the regions in the feature graph model F m The activation value and heat map of the region in the graphical model H m The activation value of and They respectively represent the activation of region m in the feature map graphical model F to region m in the heat map graphical model H, and finally the similarity between all matching regions between the two graphical models is obtained.
7. The method for fusion of unmanned aerial vehicle search and rescue images based on feature matching according to claim 6, characterized in that: The twin network also includes: exchanging mutual information between the CNN module and the Transformer module through a feature coupling unit, and sharing weights between the first L1 regularization module and the second L1 regularization module through cross attention.
8. The method for fusion of unmanned aerial vehicle search and rescue images based on feature matching according to claim 1, characterized in that: The method of obtaining the best region matching result by minimizing the global energy based on the similarity between the matching regions includes: For each candidate node Calculate global matching energy , the calculation formula is: , in, is the candidate node in the updated heat map graphical model H, is the weight of the balancing term, is the partition function; is with and The energy associated with the matching probability between: , in, is the source node in the updated feature graph model F, is the region corresponding to the source node in the feature graph model F The area corresponding to the candidate node in the heat map graphical model H The similarity between is with and The energy associated with the matching probability between pairs of parent nodes is: , in, For the heat map graphical model H The index set of the parent node of is the node index, for The area corresponding to the parent node in the heat map model H The similarity between and same, yes and The energy related to the matching probability between the child node pairs, yes and The energy related to the matching probability between pairs of neighbor nodes; By minimizing Find the node in the heat map graph model H corresponding to the best matching area , minimize the energy value The calculation formula is: , like The corresponding global matching energy Greater than the set threshold parameter , then the source region node is considered No match, otherwise it is considered and The match was successful.
9. The method for fusion of unmanned aerial vehicle search and rescue images based on feature matching according to claim 1, characterized in that: The step of extracting image features and thermal features from the feature map and the thermal map respectively, and obtaining a reconstructed feature map and a reconstructed thermal map through spatial attention operation, channel attention operation and feature reconstruction includes: Image features and thermal features are extracted from feature maps and thermal maps respectively, and a common spatial attention map of image features and thermal features is obtained through spatial attention operation. The common spatial attention map is used as the weight of image features and thermal features to calculate the spatial attention map of image features and the spatial attention map of thermal features respectively. For the spatial attention map of the image feature and the spatial attention map of the thermal feature, the channel attention map of the image feature and the channel attention map of the thermal feature are obtained respectively through channel attention operation; The spatial attention map of the image feature is multiplied by the channel attention map of the image feature to obtain the reconstructed feature map, and the spatial attention map of the thermal feature is multiplied by the channel attention map of the thermal feature to obtain the reconstructed heat map.
10. A UAV search and rescue image fusion device based on feature matching, implemented by the UAV search and rescue image fusion method based on feature matching according to any one of claims 1 to 9, characterized in that: include: Image preprocessing module, graphic model building module, feature matching and fusion module, feature reconstruction and optimization module and search and rescue image generation module; The image preprocessing module is used to preprocess the acquired RGB image and infrared image to obtain a feature map and a thermal map respectively; The graphic model construction module is used to convert the feature map and the heat map into a feature map graphic model and a heat map graphic model respectively. In each graphic model, the image area is represented by a node, and the image area is divided into multiple regional scale levels according to the size of different image areas. The inclusion relationship between nodes in different levels and the adjacency relationship between nodes in the same level are respectively constructed by directed edges and undirected edges; The feature matching fusion module is used to update the undirected edges of the feature map graphic model and the heat map graphic model through a Markov random field based on the image region, and update the directed edges through a Bayesian network based on the image region, so as to obtain an updated feature map graphic model and a heat map graphic model, calculate the similarity between all matching regions between the two updated graphic models by using a twin network, obtain the best region matching result based on the similarity between the matching regions by global energy minimization, and fuse the feature map and the heat map based on the best region matching result to obtain a first fused image; The feature reconstruction optimization module is used to extract image features and thermal features from the feature map and the thermal map respectively, and then obtain a reconstructed feature map and a reconstructed thermal map through spatial attention operation, channel attention operation and feature reconstruction, and fuse the reconstructed feature map and the reconstructed thermal map to obtain a second fused image; The search and rescue image generation module is used to merge the first fused image and the second fused image into a final generated image for search and rescue.
Citation Information
Patent Citations
RGBT target tracking method based on twin network structure and anchor frame adaptive thought
CN116563343A
Salient target detection method based on multiband visual image perception and fusion
CN117132759A