UAV relative pose estimation method and device based on multi-task image matching

By building a multi-task image learning network in the drone, semantic segmentation and feature description are achieved simultaneously, the problem of low image matching accuracy and efficiency in drone visual positioning is solved, the accuracy and efficiency of relative pose estimation is improved, and resource utilization is reduced.

CN119810200BActive Publication Date: 2025-05-16NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510280585.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-05-16
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

In the prior art, in the process of visual positioning of drones, image matching methods have problems with accuracy and low efficiency, especially in drones with limited resources, it is difficult to perform semantic segmentation and image matching tasks simultaneously.

Method used

The multi-task image matching method is adopted to build a multi-task image learning network, including feature extraction module, dual-branch network and feature fusion module, to achieve semantic segmentation and feature description simultaneously, reduce model parameters, and improve the accuracy and efficiency of image matching.

Benefits of technology

It improves the image matching accuracy and efficiency of drone relative position estimation, reduces the occupation of airborne resources, solves the problem of limited resources in drones, and improves the efficiency and accuracy of mission execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810200B_ABST
    Figure CN119810200B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for estimating relative pose of unmanned aerial vehicle based on multi-task image matching, the method comprising: constructing a multi-task image learning network, the multi-task image learning network comprising a feature extraction module, a dual-branch network and a feature fusion module arranged in sequence, the dual-branch network comprising a semantic segmentation branch network and a graphic feature description branch network; training the multi-task image learning network, using a loss function including descriptor loss, semantic segmentation loss and feature map loss during the training process; obtaining images captured by the unmanned aerial vehicle during flight, inputting them into the trained multi-task image learning model, obtaining matching pair relationships between input images, and using a semantic segmentation map to guide image matching tasks, estimating the relative pose relationship between cameras according to the matching pair relationship. The present invention can improve the image matching accuracy and efficiency of relative pose estimation of unmanned aerial vehicle, while reducing the storage occupation on the aircraft.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle visual positioning, and in particular to a method and device for estimating relative posture of an unmanned aerial vehicle based on multi-task image matching. Background Art

[0002] Image matching is the premise of visual positioning. It can enable drones to have high-precision positioning information in satellite denial or indoor conditions, help build a three-dimensional model of the scene, and realize the exploration of the environment. Image matching is to directly extract and match features of images to find corresponding feature points or feature descriptors in two or more images or data sets, and establish the corresponding relationship between feature points or feature descriptors.

[0003] Image matching methods can be divided into two categories at present: one is feature point-based image matching methods, and the other is direct image matching methods. Among them, the feature point-based image matching method is the main method of local feature matching. Its matching strategy is divided into two stages. The first stage extracts the key points and corresponding descriptors in the image, and the second stage uses the descriptors to complete the image matching task. This type of method has high accuracy and robustness, and can extract feature points in complex environments. However, when there is occlusion, partial occlusion, or the image lacks texture information or contains a large number of repeated features in the scene, it will be difficult to extract enough feature points or may produce a large number of mismatches. To solve the above problems, the current main method is to extract more feature points or improve the accuracy of the descriptor, or improve the reasoning speed of the network, but it will lead to a high complexity of implementation and reduce matching efficiency. In addition, due to the limitation of feature point detection, this type of method obtains sparse matching results. The direct image matching method directly obtains the matching relationship between the two images, eliminating the intermediate process. This method extracts the image matching problem into the correspondence problem of two point sets, that is, the pure point set matching problem. Although this type of method works better in weak texture areas or repeated areas, the inference speed is slow, and because it needs to calculate dense matching relationships, the computational complexity is high and it requires a lot of resources.

[0004] In the process of UAV visual positioning, in order to achieve relative pose estimation, multiple tasks need to be performed, including semantic segmentation tasks and image matching tasks. When the above-mentioned traditional image matching method is used to realize the relative pose estimation of the UAV, since the existing technology usually trains a separate network for each task in the process of UAV relative pose estimation, it takes up more resources, and the tasks are executed independently of each other. However, the onboard resources of the UAV are limited, and it is difficult to meet the requirements of simultaneous execution of semantic segmentation and image matching tasks, resulting in low task execution efficiency. Summary of the invention

[0005] The technical problem to be solved by the present invention is: in response to the technical problems existing in the prior art, the present invention provides a method and device for estimating the relative pose of a UAV based on multi-task image matching, which can improve the image matching accuracy and efficiency of the relative pose estimation of the UAV, while reducing the storage occupancy on the machine.

[0006] In order to solve the above technical problems, the technical solution proposed by the present invention is:

[0007] A method for estimating relative pose of a drone based on multi-task image matching, comprising the following steps:

[0008] Constructing a multi-task image learning network, the multi-task image learning network includes a feature extraction module, a dual-branch network and a feature fusion module which are arranged in sequence, the feature extraction module is used to extract a feature map from an input image and provide it to the dual-branch network to learn the common features in the input image, the dual-branch network includes a semantic segmentation branch network for performing a semantic segmentation task on the input image and a graphic feature description branch network for performing a graphic feature description task on the input image, the semantic segmentation branch network extracts a semantic segmentation descriptor and a semantic segmentation map, the graphic feature description branch network extracts a feature descriptor, and the feature fusion module is used to fuse the semantic segmentation descriptor extracted by the semantic segmentation branch network and the feature descriptor extracted by the graphic feature description branch network to form a final dense feature descriptor, so as to explicitly embed semantic global information into the feature descriptor;

[0009] The multi-task image learning network is trained, and during the training process, a loss function including a descriptor loss, a semantic segmentation loss, and a feature map loss is used to control the training process, and after the training is completed, a multi-task image learning model is obtained;

[0010] The images captured by the camera of the UAV during flight are obtained, and the obtained images are input into the trained multi-task image learning model to obtain a semantic segmentation map and a dense feature descriptor of the input image, and a matching pair relationship between images is obtained by feature matching according to the dense feature descriptors of each image, and the semantic segmentation map is used to guide the image matching task, and the relative posture relationship between cameras is estimated according to the matching pair relationship.

[0011] Furthermore, the feature extraction module is a feature encoding backbone network, and the semantic segmentation branch network and the graphic feature description branch network in the dual-branch network share the feature encoding backbone network to learn the common features in the input image, and the different features of the two tasks are learned respectively through the feature description head of the graphic feature description branch network and the semantic segmentation head of the semantic segmentation branch network.

[0012] Furthermore, the multi-task image learning network is also provided with a feature distillation network, which performs feature distillation on the feature map extracted by the feature encoding backbone network by using a semantic segmentation model as a teacher network, and distills the output result of the semantic segmentation head using the segmentation result obtained by segmenting the input image using the teacher network.

[0013] Furthermore, the feature fusion module fuses the semantic segmentation result with the feature descriptor through multiple stacked self-attention layers and cross-attention layers. The input vector of the self-attention layer includes a query vector Q, a key-value vector K and a value vector V, wherein Q, K and V of the self-attention layer are all from the same descriptor input, and Q, K and V in the cross-attention layer are from different descriptor inputs, respectively. The query vector retrieves information in the value vector by calculating the similarity with the key-value vector, and the calculation expression is:

[0014]

[0015]

[0016] in, represents the self-attention layer, represents the cross attention layer, represents the activation function, represents the first input feature, Represents the second input feature.

[0017] Furthermore, the feature fusion module is also used to perform a position encoding operation after flattening the two-dimensional image into a one-dimensional vector, and use the encoded vector x as the input source and y of the attention mechanism, and the calculation expression is:

[0018]

[0019]

[0020]

[0021] Among them, x represents the encoded vector, represents the feature descriptor obtained by the graphic feature description branch, represents a function that flattens a two-dimensional vector into one dimension, The flattened feature descriptors and semantic descriptors are respectively represents the semantic descriptor obtained by the semantic segmentation branch, Represents the input features, that is, the flattened semantic descriptor or feature descriptor, Express Features after position encoding.

[0022] Furthermore, the total loss function used during training is The calculation expression is:

[0023]

[0024]

[0025]

[0026] in, represents the semantic segmentation loss, represents the feature map loss, is the feature map of the i-th layer of the backbone network, is the intermediate feature map of the i-th layer of the semantic segmentation teacher network, represents the descriptor loss, , as well as They represent weight coefficients respectively, is the semantic segmentation result output by the teacher network, It is the semantic segmentation result predicted by the dual-branch network.

[0027] Furthermore, the descriptor loss The calculation expression is:

[0028]

[0029]

[0030]

[0031] in, represents the similarity matrix, They represent the i-th feature descriptor extracted from image A and the j-th feature descriptor extracted from image B in the image pair to be matched, respectively. represents the similarity score matrix calculated by dual softmax, represents the real matching relationship, ij is a pair of matches, i, j represent the i, j feature points extracted from image A and image B respectively, t represents the scaling factor, They represent the matching scores from image A to image B and from image B to image A, respectively. represents the index of the matching pair that matches i or j, Indicates Figure A, Represents Figure B.

[0032] Furthermore, the use of the semantic segmentation map to guide the image matching task includes: retaining feature matching pairs on the background or static objects according to the semantic segmentation map, and reducing the weight of or directly discarding matching pairs on dynamic objects.

[0033] A device for estimating relative posture of a drone based on multi-task image matching, comprising:

[0034] A model construction module, for constructing a multi-task image learning network, wherein the multi-task image learning network includes a dual-branch network, a feature fusion module, a feature extraction module, and an image matching module, which are arranged in sequence. The dual-branch network includes a semantic segmentation branch network for performing a semantic segmentation task and a graphic feature description branch network for performing a graphic feature description task. The feature fusion module fuses the semantic segmentation result extracted by the semantic segmentation branch network and the feature descriptor extracted by the graphic feature description branch network to form a fusion feature, so as to explicitly embed semantic global information into the feature descriptor. The feature extraction module extracts features from the fusion features output by the feature fusion module. The image matching module calculates the matching degree between the two images to be matched according to the features extracted from the two images to be matched.

[0035] A model training module is used to train the multi-task image learning network. During the training process, a loss function including a descriptor loss, a semantic segmentation loss, and a feature map loss is used. After the training is completed, a multi-task image learning model is obtained.

[0036] The pose estimation module is used to obtain images captured by the camera during the flight of the UAV, input the acquired images into the trained multi-task image learning model, obtain the matching pair relationship between the input images, and estimate the relative pose relationship between the cameras based on the matching pair relationship.

[0037] A computer device comprises a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to perform the above method.

[0038] A computer-readable storage medium storing a computer program, wherein the computer program implements the above method when executed by a processor.

[0039] Compared with the prior art, the advantages of the present invention are:

[0040] 1. The present invention adopts a multi-task learning method and a dual-branch network to simultaneously complete the semantic segmentation task and the feature description task. A multi-task image learning network is constructed by the dual-branch network, the feature fusion module, and the feature extraction module. The multi-task image learning network is used to learn the matching relationship between the two images in the process of relative pose estimation of the drone, thereby realizing the relative pose estimation of the drone. The use of multi-task learning can help reduce model parameters, so that there is no need to build a separate network for each task, that is, image matching can be achieved quickly and accurately, which can greatly reduce the resources required and solve the problem of limited onboard resources in the drone. At the same time, it can also improve the efficiency and precision of task execution, thereby improving the accuracy and efficiency of relative pose estimation of the drone.

[0041] 2. The present invention further utilizes semantic segmentation to implicitly and explicitly guide image matching tasks, uses a semantic segmentation model to distill the backbone network, enables the model to implicitly learn semantic descriptions, and explicitly combines semantic descriptors with feature descriptors to embed semantic information in the descriptors, thereby effectively improving the accuracy and robustness of image matching under relative pose estimation for drones. At the same time, directly using semantic information as a guide in the image matching process can also reduce the number of feature points on the moving foreground and the phenomenon of feature point spillover, thereby improving the accuracy of feature description. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a schematic diagram of the implementation flow of the method for estimating relative posture of a UAV based on multi-task image matching in this embodiment.

[0043] Figure 2 Schematic diagram of the structural principle of the multi-task image learning network in this embodiment. DETAILED DESCRIPTION

[0044] The present invention is further described below in conjunction with the accompanying drawings and specific preferred embodiments, but the protection scope of the present invention is not limited thereby.

[0045] Image matching tasks and semantic segmentation tasks are similar. Both of them construct a metric space so that the same or similar points are as close as possible, and points of different classes are as far away as possible. The only difference is the granularity of the two. The purpose of semantic segmentation is to predict the corresponding category for each pixel of the image at a coarse level of accuracy. Pixels of the same category have the same semantic segmentation label. The image feature description task is to predict a descriptor for each pixel of the image at a detailed level. It is usually a high-dimensional vector. The descriptors of different pixels under the same semantic label are different. Taking into account that onboard resources are limited, and semantic segmentation and image feature description tasks have similar characteristics, the present invention adopts a multi-task learning approach and a dual-branch network to simultaneously complete the semantic segmentation task and the feature description task. A multi-task image learning network is constructed by the dual-branch network, a feature fusion module, a feature extraction module, and an image matching module. The multi-task image learning network is used to learn the matching relationship between the two images in the process of relative pose estimation of the drone, thereby realizing the relative pose estimation of the drone. The use of a multi-task learning approach can help reduce model parameters, so that there is no need to build a separate network for each task, that is, image matching can be achieved quickly and accurately, which can greatly reduce the resources required and solve the problem of limited onboard resources in the drone. At the same time, it can also improve the efficiency and precision of task execution, thereby improving the accuracy and efficiency of relative pose estimation of the drone.

[0046] like Figure 1 As shown, the steps of the method for estimating relative pose of a drone based on multi-task image matching in this embodiment include:

[0047] Step S01. Construct a multi-task image learning network, which includes a feature extraction module, a dual-branch network and a feature fusion module which are arranged in sequence. The feature extraction module is used to extract a feature map from an input image and provide it to the dual-branch network to learn the common features in the input image. The dual-branch network includes a semantic segmentation branch network for performing a semantic segmentation task on the input image and a graphic feature description branch network for performing a graphic feature description task on the input image. The semantic segmentation branch network extracts a semantic segmentation descriptor and a semantic segmentation map, and the graphic feature description branch network extracts a feature descriptor. The feature fusion module is used to fuse the semantic segmentation descriptor extracted by the semantic segmentation branch network and the feature descriptor extracted by the graphic feature description branch network to form a final dense feature descriptor, so as to explicitly embed semantic global information into the feature descriptor.

[0048] The multi-task image learning network model constructed in this embodiment is as follows: Figure 2As shown, it includes a feature extraction module, a dual-branch network and a feature fusion module which are arranged in sequence, wherein the feature extraction module extracts a feature map from the original image as a common feature of the dual-branch network, and the dual-branch network simultaneously completes the semantic segmentation and feature description tasks through the symmetrical structure of the dual branches, realizes multi-task learning, reduces network model parameters, and makes it more convenient for use on airborne platforms; the feature fusion module improves the robustness of the feature descriptor by explicitly fusing the semantic descriptor with the feature descriptor, and then the matching pair relationship between the images can be obtained through feature matching, and then the relative posture relationship between the cameras can be estimated according to the matching pair relationship between the images to be matched.

[0049] This embodiment aims at image matching in the process of relative pose estimation of drones. By using a dual-branch network to simultaneously complete the image matching task and the semantic segmentation task, the prediction of the two tasks is completed simultaneously in one network, and the feature description and semantic segmentation of multiple tasks are realized. This can reduce the storage usage on the machine, and at the same time improve the accuracy of the feature description, thereby improving the accuracy of image matching for relative pose estimation of drones. In the dual-branch network, the semantic segmentation branch network and the graphic feature description branch network share a feature encoding backbone network to learn common features, and the different features of the two tasks are learned respectively through the feature description head of the graphic feature description branch network and the semantic segmentation head of the semantic segmentation branch network.

[0050] Considering that semantic segmentation requires that the features of the same category are as close as possible and the features of different categories are as far away as possible in the metric space it constitutes, image feature description requires that the features of the same point are as close as possible and the features of different points are as far away as possible in the metric space it constitutes, so it is necessary to construct two different feature spaces, but they have similarities, such as Figure 2 As shown, in this embodiment, the two branches are roughly aligned during design, share a feature encoding backbone network (feature extraction module) to learn the common feature f, and then learn different features of the two tasks in the feature description head and the semantic segmentation head. In order to maintain the balance of the network, the size of the backbone network and the decoding head network can also be made approximately the same so that the two decoding heads are also balanced.

[0051] In a specific application embodiment, ResNet can be used as the backbone network to extract a feature map with a resolution of 1 / 16. In the decoding part, the feature descriptor outputs a d-dimensional feature descriptor with a resolution of 1 / 4, and in the semantic segmentation part, the segmentation result with the original resolution is predicted. For example, given an image , obtained through the backbone network Feature map[ , ,..., ], whose resolutions are 1, 1 / 4, 1 / 8, and 1 / 16 respectively. The feature maps obtained by the feature description head and the backbone network are combined with upsampling to finally obtain the feature ∈ , That is, the feature descriptor obtained by the graphic feature description branch, after the semantic description head and the feature map obtained by the backbone network are combined with upsampling to finally obtain the feature ∈ , That is the semantic descriptor obtained by the semantic segmentation branch.

[0052] Since the semantic descriptor focuses on global information and is more reliable for appearance transformation or viewpoint change, the feature descriptor focuses more on local information and is discriminative. In order to improve the robustness of the feature descriptor, in this embodiment, a feature distillation network is further provided in the multi-task image learning network. The feature map extracted by the feature encoding backbone network is distilled by using the semantic segmentation model as the teacher network, and the output result of the semantic segmentation head is distilled by using the segmentation result obtained by segmenting the input image using the teacher network. The semantic segmentation result is implicitly embedded in the feature descriptor by feature distillation. Feature distillation is to improve the performance and generalization ability of the student model by extracting feature information from a large and complex teacher model and passing it to a small and more computationally efficient student model.

[0053] The traditional way of implicitly learning semantic global information through feature distillation is to distill only the intermediate features. This embodiment specifically uses a large semantic segmentation model as the teacher network. The semantic segmentation branch network, i.e., the student network, is distilled from two aspects: process features and output distillation. The feature map extracted by the teacher network is used Feature map extracted from the backbone network Feature distillation, using the segmentation results output by the teacher network Output of the segmentation head Distillation can enable the descriptor to fully learn the global semantic information, implicitly embed the semantic information into the feature descriptor, and improve the robustness of the descriptor.

[0054] Taking into account the problem of unstable feature point distribution in the feature matching pairs directly obtained by matching with feature descriptors during the relative pose estimation of the drone, that is, some matching pairs belong to the foreground part, and such feature points will change with time in the world coordinates. To address the above-mentioned problems of unstable feature point distribution and foreground interference, this embodiment explicitly embeds semantic global information into the feature descriptor and uses the semantic segmentation results to guide the feature matching task.

[0055] Compared with feature detection, semantic prediction is more complex and requires encoding and aggregation layers. Therefore, in order to embed the model with semantic information, this embodiment combines the knowledge distillation method, learns semantic information implicitly through distillation at the feature layer, enables the feature map of the image matching network to learn the feature map of semantic segmentation, and then directly fuses the feature descriptor and the semantic descriptor, so as to realize implicit and explicit semantic segmentation embedding image description, and uses semantic information to implicitly and explicitly embed into the feature descriptor, effectively improving the robustness of the descriptor, thereby improving the robustness of the image matching task, and using the semantic information output by the semantic segmentation branch to guide the image matching task. For example, the semantic segmentation results can be used to retain feature matching pairs on the background or static objects, and the matching pairs on dynamic objects can be reduced in weight or directly discarded, so that the matching pairs are more conducive to relative pose estimation, reduce the spillover of feature points and feature points of moving foregrounds, and improve the accuracy of image matching. This embodiment uses a large semantic segmentation model to make the descriptor implicitly learn semantic information, and explicitly embeds the semantic information to guide image matching, which can also solve the problem that the image matching accuracy is easily disturbed by the moving foreground during the relative pose estimation of the drone.

[0056] Specifically, the feature descriptor is obtained through multi-task learning ∈ And semantic segmentation results ∈ And the semantic descriptor in the middle of the semantic segmentation branch ∈ In order to explicitly embed the semantic results into the feature descriptor, a feature fusion module based on the attention mechanism is adopted to fuse the semantic segmentation descriptor with the feature descriptor by stacking multiple self-attention layers and cross-attention layers, converting the descriptor into a feature that is easier to match. ∈ , where the self-attention layer is used to calculate the correlation of features at different positions on the same image to obtain global information, and the cross-attention layer enables the model to effectively combine information from two different inputs. The input vector of the attention layer contains the query vector Q, the key-value vector K, and the value vector V. The Q, K, and V of the self-attention layer all come from the same descriptor input, and in the cross-attention layer, Q, K, and V come from different descriptor inputs. The query vector retrieves information in the value vector by calculating the similarity with the key-value vector. The mathematical expression is as follows:

[0057] (1)

[0058] (2)

[0059] in, represents the self-attention layer, represents the cross attention layer, represents the activation function, represents the first input feature, Represents the second input feature.

[0060] In the above formula, the similarity calculation results of the query vector and the key value vector are normalized by the softmax operation, and the output vector is obtained by weighted summing of the normalized similarity scores. Therefore, this process can explicitly embed the feature descriptor and the semantic descriptor into the final descriptor, which is also a process of information transfer. In the process of calculating the dot product of Q and K by the attention mechanism module, the amount of calculation will show a quadratic growth trend as the input increases. Since the input of this embodiment is a dense descriptor at 1 / 4 resolution, in order to reduce the amount of calculation, this embodiment further replaces formula (1) (2) with a linear Transformer to reduce the computational complexity and solve the problem of linear growth trend as the input increases.

[0061] When calculating attention, Q, K, and V are all one-dimensional vectors, so the obtained dense descriptor is flattened. Since the attention mechanism does not focus on the position information of the input, this embodiment adds a position encoding operation after flattening the two-dimensional image into a one-dimensional vector. The encoded vector x is used as the input source and y of the attention mechanism. The calculation expression is:

[0062] (3)

[0063] (4)

[0064] (5)

[0065] Among them, x represents the encoded vector, Indicates that the feature descriptor is obtained through the graph feature description branch. represents the operation of flattening a two-dimensional vector into one dimension, The flattened feature descriptors and semantic descriptors are respectively represents the semantic descriptor obtained by the semantic segmentation branch, Represents the input features, that is, the flattened features, including the flattened semantic descriptors And the flattened feature descriptor , () represents the features after position encoding.

[0066] The semantic descriptor of the flattened positional encoding Or the feature descriptor of the flattened positional encoding As The input features are flattened to encode the semantic descriptor of the position And the feature descriptor of the flattened position encoding Working together The input features of .

[0067] Specifically, in formula (1), When the input is the semantic descriptor of the flattened positional encoding Or the feature descriptor of the flattened positional encoding , in formula (2) The input is the semantic descriptor of the flattened positional encoding And the flattened position encoding feature descriptor Enter at the same time.

[0068] Through the above-mentioned dual-branch network and feature fusion module, the performance of the network can be effectively improved. However, due to the multiple stacking of transformers, the reasoning speed of the network will be affected, and the traditional backbone network has a large number of parameters and the running speed is not fast enough. This embodiment further uses FasterNet as the backbone network. Without changing the structure of the network framework, it can extract feature maps with rich information and output feature maps of different scales. , ,..., ], and finally complete the feature extraction task. FasterNet uses a new partial convolution (PConv), which applies ordinary convolution to only some of the input channels to extract spatial features, and the remaining channels remain unchanged. It can extract rich feature information while reducing redundant calculations and memory accesses, thereby extracting spatial features more efficiently. FasterNet based on Pconv can achieve faster running speeds than other networks on various devices without affecting the accuracy of various visual tasks.

[0069] Step S02: Train the multi-task image learning network. During the training process, a loss function including descriptor loss, semantic segmentation loss, and feature map loss is used to control the training process. After the training is completed, a multi-task image learning model is obtained.

[0070] In the process of training the multi-task image learning network, for image matching for relative pose estimation of drones, the loss function consists of three parts: descriptor loss, semantic segmentation loss, and feature map loss. The training process is controlled from the three aspects of descriptor, semantic segmentation, and feature map, so as to ensure the accuracy of semantic segmentation, feature description, and feature extraction, thereby ensuring the accuracy of image matching. Among them, the descriptor loss can use dualsoftmax loss, and the two input images are extracted through the network to obtain two dense feature maps. and , after extracting the descriptors corresponding to the feature points, the similarity between the feature points on the two images is calculated to obtain the similarity score matrix, as shown in the following formula:

[0071] (6)

[0072] (7)

[0073] (8)

[0074] in, represents the similarity matrix, They represent the i-th feature descriptor extracted from image A and the j-th feature descriptor extracted from image B in the image pair to be matched, respectively. represents the similarity score matrix calculated by dual softmax, represents the real matching relationship, ij is a pair of matches, i, j represent the i, j feature points extracted from image A and image B respectively, t represents the scaling factor, They represent the matching scores from image A to image B and from image B to image A, respectively. represents the index of the matching pair that matches i or j, Indicates Figure A, Represents Figure B.

[0075] In this embodiment, the semantic segmentation loss can adopt the cross entropy loss, as shown in the following formula:

[0076] (9)

[0077] in is the semantic segmentation result output by the teacher network, It is the semantic segmentation result predicted by the semantic segmentation branch of the dual-branch network.

[0078] The feature map loss can be specifically adopted The mathematical expression of loss is shown in the following formula (10), where is the feature map of the i-th layer of the backbone network, is the i-th layer intermediate feature map of the semantic segmentation teacher network.

[0079] (10)

[0080] Therefore, the total loss function is the weighted sum of the above three losses, and its mathematical expression is as follows:

[0081] (11)

[0082] in, represents the semantic segmentation loss, represents the feature map loss, is the feature map of the i-th layer of the backbone network, , as well as They represent weight coefficients respectively.

[0083] Step S03. Obtain images captured by the camera during the flight of the drone, input the acquired images into the trained multi-task image learning model, obtain the semantic segmentation map and dense feature descriptor of the input image, obtain the matching pair relationship between images through feature matching based on the dense feature descriptors of each image, and use the semantic segmentation map to guide the image matching task, and estimate the relative posture relationship between cameras based on the matching pair relationship.

[0084] By using the trained multi-task image learning model, we can extract features from the images captured by the camera during the flight of the drone to obtain semantic segmentation maps and dense feature descriptors. The extracted dense feature descriptors can be used to obtain the matching pair relationship between images through feature matching. In the image matching process, the semantic segmentation map is used to guide the image matching task. For example, according to the semantic segmentation map, feature matching pairs on the background or static objects are retained, and matching pairs on dynamic objects are reduced in weight or directly discarded. This can reduce the spillover of feature points and feature points of moving foregrounds, and improve the accuracy of image matching. After knowing the intrinsic parameters of the cameras of the two images and the relationship between the matching pairs on the two images, the essential matrix E between the images can be estimated. For example, the essential matrix can be decomposed into , where is the rotation matrix R, t is the translation vector, is the antisymmetric matrix of t, from which the relative posture relationship between the cameras is obtained.

[0085] The present invention estimates the relative posture relationship between cameras by estimating the matching pair relationship between images captured by the cameras. Through accurate image matching results, the accuracy of relative posture estimation can be effectively improved.

[0086] In summary, the present invention uses a dual-branch multi-task network to achieve the task goals of semantic segmentation and feature description at the same time, completes the task while learning, changes the network from two to one, and ensures that the performance of the feature description network does not decrease, thereby reducing the number of model parameters and storage space, and solving the problem of storage resource occupation when performing multiple tasks on drones;

[0087] At the same time, the present invention proposes to use semantic segmentation to implicitly and explicitly guide image matching tasks, use a large model to distill the backbone network, so that the model implicitly learns semantic descriptions, and then explicitly combines the semantic descriptor and the feature descriptor through a transformer to embed semantic information in the descriptor, thereby improving the robustness of the feature descriptor and effectively improving the accuracy and robustness of image matching under relative pose estimation for drones. At the same time, directly using semantic information as a guide in the image matching process can also reduce the number of feature points on the moving foreground and the phenomenon of feature point spillover, thereby improving the accuracy of feature description.

[0088] The relative pose estimation device of a drone based on multi-task image matching in this embodiment includes:

[0089] A model building module is used to build a multi-task image learning network. The multi-task image learning network includes a dual-branch network, a feature fusion module, a feature extraction module and an image matching module which are arranged in sequence. The dual-branch network includes a semantic segmentation branch network for performing a semantic segmentation task and a graphic feature description branch network for performing a graphic feature description task. The feature fusion module fuses the semantic segmentation result extracted by the semantic segmentation branch network and the feature descriptor extracted by the graphic feature description branch network to form a fusion feature, so as to explicitly embed the semantic global information into the feature descriptor. The feature extraction module extracts features from the fusion features output by the feature fusion module. The image matching module calculates the matching degree between the two images to be matched according to the features extracted from the two images to be matched.

[0090] The model training module is used to train the multi-task image learning network. During the training process, loss functions including descriptor loss, semantic segmentation loss, and feature map loss are used. After the training is completed, a multi-task image learning model is obtained.

[0091] The pose estimation module is used to obtain images captured by the camera during the flight of the UAV, input the acquired images into the trained multi-task image learning model, obtain the matching pair relationship between the input images, and estimate the relative pose relationship between the cameras based on the matching pair relationship.

[0092] The device for estimating relative pose of a UAV based on multi-task image matching in this embodiment corresponds one to one with the method for estimating relative pose of a UAV based on multi-task image matching described above, and will not be described in detail here.

[0093] This embodiment further provides a computer device, including a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to perform the above method.

[0094] It is understandable that the above method of this embodiment can be executed by a single device, such as a computer or server, etc., and can also be applied to a distributed scenario and completed by multiple devices in cooperation with each other. In the case of a distributed scenario, one of the multiple devices can only execute one or more steps in the above method of this embodiment, and multiple devices interact to complete the above method. The processor can be implemented in the form of a general-purpose CPU, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, etc., for executing related programs to implement the above method of this embodiment. The memory can be implemented in the form of a read-only memory ROM, a random access memory RAM, a static storage device, and a dynamic storage device. The memory can store an operating system and other applications. When the above method of this embodiment is implemented by software or firmware, the relevant program code is stored in the memory and called and executed by the processor.

[0095] This embodiment further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above method is implemented.

[0096] Those skilled in the art should understand that the above-mentioned embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the process Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0097] The above is only a preferred embodiment of the present invention, and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment, it is not intended to limit the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the scope of protection of the technical solution of the present invention.

Claims

1. A method for estimating relative pose of a drone based on multi-task image matching, characterized in that the steps include: Construct a multi-task image learning network, the multi-task image learning network includes a feature extraction module, a dual-branch network and a feature fusion module which are arranged in sequence, the feature extraction module is used to extract a feature map from an input image and provide it to the dual-branch network to learn the common features in the input image, the dual-branch network includes a semantic segmentation branch network for performing a semantic segmentation task on the input image and a graphic feature description branch network for performing a graphic feature description task on the input image, the semantic segmentation branch network extracts a semantic segmentation descriptor and a semantic segmentation map, the graphic feature description branch network extracts a feature descriptor, the feature fusion module is used to fuse the semantic segmentation descriptor extracted by the semantic segmentation branch network and the feature descriptor extracted by the graphic feature description branch network to form a final dense feature descriptor, so as to explicitly embed semantic global information into the feature descriptor, the feature extraction module is a feature encoding backbone network, the semantic segmentation branch network and the graphic feature description branch network in the dual-branch network share the feature encoding backbone network to learn the common features in the input image, and the different features of the two tasks are respectively learned through the feature description head of the graphic feature description branch network and the semantic segmentation head of the semantic segmentation branch network; The multi-task image learning network is trained, and during the training process, a loss function including a descriptor loss, a semantic segmentation loss, and a feature map loss is used to control the training process, and after the training is completed, a multi-task image learning model is obtained; The images captured by the camera of the UAV during flight are obtained, and the obtained images are input into the trained multi-task image learning model to obtain a semantic segmentation map and a dense feature descriptor of the input image, and a matching pair relationship between images is obtained by feature matching according to the dense feature descriptors of each image, and the semantic segmentation map is used to guide the image matching task, and the relative posture relationship between cameras is estimated according to the matching pair relationship.

2. The method for estimating relative pose of unmanned aerial vehicle based on multi-task image matching according to claim 1, characterized in that: The multi-task image learning network is also provided with a feature distillation network, which performs feature distillation on the feature map extracted by the feature encoding backbone network by using a semantic segmentation model as a teacher network, and distills the output result of the semantic segmentation head by using the segmentation result obtained by segmenting the input image using the teacher network.

3. The method for estimating relative pose of unmanned aerial vehicle based on multi-task image matching according to claim 1, characterized in that: In the feature fusion module, the semantic segmentation descriptor is fused with the feature descriptor through multiple stacked self-attention layers and cross-attention layers. The input vector of the self-attention layer includes a query vector Q, a key-value vector K and a value vector V, wherein Q, K and V of the self-attention layer are all from the same descriptor input, and Q, K and V in the cross-attention layer are from different descriptor inputs, respectively. The query vector retrieves information in the value vector by calculating the similarity with the key-value vector, and the calculation expression is: in, represents the self-attention layer, represents the cross attention layer, represents the activation function, represents the first input feature, Represents the second input feature.

4. The method for estimating relative pose of unmanned aerial vehicle based on multi-task image matching according to claim 3, characterized in that: The feature fusion module is also used to perform a position encoding operation after flattening the two-dimensional image into a one-dimensional vector, and use the encoded vector x as the input source and y of the attention mechanism. The calculation expression is: Among them, x represents the encoded vector, represents the feature descriptor obtained by the graphic feature description branch, represents the operation of flattening a two-dimensional vector into one dimension, The feature descriptors and semantic descriptors of the positional encoding are flattened respectively. represents the semantic descriptor obtained by the semantic segmentation branch, Represents the input features, that is, the flattened semantic descriptor or feature descriptor, Express Features after position encoding; The semantic descriptor of the flattened positional encoding Or the feature descriptor of the flattened positional encoding As The input features are flattened to encode the semantic descriptor of the position And the feature descriptor of the flattened position encoding Working together The input features of .

5. The method for estimating relative pose of a UAV based on multi-task image matching according to any one of claims 1 to 4, characterized in that: The total loss function used during training The calculation expression is: in, represents the semantic segmentation loss, represents the feature map loss, is the feature map of the i-th layer of the backbone network, is the i-th layer intermediate feature map of the semantic segmentation teacher network, represents the descriptor loss, , as well as They represent weight coefficients respectively, is the semantic segmentation result output by the teacher network, It is the semantic segmentation result predicted by the dual-branch network.

6. The method for estimating relative pose of unmanned aerial vehicle based on multi-task image matching according to claim 5, characterized in that: Descriptor loss The calculation expression is: in, represents the similarity matrix, They represent the i-th feature descriptor extracted from image A and the j-th feature descriptor extracted from image B in the image pair to be matched, respectively. represents the similarity score matrix calculated by dual softmax, represents the true matching relationship, i, j represents the i-th and j-th feature points extracted from image A and image B respectively, t represents the scaling factor, They represent the matching scores from image A to image B and from image B to image A, respectively. represents the index of the matching pair that matches i or j, Indicates Figure A, Represents Figure B.

7. The method for estimating relative pose of a UAV based on multi-task image matching according to any one of claims 1 to 4, characterized in that: The using the semantic segmentation map to guide the image matching task includes: retaining feature matching pairs on the background or static objects according to the semantic segmentation map, and reducing the weight of or directly discarding matching pairs on dynamic objects.

8. A device for estimating relative pose of unmanned aerial vehicle based on multi-task image matching, characterized in that: include: A model construction module, for constructing a multi-task image learning network, the multi-task image learning network includes a feature extraction module, a dual-branch network and a feature fusion module which are arranged in sequence, the feature extraction module is used to extract a feature map from an input image and provide it to the dual-branch network to learn the common features in the input image, the dual-branch network includes a semantic segmentation branch network for performing a semantic segmentation task on the input image and a graphic feature description branch network for performing a graphic feature description task on the input image, the semantic segmentation branch network extracts a semantic segmentation descriptor and a semantic segmentation map, the graphic feature description branch network extracts a feature descriptor, and the feature fusion module is used to fuse the semantic segmentation descriptor extracted by the semantic segmentation branch network and the feature descriptor extracted by the graphic feature description branch network to form a final dense feature descriptor, so as to explicitly embed semantic global information into the feature descriptor; A model training module is used to train the multi-task image learning network. During the training process, a loss function including a descriptor loss, a semantic segmentation loss, and a feature map loss is used. After the training is completed, a multi-task image learning model is obtained. The pose estimation module is used to obtain images captured by the camera during the flight of the UAV, input the acquired images into the trained multi-task image learning model, obtain a semantic segmentation map and a dense feature descriptor of the input image, obtain a matching pair relationship between images through feature matching based on the semantic segmentation map and dense feature descriptor of each image, and estimate the relative pose relationship between cameras based on the matching pair relationship.

9. A computer device comprising a processor and a memory, wherein the memory is used to store a computer program, wherein: The processor is configured to execute the computer program to perform the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Anti-transmission distortion unmanned aerial vehicle inclined image semantic information extraction method and equipment

    CN115861635A

  • Unmanned aerial vehicle image semantic segmentation method and device, equipment and storage medium

    CN116704382A