Unmanned aerial vehicle image retrieval method and device, terminal and storage medium
By generating training samples and iteratively training image search models, combining three-dimensional reconstruction and index structure, the problem of low retrieval accuracy of drone image matching pairs is solved, and higher-precision image matching pair retrieval is achieved.
Patent Information
- Application Number
- CN202510450963.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-08-26
AI Technical Summary
In the prior art, the accuracy of drone image matching pairing is low, mainly because there are significant differences in shooting perspective, observation scale, target details and background content between drone images and common data sets.
By generating training samples, the image search model is iteratively trained, and the global feature descriptor vector is output using the trained image search model, and the index structure is constructed and searched, including feature extraction, three-dimensional reconstruction, loss function optimization and the use of index structures to improve the matching pair search accuracy.
It effectively improves the accuracy of drone image matching pairing and improves the matching accuracy of image search models in different perspectives and backgrounds.
Smart Images

Figure CN120541255A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a drone image retrieval method, device, terminal and storage medium. Background Art
[0002] In the field of image retrieval, deep global features can capture the global context and semantic information of images by automatically learning feature representations from data, and can still maintain high discrimination in weak texture scenes. Therefore, they have been widely used and can be used for tasks such as instance retrieval, landmark retrieval, and visual positioning.
[0003] However, in the task of drone image pairing retrieval, deep global features have certain limitations: drone images and the datasets used in the above tasks have significant differences in shooting perspective, observation scale, target details, and background content. As a result, when directly using existing pre-trained models for drone image pairing retrieval, the accuracy is often low.
[0004] Therefore, the existing technology has defects and needs to be improved and developed. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a drone image retrieval method, device, terminal and storage medium in response to the above-mentioned defects of the prior art, aiming to solve the problem of low accuracy of drone image matching retrieval in the prior art.
[0006] The technical solutions adopted by the present invention to solve the technical problems are as follows:
[0007] In a first aspect, an embodiment of the present invention provides a method for retrieving drone images, the method comprising:
[0008] Acquire and generate training samples based on the first drone image set;
[0009] Iteratively training the image retrieval model to be trained using the training samples, calculating the loss function and reversely updating the model parameters in each round of training until an end condition is met, thereby obtaining a trained image retrieval model;
[0010] Obtain a set of drone images to be retrieved, input each drone image in the set of drone images to be retrieved into the trained image retrieval model in sequence, and obtain a global feature descriptor vector corresponding to each drone image in the set of drone images to be retrieved;
[0011] An index structure is constructed based on all the global feature descriptor vectors, and each global feature descriptor vector is sequentially input into the index structure for search, to obtain a pairing result for each drone image in the drone image set to be retrieved.
[0012] In one embodiment, obtaining and generating training samples based on a first drone image set includes:
[0013] Acquire a first drone image set, where the first drone image set includes drone image subsets of several scenes, and each drone image subset includes several drone images having a geometrically overlapping relationship;
[0014] Building a retrieval dataset based on the first drone image collection;
[0015] A training sample is generated based on the search data set.
[0016] In one embodiment, constructing a search dataset based on the first drone image set includes:
[0017] Performing feature extraction on all the drone images to obtain local features corresponding to each drone image, and performing retrieval based on all local features of each drone image subset to obtain a number of potential image matching pairs corresponding to each scene;
[0018] Perform 3D reconstruction based on all potential image matching pairs corresponding to each scene, and obtain the 3D reconstruction model corresponding to each scene and the 3D point trajectory connecting the same-name points in the images;
[0019] Calculating the geometric similarity of each potential image matching pair based on all the three-dimensional point trajectories in each scene, and selecting potential image matching pairs that meet a threshold condition as overlapping image pairs;
[0020] Taking any UAV image in each overlapping image pair as an anchor image, generating a corresponding overlapping image list, wherein the overlapping image list records the UAV images associated and matched with the anchor image in all overlapping image pairs;
[0021] Aggregate each anchor image and its corresponding overlapping image list to form a retrieval dataset.
[0022] In one embodiment, performing 3D reconstruction based on all potential image matching pairs corresponding to each scene to obtain a 3D reconstructed model corresponding to each scene includes:
[0023] Performing guided feature matching on all potential image matching pairs corresponding to each scene to obtain a plurality of matching point distribution maps, each of the matching point distribution maps being used to display the distribution of all matching points corresponding to the potential image matching pair;
[0024] Based on the distribution graph of all matching points corresponding to each scene, an undirected weighted graph corresponding to each scene is constructed;
[0025] Dividing each of the undirected weighted graphs into a plurality of subclusters, performing incremental SfM reconstruction on all subclusters in parallel, and generating a submodel corresponding to each subcluster;
[0026] After iteratively merging all sub-models of each scene, a global bundle adjustment optimization is performed to obtain the 3D reconstructed model of each scene.
[0027] In one embodiment, generating training samples based on the search data set includes:
[0028] Randomly select several scenes from all scenes covered by the retrieval dataset, and randomly select an anchor image from each scene as an anchor sample;
[0029] For each anchor sample, several drone images with geometric similarity greater than a preset threshold are selected from the overlapping image list corresponding to each anchor sample as positive samples, and positive samples in other scenes are used as negative samples of the anchor sample;
[0030] The training samples are composed of all anchor samples, as well as the positive samples and negative samples corresponding to each anchor sample.
[0031] In one embodiment, the training method of the image retrieval model includes:
[0032] Dividing the training samples into a plurality of batches, each batch having a preset number of samples;
[0033] In each round of training, all samples in each batch are input into the image retrieval model to be trained for processing, and the global feature descriptor vector of all samples in each batch is obtained. The loss function is calculated based on the global feature descriptor vector of all samples in each batch, and the model parameters are updated through backpropagation.
[0034] Iteratively training the image retrieval network model until a preset end condition is reached to obtain a trained image retrieval network;
[0035] The loss function is the sum of a first loss function and a second loss function. The first loss function is used to optimize the image retrieval model's ability to distinguish between positive samples and negative samples, and the second loss function is used to optimize the image retrieval model's ability to distinguish overlapping images.
[0036] In one embodiment, the image retrieval network model includes a backbone network and a NetVLAD network; each drone image in the set of drone images to be retrieved is sequentially input into the trained image retrieval model to obtain a global feature descriptor vector corresponding to each drone image, including:
[0037] Inputting each drone image in the set of drone images to be retrieved into the backbone network in the trained image retrieval model;
[0038] The backbone network extracts features from each drone image to obtain a depth feature map corresponding to each drone image;
[0039] Each of the depth feature maps is input into the NetVLAD network and processed by the NetVLAD network to obtain a global feature descriptor vector corresponding to each drone image.
[0040] In a second aspect, an embodiment of the present invention further provides a drone image retrieval device, the device comprising:
[0041] A training sample generation module, configured to obtain and generate training samples based on a first UAV image set;
[0042] A model training module is used to iteratively train the image retrieval model to be trained using the training samples, calculate the loss function and reversely update the model parameters in each round of training until the end condition is met, thereby obtaining a trained image retrieval model;
[0043] A vector generation module is used to obtain a set of drone images to be retrieved, input each drone image in the set of drone images to be retrieved into the trained image retrieval model in sequence, and obtain a global feature descriptor vector corresponding to each drone image;
[0044] The retrieval module is used to construct an index structure based on all the global feature descriptor vectors, and input each global feature descriptor vector into the index structure in turn for search, so as to obtain a pairing result for each drone image in the drone image set to be retrieved.
[0045] In a third aspect, an embodiment of the present invention further provides a terminal comprising: a memory, a processor, and a drone image retrieval program stored on the memory and runnable on the processor, wherein the drone image retrieval program implements the steps of the drone image retrieval method described above when executed by the processor.
[0046] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a drone image retrieval program, and the drone image retrieval program can be executed to implement the steps of the drone image retrieval method as described above.
[0047] The beneficial effects of the present invention are as follows: the present invention generates training samples; iteratively trains the image retrieval model to be trained to obtain a trained image retrieval model; obtains a set of drone images to be retrieved, and sequentially inputs each drone image in the set of drone images to be retrieved into the trained image retrieval model to obtain a global feature descriptor vector corresponding to each drone image in the set of drone images to be retrieved; constructs an index structure based on all global feature descriptor vectors, and sequentially inputs each global feature descriptor vector into the index structure for search to obtain a pairing result for each drone image in the set of drone images to be retrieved. The present invention utilizes the global feature descriptor vectors output by the trained image retrieval model to construct an index structure, and utilizes the index structure for retrieval, which can effectively improve the accuracy of drone image matching retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a flow chart of a preferred embodiment of the drone image retrieval method of the present invention.
[0049] Figure 2 This is a workflow diagram for performing three-dimensional reconstruction on potential image matching pairs in the present invention.
[0050] Figure 3 Schematic diagram of three negative samples involved in the present invention.
[0051] Figure 4 is a schematic diagram of samples being misordered among triplets.
[0052] Figure 5 This is a diagram for optimizing a sorted list consisting of a set of positive samples and a set of negative samples.
[0053] Figure 6 It is a structural diagram of a preferred embodiment of the drone image retrieval device in the present invention.
[0054] Figure 7 It is a block diagram of the terminal principle of the present invention. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0056] In the field of image retrieval, deep global features can capture the global context and semantic information of images by automatically learning feature representations from data, and can still maintain high discrimination in weak texture scenes. Therefore, they have been widely used and can be used for tasks such as instance retrieval, landmark retrieval, and visual positioning.
[0057] However, in the task of drone image pairing retrieval, deep global features have certain limitations: drone images and the datasets used in the above tasks have significant differences in shooting perspective, observation scale, target details, and background content. As a result, when directly using existing pre-trained models for drone image pairing retrieval, the accuracy is often low.
[0058] In response to the above-mentioned deficiencies in the prior art, the present invention provides a method, device, terminal, and storage medium for drone image retrieval, the method comprising: generating training samples; iteratively training an image retrieval model to be trained to obtain a trained image retrieval model; obtaining a set of drone images to be retrieved, inputting each drone image in the set of drone images to be retrieved into the trained image retrieval model in sequence, and obtaining a global feature descriptor vector corresponding to each drone image in the set of drone images to be retrieved; constructing an index structure based on all global feature descriptor vectors, and inputting each global feature descriptor vector into the index structure in sequence for search, and obtaining a pairing result for each drone image in the set of drone images to be retrieved. The present invention utilizes the global feature descriptor vectors output by the trained image retrieval model to construct an index structure, and utilizes the index structure for retrieval, which can effectively improve the accuracy of drone image matching retrieval.
[0059] See Figure 1 The drone image retrieval method according to an embodiment of the present invention includes the following steps:
[0060] Step S100: Acquire and generate training samples based on a first drone image set.
[0061] Specifically, a first drone image set is first obtained. The first drone image set contains drone image subsets of several scenes, and each drone image subset contains several drone images with geometric overlapping relationships. The number of scenes of the present invention is 30, and the scenes cover rural farmland, urban blocks, river corridors, mountainous areas, building complexes, and mixed scenes. Each scene contains 100 to 4,000 images with large geometric overlap. These images are taken by drones from multiple scales and perspectives. The first drone image set contains a total of 21,622 drone images. After obtaining the first drone image set, a retrieval dataset is constructed based on the first drone image set, and training samples are generated based on the retrieval dataset.
[0062] In one implementation, constructing a retrieval dataset based on the first drone image set includes:
[0063] Performing feature extraction on all the drone images to obtain local features corresponding to each drone image, and performing retrieval based on all local features of each drone image subset to obtain a number of potential image matching pairs corresponding to each scene;
[0064] Perform 3D reconstruction based on all potential image matching pairs corresponding to each scene, and obtain the 3D reconstruction model corresponding to each scene and the 3D point trajectory connecting the same-name points in the images;
[0065] Calculating the geometric similarity of each potential image matching pair based on all the three-dimensional point trajectories in each scene, and selecting potential image matching pairs that meet a threshold condition as overlapping image pairs;
[0066] Taking any UAV image in each overlapping image pair as an anchor image, generating a corresponding overlapping image list, wherein the overlapping image list records the UAV images associated and matched with the anchor image in all overlapping image pairs;
[0067] Aggregate each anchor image and its corresponding overlapping image list to form a retrieval dataset.
[0068] Specifically, since the first drone image set contains large-scale disordered drone images, the present invention adopts a parallel three-dimensional reconstruction method to improve the efficiency of three-dimensional reconstruction. Utilizing all local features of each drone image subset, a unified visual vocabulary is generated through clustering, the local features of each drone image are mapped to the visual vocabulary, a bow vector is generated, and then the similarity between the bow vectors is calculated, and several potential image matching pairs are obtained after sorting. Then, all potential image matching pairs corresponding to each scene are three-dimensionally reconstructed in parallel to obtain a three-dimensional reconstruction model corresponding to each scene. During the reconstruction, a three-dimensional point trajectory connecting the image points of the same name is generated. The geometric similarity of two potential image matching pairs is calculated based on the three-dimensional point trajectory. The calculation formula is as follows:
[0069]
[0070] Where P(i) represents the 3D point observed in image i, P(j) represents the 3D point observed in image j, and P(i)∩P(j) represents the common 3D point trajectory between images i and j. Then, potential image matching pairs with a geometric similarity greater than 0.2 in the 3D reconstruction model are selected as overlapping image pairs.
[0071] Traditional instance or landmark retrieval datasets, such as Oxford5k, Paris6k, GLDv1 and GLDv2, usually contain strong semantic annotations at the instance level or landmark level, focusing on single salient objects in the image. In contrast, the goal of matching pair retrieval is to find image pairs that have high spatial overlap and are likely to match, and it focuses more on geometric context and spatial relationships rather than specific object features. The present invention can effectively filter out almost all unmatched images through SfM three-dimensional reconstruction. If an image pair has significant overlap in three-dimensional space, its matching feature points will form a continuous three-dimensional point trajectory, that is, the three-dimensional point trajectory can accurately reflect the overlapping relationship of the image pair. Therefore, the present invention uses SfM three-dimensional reconstruction to realize automatic annotation of drone image pairs (that is, automatically screening out overlapping image pairs).
[0072] In one implementation, performing 3D reconstruction based on all potential image matching pairs corresponding to each scene to obtain a 3D reconstruction model corresponding to each scene includes:
[0073] Performing guided feature matching on all potential image matching pairs corresponding to each scene to obtain a plurality of matching point distribution maps, each of the matching point distribution maps being used to display the distribution of all matching points corresponding to the potential image matching pair;
[0074] Based on the distribution graph of all matching points corresponding to each scene, an undirected weighted graph corresponding to each scene is constructed;
[0075] Dividing each of the undirected weighted graphs into a plurality of subclusters, performing incremental SfM reconstruction on all subclusters in parallel, and generating a submodel corresponding to each subcluster;
[0076] After iteratively merging all sub-models of each scene, a global bundle adjustment optimization is performed to obtain the 3D reconstructed model of each scene.
[0077] Specifically, the present invention adopts a divide-and-conquer strategy in the 3D reconstruction process. The core idea is to divide the entire scene into small and compact subclusters that can be reconstructed efficiently and accurately. After feature-guided matching, a matching point distribution map is obtained for each scene. Then, a scene graph G is constructed for each scene based on the matching point distribution map. The graph is represented as an undirected weighted graph G = (V, E), where the vertex set V = {v i} represents the image {i i}, edge set E={e ij} means connecting the image {i i} and image {i j} potential image matching pairs {p ij}. In addition, the edge ij is assigned a weight value w ij, is used to measure the importance of the connection, which will affect the performance of SfM 3D reconstruction. In the present invention, the weight w ij It is calculated by the number of feature matches and their distribution on the image plane, and the calculation formula is as follows:
[0078] w ij =R ew ×w inlier +(1-R ew )×w overlap ;
[0079] In the formula, R ew is the preset weight ratio, w ivlier is the weight of the number of feature matches, w overlap is the weight of the distribution. inlier The calculation formula is as follows: w overlap The calculation formula is as follows: Where N inlier is a potential image matching pair p ij The number of matching internal points, N maxinlier is the largest N among all potential image matching pairs inlier 、A i It is an image i The plane area, A j It is an image j Plane area, CH i is a potential image matching pair p ij In the image i The convex hull area on CH j is a potential image matching pair p ij In the image i j The convex hull area on the image. Then, scene clustering is performed, and the Normalized Cut (NC) algorithm is used to divide the created scene graph into subclusters according to the weights. The small-scale subscenes corresponding to the subclusters have strong internal connections. Then, the incremental SfM 3D reconstruction of each subscene is performed in parallel to generate submodels sorted by the number of common 3D points, and merge them according to the sorting. Finally, after all submodels are iteratively merged, a global bundle adjustment optimization is performed to obtain the final reconstructed model. The workflow diagram for 3D reconstruction of potential image matching pairs is shown in the figure below. Figure 2 shown.
[0080] In one implementation, generating training samples based on the search data set includes:
[0081] Randomly select several scenes from all scenes covered by the retrieval dataset, and randomly select an anchor image from each scene as an anchor sample;
[0082] For each anchor sample, several drone images with geometric similarity greater than a preset threshold are selected from the overlapping image list corresponding to each anchor sample as positive samples, and positive samples in other scenes are used as negative samples of the anchor sample;
[0083] The training samples are composed of all anchor samples, as well as the positive samples and negative samples corresponding to each anchor sample.
[0084] Specifically, the present invention uses a retrieval dataset to construct training samples for image retrieval model training. The training samples include anchor samples q, as well as positive samples M(q) and negative samples N(q) corresponding to each anchor sample. In the prior art, when constructing global hard negative samples, a pre-trained model is usually used to first extract descriptors for all images in the dataset. Then, images with the minimum descriptor distance to the anchor sample q are selected from other scenes as hard negative samples. This process can be expressed as:
[0085] N(q)=random{n:argmin||f p -f n ||}; where f p Describes the descriptor of q, f n Descriptors representing images from other scenes. After a fixed number of training steps, the descriptors for all images are updated and the mining step is repeated. This means that the global hard negative sample mining strategy requires iteratively updating the descriptors of all images in the database and calculating the descriptor distance, which is expensive in terms of resources and time.
[0086] The present invention optimizes the training sample generation process and combines it with a batch non-trivial sample mining strategy to improve training efficiency. Specifically: First, N different scenes are randomly selected and an image is selected from each scene to form an anchor sample set {q i |i=1,…,N}; Then, randomly select the drone image from the overlapping image list {p} corresponding to each anchor sample q. This process can be expressed as:
[0087] M(q)=random{p:GS(q,p)>ε};where ε is the threshold value, which is 0.2. GS(q,p) is the geometric similarity between image q and image p. In this way, M positive samples are selected for each anchor sample. And sorted by geometric similarity, a training batch consists of N×(M+1) images, and the anchor sample q i The negative samples of are the positive samples of other anchor samples, which can be expressed as:
[0088]
[0089] Since this process ignores image descriptors, there will be many trivial samples with zero loss during training. These trivial samples weaken the contribution of non-trivial samples in gradient averaging, so they are eliminated in the subsequent loss calculation. This strategy is called batch non-trivial sample mining.
[0090] See Figure 1 The drone image retrieval method according to the embodiment of the present invention further includes the following steps:
[0091] Step S200: Iteratively train the image retrieval model to be trained using the training samples, calculate the loss function and reversely update the model parameters in each round of training until the end condition is met, and obtain a trained image retrieval model.
[0092] Specifically, the training method of the image retrieval model includes: dividing the training samples into several batches, each batch having a preset number of samples; in each round of training, inputting all samples of each batch into the image retrieval model to be trained for processing in turn, obtaining the global feature descriptor vector of all samples in each batch, calculating the loss function based on the global feature descriptor vector of all samples in each batch, and updating the model parameters through back propagation; iteratively performing the training process of the image retrieval network model until a preset end condition is reached to obtain a trained image retrieval network; wherein the loss function is the sum of a first loss function and a second loss function, the first loss function is used to optimize the image retrieval model's ability to distinguish between positive samples and negative samples, and the second loss function is used to optimize the image retrieval model's ability to distinguish between overlapping images.
[0093] In the prior art, the commonly used loss function for image retrieval is triplet loss, which is expressed as:
[0094] L triplet =max(0,D(A,P)-D(A,N)+m);
[0095] Where D(A,P) represents the distance between the descriptor of anchor sample A and the descriptor of positive sample P, D(A,N) represents the distance between the descriptor of anchor sample A and the descriptor of negative sample N, and m is a predefined boundary used to control the distance difference between positive and negative samples. It only focuses on the local similarity structure of positive and negative sample pairs and ignores the global similarity structure, resulting in limited discrimination ability of deep global features. For the selected anchor samples and positive samples, there are three possibilities for negative samples, namely simple negative samples, semi-hard negative samples and hard negative samples, such as Figure 3As shown. Simple negative samples already meet the goal of triplet loss, and the loss value during training is zero, which will not affect the gradient update. If there are too many simple negative samples in the training set, the model cannot fully learn the discriminative feature representation. Hard negative samples refer to negative samples whose distance to the anchor sample is less than or close to the distance between the positive sample and the anchor sample. If too many hard negative samples are used, the model may fall into a local optimum or training oscillation. Semi-hard negative samples refer to negative samples whose distance to the anchor sample is greater than the distance between the positive sample and the anchor sample, but does not exceed the boundary m. Semi-hard negative samples are the most commonly used type of negative samples in training, which can improve the discriminative ability of the model while ensuring training stability. Therefore, during the training process, it is necessary to continuously mine semi-hard triplets for training. As the size of the dataset grows, the number of triplets grows exponentially, and the cost of mining semi-hard triplets also increases. Since the triplet loss only optimizes the local ordering within a single triple, this may cause the samples to be sorted correctly within the triplet, but incorrectly sorted between triplets, such as Figure 4 shown.
[0096] To solve the above problems, the present invention constructs a loss function including the first loss function and the second loss function, which is expressed as: L rankedlist =L1+L2; where L1 is the first loss function and L2 is the second loss function. The expression of L1 is as follows:
[0097]
[0098] Where, [*] + Represents the function max(0,*), which automatically filters out trivial samples. P and N represent the positive and negative sample sets, respectively. α represents the negative sample boundary, and α-m represents the positive sample boundary. By optimizing L1, the positive sample set is constrained to a hypersphere with a radius of α-m, while the negative sample set is pushed outside the hypersphere with a radius of α. || represents the number of elements in the set. D(A,n) represents the distance between negative sample n and anchor sample A in feature space, and D(A,p) represents the distance between positive sample p and anchor sample A in feature space. The expression for L2 is as follows:
[0099] L2 represents the optimization of the internal sorting of the positive sample set, and D(A, p+1) represents the distance between positive sample p+1 and anchor sample A in feature space. Specifically, the positive sample set P is sorted by geometric similarity, where p is any positive sample in the positive sample set and n is any negative sample in the negative sample set. By optimizing the L2 model, we can better distinguish the differences between overlapping images and improve the discriminative ability of deep global features.
[0100] This loss function can be used to directly optimize the sorted list consisting of the positive sample set and the negative sample set, rather than optimizing the sorting of each sample pair separately, such as Figure 5 As shown in the figure, this not only avoids the resource- and time-consuming semi-hard triple mining, but also ensures the correct ordering of triplets. Furthermore, by optimizing the internal ordering of positive samples, the model can capture the differences between different positive samples, thereby improving the discriminative power of deep global features.
[0101] See Figure 1 The drone image retrieval method according to the embodiment of the present invention further includes the following steps:
[0102] Step S300: Obtain a set of drone images to be retrieved, input each drone image in the set of drone images to be retrieved into the trained image retrieval model in sequence, and obtain a global feature descriptor vector corresponding to each drone image in the set of drone images to be retrieved.
[0103] Specifically, once the trained image retrieval model is obtained, it is used to extract features and obtain the global feature descriptor vector corresponding to each drone image. Because the model undergoes multiple iterations during training, it can capture the differences between different positive samples and accurately extract deep global features.
[0104] In one implementation, the image retrieval network model includes a backbone network and a NetVLAD network; each drone image in the set of drone images to be retrieved is sequentially input into the trained image retrieval model to obtain a global feature descriptor vector corresponding to each drone image, including:
[0105] Inputting each drone image in the set of drone images to be retrieved into the backbone network in the trained image retrieval model;
[0106] The backbone network extracts features from each drone image to obtain a depth feature map corresponding to each drone image;
[0107] Each of the depth feature maps is input into the NetVLAD network and processed by the NetVLAD network to obtain a global feature descriptor vector corresponding to each drone image.
[0108] Specifically, the backbone network of the present invention can be a VGG-16 network. i,k This makes it impossible to embed it into CNN for end-to-end training. In this paper, a differentiable NetVLAD layer is used to aggregate the deep feature maps. The new allocation function is as follows:
[0109]
[0110] Where x irepresents the local feature of the i-th spatial position in the depth feature map, c k represents the kth cluster center, α is used to control the degree of attenuation of the distribution value with distance. A larger α means a harder distribution. k′ is any cluster center, that is, the numerator in the formula is the calculation of the local feature x i To the kth cluster center c k The negative exponential distance, the denominator is calculated for the local feature x i The negative exponential distances of all cluster centers are summed. Unlike VLAD, which assigns local features to only one cluster center, the distribution values of local features to different cluster centers in NetVLAD are determined by the distances between them. By expanding the formula of the distribution function and dividing both the numerator and the denominator by The following formula is obtained:
[0111]
[0112] where w k =2αc k 、b k =-α‖c k ‖ 2 The descriptors extracted by the NetVLAD network are calculated according to the following formula:
[0113]
[0114] Where j represents the jth dimension, x i,j Represents the local feature x i The value of the j-th dimension, c k,j Represents the cluster center c k The value of the jth dimension. Parameter w k 、b k and c k is trainable, which means that the cluster centers and the distribution function are learnable. However, in practical applications, in order to speed up convergence, the present invention adopts a simpler distribution function, that is, b k Fixed to 0 and w k Initialized to αc k / ‖c k ‖, which can effectively convert drone images into global feature descriptor vectors.
[0115] See Figure 1 The drone image retrieval method according to the embodiment of the present invention further includes the following steps:
[0116] Step S400: construct an index structure based on all the global feature descriptor vectors, and input each global feature descriptor vector into the index structure in turn for search, to obtain a pairing result for each drone image in the drone image set to be retrieved.
[0117] Specifically, in the prior art, after obtaining the global feature descriptor vector of the image, the similarity between the vectors is usually determined by directly calculating the Euclidean distance between the vectors, thereby reflecting the similarity between the images. However, for drone images with a large number of images, the time consumption of pairwise comparison is very large. To this end, the idea of the present invention is to use an efficient index structure to implement ANN (Approximate Nearest Neighbor) retrieval of vectors. HNSW (Hierarchical Navigable Small World) is a graph-based approximate nearest neighbor retrieval algorithm, which uses a hierarchical structure graph to construct a vector index graph. The present invention uses the global feature descriptor vector of the image as the vertex set of the graph to construct an index graph to achieve efficient retrieval.
[0118] Before searching, all calculated global feature descriptor vectors need to be established as an index structure, adding only one vector vertex at a time. The specific steps are as follows:
[0119] Step 1: Calculate the newly added vertex V i The highest layer number L projected to is calculated as follows:
[0120]
[0121] Among them, unif(0,1) is a random number between 0 and 1, m L is a normalization factor to ensure that the calculated number of layers meets actual needs.
[0122] Step 2: If there is no vertex in the index structure, add vertex V directly to layer i, i = 1, 2, ..., L i Otherwise: traverse all vertices in the top level H to find the distance vertex V i The nearest vertex V i ' as the input of H-1 layer.
[0123] Step 3: Traverse all vertices in the next layer to find the distance vertex V i 'The nearest vertex V i ” as the input of the next layer.
[0124] Step 4: Repeat step 3 until the input vertex V of the L layer is found L , traverse vertex V in layer L L Neighbors are found at distance V i The nearest M nodes, put the vertex V i Insert layer L to establish connections with these M nodes, making them neighbors.
[0125] Step 5: Take the M nodes found in the previous step as the input of the next layer, and traverse the M nodes in the next layer to find the distance V i The nearest M nodes, put the vertex V i Insert this layer to establish connections with the newly found M nodes, making them neighbors.
[0126] Step 6. Repeat step 5 until all layers are completed with vertex V i Insert.
[0127] Where M is a manually set parameter. In the present invention, M can be set freely. Generally, the larger M is, the more complex the index structure is, the longer the indexing time is, and the higher the accuracy is; the smaller M is, the simpler the index structure is, the shorter the indexing time is, and the lower the accuracy is.
[0128] Using the index structure constructed in the previous step, input the global feature descriptor vectors constructed for each image in sequence to find adjacent vectors. The specific steps are as follows:
[0129] Step 1: Input the vector vertex V to be queried i , traverse all vertices in the top level H and find the distance to the vertex V i The nearest vertex V i ' as the input of H-1 layer.
[0130] Step 2: Traverse the vertex V in the next layer i 'Neighbors find the distance from vertex V i 'The nearest vertex V i ” as the input of the next layer.
[0131] Step 3: Repeat step 2 until the bottom-level input vertex V1 is found, and then traverse all the vertices at the bottom level to find the N nearest to vertex V1. i Vertices, complete the search.
[0132] Based on the above steps, the image retrieved from each image and the image are combined into an image pair as a matching pair. That is, image i retrieves image i' j ,j=1,2,...,N i , then the image pair (i,i' j ),j=1,2,...,N i as a matching pair.
[0133] In summary, the present invention generates training samples; iteratively trains the image retrieval model to be trained to obtain a trained image retrieval model; obtains a set of drone images to be retrieved, and sequentially inputs each drone image in the set of drone images to be retrieved into the trained image retrieval model to obtain a global feature descriptor vector corresponding to each drone image in the set of drone images to be retrieved; constructs an index structure based on all global feature descriptor vectors, and sequentially inputs each global feature descriptor vector into the index structure for search, to obtain a pairing result for each drone image in the set of drone images to be retrieved. The present invention utilizes the global feature descriptor vectors output by the trained image retrieval model to construct an index structure, and utilizes the index structure for retrieval, which can effectively improve the accuracy of drone image matching retrieval.
[0134] In one embodiment, if Figure 6 As shown, based on the above-mentioned drone image retrieval method, the present invention also provides a drone image retrieval device, which includes:
[0135] A training sample generation module, configured to obtain and generate training samples based on a first UAV image set;
[0136] A model training module is used to iteratively train the image retrieval model to be trained using the training samples, calculate the loss function and reversely update the model parameters in each round of training until the end condition is met, thereby obtaining a trained image retrieval model;
[0137] A vector generation module is used to obtain a set of drone images to be retrieved, input each drone image in the set of drone images to be retrieved into the trained image retrieval model in sequence, and obtain a global feature descriptor vector corresponding to each drone image;
[0138] The retrieval module is used to construct an index structure based on all the global feature descriptor vectors, and input each global feature descriptor vector into the index structure in turn for search, so as to obtain a pairing result for each drone image in the drone image set to be retrieved.
[0139] It should be noted that the above explanation of the embodiment of the drone image retrieval method is also applicable to the drone image retrieval device of this embodiment and will not be repeated here.
[0140] Based on the above embodiment, the present invention further provides a terminal, whose structural diagram can be as follows: Figure 7As shown. The terminal includes a processor, a memory, a network interface and a display screen connected via a device bus. The processor of the terminal is used to provide computing and control capabilities. The memory of the terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating device and a drone image retrieval program. The internal memory provides an environment for the operation of the operating device and the drone image retrieval program in the non-volatile storage medium. The network interface of the terminal is used to communicate with an external terminal via a network connection. When the drone image retrieval program is executed by the processor, the steps of any one of the above-mentioned drone image retrieval methods are implemented. The display screen of the terminal can be a liquid crystal display or an electronic ink display.
[0141] Those skilled in the art will understand that Figure 7 The structural schematic diagram shown in the figure is only a schematic diagram of a partial structure related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0142] In one embodiment, a terminal is provided, which includes a memory, a processor, and a drone image retrieval program stored in the memory and runnable on the processor. When the drone image retrieval program is executed by the processor, the steps of any one of the drone image retrieval methods provided in the embodiments of the present invention are implemented.
[0143] An embodiment of the present invention further provides a computer-readable storage medium, on which a drone image retrieval program is stored. When the drone image retrieval program is executed by a processor, the steps of any one of the drone image retrieval methods provided in an embodiment of the present invention are implemented.
[0144] It should be understood that the sequence numbers of the steps in the above embodiments do not imply a specific order of execution; the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0145] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the above-mentioned device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0146] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0147] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0148] In the embodiments provided by the present invention, it should be understood that the disclosed apparatus / terminal device and method can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units described above is merely a logical functional division. In actual implementation, other division methods may be used. For example, multiple units or components may be combined or integrated into another device, or some features may be omitted or not implemented.
[0149] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A drone image retrieval method, characterized in that: The method comprises: Acquire and generate training samples based on the first drone image set; Iteratively training the image retrieval model to be trained using the training samples, calculating the loss function and reversely updating the model parameters in each round of training until an end condition is met, thereby obtaining a trained image retrieval model; Obtain a set of drone images to be retrieved, input each drone image in the set of drone images to be retrieved into the trained image retrieval model in sequence, and obtain a global feature descriptor vector corresponding to each drone image in the set of drone images to be retrieved; An index structure is constructed based on all the global feature descriptor vectors, and each global feature descriptor vector is sequentially input into the index structure for search, to obtain a pairing result for each drone image in the drone image set to be retrieved.
2. The drone image retrieval method according to claim 1, characterized in that: Obtain and generate training samples based on the first drone image set, including: Acquire a first drone image set, where the first drone image set includes drone image subsets of several scenes, and each drone image subset includes several drone images having a geometrically overlapping relationship; Building a retrieval dataset based on the first drone image collection; A training sample is generated based on the search data set.
3. The drone image retrieval method according to claim 2, characterized in that: A retrieval dataset is constructed based on the first drone image collection, including: Performing feature extraction on all the drone images to obtain local features corresponding to each drone image, and performing retrieval based on all local features of each drone image subset to obtain a number of potential image matching pairs corresponding to each scene; Perform 3D reconstruction based on all potential image matching pairs corresponding to each scene, and obtain the 3D reconstruction model corresponding to each scene and the 3D point trajectory connecting the same-name points in the images; Calculating the geometric similarity of each potential image matching pair based on all the three-dimensional point trajectories in each scene, and selecting potential image matching pairs that meet a threshold condition as overlapping image pairs; Taking any UAV image in each overlapping image pair as an anchor image, generating a corresponding overlapping image list, wherein the overlapping image list records the UAV images associated and matched with the anchor image in all overlapping image pairs; Aggregate each anchor image and its corresponding overlapping image list to form a retrieval dataset.
4. The drone image retrieval method according to claim 3, characterized in that: Based on all potential image matching pairs corresponding to each scene, 3D reconstruction is performed to obtain a 3D reconstruction model corresponding to each scene, including: Performing guided feature matching on all potential image matching pairs corresponding to each scene to obtain a plurality of matching point distribution maps, each of the matching point distribution maps being used to display the distribution of all matching points corresponding to the potential image matching pair; Based on the distribution graph of all matching points corresponding to each scene, an undirected weighted graph corresponding to each scene is constructed; Dividing each of the undirected weighted graphs into a plurality of subclusters, performing incremental SfM reconstruction on all subclusters in parallel, and generating a submodel corresponding to each subcluster; After iteratively merging all sub-models of each scene, a global bundle adjustment optimization is performed to obtain the 3D reconstructed model of each scene.
5. The drone image retrieval method according to claim 2, characterized in that: Generating training samples based on the retrieval data set includes: Randomly select several scenes from all scenes covered by the retrieval dataset, and randomly select an anchor image from each scene as an anchor sample; For each anchor sample, several drone images with geometric similarity greater than a preset threshold are selected from the overlapping image list corresponding to each anchor sample as positive samples, and positive samples in other scenes are used as negative samples of the anchor sample; The training samples are composed of all anchor samples, as well as the positive samples and negative samples corresponding to each anchor sample.
6. The drone image retrieval method according to claim 1, characterized in that: The training method of the image retrieval model includes: Dividing the training samples into a plurality of batches, each batch having a preset number of samples; In each round of training, all samples in each batch are input into the image retrieval model to be trained for processing, and the global feature descriptor vector of all samples in each batch is obtained. The loss function is calculated based on the global feature descriptor vector of all samples in each batch, and the model parameters are updated through backpropagation. Iteratively training the image retrieval network model until a preset end condition is reached to obtain a trained image retrieval network; The loss function is the sum of a first loss function and a second loss function. The first loss function is used to optimize the image retrieval model's ability to distinguish between positive samples and negative samples, and the second loss function is used to optimize the image retrieval model's ability to distinguish overlapping images.
7. The drone image retrieval method according to claim 1, characterized in that: The image retrieval network model includes a backbone network and a NetVLAD network. Each drone image in the set of drone images to be retrieved is sequentially input into the trained image retrieval model to obtain a global feature descriptor vector corresponding to each drone image, including: Inputting each drone image in the set of drone images to be retrieved into the backbone network in the trained image retrieval model; The backbone network extracts features from each drone image to obtain a depth feature map corresponding to each drone image; Each of the depth feature maps is input into the NetVLAD network and processed by the NetVLAD network to obtain a global feature descriptor vector corresponding to each drone image.
8. A drone image retrieval device, characterized in that: include: A training sample generation module, configured to obtain and generate training samples based on a first UAV image set; A model training module is used to iteratively train the image retrieval model to be trained using the training samples, calculate the loss function and reversely update the model parameters in each round of training until the end condition is met, thereby obtaining a trained image retrieval model; A vector generation module is used to obtain a set of drone images to be retrieved, input each drone image in the set of drone images to be retrieved into the trained image retrieval model in sequence, and obtain a global feature descriptor vector corresponding to each drone image; The retrieval module is used to construct an index structure based on all the global feature descriptor vectors, and input each global feature descriptor vector into the index structure in turn for search, so as to obtain a pairing result for each drone image in the drone image set to be retrieved.
9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a drone image retrieval program stored in the memory and runnable on the processor. When the drone image retrieval program is executed by the processor, the steps of the drone image retrieval method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a drone image retrieval program. When the drone image retrieval program is executed by the processor, the steps of the drone image retrieval method according to any one of claims 1 to 7 are implemented.