Image retrieval method, device and equipment and storage medium

By using a student network model and a hash network to process remote sensing images in UAV visual positioning, the problem of low efficiency in remote sensing image retrieval was solved, achieving efficient and high-precision image retrieval results.

CN115994242BActive Publication Date: 2026-01-06Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310079475.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-20
Publication Date
2026-01-06
Estimated Expiration
2043-01-20

AI Technical Summary

Technical Problem

Existing remote sensing image retrieval technologies suffer from low efficiency and accuracy in scenarios such as UAV visual positioning where storage resources are limited.

Method used

The remote sensing image to be identified is input into a student network model obtained by distillation training a pre-set loss function and a pre-trained teacher network model to obtain the feature data of the remote sensing image. Then, the hash network in the student network model is used to hash the feature data of the remote sensing image to obtain the hash code of the remote sensing image. Finally, the target reference image of the remote sensing image is retrieved based on the hash code of the remote sensing image.

Benefits of technology

It enables efficient and high-precision image retrieval in specific scenarios such as UAV visual positioning where storage resources are limited.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115994242B_ABST
    Figure CN115994242B_ABST
Patent Text Reader

Abstract

The application discloses an image retrieval method and device, equipment and a storage medium. The image retrieval method comprises the following steps: acquiring a remote sensing image to be identified; performing feature extraction on the remote sensing image based on a student network model to obtain feature data of the remote sensing image, wherein the student network model is obtained by performing distillation training on a student network model to be trained based on a pre-set loss function and a pre-trained teacher network model; performing hash processing on the feature data by using a hash network in the student network model to obtain a hash code of the remote sensing image; and searching for a target reference image of the remote sensing image according to the hash code of the remote sensing image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer technology, and in particular relates to an image retrieval method, apparatus, device and storage medium. Background Technology

[0002] Remote sensing image retrieval technology can effectively assist UAVs in image target localization and self-localization when denied by the Global Navigation Satellite System (GNSS), playing a crucial role in the stability of UAV intelligent sensing operations such as target reconnaissance, wide-area search, and target tracking in harsh environments. Currently, existing retrieval technologies are mainly based on deep learning-based symmetric retrieval, which uses the same depth model for data encoding on both the query image and the reference image in the database. To ensure retrieval accuracy, a high-performance, large model is typically chosen for feature extraction, resulting in low efficiency.

[0003] Therefore, how to achieve efficient and high-precision image retrieval in specific scenarios such as UAV visual positioning with limited storage resources is a technical problem that urgently needs to be solved. Summary of the Invention

[0004] This application provides an image retrieval method, apparatus, device, and storage medium, which can achieve efficient and high-precision image retrieval in specific scenarios such as UAV visual positioning where storage resources are limited.

[0005] In a first aspect, embodiments of this application provide an image retrieval method, the method comprising: acquiring a remote sensing image to be identified; extracting features from the remote sensing image based on a student network model to obtain feature data of the remote sensing image, wherein the student network model is obtained by distillation training of the student network model to be trained based on a pre-set loss function and a pre-trained teacher network model; hashing the feature data using a hash network in the student network model to obtain a hash code of the remote sensing image; and retrieving a target reference image of the remote sensing image based on the hash code of the remote sensing image.

[0006] Secondly, embodiments of this application provide an image retrieval device, comprising: an acquisition module for acquiring a remote sensing image to be identified; an extraction module for extracting features from the remote sensing image based on a student network model to obtain feature data of the remote sensing image, wherein the student network model is obtained by distillation training of the student network model to be trained based on a pre-set loss function and a pre-trained teacher network model; a processing module for hashing the feature data using a hash network in the student network model to obtain a hash code of the remote sensing image; and a retrieval module for retrieving a target reference image of the remote sensing image based on the hash code of the remote sensing image.

[0007] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0009] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0010] In the image retrieval method provided in this application, the remote sensing image to be identified is input into a student network model obtained by distillation training using a pre-set loss function and a pre-trained teacher network model to obtain the feature data of the remote sensing image. Then, the hash network in the student network model is used to hash the feature data of the remote sensing image to obtain the hash code of the remote sensing image. Finally, the target reference image of the remote sensing image is retrieved based on the hash code of the remote sensing image. This method can achieve efficient and high-precision image retrieval in specific scenarios such as UAV visual positioning with limited storage resources. Attached Figure Description

[0011] Figure 1 This is a schematic flowchart of an image retrieval method provided in an embodiment of this application;

[0012] Figure 2 This is a network architecture diagram provided in an embodiment of this application;

[0013] Figure 3 This is a structural diagram of a hash network provided in an embodiment of this application;

[0014] Figure 4This is a flowchart illustrating another image retrieval method provided in an embodiment of this application;

[0015] Figure 5 This is a schematic diagram of the structure of an image retrieval device provided in an embodiment of this application;

[0016] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0018] Remote sensing image retrieval technology can effectively assist UAVs in image target localization and self-localization when denied by the Global Navigation Satellite System (GNSS), playing a crucial role in the stability of UAV target reconnaissance, wide-area search, target tracking, and other intelligent perception tasks in harsh environments. Early retrieval methods mainly relied on extracting manually designed features. However, for cross-viewpoint and cross-platform visual localization problems, the significant differences in viewpoints make manually designed features difficult to adapt to, resulting in low accuracy. Currently, most advanced remote sensing image retrieval methods are based on deep learning. Some studies have specifically designed network structures or modules for visual geolocation tasks to extract powerful feature representations. Inspired by the traditional Vector of Locally Aggregated Descriptors (VLAD) method, a learnable VLAD layer called NetVLAD has been proposed, which can be embedded into convolutional neural networks for end-to-end training. Based on NetVLAD, a weighted soft-margin ranking loss function has been proposed, which significantly improves the convergence speed and accuracy of NetVLAD.

[0019] Image representations based on deep learning methods are high-dimensional, ranging from 512 to 70,000. As the flight time and distance of UAVs increase, the image search space expands dramatically. In practical localization, after a coarse screening of similar reference data, further filtering and matching are required. Moreover, when UAVs operate autonomously for extended periods, multiple remote sensing datasets are typically referenced. If the algorithm is too complex, search and matching efficiency will decrease, affecting localization results. Furthermore, remote sensing technologies based on deep hashing methods mainly include supervised and unsupervised hashing. Regardless of whether the method is supervised or unsupervised, improving retrieval accuracy relies on large models and large training datasets, which is insufficient in specific scenarios such as UAV visual localization where computational and storage resources are limited.

[0020] The following description, in conjunction with the accompanying drawings, details an image retrieval method, apparatus, device, and storage medium provided in this application through specific embodiments and application scenarios.

[0021] Figure 1 This application illustrates an embodiment of an image retrieval method, which can be executed by an electronic device, including a server and / or a terminal device. In other words, the method can be executed by software or hardware installed on the electronic device, and includes the following steps:

[0022] Step 110: Acquire the remote sensing image to be identified.

[0023] Remote sensing technology has many advantages, including a wide observation range, large information capacity, fast information acquisition, short update cycle, saving manpower and material resources, and fewer human interference factors. Remote sensing technology can conveniently, quickly, and accurately obtain remote sensing images of the area to be identified.

[0024] Specifically, the remote sensing images to be identified in the embodiments of this application can be acquired using high-resolution remote sensing satellites, such as IKONOS satellites, GeoEye satellites, QuickBird satellites, WorldView series satellites, and my country's domestically produced "Gaofen" series satellites.

[0025] It should be noted that the remote sensing images to be identified can include any type of land surface image.

[0026] Step 120: Extract features from the remote sensing image based on the student network model to obtain the feature data of the remote sensing image.

[0027] The student network model is obtained by distilling the student network model to be trained based on a pre-set loss function and a pre-trained teacher network model.

[0028] It is understood that, in this embodiment of the application, after capturing remote sensing images of a small area using a drone, the features of the remote sensing image of the area to be identified can be obtained by preprocessing the image using a trained student network model. Different feature parameters can be selected as the remote sensing image features of the area to be identified, depending on the actual situation. The specific type of remote sensing image features of the area to be identified is not specifically limited in this embodiment of the invention.

[0029] like Figure 2 The diagram shows the architecture of the teacher network model and the student network model. The student network model may include a first intermediate layer, a second intermediate layer, a third intermediate layer, and an output layer. The teacher network model also correspondingly includes a first intermediate layer, a second intermediate layer, a third intermediate layer, and an output layer. In one implementation, the distillation training of the student network model to be trained based on a pre-set loss function and a pre-trained teacher network model includes the following steps:

[0030] Step 121: Obtain a training sample set, which includes multiple sample images.

[0031] Step 122: Input the multiple sample images into the student network model and the teacher network model to train the models.

[0032] Step 123: During the training process, obtain the first feature vector output by the student network model and the second feature vector output by the teacher network model.

[0033] Wherein, the first feature vector is the feature of the remote sensing image extracted by the student network model, and the second feature vector is the feature of the remote sensing image extracted by the teacher network model.

[0034] Step 124: Input the first feature vector and the second feature vector into the pre-set loss function respectively.

[0035] Step 125: Adjust the parameters of the student network model based on the loss function.

[0036] It should be noted that both the teacher and student network models can consist of multi-scale feature extractors and network heads. The backbone of the teacher network model can be common network structures such as ResNet-50, ResNet-101, or ViT, and the network head can consist of three fully connected layers, mapping feature vectors to class probabilities or simplified binary hash codes. The backbone of the student network model can be lightweight networks such as MobileNet and ShuffleNet, or a simple network consisting of multiple stacked grouped convolutional layers, depthwise separable convolutional layers, and pooling layers. The network head can also consist of three fully connected layers, mapping feature vectors to simplified binary hash codes.

[0037] In this implementation, multiple sample images from the training sample set are input into the teacher network model and the student network model in batches. After processing through their respective first, second, and third intermediate layers, the second feature vector output by the teacher network model and the first feature vector output by the student network model are obtained. The first and second feature vectors are then input into a preset loss function, thereby continuously adjusting the student network model based on the loss function corresponding to the second feature vector output by the teacher network model, so that the difference between the first similarity matrix output by the student network model and the second similarity matrix output by the teacher network model is less than a preset threshold.

[0038] In one implementation, the loss function includes a first distillation loss function, a second distillation loss function, and a third distillation loss function.

[0039] See also Figure 2 The first intermediate layer of the student network model and the first intermediate layer of the teacher network model are respectively distilled based on the network output. The output of the student network model is processed by the first distillation loss function so that the difference between the output of the student network model and the teacher network model is less than a preset threshold.

[0040] In one implementation, the first distillation loss function is determined, wherein the first distillation loss function can be determined by the following formula:

[0041] ,

[0042] K is the number of sample images in the training sample set. The first soft label for the i-th sample image is provided by the teacher network model. This refers to the class prediction of the student network model for the i sample images. It should be noted that... and The outputs can be obtained from the fully connected layers of the teacher network model and the student network model, respectively. This can be achieved using a softmax function with a temperature parameter T, according to the following formula:

[0043]

[0044]

[0045] In this model, the fully connected layers of both the teacher and student network models output an array z. Let i be the i-th element.

[0046] See also Figure 2 The second intermediate layer of the student network model and the second intermediate layer of the teacher network model undergo knowledge distillation based on intermediate features, respectively. The feature space of the student network model is then guided by the second distillation loss function. In other words, the features of the intermediate layers of the teacher network model are used to guide the training of the student network model, ensuring that the feature spaces of the student and teacher network models have similar structures. This is mainly achieved through the L2 norm, i.e., the Mean Square Error (MSE) loss function.

[0047] In one implementation, the second distillation loss function is determined, wherein the second distillation loss function can be determined by the following formula:

[0048]

[0049] Let be the first network parameter of the student network model. The second network parameter of the student network model. This is the third network parameter of the student network model. The teacher network model depends on the first network parameters. intermediate features, The student network model depends on the second network parameters. The intermediate feature, r, is the student network model's dependence on the third network parameter. intermediate features, The sample images are input to the student network model and the teacher network model.

[0050] Alternatively, knowledge distillation based on intermediate features can also be based on feature probabilistic similarity. Specifically, pooling features are first extracted from the backbone network, and a similarity matrix is ​​obtained using cosine similarity. Then, the probabilistic similarity between the two similarities is optimized based on Kullback Leibler (KL) loss.

[0051] See also Figure 2The third intermediate layer of the student network model and the third intermediate layer of the teacher network model are subjected to knowledge distillation based on sample similarity. It is understandable that knowledge distillation based on sample similarity can preserve sample similarity. Paired sample features are extracted between the student and teacher network models and activated to generate a first similarity matrix and a second similarity matrix. Then, the second similarity matrix of the teacher network model is used as the optimization objective of the student network model.

[0052] See also Figure 2 The output layers of the student network model and the teacher network model undergo another form of knowledge distillation based on sample similarity. It can be understood that, aside from the intermediate layers of the network, the second similarity matrix reconstructed from the second feature vector output by the teacher network model at the output layer can be used to guide the construction of the first similarity matrix from the hash code output by the student network model. First, deep features are extracted using the teacher network model, and similarity information is obtained using a distance metric to construct the second similarity matrix based on deep features. On the other hand, similarity-preserving hash codes are generated in the student network model, and a first similarity matrix is ​​constructed based on these hash codes. The similarity matrices output by these two branches are used to calculate the similarity reconstruction error, serving as the third distillation loss function. The third distillation loss function is used to guide the construction of the first similarity matrix from the hash codes of the student network model.

[0053] In one implementation, the determination of the third distillation loss function is made by the following formula:

[0054]

[0055] in, The first similarity matrix is ​​constructed by the student network model based on the first feature vector through the hash network. The second similarity matrix is ​​constructed by the teacher network model based on the second feature vector.

[0056] In one implementation, the loss function is a fusion result obtained by weighted fusion of the first distillation loss function, the second distillation loss function, and the third distillation loss function.

[0057] In this implementation, output-based distillation minimizes the difference between the outputs of the student and teacher network models; feature similarity-based distillation aims to guide the student network model's learning using intermediate-level features from the teacher network model, making their feature space structures similar; and sample similarity-based knowledge distillation aims to preserve sample similarity. The student network can also incorporate a classification task, using the true labels and the student network model's output to calculate cross-entropy loss, resulting in a final loss function composed of the above multiple loss functions.

[0058] Step 130: Use the hash network in the student network model to perform hash processing on the feature data to obtain the hash code of the remote sensing image.

[0059] Understandably, the remote sensing image to be identified is processed by the student network model to obtain features; these features are then input into the hash network, i.e., the head network, to obtain the hash code of the remote sensing image. For example... Figure 3 The diagram shows the structure of a hash network, which consists of three fully connected layers. Finally, the data is activated by the Tanh activation function, which fixes the range of data within [-1, 1] as the final output.

[0060] Step 140: Retrieve the target reference image of the remote sensing image based on the hash code of the remote sensing image.

[0061] In one implementation, retrieving the target reference image of the remote sensing image based on its hash code includes: calculating the Hamming distance between the hash code of the remote sensing image and each target hash code in a preset database; and selecting the target remote sensing image corresponding to the target hash code with the smallest Hamming distance between the hash code of the remote sensing image and each target hash code as the target reference image of the remote sensing image.

[0062] It is understandable that by calculating the Hamming distance between the hash code of the remote sensing image and each target hash code in the preset database, and returning the image index with the smallest Hamming distance in ascending order according to the Hamming distance, the corresponding image in the remote sensing image database is found based on the image index, thus completing the image retrieval.

[0063] It should be noted that hash retrieval technology can map high-dimensional features similarly to hash codes, significantly reducing storage and computational overhead during the retrieval process, thus making it a highly promising and efficient retrieval method. The core idea of ​​hashing is to map high-dimensional remote sensing images into low-dimensional, compact binary codes while maintaining semantic similarity. This ensures that similar images have similar codes in the Hamming space, meaning there is a one-to-one correspondence between the high-dimensional data in the original space and the hash codes in the low-dimensional space, and similar data in the original space have a high probability of having similar hash codes in the low-dimensional space. Furthermore, Hamming distance can be quickly calculated through simple bit operations and XOR operations, achieving efficient retrieval. The main process of nearest neighbor retrieval using locality-sensitive hashing consists of three steps:

[0064] (1) Design a hash function to convert the original data into binary hash code.

[0065] (2) Use a hash function to hash the original information in the database.

[0066] (3) Use hash codes to perform nearest neighbor retrieval. Based on the hash code of the information to be retrieved, query the database for information with similar hash codes to obtain the retrieval results.

[0067] In the image retrieval method provided in this application, the remote sensing image to be identified is input into a student network model obtained by distillation training using a pre-set loss function and a pre-trained teacher network model to obtain the feature data of the remote sensing image. Then, the hash network in the student network model is used to hash the feature data of the remote sensing image to obtain the hash code of the remote sensing image. Finally, the target reference image of the remote sensing image is retrieved based on the hash code of the remote sensing image. This method can achieve efficient and high-precision image retrieval in specific scenarios such as UAV visual positioning with limited storage resources.

[0068] In one implementation, after adjusting the parameters of the student network model based on the loss function, the method further includes: if the student network model is determined to have converged, obtaining the trained student network model, wherein the trained student network model is used to retrieve a reference image of the remote sensing image.

[0069] The following is through Figure 4 A specific embodiment of this application will be described. Specifically, this embodiment includes the following steps:

[0070] Step 410: Select the teacher network model and the student network model.

[0071] Based on the characteristics of the retrieval task and the remote sensing reference images (satellite images or terrain data) to be retrieved, a more complex teacher network model and a lighter student network model were selected respectively.

[0072] Step 420: Train the teacher network model.

[0073] Using a large dataset, a teacher network model is trained using self-supervised or supervised methods. Benchmark remote sensing images, such as satellite images, are segmented and encoded, and stored in the UAV's database before the flight mission. The dataset can be selected from University-1652, ALTO, etc.

[0074] Step 430: Connect the teacher network model with the student network model.

[0075] After the teacher network model is trained, the network parameters are frozen. Using... Figure 1 The network architecture shown connects the teacher network model and the student network model, extracts intermediate features, and calculates a similarity matrix for use in calculating the loss function.

[0076] Step 440: Train the student network model based on the teacher network model.

[0077] Input remote sensing imagery into both the teacher and student network models, and calculate the distillation loss function using features corresponding to different network levels. If a network classification head is used, the cross-entropy loss L_hard is also calculated. The network is trained using paired reference remote sensing imagery and UAV imagery related to the flight area, updating the parameters of the student network model.

[0078] Step 450: Obtain the converged student network model.

[0079] After training converges, the parameters of the student network are fixed, and the retrieval on the UAV uses only the trained student network.

[0080] Step 460: Image retrieval using the student network model.

[0081] During the flight mission, images (query points) captured by the UAV are input into the student network, and the query points are hashed; the Hamming distance between the hash code and the hash code in the database is calculated to determine the nearest reference image.

[0082] The above process enables the entire process of remote sensing image retrieval.

[0083] In summary, the embodiments of this application have the advantages of high retrieval accuracy, small quantization loss, and more efficient hash encoding.

[0084] Figure 5The diagram shows a schematic of an image retrieval device for a cloud platform according to an embodiment of this application. The image retrieval device 500 may include: an acquisition module 510, an extraction module 520, a processing module 530, and a retrieval module 540.

[0085] The acquisition module 510 is used to acquire the remote sensing image to be identified. The extraction module 520 is used to extract features from the remote sensing image based on a student network model to obtain feature data of the remote sensing image. The student network model is obtained by distillation training of the student network model to be trained using a pre-set loss function and a pre-trained teacher network model. The processing module 530 is used to perform hash processing on the feature data using a hash network in the student network model to obtain the hash code of the remote sensing image. The retrieval module 540 is used to retrieve the target reference image of the remote sensing image based on the hash code of the remote sensing image.

[0086] In one implementation, the extraction module 520 is further configured to acquire a training sample set, the training sample set including multiple sample images; input the multiple sample images into the student network model and the teacher network model for model training; during the training process, acquire a first feature vector output by the student network model and a second feature vector output by the teacher network model, wherein the first feature vector is a feature of the remote sensing image extracted by the student network model, and the second feature vector is a feature of the remote sensing image extracted by the teacher network model; input the first feature vector and the second feature vector into the pre-set loss function respectively; and adjust the parameters of the student network model based on the loss function.

[0087] In one implementation, after adjusting the parameters of the student network model based on the loss function, the extraction module 520 is further configured to obtain the trained student network model when it is determined that the student network model has converged, and the trained student network model is used to retrieve the reference image of the remote sensing image.

[0088] In one implementation, the loss function in the extraction module 520 includes a first distillation loss function, a second distillation loss function, and a third distillation loss function; wherein, the first distillation loss function is used to process the output of the student network model; the second distillation loss function is used to guide the feature space of the student network model; and the third distillation loss function is used to guide the construction of a first similarity matrix from the hash codes of the student network model.

[0089] In one implementation, the loss function in the extraction module 520 is a fusion result obtained by weighted fusion of the first distillation loss function, the second distillation loss function, and the third distillation loss function.

[0090] In one implementation, the retrieval module 540 is further configured to calculate the Hamming distance between the hash code of the remote sensing image and each target hash code in the preset database; and to use the target remote sensing image corresponding to the target hash code with the smallest Hamming distance between the hash code of the remote sensing image and each target hash code as the target reference image of the remote sensing image.

[0091] In one implementation, the extraction module 520 is further configured to determine the first distillation loss function, wherein the first distillation loss function is determined by the following formula:

[0092]

[0093] The first soft label for the i-th sample image is provided by the teacher network model. The student network model makes category predictions for the i sample images.

[0094] In one implementation, the extraction module 520 is further configured to determine the second distillation loss function, wherein the second distillation loss function is determined by the following formula:

[0095]

[0096] Let be the first network parameter of the student network model. The second network parameter of the student network model. This is the third network parameter of the student network model. The teacher network model depends on the first network parameters. intermediate features, The student network model depends on the second network parameters. The intermediate feature, r, is the student network model's dependence on the third network parameter. intermediate features, The sample images are input to the student network model and the teacher network model.

[0097] In one implementation, the extraction module 520 is further configured to determine the third distillation loss function, wherein the third distillation loss function is determined by the following formula:

[0098]

[0099] in, The first similarity matrix is ​​constructed by the student network model based on the first feature vector through the hash network. The second similarity matrix is ​​constructed by the teacher network model based on the second feature vector.

[0100] The image retrieval device provided in this application embodiment can achieve... Figures 1-3 The various processes implemented in the method embodiments shown will not be described again here to avoid repetition.

[0101] The image retrieval device in the embodiments of this application can be a device, or it can be a component, integrated circuit, or chip in an electronic device. The embodiments of this application are not specifically limited.

[0102] The image retrieval device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0103] Optional, such as Figure 6 As shown, this application embodiment also provides an electronic device 600, including a processor 610, a memory 620, and a program or instructions stored in the memory 620 and executable on the processor 610. When the program or instructions are executed by the processor 610, they implement the various processes of the above-described image retrieval method embodiment and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0104] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image retrieval method embodiments and achieve the same technical effects. To avoid repetition, these will not be described again here.

[0105] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0106] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described image retrieval method embodiments. To avoid repetition, these will not be described again here.

[0107] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0108] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0109] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0110] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An image retrieval method, characterized by, The method comprises the following steps: obtaining a remote sensing image to be identified; performing feature extraction on the remote sensing image based on a student network model to obtain feature data of the remote sensing image, wherein the student network model is obtained by performing distillation training on a student network model to be trained based on a pre-set loss function and a pre-trained teacher network model; performing hash processing on the feature data by using a hash network in the student network model to obtain a hash code of the remote sensing image; retrieving a target reference image of the remote sensing image according to the hash code of the remote sensing image; wherein the distillation training of the student network model to be trained based on the pre-set loss function and the pre-trained teacher network model comprises: obtaining a training sample set, wherein the training sample set comprises a plurality of sample images; inputting the plurality of sample images into the student network model and the teacher network model for model training; during the training process, obtaining a first feature vector output by the student network model and a second feature vector output by the teacher network model, wherein the first feature vector is a feature of the remote sensing image extracted by the student network model, and the second feature vector is a feature of the remote sensing image extracted by the teacher network model; inputting the first feature vector and the second feature vector into the pre-set loss function respectively; adjusting parameters of the student network model based on the loss function; wherein the loss function comprises a first distillation loss function, a second distillation loss function and a third distillation loss function; the first distillation loss function is determined by the following formula: a first soft label of the i-th sample image for the teacher network model, a class prediction of the i-th sample image for the student network model, a number of sample images in the training sample set; the second distillation loss function is determined by the following formula: Let be the first network parameter of the student network model. The second network parameter of the student network model. This is the third network parameter of the student network model. The teacher network model depends on the first network parameters. intermediate features, The student network model depends on the second network parameters. The intermediate feature, r, is the student network model's dependence on the third network parameter. intermediate features, The sample images are input into the student network model and the teacher network model; the third distillation loss function is determined by the following formula: wherein, a first similarity matrix constructed by the hash network based on the first feature vector for the student network model, a second similarity matrix constructed by the teacher network model based on the second feature vector.

2. The method of claim 1, wherein, after adjusting the parameters of the student network model based on the loss function, further comprising: in a case where it is determined that the student network model converges, obtaining a trained student network model, wherein the trained student network model is used to retrieve a reference image of the remote sensing image.

3. The method of claim 1, wherein, The loss function is a fusion result obtained by weighting and fusing the first distillation loss function, the second distillation loss function and the third distillation loss function.

4. The method of claim 1, wherein, The retrieval of the target reference image of the remote sensing image according to the hash code of the remote sensing image comprises: calculating a Hamming distance between the hash code of the remote sensing image and each target hash code in a preset database respectively; taking a target remote sensing image corresponding to a target hash code with the smallest Hamming distance between the hash code of the remote sensing image and the target hash code as the target reference image of the remote sensing image.

5. An image retrieval apparatus characterized by comprising: The method comprises the following steps: an obtaining module configured to obtain a remote sensing image to be identified; an extracting module configured to perform feature extraction on the remote sensing image based on a student network model to obtain feature data of the remote sensing image, wherein the student network model is obtained by performing distillation training on a student network model to be trained based on a pre-set loss function and a pre-trained teacher network model; a processing module configured to perform hash processing on the feature data by using a hash network in the student network model to obtain a hash code of the remote sensing image; The retrieval module is configured to retrieve a target reference image of the remote sensing image according to a hash code of the remote sensing image. The extraction module is further configured to obtain a training sample set, the training sample set including a plurality of sample images; input the plurality of sample images into the student network model and the teacher network model to perform model training; during the training process, obtain a first feature vector output by the student network model and a second feature vector output by the teacher network model, wherein the first feature vector is a feature of the remote sensing image extracted by the student network model, and the second feature vector is a feature of the remote sensing image extracted by the teacher network model; input the first feature vector and the second feature vector into the pre-set loss function respectively; and adjust parameters of the student network model based on the loss function. The loss function includes a first distillation loss function, a second distillation loss function, and a third distillation loss function. The extraction module is further configured to determine the first distillation loss function by the following formula: a first soft label of the i-th sample image for the teacher network model, a class prediction of the i-th sample image for the student network model, a number of sample images in the training sample set; The extraction module is further configured to determine the second distillation loss function by the following formula: a first network parameter of the student network model, a second network parameter of the student network model, a third network parameter of the student network model, intermediate features of the teacher network model dependent on the first network parameter intermediate features of the student network model dependent on the second network parameter intermediate features of the student network model dependent on the third network parameter intermediate features of the student network model dependent on the third network parameter intermediate features of the student network model dependent on the third network parameter the sample imagery input to the student network model and the teacher network model; The extraction module is further configured to determine the third distillation loss function by the following formula: wherein, a first similarity matrix constructed by the hash network based on the first feature vector for the student network model, a second similarity matrix constructed by the teacher network model based on the second feature vector.

6. An electronic device, comprising: A processor, a memory, and a program or instructions stored on the memory and executable on the processor, the program or instructions being executed by the processor to implement the steps of the image retrieval method according to any one of claims 1-4.

7. A readable storage medium characterized by, The program or instructions are stored on the readable storage medium, and the program or instructions are executed by the processor to implement the steps of the image retrieval method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Cross-source image retrieval method and system, medium and equipment

    CN112860935A

  • Multi-source remote sensing image retrieval method based on deep hash

    CN113591784A