Remote sensing image retrieval method based on multi-scale feature representation

Through the dual-branch Transformer network structure and multi-scale feature interaction module, the problem of single-scale feature extraction in remote sensing image retrieval is solved, and high-precision and robust remote sensing image retrieval is realized, which is suitable for the recognition of multi-scale objects in complex scenarios.

CN120541258APending Publication Date: 2025-08-26ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510673142.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing remote sensing image retrieval methods ignore shallow features when extracting feature, resulting in the single-scale receptive field being unable to take into account details and global information. The single-scale model has limited ability to integrate context information, making it difficult to accurately identify objects of different sizes in complex scenarios, and is not robust enough.

Method used

The dual-branch Transformer network structure is adopted, and multi-level feature fusion and multi-scale feature interaction modules are combined with self-attention and cross-attention mechanisms to extract multi-scale features, and the hash code generation is optimized through joint loss functions to enhance feature representation capabilities.

Benefits of technology

It improves the accuracy and robustness of remote sensing image retrieval, can accurately identify objects of different sizes in complex scenarios, and improves the network's feature representation ability and cross-task migration potential.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541258A_ABST
    Figure CN120541258A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image retrieval method based on multi-scale feature representation, and the method is characterized in that the method comprises the following steps: S1, carrying out the preprocessing of original data, and dividing a data set into a training set and a test set for later use; s2, constructing a remote sensing image retrieval model based on multi-scale feature representation, and optimizing the remote sensing image retrieval model by adopting the proposed joint loss function; s3, training the data sets processed in the step S1 through the deep learning network constructed in the step S2, optimizing model parameters at the same time, and performing evaluation of retrieval precision, robustness and generalization ability on a plurality of data sets by comparing with a plurality of existing retrieval methods; according to the method, a main body framework of a multi-scale Transform network is used, the extraction capacity, the global semantic understanding capacity and the cross-task migration potential of advanced features of remote sensing images are enhanced, and especially in complex scene tasks such as remote sensing, the retrieval performance is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing image technology, and in particular to a remote sensing image retrieval method based on multi-scale feature representation. Background Art

[0002] Early remote sensing image retrieval technology was primarily based on handcrafted features (such as color histograms, texture filters, or spectral band statistics). While computationally efficient, this approach was limited by insufficient feature abstraction. Remote sensing images typically feature high resolution and complex scenes, making traditional handcrafted feature extraction methods significantly limited. Deep learning technology, through the training mechanism of neural networks, mimics the cognitive processes of the human brain, enabling automatic learning and extraction of features at different levels, improving the ability to express image features and bringing new breakthroughs to remote sensing image retrieval. Furthermore, deep learning networks possess excellent scalability and adaptability, adapting to the complex demands of various tasks and playing a significant role in advancing the development of remote sensing image retrieval technology. Retrieval methods based on convolutional neural networks extract global feature vectors from remote sensing images, alleviating the scarcity of annotated data through transfer learning techniques and achieving excellent retrieval results. Recently, the Transformer model (ViT) has garnered widespread attention in the computer vision community for image classification tasks. After training on large datasets, ViT outperforms traditional convolutional neural networks and demonstrates stronger generalization capabilities, leading to a wider application of Transformer networks.

[0003] Due to the way they are captured, remote sensing images often have large coverage, variable object scales, and complex backgrounds. This places higher demands on feature extraction networks. However, existing retrieval methods often only use deep features when performing feature retrieval, ignoring the potential information in shallow features. This leads to the following potential problems: First, the receptive field of a single scale cannot account for both details and global information. Shallow networks can capture fine-grained features but lack semantic abstraction; deep networks, while providing high-level semantic understanding, lose spatial detail, resulting in low accuracy in detecting small objects. Second, object scales vary significantly in complex scenes, and single-scale features are susceptible to scale sensitivity, making it difficult to accurately identify objects of different sizes simultaneously. Furthermore, single-scale models have limited ability to integrate contextual information, cannot effectively establish long-range dependencies, and lack robustness to challenges such as occlusion and deformation. Such as the literature Huang M, Dong L, Dong W, et al. Supervised contrastive learning based on fusion of global and local features for remote sensing image retrieval[J]. IEEE Transactions onGeoscience and Remote Sensing, 2023, 61: 1-13. Zhao D, Chen Y, Xiong S. Multiscalecontext deep hashing for remote sensing image retrieval[J]. IEEE Journal ofSelected Topics in Applied EarthObservations and Remote Sensing, 2023, 16: 7163-7172.

[0004] Therefore, a remote sensing image retrieval method based on multi-scale feature representation is provided to solve the above technical problems. Summary of the Invention

[0005] The purpose of the present invention is to provide a remote sensing image retrieval method based on multi-scale feature representation. The present invention extracts and utilizes multi-scale features through a dual-branch structure, and combines the loss function to improve the quality of generated hash codes. Specifically, the present invention adopts a dual-branch structure to process image block inputs of different granularities respectively: the fine-grained branch focuses on local detail features, and the coarse-grained branch captures global scene information. In the feature extraction stage, multi-level feature fusion is first implemented within each branch, and the context information modeling capability is enhanced through an adaptive fusion method; then the features of the two branches are input into a specially designed multi-scale feature interaction module, which uses a cross-attention mechanism to achieve adaptive complementarity and enhancement between features of different scales. In order to further optimize the feature representation, a joint loss function is proposed to fine-tune the model, and a compact hash code is generated while optimizing the feature space distribution.

[0006] The object of the present invention is achieved like this: A remote sensing image retrieval method based on multi-scale feature representation includes the following steps: S1. Preprocess the original data and divide the data set into training set and test set for future use. S2. Build a remote sensing image retrieval model based on multi-scale feature representation. The remote sensing image retrieval model includes the following modules: The dual-branch Transformer network architecture leverages the network's self-attention mechanism to leverage powerful global modeling capabilities and obtain effective feature representation. Multi-level feature fusion module: The multi-level feature fusion module fuses features at different depths of the network to avoid relying on a single high-level feature; Multi-scale feature fusion module: The multi-scale feature fusion module interacts the features of the two parallel branch architectures to obtain comprehensive contextual multi-scale features; Hash module: The hash module is introduced to reduce the dimension of the extracted real-valued features, reduce the storage space of the features, and improve the computing efficiency; The proposed joint loss function is used to optimize the remote sensing image retrieval model; S3. The dataset processed in step S1 is trained through the deep learning network constructed in step S2, and the model parameters are optimized. The dataset is compared with various existing retrieval methods on multiple datasets to evaluate the retrieval accuracy, robustness, and generalization ability.

[0007] The specific operations of S1 are as follows: S1.1. Perform stratified sampling on the original data; divide the data into training, validation, and test sets according to preset proportions, focusing on maintaining balance among the subsets in terms of sample size, category distribution, and scenario categories. This ensures that the entire process of model training, hyperparameter optimization, and performance verification is based on statistically significant data, thereby improving the model's generalization ability in real-world application scenarios. S1.2. Since the image sizes of the two branches of the constructed network are different, the images are scaled by the ratio of RGB and RGB, respectively, and then the RGB three-channel values ​​are normalized and standardized. Standardization reduces the instability problem caused by the difference value, that is, the channel mean is subtracted from each input channel and divided by the channel standard deviation; the present invention sets the mean and standard deviation to [0.485, 0.456, 0.406] and [0.485, 0.456, 0.406] respectively. [0.229, 0.224, 0.225]; S1.3. Implement data augmentation strategies in the training set; systematically expand the morphological spatial representation of the training data through geometric transformations such as random rotations, multi-scale scaling, and mirror flipping. This geometric invariance enhancement strategy effectively mitigates the overfitting problem caused by differences in shooting angles, object scales, and morphological diversity in remote sensing imagery, significantly improving the spatial adaptability of the feature extraction network.

[0008] The specific operation of S2 is as follows: S2.1. Resize the input data of the dual-branch structure to 224x224 and 240x240, respectively. Then, set convolution kernels of different sizes (16x16 and 12x12) to segment the input image, obtain feature information of remote sensing images at different scales, and then use the self-attention mechanism module to extract deep features of the image. S2.2. Adding a multi-level feature fusion module after the multi-layer self-attention module effectively reduces the traditional method's dependence on a single feature in the final output layer, significantly alleviating the information loss problem during feature extraction, and obtaining more comprehensive fusion features. The output features of each self-attention module in each branch are spliced, and then 1D convolution is used to achieve adaptive fusion of features at different levels; S2.3. After the multi-level feature fusion module, a multi-scale feature fusion module is added to achieve the complementarity of multi-scale features on both branches. By introducing a cross-attention mechanism, the class tokens containing global information in the dual-branch features are exchanged, and then self-attention is interacted with the local feature patch tokens of the other scale to achieve the complementarity of global feature information and obtain the attention weights between the global and local at different scales. The information of the patch tokens is also fused, and the local features of different scales are transformed into the same feature dimension to achieve the fusion of patch tokens. The image features of different scales are obtained on both branches. S2.4. Design a hash module to fuse the output features of the two branches and generate a binary hash code through a hash function. S2.5, use Represents the overall loss function, which includes the supervised contrast loss function and polarization loss function The parameters of the entire network architecture are optimized through the loss function; the expression of the supervised contrast loss function is: Where Q represents the index set of the query image; S(i) represents the index set of positive samples with the same label as the query sample; n(S(i)) represents the number of indices contained in S(i); ε is a hyperparameter set to 1e-5 to avoid n(S(i)) being 0; τ represents the temperature parameter set to 0.1; the supervised contrast loss function is generated by extending self-supervised contrastive learning to the fully supervised task, making full use of label information to bring the features of similar samples closer while separating the features of heterogeneous samples, thereby reducing the deviation between semantic similarity and Hamming distance; The expression for polarization loss is: In the formula, the margin threshold is pre-set to m=1, f i represents the output feature of the i-th sample, t i Represents the target vector; the polarization loss reduces the quantization error of the hash code through a binary distribution polarization strategy, ensuring a balance between coding compactness and retrieval efficiency.

[0009] In S2.2, formula (1) represents the process of extracting features from each level using the self-attention mechanism Where g i Represents the output features of each layer of the self-attention module, and SA represents the calculation process of the self-attention module; Formula (2) represents the multi-level feature fusion process. After the input features pass through N layers of self-attention modules, the output features are spliced ​​and the dimension is reduced through convolution operation to achieve adaptive fusion of features. g all =Conv(Concat(g1;g2;…;g i )) (2) In the formula, Concat represents the feature concatenation operation and Conv represents convolution.

[0010] In S2.3, the specific calculation process of the large-scale feature fusion module is as shown in formula (3), which realizes the transformation of global features and splicing with small-scale branch local features; Where, F L The (·) function represents the projection function, which is used to align the parameter dimensions. Class tokens representing large-scale branches, patch tokens representing small-scale branches; Formula (4) represents the introduction of a multi-head self-attention mechanism to achieve the attention weight distribution of class tokens. The features obtained by formula (3) are first aligned with the original dimension through the projection function, and then the new class token is obtained through the residual connection; Where, F L The (·) function represents the projection function, and MAtt represents the calculation process of the multi-head self-attention mechanism. Formula (5) represents the process of splicing the patch tokens of the two branches to achieve the fusion process of patch tokens; Where, F L (·) function represents the projection function, and Concat represents the feature concatenation operation; Formula (6) represents the reverse projection of the fused patch tokens to the input dimension through projection transformation; Where g L (·) represents the inverse projection function; Formula (7) represents the concatenation of the fused class tokens and patch tokens to obtain the new feature Z L ; Where g L (·) represents the back-projection function, which is used to align the fused class token and patch token dimensions. In S2.4, formula (8) projects the real-valued features output by the multi-scale feature fusion module to a dimension consistent with the preset hash code length for subsequent hash code generation; Where U(·) represents the mapping function, and Token-like features representing the output of the dual branch; Formula (9) represents the fusion process of real-valued features. The feature values ​​of the corresponding dimensions of the two-branch features are added and averaged to obtain the final feature. Where L represents the length of the preset hash code, and Respectively represent the eigenvalues ​​at the i-th position; Formula (10) represents the process of generating binary hash codes. The real-valued features are compressed by the hash function to generate the corresponding binary hash codes. Where, Represents a real-valued feature of length L, Sign represents a sign function used to generate a hash code; Since the class tokens contain global information, only the class tokens are used for fusion in this process to generate binary hash codes.

[0011] The beneficial effects of the present invention are: 1. The present invention uses the main architecture of the multi-scale Transformer network to enhance the ability to extract high-level features of remote sensing images, the ability to understand global semantics, and the potential for cross-task migration, especially in complex scene tasks such as remote sensing, which brings a significant improvement in retrieval performance. 2. To address the problem of information loss between different levels during feature extraction, this invention alleviates the information loss during deep propagation by splicing and adaptively fusing features at each level; 3. To address the problem of incomplete expression of single-scale features in existing methods, the present invention introduces a cross-attention mechanism to interact with multi-scale information, effectively realizing the complementarity between global and local information and improving the feature representation ability of the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is the overall framework diagram of the present invention; Figure 2 This is a multi-level feature fusion module diagram of the present invention; Figure 3 This is a diagram of the multi-scale feature fusion module of the present invention; Figure 4 UC Merced dataset search result diagram of the present invention; a. Precision@k curve diagram; b. Recall@k curve diagram; cP-R curve diagram; Figure 5 The AID dataset retrieval result diagram of the present invention; a. Precision@k curve diagram; b. Recall@k curve diagram; cP-R curve diagram; Figure 6 These are the retrieval result diagrams of the NWPU-RESISC45 dataset of the present invention; a. Precision@k curve diagram; b. Recall@k curve diagram; cP-R curve diagram. DETAILED DESCRIPTION

[0013] The present invention will be further described below with reference to the accompanying drawings and examples.

[0014] A remote sensing image retrieval method based on multi-scale feature representation, such as Figure 1 As shown, the following steps are included: S1. Preprocess the raw data and divide the dataset into training and test sets for future use. Collect commonly used remote sensing image retrieval datasets such as UC Merced, AID, and NWPU-RESISC45. Randomly sample samples from each class according to a fixed ratio to ensure balanced data within each class. Furthermore, to allow the network to better focus on semantic features within the image, scale, normalize, and scale the image.

[0015] The specific operations of S1 are as follows: S1.1. Perform stratified sampling on the original data; divide the data into training, validation, and test sets according to preset proportions, focusing on maintaining the balance of each subset in terms of sample size, category distribution, and scenario category; to ensure that the entire process of model training, hyperparameter optimization, and performance verification is based on statistically significant data, thereby improving the model's generalization ability in real application scenarios.

[0016] S1.2. Since the image sizes of the two branches of the constructed network are different, the images are scaled by the ratio of RGB and RGB, respectively, and then the RGB three-channel values ​​are normalized and standardized. Standardization reduces the instability problem caused by the difference value, that is, the channel mean is subtracted from each input channel and divided by the channel standard deviation; the present invention sets the mean and standard deviation to [0.485, 0.456, 0.406] and [0.485, 0.456, 0.406] respectively. [0.229, 0.224, 0.225].

[0017] S1.3. Implement data augmentation strategies in the training set; systematically expand the morphological spatial representation of the training data through geometric transformations such as random rotations, multi-scale scaling, and mirror flipping. This geometric invariance enhancement strategy effectively mitigates the overfitting problem caused by differences in shooting angles, object scales, and morphological diversity in remote sensing imagery, significantly improving the spatial adaptability of the feature extraction network.

[0018] S2. Build a remote sensing image retrieval model based on multi-scale feature representation. The remote sensing image retrieval model includes the following modules: The dual-branch Transformer network architecture leverages the network's self-attention mechanism to leverage powerful global modeling capabilities and obtain effective feature representation. Multi-level feature fusion module: The multi-level feature fusion module fuses features at different depths of the network to avoid relying on a single high-level feature; Multi-scale feature fusion module: The multi-scale feature fusion module interacts the features of the two parallel branch architectures to obtain comprehensive contextual multi-scale features; Hash module: The hash module is introduced to reduce the dimension of the extracted real-valued features, reduce the storage space of the features, and improve the computing efficiency; The proposed joint loss function is used to optimize the remote sensing image retrieval model.

[0019] The specific operation of S2 is as follows: S2.1. Resize the input data of the dual-branch structure to 224x224 and 240x240, respectively. Then, set convolution kernels of different sizes (16x16 and 12x12) to segment the input image, obtain feature information of remote sensing images at different scales, and then use the self-attention mechanism module to extract deep features of the image. S2.2. Adding a multi-level feature fusion module after the multi-layer self-attention module effectively reduces the traditional method’s dependence on a single feature in the final output layer, significantly alleviates the information loss problem in the feature extraction process, and obtains more comprehensive fusion features. The output features of each self-attention module in each branch are spliced, and then 1D convolution is used to achieve adaptive fusion of features at different levels, such as Figure 2 shown.

[0020] In S2.2, formula (1) represents the process of extracting features from each level using the self-attention mechanism Where g i Represents the output features of each layer of the self-attention module, and SA represents the calculation process of the self-attention module.

[0021] Formula (2) represents the multi-level feature fusion process. After the input features pass through N layers of self-attention modules, the output features are spliced ​​and the dimension is reduced through convolution operation to achieve adaptive fusion of features. g all =Conv(Concat(g1;g2;...;g i )) (2) In the formula, Concat represents the feature concatenation operation and Conv represents convolution.

[0022] S2.3. After the multi-level feature fusion module, a multi-scale feature fusion module is added to achieve the complementarity of multi-scale features on the two branches; by introducing the cross-attention mechanism, the class tokens containing global information in the dual-branch features are exchanged, and then the self-attention interaction is performed with the local feature patch tokens of another scale to achieve the complementarity of global feature information and obtain the attention weights between the global and local at different scales; the information of the patch tokens is also fused, and the local features of different scales are transformed into the same feature dimension to achieve the fusion of patch tokens; the image features of different scales are obtained on both branches, such as Figure 3 shown.

[0023] In S2.3, the specific calculation process of the large-scale feature fusion module is as shown in formula (3), which realizes the transformation of global features and splicing with small-scale branch local features; Where, F L The (·) function represents the projection function, which is used to align the parameter dimensions. Class tokens representing large-scale branches, patch tokens representing small-scale branches; Formula (4) represents the introduction of a multi-head self-attention mechanism to achieve the attention weight distribution of class tokens. The features obtained by formula (3) are first aligned with the original dimension through the projection function, and then the new class token is obtained through the residual connection; Where, F L (·) function represents the projection function, MAtt represents the calculation process of the multi-head self-attention mechanism; Formula (5) represents the process of splicing the patch tokens of the two branches to achieve the fusion process of patch tokens; Where, F L (·) function represents the projection function, and Concat represents the feature concatenation operation; Formula (6) represents the reverse projection of the fused patch tokens to the input dimension through projection transformation; Where g L (·) represents the inverse projection function; Formula (7) represents the concatenation of the fused class tokens and patch tokens to obtain the new feature Z l . Where g L(·) represents the inverse projection function, which is used to align the dimensions of the fused class tokens and patch tokens. In S2.4, Equation (8) projects the real-valued features output by the multi-scale feature fusion module to a dimension consistent with the preset hash code length for subsequent hash code generation; Where U(·) represents the mapping function, and Token-like features representing the output of the dual branch; Formula (9) represents the fusion process of real-valued features. The feature values ​​of the corresponding dimensions of the two-branch features are added and averaged to obtain the final feature. Where L represents the length of the preset hash code, and Represent the eigenvalues ​​at the i-th position respectively. Formula (10) represents the process of generating binary hash codes. The real-valued features are compressed by the hash function to generate the corresponding binary hash codes. Where, Represents a real-valued feature of length L, Sign represents a sign function used to generate a hash code; Since the class tokens contain global information, only the class tokens are used for fusion in this process to generate binary hash codes.

[0024] S2.4. Design a hash module, fuse the output features of the two branches in the hash module, and generate a binary hash code through the hash function.

[0025] S2.5, use Represents the overall loss function, which includes the supervised contrast loss function Polarization loss function The parameters of the entire network architecture are optimized through the loss function; the expression of the supervised contrast loss function is: Where Q represents the index set of the query image; S(i) represents the index set of positive samples with the same label as the query sample; n(S(i)) represents the number of indices contained in S(i); ε is a hyperparameter set to 1e-5 to avoid n(S(i)) being 0; τ represents the temperature parameter set to 0.1; the supervised contrast loss function is generated by extending self-supervised contrastive learning to the fully supervised task, making full use of label information to bring the features of similar samples closer while separating the features of heterogeneous samples, thereby reducing the deviation between semantic similarity and Hamming distance; The expression for polarization loss is: In the formula, the margin threshold is pre-set to m=1, f i represents the output feature of the i-th sample, t i Represents the target vector; the polarization loss reduces the quantization error of the hash code through a binary distribution polarization strategy, ensuring a balance between coding compactness and retrieval efficiency.

[0026] S3. The dataset processed in step S1 is trained through the deep learning network constructed in step S2, and the model parameters are optimized. The dataset is compared with various existing retrieval methods on multiple datasets to evaluate the retrieval accuracy, robustness, and generalization ability.

[0027] First, it is necessary to screen the multiple models or parameter configurations obtained in the training and validation stages through unified evaluation on the test set. The data in the test set has not participated in model training or parameter adjustment, and its diversity and complexity can more realistically reflect the generalization performance of the model in actual scenarios. In order to verify the superiority of this estimation model, it is compared with several of the most widely used deep learning models. The data sets use the public UC Merced, AID and NWPU-RESISC45 data sets, which contain a variety of remote sensing scenes such as airplanes, houses, rivers, roads, baseball fields, etc. The experiment of the present invention randomly selects 70% of the images of each category as the training set. The batch size set in each iteration is 32. The Adam optimizer is used for training, the learning rate is set to 5e-5, the margin threshold m is initially set to 1, and the temperature parameter τ is initially set to 0.1. The balance weight parameters α and β are set to 100 and 1.0 respectively, and ε is set to 10 -5 All experiments are implemented in Pytorch framework with Python version 3.10, using a 13th Gen Intel(R) Core(TM) i7-13700H CPU and an NVIDIA GeForce RTX4070Laptop GPU. Table 1 Different retrieval methods mAP value

[0028] To ensure fairness in comparison, all estimation models were trained and tested under the same hardware and software environment. Table 1 shows the average retrieval accuracy of all methods, which can reflect the overall performance of different methods. By analyzing Table 1, it can be seen that this method can achieve the best retrieval accuracy and higher mAP values ​​under the hash code length of 32 bits, 64 bits, and 128 bits, indicating the excellent performance of this method. Figure 4 Figure 5 Figure 6The accuracy, recall and accuracy-recall interaction curves of the models on different data sets are respectively shown. It can be seen from the figure that the IDHN model performs relatively poorly and performs poorly on all three data sets. Although the CSQ method, which ranks second in accuracy, is not much different from the present invention in terms of accuracy, its recall rate decreases significantly compared with the present method as the number of query images increases. This shows that the present invention has good adaptability and robustness in dealing with the large scale changes of remote sensing images, which confirms the results obtained in Table 1. In summary, in the field of remote sensing image retrieval, the remote sensing image retrieval method based on multi-scale feature representation proposed by the present invention is significantly better than other mainstream methods. As the number of data sets increases, the performance of each method will decline to varying degrees. However, in this case, the experimental results of the present invention are still better than those of other methods, which shows that the present invention can effectively improve the accuracy of retrieval, and exhibits excellent retrieval performance and wide application potential.

Claims

1. A remote sensing image retrieval method based on multi-scale feature representation, characterized in that: The following steps are involved: S1. Preprocess the original data and divide the data set into training set and test set for future use; S2. Build a remote sensing image retrieval model based on multi-scale feature representation. The remote sensing image retrieval model includes the following modules: The dual-branch Transformer network architecture leverages the network's self-attention mechanism to leverage powerful global modeling capabilities and obtain effective feature representation. Multi-level feature fusion module: The multi-level feature fusion module fuses features at different depths of the network to avoid relying on a single high-level feature; Multi-scale feature fusion module: The multi-scale feature fusion module interacts the features of the two parallel branch architectures to obtain comprehensive contextual multi-scale features; Hash module: The hash module is introduced to reduce the dimension of the extracted real-valued features, reduce the storage space of the features, and improve the computing efficiency; The proposed joint loss function is used to optimize the remote sensing image retrieval model; S3. The dataset processed in step S1 is trained through the deep learning network constructed in step S2, and the model parameters are optimized. The dataset is compared with various existing retrieval methods on multiple datasets to evaluate the retrieval accuracy, robustness, and generalization ability.

2. The remote sensing image retrieval method based on multi-scale feature representation according to claim 1, characterized in that: The specific operations of S1 are as follows: S1.

1. Perform stratified sampling on the original data; divide the data into training, validation, and test sets according to the preset ratios, focusing on maintaining the balance of sample size, category distribution, and scenario categories among the subsets; This ensures that the entire process of model training, hyperparameter optimization, and performance verification is based on statistically significant data, thereby improving the model's generalization capabilities in real-world application scenarios. S1.

2. The constructed network has two branches with different input image sizes. The images are scaled by the ratio of RGB and RGB respectively, and then the RGB channel values ​​are normalized and standardized. Standardization reduces the instability caused by the difference value. That is, the channel mean is subtracted from each input channel and divided by the channel standard deviation. S1.

3. Implement data augmentation strategies on the training set; systematically expand the morphological space representation of the training data through geometric deformation operations such as random rotation transformations, multi-scale scaling, and mirror flipping.

3. The remote sensing image retrieval method based on multi-scale feature representation according to claim 1, characterized in that: The specific operation of S2 is as follows: S2.

1. Resize the input data of the dual-branch structure to 224x224 and 240x240, respectively. Then, set convolution kernels of different sizes (16x16 and 12x12) to segment the input image, obtain feature information of remote sensing images at different scales, and then use the self-attention mechanism module to extract deep features of the image. S2.

2. Add a multi-level feature fusion module after the multi-level self-attention module, concatenate the output features of each self-attention module in each branch, and then use 1D convolution to achieve adaptive fusion of features at different levels; S2.

3. After the multi-level feature fusion module, a multi-scale feature fusion module is added to achieve the complementarity of multi-scale features on both branches. By introducing a cross-attention mechanism, the class tokens containing global information in the dual-branch features are exchanged, and then self-attention is interacted with the local feature patch tokens of the other scale to achieve the complementarity of global feature information and obtain the attention weights between the global and local at different scales. The information of the patch tokens is also fused, and the local features of different scales are transformed into the same feature dimension to achieve the fusion of patch tokens. The image features of different scales are obtained on both branches. S2.

4. Design a hash module to fuse the output features of the two branches and generate a binary hash code through a hash function. S2.5, use Represents the overall loss function, which includes the supervised contrast loss function and polarization loss function The parameters of the entire network architecture are optimized through the loss function; the expression of the supervised contrast loss function is: Where Q represents the index set of the query image; S(i) represents the index set of positive samples with the same label as the query sample; n(S(i)) represents the number of indices contained in S(i); ε is a hyperparameter set to 1e-5 to avoid n(S(i)) being 0; τ represents the temperature parameter set to 0.1; the supervised contrast loss function is generated by extending self-supervised contrastive learning to the fully supervised task, making full use of label information to bring the features of similar samples closer while separating the features of heterogeneous samples, thereby reducing the deviation between semantic similarity and Hamming distance; The expression for polarization loss is: In the formula, the margin threshold is pre-set to m=1, f i represents the output feature of the i-th sample, t i Represents the target vector; the polarization loss reduces the quantization error of the hash code through a binary distribution polarization strategy, ensuring a balance between coding compactness and retrieval efficiency.

4. The remote sensing image retrieval method based on multi-scale feature representation according to claim 3, characterized in that: In S2.2, formula (1) represents the process of extracting features from each level using the self-attention mechanism Where g i Represents the output features of each layer of the self-attention module, and SA represents the calculation process of the self-attention module; Formula (2) represents the multi-level feature fusion process. After the input features pass through N layers of self-attention modules, the output features are spliced ​​and the dimension is reduced through convolution operation to achieve adaptive fusion of features. g all =Conv(Concat(g1;g2;...;g i )) (2) In the formula, Concat represents the feature concatenation operation and Conv represents convolution.

5. The remote sensing image retrieval method based on multi-scale feature representation according to claim 3 is characterized in that: In S2.3, the specific calculation process of the large-scale feature fusion module is as shown in formula (3), which realizes the transformation of global features and splicing with small-scale branch local features; Where, F L The (·) function represents the projection function, which is used to align the parameter dimensions. Class tokens representing large-scale branches, patch tokens representing small-scale branches; Formula (4) represents the introduction of a multi-head self-attention mechanism to achieve the attention weight distribution of class tokens. The features obtained by formula (3) are first aligned with the original dimension through the projection function, and then the new class token is obtained through the residual connection; Where, F L (·) function represents the projection function, MAtt represents the calculation process of the multi-head self-attention mechanism; Formula (5) represents the process of splicing the patch tokens of the two branches to achieve the fusion process of patch tokens; Where, F L (·) function represents the projection function, and Concat represents the feature concatenation operation; Formula (6) reverse projects the fused patch tokens to the input dimension through projection transformation; Where g L (·) represents the inverse projection function; Formula (7) represents the concatenation of the fused class tokens and patch tokens to obtain the new feature Z L ; Where g L (·) represents the back-projection function, which is used to align the fused class token and patch token dimensions.

6. The remote sensing image retrieval method based on multi-scale feature representation according to claim 3, characterized in that: In S2.4, formula (8) projects the real-valued features output by the multi-scale feature fusion module to a dimension consistent with the preset hash code length for subsequent hash code generation; Where U(·) represents the mapping function, and Token-like features representing the output of the dual branch; Formula (9) represents the fusion process of real-valued features. The feature values ​​of the corresponding dimensions of the two-branch features are added and averaged to obtain the final feature. Where L represents the length of the preset hash code, and Respectively represent the eigenvalues ​​at the i-th position; Formula (10) represents the process of generating binary hash codes. The real-valued features are compressed by the hash function to generate the corresponding binary hash codes. Where, represents a real-valued feature of length L, and Sign represents a sign function, which is used to generate a hash code. Since the class token contains global information, only the class token is used for fusion in this process to generate a binary hash code.