Lightweight remote sensing image retrieval method based on asymmetric Hash

By improving the FasterNet network FSNet, using partial convolution PConv and structural embedding module, the computing efficiency and memory access problems in remote sensing image feature extraction are solved, and efficient and low resource consumption are achieved, and the retrieval accuracy and efficiency are improved.

CN120541259APending Publication Date: 2025-08-26ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510673175.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing asymmetric hashing method has insufficient computational efficiency and memory access performance during the remote sensing image feature extraction stage, which affects the model's retrieval efficiency and accuracy in complex scenarios.

Method used

The improved FasterNet lightweight network FSNet is constructed, and the partially convolutional PConv is used to reduce computational redundancy, and a structural embedding module is embedded in the backbone network, combining linear classifiers and asymmetric similarity information to generate discriminant hash codes.

Benefits of technology

While maintaining high retrieval accuracy, the efficiency of remote sensing image retrieval is significantly improved, the consumption of computing resources is reduced, and a more discriminant hash code is generated, which is suitable for rapid retrieval of large-scale remote sensing images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541259A_ABST
    Figure CN120541259A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight remote sensing image retrieval method based on asymmetric Hash, and the method comprises the steps: constructing a lightweight network FSNet of an improved FasterNet in a feature extraction stage, employing a partial convolution PConv to carry out the convolution operation of a part of channels of an input feature map, keeping the characteristics of the remaining part of channels unchanged, effectively reducing the calculation redundancy, and improving the retrieval efficiency. And the calculation efficiency is improved. According to the method, the optimized structure embedding module is embedded in the backbone network, the expression of the structure information in the image is enhanced by capturing the internal structure of the image while efficient extraction of deep features is realized, and the retrieval precision is further improved on the basis of improving the retrieval efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing image technology, and in particular to a lightweight remote sensing image retrieval method based on asymmetric hashing. Background Art

[0002] With the rapid development of remote sensing technology, the amount of remote sensing data acquired has increased exponentially. Covering multiple layers and dimensions, from surface coverage to atmospheric conditions, these data contain rich spatial, spectral, and structured information. These data provide critical support for target recognition, object classification, and environmental monitoring. However, faced with the massive and diverse volume of remote sensing imagery, achieving efficient and accurate target retrieval has become a core challenge in remote sensing data processing.

[0003] Deep learning provides a revolutionary technological path for remote sensing image retrieval. By building deep neural networks that mimic the human brain's hierarchical cognitive mechanisms, it can directly learn multi-level abstract features from massive amounts of data, overcoming the limitations of traditional handcrafted features, which suffer from weak expressiveness and a significant semantic gap. The synergistic effect of convolutional architectures and attention mechanisms allows for the simultaneous extraction of local details and global semantics within an image, significantly enhancing the ability to identify complex features. This paradigm shift from rule-driven to data-driven approaches has significantly propelled remote sensing retrieval systems toward intelligent and adaptive development.

[0004] Remote sensing images typically feature high resolution and complex scenes, and traditional manual feature extraction methods have significant limitations. Early retrieval methods primarily relied on low-level visual features and specific image descriptors (such as color and texture), combined with metric learning or indexing structures. However, these methods lacked sufficient feature representation and were computationally inefficient. Deep learning technology, through the training mechanism of neural networks, mimics the cognitive processes of the human brain, automatically learning and extracting features at different levels, improving the representation of image features and bringing new breakthroughs to remote sensing image retrieval. Retrieval methods based on convolutional neural networks (CNNs) extract global feature vectors from remote sensing images and, through transfer learning techniques, alleviate the scarcity of annotated data, achieving excellent retrieval results. Building on this foundation, deep hashing technology further compresses high-dimensional data into low-dimensional binary hash codes, significantly improving retrieval efficiency and accuracy. Mainstream symmetric deep hashing methods optimize hash codes by combining data features with label information, but their semantic representation and feature discrimination remain limited in complex scenes. To this end, researchers have gradually improved the semantic representation capabilities of hash codes by fusing multi-scale features, introducing contrastive learning mechanisms, and optimizing classification loss functions.

[0005] Current asymmetric hashing methods rely primarily on dense convolutional computations during the feature extraction phase, which still suffers from significant deficiencies in computational efficiency and memory access performance. In recent years, to improve model computational efficiency, lightweight convolutional networks (CNNs) such as MobileNet, ShuffleNet, and GhostNet have primarily adopted depthwise separable convolution (DWConv) or group convolution (GConv) to reduce computational overhead. However, while these methods reduce FLOPs, they often increase memory access overhead, resulting in low floating-point operations per second (FLOPS), limiting the actual acceleration effect. FasterNet, a new lightweight CNN, utilizes partial convolution (PConv) to perform convolution computations on only a portion of the input channels and subsequently fuses information through pointwise convolution (PWConv), effectively reducing computational complexity. However, FasterNet still has certain limitations in expressing the structural features of images. PConv only utilizes partial channel information during feature extraction, which weakens the image's structural information and thus affects its ability to model complex spatial relationships. Summary of the Invention

[0006] The purpose of this invention is to provide a lightweight remote sensing image retrieval method based on asymmetric hashing. During the feature extraction phase, the present invention constructs a lightweight network, FSNet, which improves on FasterNet. Using partial convolution (PConv), the method applies convolution operations to a subset of channels in the input feature map, while leaving the characteristics of the remaining channels unchanged. This effectively reduces computational redundancy and improves computational efficiency. By embedding an optimized structural embedding module within the backbone network, the present invention achieves efficient deep feature extraction while capturing the internal structure of the image to enhance the representation of structural information within the image. This improves retrieval accuracy while also increasing retrieval efficiency.

[0007] The object of the present invention is achieved like this: A lightweight remote sensing image retrieval method based on asymmetric hashing includes the following steps: S1. Preprocess the original data and divide the data set into training set and test set for future use; S2. Construct a fast asymmetric hash remote sensing image retrieval model based on the improved FasterNet, as follows: Construct a lightweight feature extraction network FSNet based on the improved FasterNet. The lightweight feature extraction network FSNet uses partial convolution PConv to perform convolution in specific areas of the input channel to reduce redundant calculations, and embeds the optimized structure embedding module in the FasterNet backbone network to achieve efficient extraction of deep features while capturing the internal structure of the image to enhance the expression of structural information in the image and improve the feature expression ability of the hash code; introduce a linear classifier to capture the semantic features of the image label and combine it with asymmetric similarity information to generate a discriminative image hash code to achieve effective improvement in similarity sorting. S3: Train and evaluate the dataset preprocessed by S1 using the model constructed by S2, select the optimal model, and compare and analyze the performance of each model under different indicators.

[0008] Said S1 comprises the following steps: S1.1. Perform stratified sampling on the original data, dividing it into training, validation, and test sets according to preset proportions. Maintain balance among the subsets in terms of sample size, category distribution, and scenario categories. This ensures that the entire process of model training, hyperparameter optimization, and performance verification is based on statistically significant data, thereby improving the model's generalization capabilities in real-world application scenarios. S1.2. All 256×256 and 600×600 pixel input images are scaled to 224x224, and then the RGB channel values ​​are normalized to reduce the instability caused by the difference values. S1.3. Implement data augmentation strategies in the training set; systematically expand the morphological space expression of the training data through geometric deformation operations such as random rotation transformation, multi-scale scaling, and mirror flipping; this geometric invariance enhancement strategy effectively alleviates the model overfitting problem caused by shooting angles, target scale differences, and diversity of ground object morphology in remote sensing images, and improves the spatial adaptability of the feature extraction network.

[0009] The S2 comprises the following steps: S2.1. Replace the original ResNet network with the FasterNet-TO backbone network of the original FSNet model. This network uses convolution PConv to apply conventional Conv to a portion of the input channels for spatial feature extraction, while leaving the other channels unchanged. This reduces computational redundancy and improves computational efficiency. S2.2. In the spatial feature extraction stage, an optimized structure embedding module is added to encode high-dimensional self-similarity into a compact self-similar descriptor. At the same time, different geometric structures are learned and analyzed from different images, thereby enhancing the network's ability to understand image details and geometric structures and improving the feature expression ability of hash codes. The structure embedding network module is mainly composed of three sub-modules, namely the self-similarity calculation module, the self-similarity encoding module, and the feature fusion module. Among them, the self-similarity calculation module generates a non-negative self-similarity descriptor for each pixel position and its surrounding area through linear layer dimensionality reduction and cosine similarity calculation; the self-similarity encoding module uses a sequence of convolution blocks to encode the descriptor into a dense descriptor of the same size as the initial feature map, and adjusts the channel dimension through a linear layer; the feature fusion module adaptively adjusts the contribution ratio of the initial features and self-similarity descriptors through the dynamic weight module (DWG) to achieve more optimized feature fusion, thereby enhancing the expression ability of global visual content while retaining local structural information; S2.3 performs adaptive average pooling and 1×1 convolution on the output features fused in S2.2. AAP maps input features of any size to a fixed dimension through spatial dimensionality reduction, ensuring that the model has better adaptability when processing inputs of different resolutions. Conv1×1 is applied to the final output layer to integrate and compress feature dimensions to improve the compactness of feature expression while reducing computational overhead, providing more recognizable feature representations for subsequent image retrieval tasks. S2.4. Construct a hash code learning strategy based on joint loss to enhance the discriminative ability of hash codes by combining the semantic features captured by the classifier with the asymmetric similarity constraints of images.

[0010] In S2.1, the FLOPs and memory access of common convolution and PConv are shown below: h×w×k 2 ×c 2 h×w×2c+k 2 ×c 2 ≈h×w×2c Among them, h, w, c are the height, width and number of channels of the input image respectively, and k is the convolution kernel size.

[0011] When the ratio of channels involved in convolution c pWhen the number of FLOPs and memory access of PConv is 1 / 4, the FLOPs and memory access of PConv are 1 / 16 and 1 / 4 of those of conventional convolution respectively; the MLPBlock structure of FasterNet-TO is optimized, and the stacking structure analysis is adopted to eliminate redundant modules; the present invention further designs a lightweight FasterNet backbone network, aiming to reduce computational overhead while improving the ability to extract remote sensing image features. The optimization strategy includes: (1) reserving more computing resources for deep key layers; (2) improving the settings of embedding layers and convolution layers to adapt them to the characteristics of remote sensing image data; by comparing the performance of different backbone network structures on the remote sensing image datasets of UCMD and AID, the experimental results show that the blocks in the last two stages consume less memory access, so more computing resources are allocated to these two stages to improve feature extraction efficiency. In the process of multiple experiments and optimization iterations, the design of the FasterNet backbone network is continuously verified and adjusted, and finally an efficient lightweight backbone network is constructed.

[0012] In S2.2, the feature expression containing the image structure information is calculated by the formula as follows: D=W d L(Conv N (S))+b d F * =α·F(x)+(1-α)·D(x) F S =max(0,W1F * +b1)W2+b2 Where: u∈[1, C′] refers to the index of the channel, v∈[-v P , v P ]×[-v P , v P ] is the relative position of each pixel x in the surrounding area of ​​size P×P, where v p =(P-1) / 2.

[0013] While structural information reflects the local content of an image, to better reflect the global content of remote sensing images, attention must also be paid to the image's visual information. Therefore, the self-similarity code D and the initial features F need to be fused. The current module adds the original feature map and the self-similarity descriptor and performs a linear transformation without assigning weights based on their importance. This can lead to insufficient or excessive information fusion in certain scenarios. A dynamic weight module (DWG) is constructed within the feature fusion submodule within the structural embedding module. This module dynamically generates weights o and 1-o to adjust the contribution ratio of the original features F and the self-similarity descriptor D, achieving adaptive weighted feature fusion.

[0014] In S2.4, a joint loss function is constructed to effectively balance the weights between different objectives during training and ensure that the learned hash codes perform well in terms of similarity ranking and classification accuracy. The joint loss function is as follows: Where λ and γ are hyperparameters, represents asymmetric loss, represents the classification loss; Asymmetric hashing effectively improves the accuracy and robustness of image retrieval by introducing similarity constraints in the hash code generation process. n database images are represented as The similarity information of pairwise supervision is obtained by multiplying the single-hot encoding and denoted as S = {-1, +1} m×n , s ij =1 indicates q i and d j are similar, whereas s ij =-1 means dissimilarity, forming the semantic similarity matrix between the two; Represents the hash code matrix of the database image D; its formula is expressed as: In the formula, F(x i ; θ) represents an efficient feature extraction network, x d is the database image hash code, I is the hash code length, s ij is the image similarity information, and its value is {1, -1}.

[0015] The classification loss is used to measure the consistency between the generated hash code and the actual semantic label to ensure that the semantic information of each image can be accurately reflected; the hash code is mapped to the category prediction space through a linear classifier and combined with the real semantic label for supervised optimization. Its formula is expressed as: Among them, yi represents the true label (0 or 1), pi represents the predicted probability of the model, represents the classification loss, and n represents the number of samples.

[0016] The present invention has the following beneficial effects: It designs a lightweight feature extraction network, FSNet. By introducing a lightweight design and structural embedding module, it significantly improves retrieval efficiency while maintaining high retrieval accuracy. It also utilizes a linear classifier to capture label semantic information and combines it with asymmetric similarity to generate more discriminative hash codes, providing an efficient and low-resource solution for the rapid retrieval of large-scale remote sensing imagery. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1This is the overall network architecture diagram of the present invention; Figure 2 It is a module diagram embedded in the structure of the present invention; Figure 3 This is a diagram of the structural embedding feature fusion module of the present invention; Figure 4 This is an example of the top 30 retrieval results of the comparison methods in the three datasets; Figure 5 Comparison chart for 64-bit hash code retrieval on the WHURS-19 dataset. DETAILED DESCRIPTION

[0018] The present invention will be further described below with reference to the accompanying drawings and examples.

[0019] A lightweight remote sensing image retrieval method based on asymmetric hashing, such as Figure 1 As shown, the following steps are included: S1. Preprocess the raw data and divide the dataset into training and test sets for future use. Three publicly available remote sensing image retrieval datasets, WHU-RS19, UC Merced, and AID, were used. Samples from each class were randomly sampled at a fixed ratio to ensure balanced data within each class. To ensure the network focuses more on semantic features within the images, the images were uniformly preprocessed using scaling, normalization, and standardization.

[0020] Said S1 comprises the following steps: S1.1. Perform stratified sampling on the original data and divide the data into training, validation, and test sets according to preset proportions. Maintain the balance of each subset in terms of sample size, category distribution, and scenario category to ensure that the entire process of model training, hyperparameter optimization, and performance verification is based on statistically significant data, thereby improving the model's generalization ability in real application scenarios.

[0021] S1.2. All 256×256 and 600×600 pixel input images are scaled to 224x224, and then the RGB three-channel values ​​are normalized and calculated to reduce the instability caused by the difference value; that is, the channel mean is subtracted from each input channel and divided by the channel standard deviation. In this paper, the mean and standard deviation are set to [0.485, 0.456, 0.406] and [0.229, 0.224, 0.225].

[0022] S1.3. Implement data augmentation strategies in the training set; systematically expand the morphological space expression of the training data through geometric deformation operations such as random rotation transformation, multi-scale scaling, and mirror flipping; this geometric invariance enhancement strategy effectively alleviates the model overfitting problem caused by shooting angles, target scale differences, and diversity of ground object morphology in remote sensing images, and improves the spatial adaptability of the feature extraction network.

[0023] S2. Build a fast asymmetric hash remote sensing image retrieval model based on improved FasterNet, such as Figure 2 As shown in the figure, the details are as follows: a lightweight feature extraction network FSNet based on the improved FasterNet is constructed. The lightweight feature extraction network FSNet uses partial convolution PConv to perform convolution in specific areas of the input channel to reduce redundant calculations, and embeds the optimized structure embedding module in the FasterNet backbone network to achieve efficient extraction of deep features while capturing the internal structure of the image to enhance the expression of structural information in the image and improve the feature expression ability of the hash code; a linear classifier is introduced to capture the semantic features of the image label and combined with asymmetric similarity information to generate a discriminative image hash code to achieve effective improvement in similarity sorting.

[0024] The S2 comprises the following steps: S2.1. Use the FasterNet-TO backbone network of the original FSNet model FasterNet to replace the original ResNet network. The FasterNet-TO backbone network uses convolution PConv to apply conventional Conv on a part of the input channel for spatial feature extraction, while keeping other channels unchanged; to reduce computational redundancy and improve computational efficiency.

[0025] In S2.1, the FLOPs and memory access of common convolution and PConv are shown below: h×w×k 2 ×c 2 h×w×2c+k 2 ×c 2 ≈h×w×2c Among them, h, w, c are the height, width and number of channels of the input image respectively, and k is the convolution kernel size.

[0026] When the ratio of channels involved in convolution c pWhen the number of FLOPs and memory access of PConv is 1 / 4, the FLOPs and memory access of PConv are 1 / 16 and 1 / 4 of those of conventional convolution respectively; the MLPBlock structure of FasterNet-TO is optimized, and the stacking structure analysis is adopted to eliminate redundant modules; the present invention further designs a lightweight FasterNet backbone network, aiming to reduce computational overhead while improving the ability to extract remote sensing image features. The optimization strategy includes: (1) reserving more computing resources for deep key layers; (2) improving the settings of embedding layers and convolution layers to adapt them to the characteristics of remote sensing image data; by comparing the performance of different backbone network structures on the remote sensing image datasets of UCMD and AID, the experimental results show that the blocks in the last two stages consume less memory access, so more computing resources are allocated to these two stages to improve feature extraction efficiency. In the process of multiple experiments and optimization iterations, the design of the FasterNet backbone network is continuously verified and adjusted, and finally an efficient lightweight backbone network is constructed.

[0027] S2.2. An optimized structure embedding module is added to the spatial feature extraction stage to encode high-dimensional self-similarity into compact self-similar descriptors. At the same time, different geometric structures are learned and analyzed from different images, thereby enhancing the network's understanding of image details and geometric structures and improving the feature expression ability of hash codes. The structure embedding network module mainly consists of three submodules: the self-similarity calculation module, the self-similarity encoding module, and the feature fusion module. The self-similarity calculation module generates a non-negative self-similarity descriptor for each pixel position and its surrounding area through linear layer dimensionality reduction and cosine similarity calculation. The self-similarity encoding module encodes this descriptor into a dense descriptor of the same size as the initial feature map using a sequence of convolutional blocks and adjusts the channel dimension through a linear layer. The feature fusion module adaptively adjusts the contribution ratio of the initial features and self-similarity descriptors through a dynamic weight module (DWG) to achieve more optimized feature fusion, thereby enhancing the expression ability of global visual content while preserving local structural information.

[0028] In S2.2, the feature expression containing the image structure information is calculated by the formula as follows: D=W d L(Conv N (S))+b d F * =α·F(x)+(1-α)·D(x) F S =max(0,W1F * +b1)W2+b2 Where: u∈[1, C′] refers to the index of the channel, v∈[-vP , v P ]×[-v P , v P ] is the relative position of each pixel x in the surrounding area of ​​size P×P, where v p =(P-1) / 2.

[0029] Structural information reflects the local content of an image. To better reflect the global content of a remote sensing image, it is also necessary to pay attention to the visual information of the image. Therefore, it is necessary to fuse the self-similarity code D and the initial feature F. The current module adds the original feature map and the self-similarity descriptor and performs a linear transformation without assigning weights to the importance of the two. This may lead to insufficient or excessive information fusion in certain scenarios. For example, Figure 3 As shown in the figure, a dynamic weight module DWG is constructed in the feature fusion submodule in the structure embedding module to dynamically generate weights o and 1-o to adjust the contribution ratio of the original feature F and the self-similarity descriptor D to achieve adaptive weighted fusion of features.

[0030] S2.3, adaptive average pooling and 1×1 convolution are performed on the output features fused by S2.2; AAP maps input features of any size to a fixed dimension through spatial dimensionality reduction, ensuring that the model has better adaptability when processing inputs of different resolutions; Conv1×1 acts on the final output layer to integrate and compress feature dimensions to improve the compactness of feature expression while reducing computational overhead, providing more recognizable feature representations for subsequent image retrieval tasks.

[0031] S2.4. Construct a hash code learning strategy based on joint loss to enhance the discriminative ability of hash codes by combining the semantic features captured by the classifier with the asymmetric similarity constraints of images.

[0032] In S2.4, a joint loss function is constructed to effectively balance the weights between different objectives during the training process and ensure that the learned hash code performs well in terms of similarity ranking and classification accuracy. The joint loss function is as follows: Where λ and γ are hyperparameters, represents asymmetric loss, represents the classification loss; Asymmetric hashing effectively improves the accuracy and robustness of image retrieval by introducing similarity constraints in the hash code generation process. n database images are represented as The similarity information of pairwise supervision is obtained by multiplying the single-hot encoding and denoted as S = {-1, +1} m×n , s ij =1 indicates qi and d j are similar, whereas s ij =-1 means dissimilarity, forming the semantic similarity matrix between the two; Represents the hash code matrix of the database image D; its formula is expressed as: In the formula, F(x i ; θ) represents an efficient feature extraction network, x d is the database image hash code, l is the hash code length, s ij is the image similarity information, and its value is {1, -1}.

[0033] The classification loss is used to measure the consistency between the generated hash code and the actual semantic label to ensure that the semantic information of each image can be accurately reflected; the hash code is mapped to the category prediction space through a linear classifier and combined with the real semantic label for supervised optimization. Its formula is expressed as: Among them, y i represents the true label (0 or 1), p i represents the predicted probability of the model, represents the classification loss, and n represents the number of samples.

[0034] S3: Train and evaluate the dataset preprocessed by S1 using the model constructed by S2, select the optimal model, and compare and analyze the performance of each model under different indicators.

[0035] Specifically, the multiple models or parameter configurations obtained during the training and validation phases are first screened through unified evaluation on a test set. The data in the test set has not been used in model training or parameter tuning, and its diversity and complexity more accurately reflect the model's generalization performance in real-world scenarios. To validate the superiority of this estimation model, it is compared with several widely used deep learning models. For the WHU-RS19, UCMD, and AID datasets, 60%, 80%, and 70% of the datasets, respectively, are randomly selected as training sets, and the remaining datasets are used as test sets. All input images are resized to 224x224 pixels. The hyperparameter settings include an initial learning rate of 0.001, a weight decay exponent of 0.0001, a batch size of 64, and hyperparameters Y and λ in the loss function of 20 and 200, respectively. In addition, the experiments were implemented in the Pytorch framework with Python version 3.10. The processor was a 13th Gen Intel(R) Core(TM) i7-13700H CPU, the graphics card was an NVIDIA GeForce RTX 4070 Laptop GPU, 64G memory, and the CUDA version was 12.2.91. The model parameters with the best results during training were saved. Table 1. Comparison methods on the three datasets on the left mAP Quantitative comparison results table

[0036] To ensure fair comparison, all estimation models are trained and tested in the same hardware and software environment.

[0037] Analysis results Table 1 and results Figure 4 、 Figure 5 It can be seen that in the field of remote sensing image retrieval, the fast asymmetric hashing method FAHM proposed in this paper demonstrates high retrieval accuracy across all datasets and hash code lengths, demonstrating that the present invention can maintain high retrieval accuracy and even achieve improved accuracy in some experiments. This performance is due to the fast asymmetric hashing method FAHM combining the lightweight FasterNet and structure embedding modules in the feature extraction stage, thereby enhancing the ability to express image structural information while maintaining computational efficiency.

[0038] To verify the effectiveness of the fast asymmetric hashing method (FAHM) proposed in this paper in improving retrieval efficiency, we conducted a time-comparison experiment on three public benchmark datasets (WHURS19, UCMD, and AID) comparing various methods in terms of both training time and retrieval time. In the experimental setup, all parameters of each method were set to their default configurations. Retrieval time was the time it took to perform similarity sorting between the query image and the database hash code saved after training. The comparison results are shown in Table 2: Table 2. Quantitative comparison results of time indicators of the comparison methods on three datasets (unit: seconds)

[0039] As can be seen from Table 2, in terms of training time, the training time of the fast asymmetric hashing method FAHM of the present invention on the three data sets is significantly lower than that of other methods. The significant advantage of the fast asymmetric hashing method FAHM of the present invention in training time is attributed to its combination with the lightweight FasterNet feature extraction module, which not only reduces the number of parameters but also speeds up the training speed. At the same time, the optimized structure embedding module can efficiently capture the internal structural information of the image, avoiding the time consumption of redundant calculations of traditional methods. In terms of retrieval time, the fast asymmetric hashing method FAHM of the present invention also shows certain advantages, which is due to the compact and highly discriminative hash code generated by the asymmetric hashing strategy, and the introduction of the linear classifier also helps to generate more discriminative hash codes, indirectly reducing the similarity calculation overhead between the database image and the query image during the retrieval process.

[0040] Experimental results show that the fast asymmetric hashing method FAHM proposed in this paper provides a solution that combines efficiency and accuracy for remote sensing image retrieval with its efficient feature extraction capability, accurate structural information expression and balanced retrieval performance.

Claims

1. A lightweight remote sensing image retrieval method based on asymmetric hashing, characterized in that: The following steps are involved: S1. Preprocess the original data and divide the data set into training set and test set for future use; S2. Construct a fast asymmetric hash remote sensing image retrieval model based on the improved FasterNet. Specifically, a lightweight feature extraction network FSNet based on the improved FasterNet is constructed. The lightweight feature extraction network FSNet uses partial convolution PConv to perform convolution in specific areas of the input channel to reduce redundant calculations, and embeds an optimized structure embedding module in the backbone network to achieve efficient extraction of deep features while capturing the internal structure of the image to enhance the expression of structural information in the image and improve the feature expression ability of the hash code. A linear classifier is introduced to capture the semantic features of the image label and combined with asymmetric similarity information to generate a discriminative image hash code to achieve effective improvement in similarity sorting. S3: Train and evaluate the dataset preprocessed by S1 using the model constructed by S2, select the optimal model, and compare and analyze the performance of each model under different indicators.

2. The lightweight remote sensing image retrieval method based on asymmetric hashing according to claim 1 is characterized in that: Said S1 comprises the following steps: S1.

1. Perform stratified sampling on the original data, dividing it into training, validation, and test sets according to preset proportions. Maintain balance among the subsets in terms of sample size, category distribution, and scenario categories. This ensures that the entire process of model training, hyperparameter optimization, and performance verification is based on statistically significant data, thereby improving the model's generalization capabilities in real-world application scenarios. S1.

2. All 256×256 and 600×600 pixel input images are scaled to 224x224, and then the RGB channel values ​​are normalized to reduce the instability caused by the difference values. S1.

3. Implement data augmentation strategies in the training set; systematically expand the morphological space representation of the training data through geometric deformation operations such as random rotation transformation, multi-scale scaling, and mirror flipping; and improve the spatial adaptability of the feature extraction network.

3. The lightweight remote sensing image retrieval method based on asymmetric hashing according to claim 1 is characterized in that: The S2 comprises the following steps: S2.

1. Replace the original ResNet network with the FasterNet-T0 backbone network of the original FSNet model FasterNet. The FasterNet-T0 backbone network uses convolution PConv to apply conventional Conv to a portion of the input channels for spatial feature extraction, while leaving the other channels unchanged, to reduce computational redundancy and improve computational efficiency. S2.

2. In the spatial feature extraction stage, an optimized structure embedding module is added to encode high-dimensional self-similarity into a compact self-similar descriptor. At the same time, different geometric structures are learned and analyzed from different images, thereby enhancing the network's understanding of image details and geometric structures and improving the feature expression ability of hash codes. The structure embedding network module consists of three submodules, namely the self-similarity calculation module, the self-similarity encoding module and the feature fusion module. Among them, the self-similarity calculation module generates a non-negative self-similarity descriptor for each pixel position and its surrounding area through linear layer dimensionality reduction and cosine similarity calculation. The self-similarity encoding module uses a sequence of convolutional blocks to encode the descriptor into a dense descriptor with the same size as the initial feature map, and adjusts the channel dimension through a linear layer. The feature fusion module adaptively adjusts the contribution ratio of the initial features and self-similarity descriptors through the dynamic weight module DWG to achieve more optimized feature fusion, thereby enhancing the expression ability of global visual content while retaining local structural information. S2.3 performs adaptive average pooling and 1×1 convolution on the output features fused in S2.

2. AAP maps input features of any size to a fixed dimension through spatial dimensionality reduction, ensuring that the model has better adaptability when processing inputs of different resolutions. Conv1×1 is applied to the final output layer to integrate and compress feature dimensions to improve the compactness of feature expression while reducing computational overhead, providing more recognizable feature representations for subsequent image retrieval tasks. S2.

4. Construct a hash code learning strategy based on joint loss to enhance the discriminative ability of hash codes by combining the semantic features captured by the classifier with the asymmetric similarity constraints of images.

4. The lightweight remote sensing image retrieval method based on asymmetric hashing according to claim 1 is characterized in that: In S2.1, the FLOPs and memory access of common convolution and PConv are shown below: h×w×k 2 ×c 2 Among them, h, w, c are the height, width and number of channels of the input image respectively, and k is the convolution kernel size; When the ratio of channels involved in convolution c p When the kernel size is 1 / 4, the FLOPs and memory access of PConv are 1 / 16 and 1 / 4 of those of conventional convolution respectively. The MLPBlock structure of FasterNet-T0 is optimized, and stacking structure analysis is used to eliminate redundant modules. The optimization strategies include: (1) reserving more computing resources for deep key layers; (2) improving the settings of embedding layers and convolution layers to adapt them to the characteristics of remote sensing image data.

5. The lightweight remote sensing image retrieval method based on asymmetric hashing according to claim 1 is characterized in that: In S2.2, the feature expression containing the image structure information is calculated by the formula as follows: D=W d ·L(Conv N (S))+b d F * =α·F(x)+(1-α)·D(x) F S =max(0,W1F * +b1)W2+b2 Where: u∈[1, C′] refers to the index of the channel, v∈[-v P , v P ]×[-v P , v P ] is the relative position of each pixel x in the surrounding area of ​​size P×P, where v p =(P-1) / 2; A dynamic weight module DWG is constructed in the feature fusion submodule in the structure embedding module to dynamically generate weights α and 1-α to adjust the contribution ratio of the original feature F and the self-similarity descriptor D to achieve adaptive weighted fusion of features.

6. The lightweight remote sensing image retrieval method based on asymmetric hashing according to claim 1 is characterized in that: In S2.4, a joint loss function is constructed to effectively balance the weights between different objectives during training and ensure that the learned hash codes perform well in terms of similarity ranking and classification accuracy. The joint loss function is as follows: Where λ and γ are hyperparameters, represents asymmetric loss, represents the classification loss; Asymmetric hashing effectively improves the accuracy and robustness of image retrieval by introducing similarity constraints in the hash code generation process. n database images are represented as The similarity information of pairwise supervision is obtained by multiplying the single-hot encoding and recorded as S = {-1, +1} m×n , s ij =1 indicates q i and d j are similar, whereas s ij =-1 means dissimilarity, forming the semantic similarity matrix between the two; Represents the hash code matrix of the database image D; its formula is expressed as: In the formula, F(x i ; θ) represents an efficient feature extraction network, x d is the database image hash code, l is the hash code length, s ij is the image similarity information, with a value of {1, -1}; The classification loss is used to measure the consistency between the generated hash code and the actual semantic label to ensure that the semantic information of each image can be accurately reflected; the hash code is mapped to the category prediction space through a linear classifier and combined with the real semantic label for supervised optimization. Its formula is expressed as: Among them, y i represents the true label, y i is 0 or 1, p i represents the predicted probability of the model, represents the classification loss, and n represents the number of samples.