Small sample remote sensing image target detection algorithm based on multilevel information interaction
By adopting a multi-level information interaction algorithm in the detection of small-sample remote sensing image targets, the feature alignment is used using dual-branch coding and dynamic attention mechanisms, and high-level semantic alignment is achieved through the region-level semantic similarity matrix, the cross-scale semantic mismatch problem is solved and the accuracy and stability of the detection is improved.
Patent Information
- Application Number
- CN202510392371.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-05-16
AI Technical Summary
In the detection of small sample object of remote sensing images, cross-scale semantic mismatch often occurs, resulting in semantic drift of the model during feature migration, weakening the model's discrimination ability.
A small sample remote sensing image object detection algorithm based on multi-level information interaction is adopted to realize feature alignment and fusion of key regions in the shallow feature stage through dual-branch encoding and dynamic attention mechanism, and a region-level semantic similarity matrix is constructed in the deep feature stage, and a region-level semantic similarity matrix is calculated in combination with self-distinguishing graph weighting to calculate the similarity between regions to achieve high-level semantic alignment.
It significantly improves the consistency of the semantic level, enhances the discrimination ability of the target area, and improves the accuracy and stability of remote sensing image object detection under small sample conditions.
Smart Images

Figure CN120014247A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image target detection, and in particular to a small sample remote sensing image target detection algorithm based on multi-level information interaction. Background Art
[0002] Object detection in remote sensing images has a wide range of applications in urban planning, disaster response, environmental monitoring, and military reconnaissance. Traditional deep learning object detection methods, such as Faster R-CNN, YOLO, and RetinaNet, have made significant progress with the support of large-scale labeled data. However, in the field of remote sensing, labeling is expensive and time-consuming, especially for small objects and rare categories, it is more difficult to obtain sufficient training samples. For this reason, Few-Shot Object Detection (FSOD) was proposed, aiming to achieve effective detection of new categories with only a very small number of labeled samples (usually 1 to 10 images per category).
[0003] The prior art has the following deficiencies: In small-sample target detection in remote sensing images, the problem of cross-scale semantic mismatch often occurs. This problem mainly stems from the serious inconsistency between the support sample and the query sample in terms of spatial resolution, target scale and semantic level. Even after a unified feature extraction network and spatial alignment processing, the high-level semantic features extracted from the support image may still show semantic drift with similar targets in the query image when the target scale is small, resulting in negative transfer. At the same time, when the model transfers the features of the support sample to the query sample, it not only fails to enhance the recognition ability, but also weakens the model's discrimination ability due to incorrect semantic associations. Summary of the invention
[0004] The purpose of the present invention is to provide a small sample remote sensing image target detection algorithm based on multi-level information interaction to overcome the shortcomings of the background technology.
[0005] In order to achieve the above object, the present invention provides the following technical solution: a small sample remote sensing image target detection algorithm based on multi-level information interaction, comprising: Obtain a query image and multiple support images, and input them into a parallel feature extraction backbone network to extract multi-level feature representations; In the shallow feature stage, a dual-branch encoding method is used to perform key-value mapping on query features and support features respectively, and dynamic attention weights are calculated based on feature similarity to achieve feature alignment and fusion in key areas; In the deep feature stage, for the positive support images of the same type as the query target, a region-level semantic similarity matrix is constructed, and the similarity between regions is weightedly calculated by combining the self-distinguishing graph to achieve high-level semantic alignment; Semantic enhancement technology is used to further filter out the distinguishing feature areas in the supporting features and suppress background interference information; The query features updated through multi-level interactions are input into the target detection head to complete target classification and bounding box regression, and the detection results in the remote sensing image are output.
[0006] Preferably, the remote sensing image to be detected is used as the query image And select K labeled sample images of each category in the M categories from the training set as the support image set ;in, ; Constructing a dual-branch backbone network with the same structure, which is used to process the query image and the support image respectively, wherein the backbone network is constructed based on a convolutional neural network; The query image Input to the query branch network, pass through multiple convolutional layers in sequence, and extract shallow features , middle-level features and deep features ; Similarly, each supporting image Input to the support branch network to extract the corresponding multi-scale features ; Perform channel normalization on all output feature maps.
[0007] Preferably, in the shallow feature stage, the extracted shallow query image features are expressed as , supporting image feature representation as , where H and W are the spatial sizes of the feature map, and C is the number of channels; Use two independent convolutional encoders to project the query features and support image features to generate key and value mappings respectively: query image key, value mapping: ; Support image key and value mapping: ;in, and All are 1×1 convolution operations. and Flatten to a 2D matrix ;in, ; Calculate the point-to-point similarity weights between query features and support features .
[0008] Preferably, based on similarity weight , weighted sum of the value mappings of the support images is fused into the query features; Map the values of the original query image Image features after interaction Concatenate and integrate via convolution: ;in: represents splicing in the channel dimension, is a 1×1 convolutional layer, It is the query feature expression after the support graph information is updated, and the final output is It is the query feature after interaction enhancement.
[0009] Preferably, in the deep feature stage, the deep features of the image are queried, the query features and each positive class support feature are convolutionally encoded, and the key feature representation required for semantic alignment is extracted: Query key mapping: ; Each supported image key mapping: ;in: represents a 1×1 convolutional encoder, is the number of channels after dimensionality reduction, HW and hw represent the flattened spatial dimensions of the query image and the support image respectively; for each positive class support image, calculate the similarity score matrix of all regions between it and the query image .
[0010] Preferably, in order to increase the attention to semantically significant regions, a self-distinguishing map of the support image is introduced: for each support image feature The discriminative response map is extracted through convolution operation; the similarity matrix is weighted modulated to obtain the final regional similarity; all positive support image features are summarized into an enhanced representation and fused with the query image to weighted aggregate support features; the final output weighted aggregate support features contain the enhanced information after the original semantics of the query image is aligned with the semantics of the support map.
[0011] Preferably, the deep features of the support image are represented as: ; Compress the channel through 1×1 convolution operation to obtain semantic summary feature map and highlight the key semantic area ;in, ; based on Generate a semantic attention map to represent the semantic importance of each spatial position , the expression is: ;in: is a convolution operation with nonlinear activation, Ensure that the output weight value is in the range of (0,1); use the semantic attention map as the weight and act on the original supporting image features , highlighting discriminative regions and suppressing low weights: ;in: represents position-wise channel broadcast multiplication; is the enhanced support graph semantic feature; the enhanced support feature is normalized to obtain , the final output It will serve as input for subsequent deep semantic interaction or feature matching.
[0012] Preferably, the query image features after shallow interaction and deep semantic alignment are processed, and a region proposal network is used to generate multiple candidate target regions, i.e., candidate box regions, on the feature map; For each candidate box area, extract the features of the corresponding area in the feature map, perform spatial alignment and unify the size, and input each RoI feature into the target detection head, including the classification branch and the regression branch, to perform target category prediction and position refinement respectively; The position and size of the candidate box are adjusted according to the regression offset to obtain the predicted bounding box; non-maximum suppression is performed on all prediction results to remove duplicate predictions, and the detection result with the highest confidence is retained, which is output as the remote sensing image target detection result.
[0013] In the above technical solution, the technical effects and advantages provided by the present invention are: 1. Aiming at the semantic mismatch and negative transfer problems caused by the significant differences in scale, semantics and spatial structure between the support image and the query image in the prior art, the present invention constructs a multi-stage feature interaction mechanism from shallow to deep layers. By introducing dual-branch encoding and dynamic attention mechanism in the shallow feature stage, accurate alignment of key areas is achieved; in the deep stage, a regional semantic similarity matrix is constructed, and combined with weighted region matching of the support image self-differentiating graph, the consistency of the semantic level is significantly improved. At the same time, the semantic importance of the supporting features is modeled with the help of the semantic enhancement module, which effectively suppresses background noise interference and enhances the discrimination ability of the target area.
[0014] 2. The present invention achieves high-precision detection of complex and diverse targets in remote sensing images under small sample conditions, and is particularly suitable for remote sensing application scenarios with small target scales, unbalanced categories, and difficult labeling. This method not only improves the effectiveness of feature migration, but also enhances the detection model's ability to focus on key areas and its robustness to complex backgrounds. While reducing dependence on large-scale labeled data, it significantly improves the generalization ability and practicality of the model, and has high engineering value and application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0016] Figure 1 It is the algorithm flow chart of the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0018] For examples, see Figure 1 As shown, the small sample remote sensing image target detection algorithm based on multi-level information interaction described in this embodiment includes: Obtain a query image and multiple support images, and input them into a parallel feature extraction backbone network to extract multi-level feature representations; In the shallow feature stage, a dual-branch encoding method is used to perform key-value mapping on query features and support features respectively, and dynamic attention weights are calculated based on feature similarity to achieve feature alignment and fusion in key areas; In the deep feature stage, for the positive support images of the same type as the query target, a region-level semantic similarity matrix is constructed, and the similarity between regions is weightedly calculated by combining the self-distinguishing graph to achieve high-level semantic alignment; Semantic enhancement technology is used to further filter out the distinguishing feature areas in the supporting features and suppress background interference information; The query features updated through multi-level interactions are input into the target detection head to complete target classification and bounding box regression, and the detection results in the remote sensing image are output.
[0019] The remote sensing image to be detected is used as the query image And select K labeled sample images of each category in the M categories from the training set as the support image set ;in, .
[0020] Constructing a dual-branch backbone network with the same structure, which is used to process query images and support images respectively. The backbone network is built based on a convolutional neural network (such as ResNet) and has the ability to output multi-scale features; The query image Input to the query branch network, pass through multiple convolutional layers in sequence, and extract shallow features , middle-level features and deep features ; Similarly, each supporting image Input to the support branch network to extract the corresponding multi-scale features ; Perform channel normalization on all output feature maps to ensure that the feature expressions between different images maintain scale consistency; Using the upsampling or downsampling method, the feature maps at different scales are adjusted to a unified spatial dimension, providing a matching basis for subsequent cross-image feature interaction operations.
[0021] In the shallow feature stage, the shallow query image features extracted are represented as , and the support image features are represented as , where H and W are the spatial dimensions of the feature map, and C is the number of channels.
[0022] Use two independent convolutional encoders to perform projection transformation on the query features and support image features, generating key (Key) and value (Value) mappings respectively: Query image key and value mappings: ; Support image key and value mappings: ; Among them, and are both 1×1 convolutional operations, with an output dimension of C′ < C, used for dimensionality reduction and compression to improve computational efficiency.
[0023] Flatten and into a two-dimensional matrix ; among them, ; calculate the pairwise similarity weights between the query features and the support features ; where: is to normalize each row so that the sum of the similarity weights at each query position is 1, and T is the matrix transpose.
[0024] According to the similarity weights , weight-sum the value mapping of the support image and fuse it into the query features: ; among them, is the query feature expression updated based on the support map information.
[0025] Concatenate the value mapping of the original query image with the interacted image features and integrate them through convolution: ; where: represents concatenation in the channel dimension, is a 1×1 convolutional layer, used for feature fusion and restoring the original channel dimension; the final output is the query feature after interaction enhancement.
[0026] In the deep feature stage, the deep features of the query image: ; The deep features of K positive class support images: ; Convolutional encoding is performed on the query features and each positive class support feature to extract the key feature representation required for semantic alignment: Query key mapping: ; Each supports image key mappings: ; in: represents a 1×1 convolutional encoder, is the number of channels after dimensionality reduction, HW and hw represent the flattened spatial dimensions of the query image and support image, respectively.
[0027] For each positive support image, calculate the similarity score matrix of all regions between it and the query image: ;in: , is a learnable linear transformation (such as 1×1 convolution or full connection); softmax normalizes the similarity score of each query region. Represents the regional similarity matrix between the k-th positive class support image and the query image.
[0028] In order to increase the attention to semantically significant areas, the self-differentiation map of the support image is introduced: for each support image feature Extract discriminative response maps through convolution operations: ; Perform weighted modulation on the similarity matrix to obtain the final regional similarity: ;in: represents the category significance score of each position in the support graph; Weighted It not only considers the similarity between the two images, but also integrates the importance of the target area in the support image.
[0029] All positive support image features are aggregated into an enhanced representation and fused with the query image to weighted aggregate support features: ; If there are multiple supporting images, they can be averaged or attention-weighted fused and then spliced.
[0030] Final Output Contains enhanced information after the original semantics of the query image and the semantics of the support graph are aligned, which is used for subsequent classification and positioning.
[0031] The deep features of the support image are represented as: ; Compress the channel through 1×1 convolution operation to obtain the semantic summary feature map and highlight the key semantic area: ;in, , which is used for dimensionality reduction to improve the response efficiency of the subsequent attention mechanism.
[0032] based on , generate a semantic attention map to represent the semantic importance of each spatial position , the expression is: ;in: is a convolution operation with nonlinear activation (such as ReLU), Ensure that the output weight value is in the range of (0,1), representing the significance score of each region; higher Indicates that the location is more likely to belong to the target key area.
[0033] Use the semantic attention map as weights to act on the original support image features , highlighting discriminative regions and suppressing low weights (background or noise regions): ;in: represents position-wise channel broadcast multiplication; It is the enhanced semantic feature of the support image; the background area is weakened because its attention value is close to 0, which effectively filters the background noise.
[0034] The enhanced support features are normalized to obtain , to avoid numerical amplification affecting model stability: the final output It will serve as the input for subsequent deep semantic interaction or feature matching, improving the model's attention to key semantic areas and its robustness to background interference.
[0035] In this application, the semantic enhancement process not only effectively screens out areas that are highly correlated with the target category by explicitly modeling the semantic significance of each area in the feature map, but also adaptively suppresses common complex background interference in remote sensing images, such as vegetation, water bodies, road textures, etc., thereby improving the accuracy and stability of target detection under small sample conditions.
[0036] The query features updated through multi-level interactions are input into the target detection head to complete target classification and bounding box regression, and the detection results in the remote sensing image are output.
[0037] The query image features after shallow interaction and deep semantic alignment are expressed as: ; This feature contains the original information of the query image and the semantic enhancement information of the fused self-support image, and is the high-quality representation ultimately used for detection.
[0038] Using region proposal network in feature map Generate multiple candidate target regions: ; Among them: each represents a candidate box, N is the number of proposed regions, and the proposed regions are sampled at different scales and aspect ratios to cover possible target locations.
[0039] For each candidate box area , extract the features of the corresponding area in the feature map, align the space and unify the size (e.g., 7×7): ; Used to ensure consistency in feature representation for target regions of different sizes.
[0040] Each RoI feature Input the target detection head, including the classification branch and the regression branch, which perform target category prediction and position refinement respectively: Classification branch output: ;Regression branch output: ;in: is the category probability distribution of each RoI, is the regression offset value of the bounding box, , , , are all learnable parameters.
[0041] According to the regression shift Adjust the candidate box The position and size of the predicted bounding box are: ; For all prediction results Non-maximum suppression is performed to remove duplicate predictions and retain the detections with the highest confidence.
[0042] Output remote sensing image target detection results, the final output is: a set of accurately located target bounding boxes ; The target category and confidence level corresponding to each bounding box These results constitute the target detection output in remote sensing images, which is suitable for high-precision target recognition and localization tasks in small sample environments.
[0043] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.
Claims
1. A small sample remote sensing image target detection algorithm based on multi-level information interaction, characterized by: include: Obtain a query image and multiple support images, and input them into a parallel feature extraction backbone network to extract multi-level feature representations; In the shallow feature stage, a dual-branch encoding method is used to perform key-value mapping on query features and support features respectively, and dynamic attention weights are calculated based on feature similarity to achieve feature alignment and fusion in key areas; In the deep feature stage, for the positive support images of the same type as the query target, a region-level semantic similarity matrix is constructed, and the similarity between regions is weightedly calculated by combining the self-distinguishing graph to achieve high-level semantic alignment; Semantic enhancement technology is used to further filter out the distinguishing feature areas in the supporting features and suppress background interference information; The query features updated through multi-level interactions are input into the target detection head to complete target classification and bounding box regression, and the detection results in the remote sensing image are output.
2. The small sample remote sensing image target detection algorithm based on multi-level information interaction according to claim 1 is characterized by: The remote sensing image to be detected is used as the query image And select K labeled sample images of each category in the M categories from the training set as the support image set ;in, ; Constructing a dual-branch backbone network with the same structure, which is used to process the query image and the support image respectively, wherein the backbone network is constructed based on a convolutional neural network; The query image Input to the query branch network, pass through multiple convolutional layers in sequence, and extract shallow features , middle-level features and deep features ; Similarly, each supporting image Input to the support branch network to extract the corresponding multi-scale features ; Perform channel normalization on all output feature maps.
3. The small sample remote sensing image target detection algorithm based on multi-level information interaction according to claim 1 is characterized by: In the shallow feature stage, the extracted shallow query image features are expressed as , supporting image feature representation as , where H and W are the spatial sizes of the feature map, and C is the number of channels; Use two independent convolutional encoders to project the query features and support image features to generate key and value mappings respectively: query image key, value mapping: ; Support image key and value mapping: ;in, and All are 1×1 convolution operations. and Flatten to a 2D matrix ;in, ; Calculate the point-to-point similarity weights between query features and support features .
4. The small sample remote sensing image target detection algorithm based on multi-level information interaction according to claim 3 is characterized by: Based on similarity weight , weighted sum of the value mappings of the support images is fused into the query features; Map the values of the original query image Image features after interaction Concatenate and integrate via convolution: ;in: represents concatenation in the channel dimension, is a 1×1 convolutional layer, It is the query feature expression after the support graph information is updated, and the final output is It is the query feature after interaction enhancement.
5. The small sample remote sensing image target detection algorithm based on multi-level information interaction according to claim 1 is characterized by: In the deep feature stage, the deep features of the query image are queried, and the query features and each positive class support feature are convolutionally encoded to extract the key feature representation required for semantic alignment: query key mapping: ; Each supported image key mapping: ;in: represents a 1×1 convolutional encoder, is the number of channels after dimensionality reduction, HW and hw represent the flattened spatial dimensions of the query image and the support image respectively; for each positive class support image, calculate the similarity score matrix of all regions between it and the query image .
6. The small sample remote sensing image target detection algorithm based on multi-level information interaction according to claim 5 is characterized by: In order to increase the attention to semantically significant areas, the self-differentiation map of the support image is introduced: for each support image feature Extract discriminative response maps through convolution operations; perform weighted modulation on the similarity matrix to obtain the final regional similarity; summarize all positive class support image features into an enhanced representation and fuse them with the query image to weighted aggregate support features; The final output weighted aggregate support features contain the enhanced information after the original semantics of the query image is aligned with the semantics of the support graph.
7. The small sample remote sensing image target detection algorithm based on multi-level information interaction according to claim 1 is characterized by: The deep features of the support image are represented as: ; Compress the channel through 1×1 convolution operation to obtain semantic summary feature map and highlight the key semantic area ;in, ; based on Generate a semantic attention map to represent the semantic importance of each spatial position , the expression is: ;in: is a convolution operation with nonlinear activation, Ensure that the output weight value is in the range of (0,1); use the semantic attention map as the weight and act on the original supporting image features , highlighting discriminative regions and suppressing low weights: ;in: represents position-wise channel broadcast multiplication; is the enhanced support graph semantic feature; the enhanced support feature is normalized to obtain , the final output It will serve as input for subsequent deep semantic interaction or feature matching.
8. The small sample remote sensing image target detection algorithm based on multi-level information interaction according to claim 7 is characterized by: The query image features after shallow interaction and deep semantic alignment are processed, and a region proposal network is used to generate multiple candidate target regions, i.e., candidate box regions, on the feature map; For each candidate box area, extract the features of the corresponding area in the feature map, perform spatial alignment and unify the size, and input each RoI feature into the target detection head, including the classification branch and the regression branch, to perform target category prediction and position refinement respectively; The position and size of the candidate box are adjusted according to the regression offset to obtain the predicted bounding box; non-maximum suppression is performed on all prediction results to remove duplicate predictions, and the detection result with the highest confidence is retained, which is output as the remote sensing image target detection result.
Citation Information
Patent Citations
Metalearning-based small sample remote sensing target detection method
CN116012716A
Remote sensing image target detection method and system based on multi-scale semantic features
CN117079139A
SAM model-based small sample learning remote sensing image detection method
CN117690031A
Cited By
Laser radar and vision fusion target detection method based on space-time registration
CN120491093A