Remote sensing small target detection method based on position information transmission and feature decoupling

By using technical means of position information transmission, feature decoupling and global context aggregation in remote sensing small object detection, the problem of low detection accuracy of small object in visible light remote sensing images is solved, and more efficient knowledge transmission and feature learning are achieved.

CN120219960APending Publication Date: 2025-06-27SHENZHEN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510287297.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, when detecting small objects in visible light remote sensing images, it is difficult to effectively transmit detailed information and balance the feature learning of the target and background, resulting in low detection accuracy.

Method used

The remote sensing small object detection method based on position information transmission and feature decoupling is adopted. The position information is recursively transmitted through the position information transmission module. The feature decoupling distillation module uses the label information to generate a mask for feature separation, and combines with the global context polymerization distillation module to improve detection accuracy.

Benefits of technology

While keeping the calculation cost unchanged, the accuracy and performance of remote sensing small object detection are significantly improved, effectively solving the problem of difficult transfer of features in knowledge distillation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219960A_ABST
    Figure CN120219960A_ABST
Patent Text Reader

Abstract

The invention provides a remote sensing small target detection method based on position information transmission and feature decoupling. The method comprises the following steps: acquiring a to-be-recognized remote sensing image; inputting a to-be-recognized remote sensing image into the trained target detection model, and reasoning the to-be-recognized remote sensing image by a student model of the trained target detection model to obtain a small target recognition result of the to-be-recognized remote sensing image; wherein in the training stage, knowledge migration is achieved through a teacher-student cooperative training framework, image features are transmitted and fused through position information transmission and a feature decoupling mechanism, and the spatial perception ability of a teacher model is inherited. The student model after knowledge distillation optimization does not need to depend on a complex feature fusion structure in the reasoning stage, and the remote sensing small target detection precision can be improved directly through a lightweight network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target recognition, and particularly to a remote sensing small target detection method and device based on position information transmission and feature decoupling. Background Art

[0002] Visible light remote sensing images present characteristics such as diverse scenes, various noises, and large variations in target sizes, posing great challenges to the target detection task. Especially for the recognition and detection of small targets in visible light remote sensing images, more difficulties are added. Currently, in order to better detect small targets, target detection algorithms based on deep convolutional neural networks usually adopt two strategies: one is to introduce complex structures such as attention mechanisms and feature pyramid enhancement modules, and strengthen the semantic representation of small targets by explicitly modeling multi-scale feature associations; the other is to directly increase the resolution of the input image to retain more detailed information. However, such methods often significantly increase the number of model parameters and computational complexity, resulting in a decrease in the inference speed and making it difficult to meet the requirements of real-time detection scenarios. Improving the performance of small target detection without significantly increasing the computational cost has become a key problem to be solved urgently. In recent years, knowledge distillation technology has provided a new idea for solving this contradiction. By constructing a "teacher-student" model framework, the knowledge of small target detection hidden in a high-precision large model is transferred to a lightweight student model, enabling the student model to inherit the sensitivity of the teacher model to small targets without changing the original architecture, thereby indirectly improving the performance of small target detection while maintaining the inference efficiency. However, when knowledge distillation is used for small target detection in visible light remote sensing images, it will face unique challenges. Since small targets in visible light remote sensing images are usually difficult to distinguish in complex backgrounds and the target sizes are small, traditional distillation methods often cannot effectively transmit the captured detailed information, and due to the sample imbalance problem between small targets and backgrounds, the background features overwhelm the target features, making it difficult to learn effective target features. Summary of the Invention

[0003] To solve the above technical problems, the present invention proposes a remote sensing small target detection method based on position information transmission and feature decoupling.

[0004] According to the first aspect of the present invention, there is provided a remote sensing small target detection method based on position information transmission and feature decoupling, the method comprising the following steps:

[0005] Step S1: Obtain a remote sensing image to be recognized;

[0006] Step S2: Input the remote sensing image to be recognized into a trained target detection model, and the student model of the trained target detection model performs inference on the remote sensing image to be recognized to obtain the small target recognition result of the remote sensing image to be recognized;

[0007] The object detection model includes a teacher model, a student model, a position information transmission module, a feature decoupling distillation module, and a global context aggregation distillation module; the teacher model, the student model, the position information transmission module, the feature decoupling distillation module, and the global context aggregation distillation module participate in training during the training phase; during the training phase, the backbone network in the student model extracts the image features of the remote sensing image to be recognized, and the neck network in the student model fuses the image features to obtain multiple second fused features; during the inference phase, the student model only extracts the image features through the backbone network and directly recognizes the small target to be recognized based on the extracted image features.

[0008] Preferably, during the training phase, the backbone network of the teacher model extracts the first image features of the training samples, and the neck network of the teacher model fuses the first image features to obtain multiple first fused features; the first fused features are input into the feature decoupling distillation module and the global context aggregation distillation module;

[0009] The backbone network in the student model extracts the second image features of the training samples, and the neck network in the student model fuses the second image features to obtain multiple second fused features, with each second fused feature corresponding to a first fused feature; the second fused features are input into the position information transmission module, and the position information transmission module recursively transmits the second fused features on multiple feature distillation layers, and enhances and injects the position information of the small targets in the training samples during the recursive transmission process to obtain enhanced features corresponding to the second fused features, and the enhanced features are input into the feature decoupling distillation module and the global context aggregation distillation module;

[0010] The feature decoupling distillation module generates a foreground mask and a background mask based on the annotation information in the training samples, and uses the foreground mask and the background mask to perform separation distillation on the first fused features and the enhanced features corresponding to the first fused features to obtain background features and target features;

[0011] The global context aggregation distillation module fuses the global context information of the first fused features with the enhanced features corresponding to the first fused features to obtain fused enhanced features;

[0012] A global loss function is constructed, and the global loss function includes a detection loss function constructed by the student model based on the enhanced features, a distillation loss function constructed by the feature decoupling distillation module based on the foreground mask and the background mask, and a distillation loss function constructed by the global context aggregation distillation module; based on the global loss function, the parameters in the student model, the position information transmission module, the feature decoupling distillation module, and the global context aggregation distillation module are adjusted through gradient backpropagation.

[0013] Preferably, the position information transmission module includes N branches, and each branch corresponds to a feature distillation layer; each branch is used to receive a second fusion feature, and the second fusion features are all distillation features; each branch includes a first convolutional layer, a position information enhancement module, and a second convolutional layer connected in sequence, where the input of the position information enhancement module of the Nth branch is the feature obtained by adding the output feature of the first convolutional layer of this branch to the output feature of the first convolutional layer of this branch; the input of the position information enhancement module of the numth branch is the feature obtained by adding the output feature of the first convolutional layer of this branch to the output feature of the position information enhancement module of the (num + 1)th branch, where 1 ≤ num ≤ N - 1; the output features of the position information enhancement modules of each branch are subjected to convolutional processing through the second convolutional layer, and the results after the convolutional processing of the second convolutional layer of each branch are used as the enhanced features output by this branch.

[0014] Preferably, the position information enhancement module includes a first convolutional branch and a second convolutional branch. The first convolutional branch includes a convolutional layer with a convolution kernel size of 3×3 and a first pooling layer. After convolutional and pooling processing, the first convolutional branch obtains first spatial information and first channel attention weights corresponding to the input of this position information enhancement module; the second convolutional branch includes a convolutional layer with a convolution kernel size of 1×1 and a second pooling layer. After convolutional and pooling processing, the second convolutional branch obtains second spatial information and second channel attention weights corresponding to the input of this position information enhancement module; the first spatial information and the second spatial information are cross-weighted, and the feature after cross-weighting is matrix-multiplied with the input of this position information enhancement module, and the obtained result is used as the output feature of the position information enhancement module.

[0015] Preferably, the feature decoupling and distillation module generates a foreground mask and a background mask based on the annotation information in the training samples, and uses the foreground mask and the background mask to perform separation and distillation on the first fusion feature and the enhanced feature corresponding to the first fusion feature, so as to obtain a background feature and a target feature, where:

[0016] The calculation formula for the foreground mask is:

[0017] M fg_k (i,j)=MaxPool HW_k (M fg (i,j))

[0018]

[0019] where M fg_k(i, j) is the foreground mask corresponding to the k-th first fusion feature and the k-th second fusion feature obtained based on the training samples. i is the position index of the pixel point in the image height direction of the training sample, j is the position index of the pixel point in the image width direction of the position information enhancement module, MaxPool is the max pooling, H and W are the height and width of the image of the training sample, respectively, and M fg (i, j) is the normalized foreground mask, N fg is the number of pixel points included in all target boxes in the annotation information, S fg is the set of annotation points of the target box, and M(i, j) is the foreground mask before normalization;

[0020] The calculation formula of the background mask is:

[0021]

[0022] where M bg_k (i, j) is the background mask corresponding to the k-th first fusion feature and the k-th second fusion feature obtained based on the training samples, N bg is the number of background pixel points.

[0023] Preferably, the global context aggregation and distillation module fuses the global context information of the first fusion feature with the enhanced feature corresponding to the first fusion feature, where:

[0024] The global context aggregation and distillation module includes three dilated convolutional layers in parallel with different dilation ratios, all of which are used to receive the first fusion feature and perform dilated convolution operations on the first fusion feature, and fuse the results of the three dilated convolutions as the global context information of the first fusion feature.

[0025] Preferably, the global loss function L is:

[0026] L = λ1L dis + λ2L g + L det

[0027] where L dis is the distillation loss function of the feature decoupling and distillation module, L g is the distillation loss function of the global context aggregation and distillation module, L det is the detection loss function of the student model, and λ1 and λ2 are weight coefficients;

[0028]

[0029] L feature_k = (f align (f MLAP (F′ S_k )) - FT_k ) 2

[0030] Among them, L feature_k represents the k-th second fusion feature obtained based on the training samples, and f align (·) represents channel and scale alignment, and f MLAP (·) is the mapping function corresponding to the position information transfer module. α and β are the background feature distillation loss weight and the foreground feature distillation loss weight respectively, represents the element-wise multiplication of matrices, and F′ S_k is the k-th second fusion feature of the student model, and F T_k is the k-th first fusion feature of the teacher model, and M bg_k is all the background masks, and M fg_k is all the foreground masks;

[0031]

[0032] Among them, represents the mapping function of the global context aggregation distillation module for the k-th first fusion feature; F″ S_k is the global context information feature obtained by the student model;

[0033] L det combines the classification loss, localization loss, and confidence loss of the student module.

[0034] According to the second aspect of the present invention, a remote sensing small target detection device based on position information transfer and feature decoupling is provided. The device includes:

[0035] An image acquisition module: configured to acquire a remote sensing image to be recognized;

[0036] A recognition module: configured to input the remote sensing image to be recognized into a trained target detection model, and the student model of the trained target detection model performs inference on the remote sensing image to be recognized to obtain the small target recognition result of the remote sensing image to be recognized;

[0037] The target detection model includes a teacher model, a student model, a position information transfer module, a feature decoupling distillation module, and a global context aggregation distillation module; the teacher model, the student model, the position information transfer module, the feature decoupling distillation module, and the global context aggregation distillation module participate in training during the training phase; during the training phase, the backbone network in the student model extracts the image features of the remote sensing image to be recognized, and the neck network in the student model fuses the image features to obtain multiple second fusion features; during the inference phase, the student model only extracts the image features through the backbone network and directly recognizes the small target to be recognized according to the extracted image features.

[0038] According to a third aspect of the present invention, there is provided an electronic device, including:

[0039] A processor for executing a plurality of instructions;

[0040] A memory for storing a plurality of instructions;

[0041] Wherein, the plurality of instructions are used to be stored by the memory and loaded and executed by the processor for the method as described above.

[0042] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium, in which a plurality of instructions are stored; the plurality of instructions are used to be loaded and executed by a processor for the method as described above.

[0043] The present invention solves the problem of applying knowledge distillation to small target detection in visible light remote sensing images, enabling small target detection to further improve the detection accuracy while keeping the number of parameters and the amount of calculation unchanged. Aiming at the problem that the same-level feature distillation between the teacher model and the student model in the feature distillation process cannot fully utilize the knowledge in the distillation process, a position information transfer module is designed. This module recursively fuses and transfers the position information in knowledge distillation on feature maps at multiple different levels, enabling the student model to more efficiently learn the spatial position knowledge transmitted by the teacher model and improving the target detection performance of the student model. Then, aiming at the problem that during the knowledge distillation process, small targets contribute less loss, resulting in their feature knowledge being masked by background knowledge, making it impossible for the student model to better improve the detection performance of small targets, this chapter proposes a feature decoupling distillation module. This module generates target and background masks respectively through label information, and uses the masks to achieve decoupled knowledge distillation at the feature level, effectively balancing the loss contribution of small targets in the distillation process. Finally, to solve the problem of context information loss that occurs during feature decoupling distillation, a global context aggregation distillation module is proposed, which perceives and aggregates context information globally through dilated convolutions with multiple different dilation rates, and uses this information to guide the student model.

[0044] The present invention has the following technical effects:

[0045] 1. The present invention proposes an efficient knowledge transfer mechanism, allowing the student model to more comprehensively and effectively absorb the knowledge of the teacher model without increasing additional computational costs, thus significantly improving the overall performance of the detection model during the inference stage and enhancing the detection accuracy of remote sensing small targets.

[0046] 2. The present invention designs a position information transfer module for transferring position information between multiple feature layers, enabling the student model to absorb and integrate the position knowledge of the teacher model between different feature layers and further strengthening the distillation effect.

[0047] 3. The present invention designs a feature decoupling and distillation module, which effectively balances the loss of target and background information by distilling the features of the target and background respectively, and further improves the detection and recognition ability of the student network for small targets.

[0048] 4. To solve the problem of context information loss that may be caused by feature decoupling and distillation, the present invention introduces a global context aggregation and distillation module, which works within the global context range and effectively guides the student model to capture and utilize context knowledge.

[0049] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly and implement it in accordance with the content of the specification, the following takes the preferred embodiments of the present invention and combines with the drawings to describe in detail as follows. Brief Description of the Drawings

[0050] The accompanying drawings that form a part of the present invention are used to provide a further understanding of the present invention, and the present invention provides the following accompanying drawings for description. In the accompanying drawings:

[0051] Figure 1 It is a schematic flow chart of a remote sensing small target detection method based on position information transmission and feature decoupling according to an embodiment of the present invention;

[0052] Figure 2 It is a schematic structural diagram of an object detection model according to an embodiment of the present invention;

[0053] Figure 3 It is a schematic structural diagram of a position information transmission module and a position information enhancement module according to an embodiment of the present invention;

[0054] Figure 4 It is a schematic structural diagram of a global context aggregation and distillation module according to an embodiment of the present invention;

[0055] Figure 5 It is a structural block diagram of a remote sensing small target detection device based on position information transmission and feature decoupling of the present invention. Detailed Embodiments

[0056] First, in combination with Figure 1 - Figure 2 describe a remote sensing small target detection method based on position information transmission and feature decoupling according to an embodiment of the present invention. As Figure 1 - Figure 2 shown, the method includes the following steps:

[0057] Step S1: Obtain the remote sensing image to be recognized;

[0058] Step S2: Input the remote sensing image to be recognized into the trained object detection model, and the student model of the trained object detection model performs inference on the remote sensing image to be recognized to obtain the small target recognition result of the remote sensing image to be recognized;

[0059] The object detection model includes a teacher model, a student model, a location information transmission module, a feature decoupling and distillation module, and a global context aggregation and distillation module; the teacher model, the student model, the location information transmission module, the feature decoupling and distillation module, and the global context aggregation and distillation module participate in training during the training phase; during the training phase, the backbone network in the student model extracts the image features of the remote sensing image to be recognized, and the neck network in the student model fuses the image features to obtain multiple second fused features; during the inference phase, the student model only extracts the image features through the backbone network and directly recognizes the small target to be recognized based on the extracted image features.

[0060] In the present invention, when recognizing a small target in the remote sensing image to be recognized, only the student model participates in the inference, and the teacher model, the location information transmission module, the feature decoupling and distillation module, and the global context aggregation and distillation module do not participate in the inference.

[0061] Further, during the training phase, the backbone network of the teacher model extracts the first image features of the training samples, and the neck network of the teacher model fuses the first image features to obtain multiple first fused features; the first fused features are input into the feature decoupling and distillation module and the global context aggregation and distillation module;

[0062] The backbone network in the student model extracts the second image features of the training samples, and the neck network in the student model fuses the second image features to obtain multiple second fused features, and each second fused feature corresponds to a first fused feature; the second fused features are input into the location information transmission module, and the location information transmission module recursively transmits the second fused features on multiple feature distillation layers, and enhances and injects the location information of the small target in the training samples during the recursive transmission process to obtain enhanced features corresponding to the second fused features, and the enhanced features are input into the feature decoupling and distillation module and the global context aggregation and distillation module;

[0063] The feature decoupling and distillation module generates a foreground mask and a background mask based on the annotation information in the training samples, and uses the foreground mask and the background mask to perform separation distillation on the first fused features and the enhanced features corresponding to the first fused features to obtain background features and target features;

[0064] The global context aggregation and distillation module fuses the global context information of the first fused features with the enhanced features corresponding to the first fused features to obtain fused enhanced features;

[0065] Construct a global loss function, which includes a detection loss function constructed by the student model based on enhanced features, a distillation loss function constructed by the feature decoupling distillation module based on the foreground mask and the background mask, and a distillation loss function constructed by the global context aggregation distillation module; based on the global loss function, adjust the parameters in the student model, the position information transmission module, the feature decoupling distillation module, and the global context aggregation distillation module through gradient backpropagation.

[0066] In the present invention, YOLOv5-X with more parameters and computational complexity is used as the teacher model, and YOLOv5-S is used as the student model. During the training phase, the teacher model adopts frozen training and does not perform gradient backpropagation to optimize the weights.

[0067] The student model extracts features through the backbone network and fuses the features through the neck network. The fused output features are mapped to intermediate features with the same number of channels by the position information transmission module MLAP. The position information transmission module recursively transmits the fused output features on multiple feature distillation layers and injects enhanced position information during the transmission process. During this process, the student model effectively transmits the position information on multiple distillation features at different scales with the help of the position information taught by the teacher model, and thus better learns the target position knowledge taught by the teacher, thereby significantly enhancing the detection performance of the student model for small targets. Then, the output features of the fused output features and the output features of the teacher network are separated and distilled for small targets and the background through the feature decoupling distillation module FDD. The feature decoupling distillation module constructs a mask using the label information of the training samples and uses the mask to decouple the target and the background of the distillation features. At the same time, the global context information of the teacher model is transmitted to the student model through the feature decoupling distillation module to alleviate the side effect of context loss brought by the feature decoupling distillation module. During the training process, calculate the detection loss of the student model and the distillation losses of the feature decoupling distillation module and the global context aggregation distillation module MGCAD, and optimize the student model through gradient backpropagation.

[0068] As Figure 3 shown, the position information transmission module includes a plurality of serially connected position information enhancement modules LIE to achieve position information enhancement on multiple feature distillation layers and adjacent-level position knowledge sharing.

[0069] The position information transmission module includes N branches, and each branch corresponds to a feature distillation layer; each branch is used to receive a second fused feature, and the second fused features are all distillation features; each branch includes a first convolutional layer, a position information enhancement module, and a second convolutional layer connected in sequence. Among them, the input of the position information enhancement module of the Nth branch is the feature obtained by adding the output feature of the first convolutional layer of this branch to the output feature of the first convolutional layer of this branch; the input of the position information enhancement module of the numth branch is the output feature of the first convolutional layer of this branch plus the output feature of the position information enhancement module of the (num + 1)th branch, where 1 ≤ num ≤ N - 1; the output feature of the position information enhancement module of each branch is subjected to convolutional processing through the second convolutional layer, and the result after convolutional processing of the second convolutional layer of each branch is used as the enhanced feature output by this branch.

[0070] Furthermore, each first convolutional layer maps the second fused feature received by its corresponding branch, and converts the number of channels of the second fused feature to the intermediate transmission channel dimension.

[0071] The position information enhancement module includes a first convolutional branch and a second convolutional branch. The first convolutional branch includes a convolutional layer with a kernel size of 3×3 and a first pooling layer. After convolutional and pooling processing, the first convolutional branch obtains the first spatial information and the first channel attention weight corresponding to the input of this position information enhancement module; the second convolutional branch includes a convolutional layer with a kernel size of 1×1 and a second pooling layer. After convolutional and pooling processing, the second convolutional branch obtains the second spatial information and the second channel attention weight corresponding to the input of this position information enhancement module; the first spatial information and the second spatial information are cross-weighted, and the feature after cross-weighting is matrix-multiplied with the input of this position information enhancement module, and the obtained result is used as the output feature of the position information enhancement module.

[0072] In the present invention, the core purpose of the position information enhancement module is to comprehensively capture and enhance the position information in the remote sensing image, so that more abundant feature representations can be transmitted on multiple different distillation features, and more feature knowledge of the teacher model can be learned more fully during the model distillation process. The position information enhancement module is a parallel dual-branch structure, divided into a 3×3 convolutional branch and a 1×1 convolutional branch. By extracting visual information and channel information with different receptive fields through the two branches respectively, and cross-weighting the two branches finally, the complementarity and enhancement of information can be well realized, and the ability to process detailed information and position information can be improved. Specifically, for the input feature F input_S ∈R C×H×W of the position information enhancement module, in the 3×3 convolutional branch, the feature is first passed through a 3×3 convolution, and the first spatial information F s1, at the same time, average pooling is used to obtain the first-channel attention weight W1. It can be expressed by the formula:

[0073]

[0074] where Softmax(·) represents the Softmax function, which is used to calculate the spatial information weight, and Conv 3×3 (·) represents the convolution with a convolution kernel size of 3×3, and AvgPool(·) represents average pooling. For the 1×1 convolution branch, first, F input_S is passed through max pooling and average pooling to extract the detailed information and global information in the features, and they are concatenated in the channel dimension to construct an informative feature map.

[0075] Then, 1×1 convolution is used to restore the number of channels of the features, and at the same time, the position information weight W covering the entire feature map is generated p :

[0076]

[0077] where f sigmoid (·) represents the Sigmoid function, Conv 1×1 (·) represents the convolution with a convolution kernel size of 1×1, and Concat(·) represents the concatenation operation in the channel dimension. After that, the weight W p is batch-normalized and then multiplied by F input_s , and the second spatial information F s2 and the second-channel attention weight W2 are calculated using the same method as the 3×3 convolution branch. It can be expressed by the formula:

[0078]

[0079] where Softmax(·) represents the Softmax function, BN(·) represents batch normalization, represents the element-wise multiplication of matrices. Finally, in order to better extract spatial information and cross-scale information interaction, the two branches are cross-enhanced and added to obtain the final output feature F output_S :

[0080]

[0081] where represents the element-wise multiplication of matrices.

[0082] In the position information transfer module, assume that the number of feature layers that the student model needs to distill is N, and each distillation feature is F S_k ∈R C×H×W, where k = 1, 2, …, N, representing the k-th distilled feature. For the distilled feature F of each layer S_k , first, use a 1×1 convolution to map F S_k for feature mapping, and convert its number of channels to the intermediate transfer channel dimension C mid , obtaining a feature with a shape of C mid × H × W The purpose of doing this is to facilitate subsequent position information transfer. Then, upsample the residual output feature F of the adjacent layer (k + 1) res_(k+1) and add it to F S_k after aligning the size dimensions to obtain the transfer feature for information transfer For the last position information transfer module, since it has no residual output feature of the adjacent layer as input, the distilled feature F S_N is used as the residual output feature of the adjacent layer. This process can be expressed as:

[0083]

[0084] where Upsample(·) represents the nearest neighbor interpolation function for upsampling, and Conv 1×1 (·) represents the convolution with a convolution kernel size of 1×1. Finally, the transfer feature is enhanced with position information through the position information transfer module, and the number of channels is restored using a 1×1 convolution to obtain the k-th output F′ of the position information transfer module S_k , and at the same time, the output of the position information transfer module is used as the residual output feature F of the current layer res_k , thus realizing cross-layer position information transfer. This process can be expressed as:

[0085]

[0086] where f LIE_k (·) represents the mapping function of the k-th LIE module. By using this method of gradually transferring position information on the distilled feature and performing feature enhancement in each layer, the transfer of position knowledge in knowledge distillation can be effectively improved, and the student model can learn the relationships between different distilled layers, thereby comprehensively improving the student model's ability to extract and fuse target features, and further improving its detection performance.

[0087] Further, the feature decoupling and distillation module generates a foreground mask and a background mask based on the annotation information in the training samples, and uses the foreground mask and the background mask to separately distill the first fused feature and the enhanced feature corresponding to the first fused feature to obtain a background feature and a target feature, where:

[0088] The calculation formula for the foreground mask is:

[0089] M fg_k (i,j) = MaxPool HW_k (M fg (i,j))

[0090]

[0091] Among them, M fg_k (i,j) is the foreground mask corresponding to the k-th first fusion feature and the k-th second fusion feature obtained based on the training samples. i is the position index of the pixel point in the image height direction of the training sample, j is the position index of the pixel point in the image width direction of the position information enhancement module, MaxPool is max pooling, H and W are the height and width of the image of the training sample respectively, and M fg (i,j) is the normalized foreground mask, N fg is the number of pixel points included in all target boxes in the annotation information, S fg is the set of annotation points of the target box, and M(i,j) is the foreground mask before normalization;

[0092] The calculation formula of the background mask is:

[0093]

[0094] Among them, M bg_k (i,j) is the background mask corresponding to the k-th first fusion feature and the k-th second fusion feature obtained based on the training samples, and N bg is the number of background pixel points.

[0095] In remote sensing images, small targets such as vehicles, ships, or crops are often difficult to identify in the vast background. The background area in object detection is usually much larger than the foreground area. Using the knowledge distillation method based on feature distillation to simply transfer features from the teacher model to the student model may lead to overemphasis on background features, thus masking the important features of small targets. The feature decoupling distillation module can balance feature learning and prevent background features from dominating the learning process by separating the learning processes of the foreground and background and processing foreground and background features differently.

[0096] The feature decoupling distillation module needs to generate a foreground mask and a background mask according to the annotation information of small targets. Since the size of the mask is different from the size of the distilled features, M i,j cannot be directly used as the mask for distilled features. To make the proportion of the foreground mask of small targets larger so that the small target features in the network can be assigned more distilled weights, M(i,j) is downsampled by the method of max pooling.

[0097] Furthermore, as Figure 4As shown, the global context aggregation and distillation module fuses the global context information of the first fused feature with the enhanced feature corresponding to the first fused feature, where:

[0098] The global context aggregation and distillation module includes three parallel atrous convolution layers with different dilation ratios, all of which are used to receive the first fused feature, perform atrous convolution operations on the first fused feature, and fuse the three atrous convolution results as the global context information of the first fused feature.

[0099] In the present invention, in remote sensing images, context information is crucial because each pixel not only contains its inherent features but also forms a rich association network with surrounding pixels. The detection of small targets such as ships, vehicles, or oil tanks benefits from their relationship with the surrounding environmental features because these targets are small in size and difficult to distinguish themselves. Environmental cues such as the surrounding terrain and man-made structures can make up for the deficiencies of direct visual features and enhance the recognition ability of the detection algorithm. The feature decoupling and distillation module of the present invention separates the target and the background using masks to solve the problem of unbalanced distillation weights between the target and the background, which inevitably weakens these associations. To solve this problem, the present invention proposes a global context aggregation and distillation module.

[0100] In the global context aggregation and distillation module, first, atrous convolutions with three different dilation ratios are used to calculate the distillation features of the teacher model and the student model respectively. This can not only maintain the fineness of local features but also expand the receptive field and capture more extensive context information. Then, the features calculated by the three atrous convolutions are added together to converge into a feature containing rich context information, which can be expressed by the formula:

[0101] F′ T_k = DConv1(F T_k ) + DConv3(F T_k ) + DConv5(F T_k )

[0102] F″ S_k = DConv1(F′ S_k ) + DConv3(F′ S_k ) + DConv5(F′ S_k )

[0103] where DConv r represents an atrous convolution with a dilation rate of r, where r is selected as 1, 3, and 5, F T_k represents the teacher distillation feature, F′ S_k represents the feature after passing the position information, F′ T_k and F″ S_krespectively represent the global context information features aggregated by the teacher module and the student module.

[0104] The present invention enables it to understand the image content at different scales, thereby more effectively simulating and transmitting the context information in the teacher model to the student model, reducing the loss of key context information of small target features during the distillation process.

[0105] Furthermore, the global loss function L is:

[0106] L = λ1L dis + λ2L g + L det

[0107] where, L dis is the distillation loss function of the feature decoupling distillation module, L g is the distillation loss function of the global context aggregation distillation module, L det is the detection loss function of the student model, and λ1, λ2 are weight coefficients.

[0108]

[0109] L feature_k = (f align (f MLAP (F′ S_k )) - F T_k ) 2

[0110] where, L feature_k represents the k-th second fusion feature obtained based on the training samples, f align (·) represents channel and scale alignment, f MLAP (·) is the mapping function corresponding to the position information transmission module, α and β are the background feature distillation loss weight and the foreground feature distillation loss weight respectively, represents element-wise multiplication of matrices, F′ S_k is the k-th second fusion feature of the student model, F T_k is the k-th first fusion feature of the teacher model, M bg_k is all background masks, M fg_k is all foreground masks.

[0111] The present invention realizes the transmission of context knowledge by minimizing the Euclidean distance between the context aggregation features of the teacher model and the student model, that is:

[0112]

[0113] where, A mapping function representing the global context aggregation and distillation module for the k-th first fusion feature, F″ S_k The global context information feature obtained by the student model.

[0114] In the present invention, L det Combining the classification loss, localization loss, and confidence loss of the student module, both the first fusion feature and the second fusion feature are distillation features. The classification loss, localization loss, and confidence loss use conventional calculation methods in the art and will not be elaborated here.

[0115] The present invention significantly improves the detection effect of YOLOv5-S, can well distill the features of small targets in remote sensing images, effectively solves the problem that features are difficult to transfer during knowledge distillation, and its experimental indicators on the AI-TOD dataset are significantly ahead of current mainstream detectors. Table 1 shows the comparison experimental results of the present invention and other detection methods on the AI-TOD dataset.

[0116] Table 1 Comparison Table

[0117] Method AP <![CDATA[AP 50 > <![CDATA[AP 75 > <![CDATA[AP t > <![CDATA[AP s > <![CDATA[AP m > FCOS 0.131 0.307 0.088 0.151 0.161 0.184 YOLOv5 0.215 0.523 0.137 0.215 0.281 0.301 Cascade R-CNN 0.196 0.465 0.133 0.197 0.246 0.319 YOLOX 0.216 0.507 0.146 0.214 0.277 0.334 DetectoRS 0.225 0.526 0.156 0.233 0.273 0.340 YOLOv6 0.179 0.402 0.135 0.164 0.249 0.329 DAMO-YOLO 0.165 0.367 0.124 0.158 0.223 0.303 RTMDet 0.201 0.479 0.131 0.187 0.270 0.370 YOLOv8 0.207 0.462 0.155 0.190 0.289 0.364 The present invention 0.288 0.635 0.218 0.291 0.346 0.380

[0118] As Figure 5 shown, the present invention provides a remote sensing small target detection device based on position information transmission and feature decoupling. The device includes:

[0119] An image acquisition module: configured to acquire a remote sensing image to be recognized;

[0120] A recognition module: configured to input the remote sensing image to be recognized into a trained target detection model, and the student model of the trained target detection model performs inference on the remote sensing image to be recognized to obtain the small target recognition result of the remote sensing image to be recognized;

[0121] The target detection model includes a teacher model, a student model, a position information transmission module, a feature decoupling and distillation module, and a global context aggregation and distillation module; the teacher model, the student model, the position information transmission module, the feature decoupling and distillation module, and the global context aggregation and distillation module participate in training during the training phase; during the training phase, the backbone network in the student model extracts the image features of the remote sensing image to be recognized, and the neck network in the student model fuses the image features to obtain multiple second fusion features; during the inference phase, the student model only extracts the image features through the backbone network and directly recognizes the small target to be recognized according to the extracted image features.

[0122] An embodiment of the present invention further provides an electronic device, including:

[0123] A processor for executing multiple instructions;

[0124] A memory for storing multiple instructions;

[0125] Among them, the multiple instructions are used to be stored by the memory and loaded and executed by the processor for the method as described above.

[0126] An embodiment of the present invention further provides a computer-readable storage medium, in which multiple instructions are stored; the multiple instructions are used to be loaded and executed by a processor for the method as described above.

[0127] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0128] In several embodiments provided by the present invention, it should be understood that the disclosed system, device and method may be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other may be through some interfaces, and the indirect coupling or communication connection of the device or unit may be in an electrical, mechanical or other form.

[0129] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0130] In addition, each functional unit in each embodiment of the present invention may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0131] The integrated unit implemented in the form of software functional units can be stored in a computer-readable storage medium. The above-mentioned software functional units are stored in a storage medium, including several instructions for causing a computer device (which can be a personal computer, a physical machine server, or a network cloud server, etc., and needs to install the Ubuntu operating system) to execute some steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

Claims

1. A remote sensing small target detection method based on position information transmission and feature decoupling, characterized in that: The method comprises the following steps: Step S1: Acquire a remote sensing image to be identified; Step S2: inputting the remote sensing image to be identified into the trained target detection model, and using the student model of the trained target detection model to infer the remote sensing image to be identified, to obtain a small target recognition result of the remote sensing image to be identified; The target detection model includes a teacher model, a student model, a position information transfer module, a feature decoupling distillation module and a global context aggregation distillation module; the teacher model, the student model, the position information transfer module, the feature decoupling distillation module and the global context aggregation distillation module participate in training in the training stage; in the training stage, the backbone network in the student model extracts the image features of the remote sensing image to be identified, and the neck network in the student model fuses the image features to obtain multiple second fusion features; in the reasoning stage, the student model only extracts image features through the backbone network, and directly identifies the small target to be identified based on the extracted image features.

2. The method according to claim 1, characterized in that In the training phase, the backbone network of the teacher model extracts the first image features of the training samples, and the neck network of the teacher model fuses the first image features to obtain multiple first fused features; the first fused features are input into the feature decoupling distillation module and the global context aggregation distillation module; The backbone network in the student model extracts the second image features of the training sample, and the neck network in the student model fuses the second image features to obtain multiple second fused features, each of which corresponds to one first fused feature; The second fused feature is input into the position information transmission module, which recursively transmits the second fused feature on multiple feature distillation layers, and in the recursive transmission process, the position information of the small target in the training sample is enhanced and injected to obtain the enhanced feature corresponding to the second fused feature, and the enhanced feature is input into the feature decoupling distillation module and the global context aggregation distillation module; The feature decoupling and distillation module generates a foreground mask and a background mask based on the annotation information in the training sample, and uses the foreground mask and the background mask to separate and distill the first fused feature and the enhanced feature corresponding to the first fused feature to obtain the background feature and the target feature; The global context aggregation and distillation module fuses the global context information of the first fused feature with the enhanced feature corresponding to the first fused feature to obtain a fused enhanced feature; A global loss function is constructed, which includes the detection loss function constructed by the student model based on enhanced features, the distillation loss function constructed by the feature decoupling distillation module based on foreground mask and background mask, and the distillation loss function constructed by the global context aggregation distillation module. Based on the global loss function, the parameters in the student model, position information transfer module, feature decoupling distillation module and global context aggregation distillation module are adjusted through gradient backpropagation.

3. The method according to claim 2, characterized in that The position information transmission module includes N branches, each branch corresponds to a feature distillation layer; each branch is used to receive a second fusion feature, and the second fusion features are all distillation features; Each branch includes a first convolutional layer, a position information enhancement module and a second convolutional layer connected in sequence, wherein the input of the position information enhancement module of the Nth branch is the output feature of the first convolutional layer of the branch plus the output feature of the first convolutional layer of the branch; the input of the position information enhancement module of the numth branch is the output feature of the first convolutional layer of the branch plus the output feature of the position information enhancement module of the num+1th branch, wherein 1≤num≤N-1; the output feature of the position information enhancement module of each branch is convolved through the second convolutional layer, and the result of the convolution processing of the second convolutional layer of each branch is used as the enhanced feature output of the branch.

4. The method according to claim 3, characterized in that The position information enhancement module includes a first convolution branch and a second convolution branch. The first convolution branch includes a convolution layer with a convolution kernel size of 3×3 and a first pooling layer. After convolution and pooling processing, the first convolution branch obtains the first spatial information and the first channel attention weight corresponding to the input of the position information enhancement module. The second convolution branch includes a convolution layer with a convolution kernel size of 1×1 and a second pooling layer. After convolution and pooling processing, the second convolution branch obtains second spatial information and a second channel attention weight corresponding to the input of the position information enhancement module; The first spatial information and the second spatial information are cross-weighted, and the cross-weighted features are matrix-multiplied with the input of the position information enhancement module, and the obtained results are used as the output features of the position information enhancement module.

5. The method according to claim 4, characterized in that The feature decoupling and distillation module generates a foreground mask and a background mask based on the annotation information in the training sample, and uses the foreground mask and the background mask to separate and distill the first fused feature and the enhanced feature corresponding to the first fused feature to obtain the background feature and the target feature, wherein: The foreground mask calculation formula is: M fg_k (i,j)=MaxPool HW_k (M fg (i,j)) Among them, M fg_k (i, j) is the foreground mask corresponding to the kth first fusion feature and the kth second fusion feature obtained based on the training sample, i is the position index of the pixel in the image height direction of the training sample, j is the position index of the pixel in the image width direction of the position information enhancement module, MaxPool is the maximum pooling, H, W are the height and width of the image of the training sample, respectively, M fg (i,j) is the normalized foreground mask, N fg is the number of pixels contained in all target boxes in the annotation information, S fg is the set of annotation points of the target box, and M(i,j) is the foreground mask before normalization; The background mask is calculated as: Among them, M bg_k (i, j) is the background mask corresponding to the kth first fusion feature and the kth second fusion feature obtained based on the training sample, N bg is the number of background pixels.

6. The method according to claim 1, characterized in that The global context aggregation distillation module fuses the global context information of the first fused feature with the enhanced feature corresponding to the first fused feature, where: The global context aggregation distillation module includes three parallel atrous convolution layers with different atrous ratios, each of which is used to receive the first fusion feature and perform a atrous convolution operation on the first fusion feature, and fuse the three atrous convolution results as the global context information of the first fusion feature.

7. The method according to claim 2, characterized in that The global loss function L is: L=λ1L dis +λ2L g +L det Among them, L dis is the distillation loss function of the feature decoupling distillation module, L g is the distillation loss function of the global context aggregation distillation module, L det is the detection loss function of the student model, λ1 and λ2 are weight coefficients; L feature_k =(f align (in MLAP (F′ S_k ))-F T_k ) 2 Among them, L feature_k represents the kth second fusion feature obtained based on the training sample, f align (·) represents the alignment of channels and scales, f MLAP (·) is the mapping function corresponding to the position information transfer module, α and β are the background feature distillation loss weights and foreground feature distillation loss weights, respectively. Represents the multiplication of corresponding elements of the matrix, F′ S_k is the kth second fusion feature of the student model, F T_k is the kth first fusion feature of the teacher model, M bg_k For all background masks, M fg_k for all foreground masks; in, represents the mapping function of the global context aggregation distillation module for the kth first fusion feature; F″ S_k The global context information features obtained for the student model; L det Comprehensive classification loss, localization loss and confidence loss of the student module.

8. A remote sensing small target detection device based on position information transmission and feature decoupling, characterized in that: The device comprises: Image acquisition module: configured to acquire remote sensing images to be identified; Recognition module: configured to input the remote sensing image to be identified into the trained target detection model, and the student model of the trained target detection model performs inference on the remote sensing image to be identified to obtain the small target recognition result of the remote sensing image to be identified; The target detection model includes a teacher model, a student model, a position information transfer module, a feature decoupling distillation module and a global context aggregation distillation module; the teacher model, the student model, the position information transfer module, the feature decoupling distillation module and the global context aggregation distillation module participate in training in the training stage; in the training stage, the backbone network in the student model extracts the image features of the remote sensing image to be identified, and the neck network in the student model fuses the image features to obtain multiple second fusion features; in the reasoning stage, the student model only extracts image features through the backbone network, and directly identifies the small target to be identified based on the extracted image features.

9. An electronic device, comprising: A processor, which is used to execute multiple instructions; A memory for storing a plurality of instructions; The plurality of instructions are used to be stored in the memory and loaded and executed by the processor according to any one of claims 1 to 7.

10. A computer-readable storage medium, wherein a plurality of instructions are stored in the storage medium; the plurality of instructions are used for a processor to load and execute the method as claimed in any one of claims 1 to 7.

Citation Information

Cited By

  • Remote sensing visual intelligent reasoning method and device based on big language model thinking chain

    CN121094140A

  • On-satellite application-oriented lightweight remote sensing image learning type compression method and system

    CN121711490A

  • Small target image generation method and device based on diffusion model, equipment and medium

    CN122067072A