Remote Sensing Image Aircraft Target Detection Method and Device Based on Feature Dynamic Calibration Optimization
Through the dynamic feature calibration optimization method, the accuracy and robustness of aircraft target detection in remote sensing images are improved, resource requirements are reduced, and the detection problem of traditional methods in complex backgrounds and low-resolution images is solved.
Patent Information
- Application Number
- CN202510032989.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-01-09
AI Technical Summary
Traditional remote sensing image aircraft object detection methods have low accuracy in complex backgrounds and low-resolution images, limited generalization capabilities, insufficient resource consumption and real-time performance, existing deep learning methods lack dynamic feature calibration capabilities, and high computing resource requirements.
Using a dynamic feature calibration optimization method, the remote sensing image features are dimensionally compressed, prompt weights generated, calibration and interactive optimization are performed through the dynamic calibration module. Combined with efficient feature calibration and optimization strategies, global average pooling layer, downsampled convolution layer, multi-head attention layer and mobile convolution blocks are used to reduce the computational complexity.
It improves detection accuracy and robustness, reduces model resource requirements, adapts to diverse computing environments and complex scenarios, and realizes efficient remote sensing image processing.
Smart Images

Figure CN119963812B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection, and more particularly to a method and device for detecting aircraft targets in remote sensing images based on feature dynamic calibration optimization. Background Art
[0002] The detection of aircraft targets in remote sensing images is to use the image data obtained by remote sensing technology to identify and locate aircraft targets in the image through an automated method. This technology has wide application value in the fields of military reconnaissance, air traffic management, disaster assessment, and environmental monitoring. Traditional aircraft target detection methods mainly rely on manual feature extraction and classical machine learning algorithms, such as support vector machines (SVMs), random forests, and sliding window methods based on feature pyramids. However, these methods face the following problems in practical applications:
[0003] 1. Insufficient adaptability to complex backgrounds and low-resolution images. Remote sensing images usually contain complex backgrounds, such as terrain undulations, buildings, and vegetation, and the target objects are usually sparsely distributed and small in size, resulting in low detection accuracy. In addition, due to imaging conditions (such as satellite resolution and weather conditions), remote sensing images often have problems of low resolution or severe noise interference, further increasing the detection difficulty.
[0004] 2. Limited generalization ability of the model. Traditional methods rely on fixed feature extraction strategies and cannot adapt to various target shapes and distribution characteristics. For example, there are significant differences in the appearance features of different aircraft types, and traditional methods have poor robustness to these changes and often require artificial design and adjustment of specific feature extraction rules.
[0005] 3. Insufficient resource consumption and real-time performance. With the rapid increase in the amount of remote sensing image data, higher requirements are put forward for the real-time performance and computational efficiency of detection algorithms. Traditional detection algorithms usually rely on the sliding window strategy for target search, and their computational complexity increases exponentially with the increase in image resolution and the number of targets, making it difficult to meet the actual needs of large-scale remote sensing data processing.
[0006] In recent years, deep learning-based object detection algorithms, such as YOLO, Faster R-CNN, and EfficientDet, have gradually become research hotspots in the field of remote sensing image object detection due to their automated feature extraction capabilities and high detection accuracy. However, these algorithms still face the following challenges in remote sensing applications:
[0007] 1. Insufficiency of dynamic feature calibration. In existing deep learning methods, a fixed feature extraction and processing process is usually adopted for all input images, lacking the dynamic adaptation ability to specific input conditions (such as different degradation types or target scales), resulting in a significant performance decline of the model in complex or unseen scenarios.
[0008] 2. Demand for efficient model design. Deep learning models usually require high computational resources, especially in the process of processing high-resolution remote sensing images. The video memory occupancy and computational time become the main bottlenecks in model deployment. How to design a lightweight and modular detection framework to reduce computational overhead and improve inference efficiency has become an urgent problem to be solved. Summary of the Invention
[0009] In view of the above technical problems, the present invention proposes a method and device for detecting aircraft targets in remote sensing images based on dynamic feature calibration optimization. By introducing a dynamic calibration module and combining efficient feature calibration and optimization strategies, the detection accuracy and robustness are significantly improved, while the resource requirements of the model are reduced. This method can adapt to diverse computational environments and remote sensing image processing tasks in complex scenarios, providing a more efficient and reliable solution for practical applications in related fields.
[0010] To achieve the above object, the present invention adopts the following technical solutions:
[0011] First, the present application discloses a method for detecting aircraft targets in remote sensing images based on dynamic feature calibration optimization, which is characterized in that a remote sensing image is acquired for feature extraction, and after the extracted features are dynamically calibrated and optimized, aircraft target detection is performed; the dynamic calibration and optimization steps include:
[0012] Compress the dimensions of the extracted features;
[0013] Generate hint weights according to the compressed features;
[0014] Calibrate the extracted features using the hint weights;
[0015] Make the calibrated features interact and optimize with the extracted features.
[0016] According to an embodiment of the present invention, the dimensions of the extracted features are compressed by a global average pooling layer and a downsampling convolutional layer in sequence.
[0017] According to an embodiment of the present invention, an activation function layer is used to generate hint weights according to the compressed features. The function activation layer is preferably softmax.
[0018] According to an embodiment of the present invention, acquiring a remote sensing image for feature extraction includes:
[0019] Segment the remote sensing image into several small pieces;
[0020] Linearly embed each small piece and add positional encoding to obtain a sequence of vectors;
[0021] Use the encoder to extract features from the sequence of vectors.
[0022] According to an embodiment of the present invention, multiple encoders are connected in series, and each encoder includes a first normalization layer, a multi-head attention layer, a first residual connection, a second normalization layer, a multi-layer perceptron, and a second residual connection.
[0023] According to an embodiment of the present invention, aircraft target detection includes: inputting the calibrated and optimized features into a decoder for decoding to output the target detection result;
[0024] The decoder includes multiple mobile convolutional blocks, a global average pooling layer, and a fully connected layer;
[0025] The mobile convolutional block includes a depthwise separable convolution, a pointwise convolution, and a residual link.
[0026] According to an embodiment of the present invention, the following loss function is used to train the aircraft target detection process;
[0027] L = λ1L c +λ2L s
[0028]
[0029] where λ1 and λ2 are hyperparameters for balancing the contribution weights of different loss functions, and L c represents the cross-entropy loss, y i represents the predicted output, that is, the probability that sample i belongs to each class, is the true label of sample i, and L s represents the smoothness loss, is the difference between the predicted value and the true value, and δ is a hyperparameter, usually taken as 1.0.
[0030] Secondly, the present application discloses a remote sensing image aircraft target detection device based on feature dynamic calibration optimization. This device applies the remote sensing image aircraft target detection method described in any one of the foregoing, and includes:
[0031] A feature extraction module for extracting features from the remote sensing image;
[0032] A feature dynamic calibration optimization module for dynamically calibrating and optimizing the extracted features; including:
[0033] A feature vector compression unit for compressing the dimensions of the extracted features;
[0034] A hint weight generation unit, configured to generate hint weights according to compressed features;
[0035] A feature calibration unit, configured to calibrate the extracted features by using the hint weights;
[0036] A feature optimization unit, configured to interactively optimize the calibrated features and the extracted features;
[0037] An aircraft target detection module, configured to perform aircraft target detection according to the calibrated and optimized features.
[0038] According to an embodiment of the present invention, it further includes an image processing module, configured to segment a remote sensing image into several small blocks; linearly embed each small block and add position encoding to obtain a vector sequence
[0039] The present invention discloses a method and device for aircraft target detection in remote sensing images based on feature dynamic calibration and optimization, which not only improves the accuracy of segmentation, but also reduces the resource requirements for model training and deployment. Compared with the prior art, the advantages include:
[0040] 1) By dynamically calibrating and optimizing the extracted features, the present application can adaptively adjust to different degradation types and target scales, thereby effectively improving the detection accuracy and robustness of the model in complex scenarios and avoiding the problem of performance degradation of traditional methods in unseen scenarios.
[0041] 2) The present application adopts efficient parameter fine-tuning and adapters to implement a lightweight dynamic calibration and optimization strategy, greatly reducing the model's demand for hardware resources and adapting to various computing environments. The design method proposed by the present invention makes the model easier to deploy, especially in high-resolution remote sensing image processing, significantly reducing the computational cost.
[0042] In summary, while improving the detection accuracy and robustness of aircraft targets in remote sensing images, the present application realizes the optimization of computing resources, providing a cost-effective solution for actual engineering deployment. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0044] Figure 1 It is a schematic flowchart of a method for aircraft target detection in remote sensing images based on feature dynamic calibration and optimization;
[0045] Figure 2 It is a process diagram for feature dynamic calibration optimization. Specific implementation manners
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0047] The embodiments of the present invention disclose a method and device for detecting aircraft targets in remote sensing images based on feature dynamic calibration optimization, aiming to improve the accuracy and robustness of remote sensing image target detection by introducing a dynamic calibration and optimization module, and at the same time reduce the resource requirements for model training and deployment, so that it can adapt to diverse computing environments and remote sensing image processing tasks in complex scenarios.
[0048] The aircraft target detection method in this embodiment includes:
[0049] Obtain a remote sensing image for feature extraction, perform dynamic calibration optimization on the extracted features, and then perform aircraft target detection.
[0050] In one embodiment, obtaining a remote sensing image for feature extraction includes:
[0051] (1) Segment the remote sensing image into several small blocks; linearly embed each small block and add position encoding to retain spatial information, obtaining a vector sequence;
[0052] In this embodiment, the input image X is segmented into several small blocks; assume the size of the image X is H×W×C, where H is the height, W is the width, and C is the number of channels. The image is segmented into N small blocks, and the size of each small block is p×p. Therefore, the image is segmented into small blocks, and there are a total of small blocks;
[0053] Furthermore, perform linear embedding on each small block to map each small block to a vector of a fixed dimension. To retain the spatial information in the image, add position encoding to each embedded vector to obtain a complete vector sequence.
[0054] (2) Use the encoder to perform feature extraction on the vector sequence to obtain the final feature representation.
[0055] In this application, there are L encoders connected end to end, and each encoder layer includes a first normalization layer, a multi-head attention layer, a first residual connection, a second normalization layer, a multi-layer perceptron, and a second residual connection. For details, refer to Figure 1After being processed by the L-layer encoder, the feature representation Φ is obtained.
[0056] In one embodiment, dynamic calibration optimization enhances the image restoration ability by generating conditional prompt information, encoding degraded context knowledge, and dynamically adjusting network behavior to restore images from different degradations. The dynamic calibration optimization in this embodiment mainly includes three stages: prompt weight generation, conditional prompt adjustment, and feature interaction optimization. Specifically:
[0057] Compress the dimension of the extracted features;
[0058] Generate prompt weights based on the compressed features;
[0059] Calibrate the extracted features using the prompt weights;
[0060] Perform interactive optimization on the calibrated features and the extracted features.
[0061] In an exemplary embodiment, first, generate prompt weights through global average pooling and channel downsampling to dynamically adjust the input features; subsequently, generate conditional prompts through convolutional processing and adapt to different resolutions through bilinear upsampling; finally, the prompt information interacts with the input features to optimize the decoder behavior, effectively removing degradations and restoring image details.
[0062] Specifically refer to Figure 2 , Figure 2 which is an instance diagram of the dynamic calibration optimization process; the corresponding steps are as follows:
[0063] 1. Feature vector compression: Define the extracted feature as Φ. In this embodiment, first perform global average pooling operation on it. The global average pooling operation compresses the spatial information of each channel into a scalar, thereby generating a compact feature vector. Furthermore, the feature vector v is processed through a channel downsampling convolutional layer (1*1 convolution) to generate a compact feature vector. The role of the channel downsampling convolutional layer is to reduce the dimension of the feature vector and make it more compact.
[0064] 2. Generate prompt weights: In this embodiment, apply an activation function (softmax) operation to the downsampled feature vector v′ to generate prompt weights for dynamically adjusting the prompt component to enable conditional adjustment according to the input features. The expression is as follows:
[0065] ω = Softmax(Conυ 1×1 (GlobalAveragePooling(Φ)))
[0066] Among them, Conv represents the convolution operation, and GlobalAveragePooling represents the global average pooling operation.
[0067] 3. Feature calibration: Use the generated prompt weight w to adjust the input feature Φ to generate the conditional prompt φ. By adjusting the prompt weight, the input feature can be dynamically adjusted. The adjusted prompt φ is further processed through a convolutional layer to help maintain the spatial consistency of the input feature: Since a smaller convolutional kernel can avoid excessive loss of spatial information and ensure that the context relationship of each feature position can be well preserved in the further processing, in this embodiment, a 3×3 convolution is adopted to improve the feature representation without sacrificing spatial accuracy; The formula is expressed as:
[0068]
[0069] where i represents the prompt index.
[0070] 4. Feature optimization: To process images of different resolutions, in this embodiment, the adjusted conditional prompt φ is upsampled to the same size as the input feature representation Φ through bilinear upsampling; Then, the upsampled prompt component is interacted with the input feature Φ to accurately remove degradation and restore image details to obtain the optimized Φ′.
[0071] The dynamic calibration and optimization process of the present invention can capture and encode various degradation feature information, thereby significantly improving the generalization ability for unknown or unseen degradation types. In addition, the dynamic calibration and optimization decouples the encoding of the degradation type from the prompt generation, avoiding relying on contrastive learning or multi-stage training, simplifies the model design and training process, and can be directly embedded into the existing network as a lightweight module.
[0072] In one embodiment, the calibrated and optimized features are input into the decoder for decoding to output the target detection result; In this embodiment, the decoder completes the feature decoding and classification of the target detection, uses the mobile convolutional block to gradually restore the spatial resolution of the feature map, and extracts features through global average pooling and fully connected layers to output the target detection result (category and location information).
[0073] In this embodiment, the decoder consists of multiple Mobile Inverted Bottleneck Convolution (MBConv) blocks, and each Mobile Inverted Bottleneck Convolution block contains depthwise separable convolution, pointwise convolution, and residual connection. Specifically, first, perform depthwise separable convolution operation on the input feature Φ′ to extract spatial features. The depthwise separable convolution reduces the computational complexity by independently applying the convolution kernel on each channel; then, perform pointwise convolution operation on the output of the depthwise separable convolution to mix channel information. The pointwise convolution is a 1×1 convolution operation; finally, perform residual connection between the input feature Φ′ and the output of the pointwise convolution to improve the expressive power and generalization ability of the network, and obtain Φ out :
[0074]
[0075] Among them, L represents the number of convolution blocks, DepthwiseConv represents depthwise separable convolution, and the spatial resolution of the feature map is restored layer by layer through L Mobile Inverted Bottleneck Convolution blocks.
[0076] At the last layer of the decoder, perform global average pooling operation on the restored feature map to generate a compact feature vector. Input the feature vector after global average pooling into the fully connected layer for final feature extraction and classification. The output of the fully connected layer is the aircraft target detection result y.
[0077] In one embodiment, the entire detection process needs to be trained before aircraft target detection. In this embodiment, to improve the generalization ability of the detection model, the training dataset is expanded, including image enhancement, denoising, cropping, size normalization, data annotation, and data augmentation, etc.
[0078] That is, first, improve the image quality and reduce noise interference through image enhancement and denoising processing; then, crop and normalize the size of the image to make it meet the model input requirements; next, perform data annotation to mark the position and boundary of the aircraft target; finally, expand the training dataset through data augmentation techniques (such as rotation, flipping, scaling, etc.).
[0079] During the model training process, optimize the classification task through cross-entropy loss and optimize the regression task through smooth L1 loss to improve the accuracy and robustness of the overall model and achieve accurate target recognition and positioning.
[0080] Among them, the cross-entropy loss needs to calculate the difference between the predicted class y and the true class \hat{y}; the formula is expressed as:
[0081]
[0082] Among them, y i is the predicted output of the model, representing the probability that sample i belongs to each class; is the true label of sample i, usually in one-hot encoding. The cross-entropy loss guides the model to adjust the weights to improve the classification accuracy by quantifying the gap between the predicted probability and the true label.
[0083] In addition, to better optimize the model parameters in the regression task, the present invention adopts the smooth L1 loss function. The smooth L1 loss is similar to the L2 loss when the error is small, and approaches the L1 loss when the error is large, reducing the impact of outliers on model training. The calculation formula of the smooth L1 loss is as follows:
[0084]
[0085] where is the difference between the predicted value and the true value, and δ is a hyperparameter, usually taken as 1.0. Through the smooth L1 loss, the model can balance accuracy and robustness during training.
[0086] The total loss of the present invention that comprehensively uses the cross-entropy loss and the smooth L1 loss is:
[0087] L = λ1L c + λ2L s
[0088] where λ1 and λ2 are two hyperparameters used to balance the contribution weights of different loss functions.
[0089] The present invention can simultaneously optimize the performance of the model in classification tasks and regression tasks, so that the model has higher accuracy and stability in practical applications.
[0090] Finally, the target detection result y output by the decoder is the final detection result, including the category and location information of the target.
[0091] In another embodiment, the present application provides a remote sensing image aircraft target detection device based on feature dynamic calibration optimization. The device executes the remote sensing image aircraft target detection method based on feature dynamic calibration optimization described in any of the above embodiments, and the structure includes:
[0092] A feature extraction module for extracting features from the remote sensing image;
[0093] A feature dynamic calibration optimization module for dynamically calibrating and optimizing the extracted features; including:
[0094] A feature vector compression unit for compressing the dimensions of the extracted features;
[0095] A hint weight generation unit for generating hint weights according to the compressed features;
[0096] A feature calibration unit for calibrating the extracted features using the prompt weights;
[0097] A feature optimization unit for interacting and optimizing the calibrated features with the extracted features;
[0098] An aircraft target detection module for detecting aircraft targets based on the calibrated and optimized features.
[0099] It further includes an image processing module for segmenting the remote sensing image into several small blocks; linearly embedding each small block and adding position encoding to obtain a vector sequence.
[0100] The detection device of this application is designed to be lightweight, which can avoid complex training processes and significantly improve the generalization ability of the model to unknown degradations.
[0101] The advantages of this application compared with the prior art include:
[0102] 1) Significantly improve the detection accuracy and robustness through the dynamic calibration and optimization process. The dynamic calibration and optimization module introduced in the present invention can dynamically generate conditional prompt information, capture and encode various complex degradation features, thereby effectively adjusting the network behavior and improving the adaptability of the model to different degradation types. Through the deep interaction between the prompt interaction module and the features, the decoder can more accurately remove degradations and restore image details, significantly improving the accuracy of aircraft target detection in remote sensing images. In addition, this module also decouples the degradation feature encoding and prompt generation, avoiding relying on cumbersome contrast learning or multi-stage training, greatly reducing the complexity of model design and optimization, and enhancing the robustness of the model in the actual environment.
[0103] 2) The lightweight design reduces resource requirements and adapts to various computing environments. The present invention adopts an efficient parameter fine-tuning and adapter module. Through a lightweight dynamic calibration and optimization strategy, it significantly reduces the hardware resources required for model training and inference. For example, the mechanism of dynamically adjusting the prompt weights can accurately adapt to the input features without significantly increasing the computational burden. The design of the mobile convolutional block (MBConv) in the decoder further optimizes the computational efficiency and significantly reduces the dependence of the model on high-performance hardware. At the same time, by combining data augmentation and model optimization, the present invention can balance accuracy and computational cost, enabling it to flexibly adapt to remote sensing image processing tasks in various computing environments and providing a cost-effective solution for actual engineering deployment.
[0104] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0105] The foregoing description of the disclosed embodiments enables those skilled in the art to make or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A remote sensing image aircraft target detection method based on feature dynamic calibration optimization, characterized in that Obtain remote sensing images for feature extraction, perform dynamic calibration and optimization on the extracted features, and then conduct aircraft target detection; Among them, the dynamic calibration and optimization steps include: Perform dimensional compression on the extracted features and generate hint weights based on the compressed features; including applying the activation function softmax operation to generate hint weights Used to dynamically adjust the hint component so that it can be conditionally adjusted according to the input features. The expression is as follows: w=Softmax(Conv 1×1 (GlobalAoeragePooling(Φ))) Among them, Conv represents the convolution operation, GlobalAveragePooling represents the global average pooling operation, and Φ represents the extracted features; Calibrate the extracted features using the prompt weights; Enable the calibrated features to interact and optimize with the extracted features.
2. The remote sensing image aircraft target detection method according to claim 1, wherein Successively compress the dimensions of the extracted features through the global average pooling layer and the downsampling convolutional layer.
3. The remote sensing image aircraft target detection method according to claim 1, characterized in that, Use the activation function layer to generate prompt weights based on the compressed features.
4. The remote sensing image aircraft target detection method according to claim 1, wherein, Obtain remote sensing images for feature extraction, including: Segment the remote sensing image into several small blocks; Linearly embed each small block and add position encoding to obtain a vector sequence; Use the encoder to extract features from the vector sequence.
5. The remote sensing image aircraft target detection method according to claim 4, characterized in that There are multiple encoders in series, and each encoder includes a first normalization layer, a multi-head attention layer, a first residual connection, a second normalization layer, a multi-layer perceptron, and a second residual connection.
6. The remote sensing image aircraft target detection method according to claim 1, wherein Conducting aircraft target detection includes: inputting the calibrated and optimized features into the decoder for decoding and outputting the target detection result; The decoder includes multiple mobile convolutional blocks, a global average pooling layer, and a fully connected layer; The mobile convolutional block includes a depthwise separable convolution, a point convolution, and a residual link.
7. The remote sensing image aircraft target detection method according to claim 1, characterized in that, Use the following loss function to train the aircraft target detection process; L = λ1L c + λ2L s where λ1 and λ2 are hyperparameters for balancing the contribution weights of different loss functions, and L c represents the cross-entropy loss, y i represents the predicted output, is the true label of sample i, and L s represents the smoothing loss, is the difference between the predicted value and the true value, and δ is a hyperparameter.
8. A remote sensing image aircraft target detection device based on feature dynamic calibration optimization, which applies the remote sensing image aircraft target detection method based on feature dynamic calibration optimization according to any one of claims 1-7, characterized in that, Including: A feature extraction module for extracting features from remote sensing images; A feature dynamic calibration and optimization module for dynamically calibrating and optimizing the extracted features; including: A feature vector compression unit for compressing the dimensions of the extracted features; A prompt weight generation unit for generating prompt weights based on the compressed features; A feature calibration unit for calibrating the extracted features using the prompt weights; A feature optimization unit for enabling the calibrated features to interact and optimize with the extracted features; A target detection module for conducting aircraft target detection based on the calibrated and optimized features.
9. The remote sensing image aircraft target detection device according to claim 8, wherein, It also includes an image processing module for segmenting the remote sensing image into several small blocks; linearly embedding each small block and adding position encoding to obtain a vector sequence.
Citation Information
Patent Citations
Gesture image feature extraction method based on dynamic fusion mechanism
CN112836651A
Target detection method and system based on global feature perception
CN113673420A