Remote sensing image target detection and recognition method based on multi-model integration and progressive prediction
Through the multi-model integration and progressive prediction methods, combined with the diversity training augmentation strategy and multi-scale feature fusion, the problem of improving the accuracy of remote sensing image object detection is solved, and high-precision object detection and recognition is achieved.
Patent Information
- Application Number
- CN202411593334.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-11-08
AI Technical Summary
The existing remote sensing image object detection and recognition methods are difficult to improve the object detection and recognition accuracy in complex remote sensing scenarios, especially in the case of high intra-class differences and low intra-class differences.
Using multi-model integration and progressive prediction methods, the augmentation strategy and multi-scale feature fusion are used to combine the advantages of Transformer and CNN structures to detect and identify targets.
The accuracy of remote sensing image object detection and recognition is significantly improved, the generalization ability of the model and the accuracy of target extraction in complex scenarios are improved. The experimental results show that mAP can be improved by more than 10%.
Smart Images

Figure CN119559499B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a remote sensing image target detection and recognition method based on multi-model integration and progressive prediction, and belongs to the technical field of target detection and recognition. Background Art
[0002] With the development of high-resolution remote sensing imaging technology, images have more refined texture detail features, making it possible to detect and fine-grainedly identify typical targets. Among them, specific targets represented by aircraft and ships have always been important research targets in the field of remote sensing image processing due to their great value in the military and civilian fields. Rapid detection and identification of high-value targets through remote sensing images is of great significance to military dynamic analysis, battlefield defense warning, civil route planning, safety search and rescue, etc.
[0003] Early detection and recognition methods required artificially designed image feature extraction operators for different types of targets. Common features for surface targets include geometric contour features, texture edge features, symmetry, etc. However, such methods have limited ability to represent target features and are only applicable to target detection in specific scenarios. With the improvement of hardware computing power and the surge in data samples, deep learning-based methods have emerged and are gradually being applied to remote sensing image target detection and recognition tasks. Such methods can adaptively learn more robust and powerful features of the target.
[0004] SCMask RCNN uses an improved self-calibrated convolution (SC-Conv) in the backbone feature extraction network of Mask RCNN to obtain more discriminative features of the target. GLF-Net adopts an encoder-decoder architecture and uses global and local multi-scale feature fusion to improve the accuracy of target detection. HFF-YOLO proposes a manual feature fusion to enhance the detection effect by combining prior knowledge of the target with deep features.
[0005] Existing methods focus on designing feature enhancement and feature fusion structures to improve the network's perception of remote sensing image targets. However, due to factors such as complex remote sensing scenes and special target characteristics, existing methods have limited improvement in target detection and recognition accuracy. Specifically, due to the special bird's-eye view of satellite images, the scene will contain targets from multiple angles and a variety of false alarm sources; and specific targets such as aircraft themselves have similar geometric structures, showing small discernible differences between different types of targets. At the same time, due to interference from imaging, environment and other factors, the same type of target may present completely different characteristics in the image. The high intra-class differences and low inter-class differences of such targets make it difficult for the network to extract identifiable features of the target, affecting the target recognition accuracy.
[0006] In view of the above-mentioned difficult problems, the present invention proposes a detection and recognition method for remote sensing image targets based on the ideas of multi-model integration and progressive prediction. Summary of the invention
[0007] In order to solve the problems existing in the background technology, the present invention provides a remote sensing image target detection and recognition method based on multi-model integration and progressive prediction.
[0008] To achieve the above object, the present invention adopts the following technical solution: a remote sensing image target detection and recognition method based on multi-model integration and progressive prediction, the method comprising the following steps:
[0009] S1: Divide the data set into a training set and a validation set in proportion, input the training set into the network for training; perform online augmentation of the training set through a training augmentation strategy to obtain augmented input data;
[0010] The online augmentation described in S1 includes the following steps:
[0011] S1-1: Use the Mosaic augmentation strategy to augment the training set online: randomly crop the images used for training, and splice the cropped images into one image as training data for training;
[0012] S1-2: Use the Copy-paste augmentation strategy to augment the training set online: randomly copy the target to the background to generate an augmented image;
[0013] S1-3: Use the Mixup augmentation strategy to augment the training set online: randomly select two images and add them in proportion.
[0014] S2: Input the augmented input data into the backbone feature extraction network composed of Swin-Transformer Blocks to obtain feature maps of different levels, and input the feature maps of different levels into the basic detection and recognition network composed of the rotating box multi-stage progressive prediction head to perform target positioning and classification training to obtain the trained model A;
[0015] The acquisition of feature maps of different levels described in S2 includes the following steps:
[0016] S2-1-1: Divide each image into multiple non-overlapping batches, and perform linear transformation on the channel data of each pixel;
[0017] S2-1-2: Apply a linear transformation layer to the original value features, project them to an arbitrary dimension, and then apply two Swin Transformer Blocks with multi-head self-attention calculations on all batches;
[0018] The Swin Transformer Block described in S2-1-2 includes a multi-head self-attention module based on a regular window / offset window and a double-layer multi-layer perception module thereafter, a layer normalization is applied before the multi-head self-attention module based on a regular window / offset window and the double-layer multi-layer perception module, and a residual connection auxiliary training is applied after the multi-head self-attention module based on a regular window / offset window and the double-layer multi-layer perception module;
[0019] In the conventional window-based multi-head self-attention module, the M×M batches are merged into a set of non-overlapping windows, and then each batch is expanded into a one-dimensional vector x p , as the input of Transformer, self-attention calculation is performed in different windows;
[0020] In the multi-head self-attention module based on the offset window, the window position is moved on the basis of the window division of the multi-head self-attention module based on the regular window in the previous layer Swin Transformer Block to calculate the self-attention. The calculation process is as follows:
[0021]
[0022] In formula (1):
[0023] Represents the output features of the multi-head self-attention module based on regular window / offset window of the lth Swin Transformer Block;
[0024] z l Represents the output features of the double-layer multi-layer perception module of the lth Swin Transformer Block;
[0025] W-MSA means multi-head self-attention using a multi-head self-attention module with regular windows;
[0026] SW-MSA stands for Multi-Head Self-Attention using Multi-Head Self-Attention Module with offset window;
[0027] MLP stands for a two-layer multi-layer perception module operation;
[0028] LN represents the layer normalization operation.
[0029] The target positioning and classification training described in S2 includes the following steps:
[0030] S2-2-1: Initial candidate box generation;
[0031] S2-2-2: Multi-stage training of cascaded prediction heads;
[0032] S2-2-3: Optimize step by step.
[0033] S2-1-3: Reduce the number of batches through batch merging layers to produce a hierarchical feature map representation;
[0034] S2-1-4: Using the feature map obtained in S2-1-3 as input, repeat the operation of S2-1-3 twice to obtain different feature maps respectively, thereby obtaining feature map representations at different levels.
[0035] S3: Input the augmented input data into the feature extraction network with ReResNet as the backbone to obtain a high-level feature map, and input the high-level feature map into the basic detection and recognition network composed of the rotating box multi-stage progressive prediction head to perform target positioning and classification training to obtain the trained model B;
[0036] The ReResNet backbone feature extraction network described in S3 reimplements all layers in the ResNet backbone network based on the rotational equivariant network of e2cnn.
[0037] S4: Use Model A and Model B to build a multi-model integrated prediction framework and output target detection and recognition results.
[0038] S4-1: Upsample and downsample the validation set images of model A and model B in the inference phase to obtain multi-scale input images with scaling factors of 0.5, 1.0, and 1.5;
[0039] S4-2: Use model A and model B to infer the validation set and obtain their respective inference results;
[0040] S4-2-1: Use model A to infer the multi-scale input image, and obtain the multi-scale input image inference results with scale factors of 0.5, 1.0 and 1.5 respectively, and reverse sample the inference results to restore them to the original size inference results, and finally obtain image inference results A-1, A-2 and A-3;
[0041] S4-2-2: Use model B to infer the multi-scale input image, and obtain the multi-scale input image inference results with scale scaling factors of 0.5, 1.0 and 1.5 respectively, and reverse sample the inference results to restore them to the original size inference results, and finally obtain image inference results B-1, B-2 and B-3;
[0042] S4-3: Use WBF to perform decision-level fusion on the inference results, adjust the coordinates and confidence scores based on the inference results of the two models, and obtain the final target recognition result, realizing the complementary advantages of different models.
[0043] S4-3-1: Create three empty lists:
[0044] Prediction box list B is used to place the prediction boxes to be processed;
[0045] The box cluster list L is used to place the set of target prediction boxes;
[0046] The box fusion list F is used to place the fusion boxes processed based on the predicted boxes.
[0047] That is, each position in L contains a set of all prediction boxes of the same target, while each position in F contains only one fused prediction box;
[0048] S4-3-2: Put the multi-scale input image inference results A-1, A-2, A-3, B-1, B-2 and B-3 of different models obtained in S4-2-1 and S4-2-2 into the prediction box list B;
[0049] S4-3-3: Sort the prediction box list B in descending order of confidence score C;
[0050] S4-3-4: Traverse the prediction box list B in order, and for each prediction box B i Find the matching box in the box fusion list F, and define the match as the intersection over union ratio (IoU) > Th between the predicted box and the fused box. IoU is the ratio of the intersection area of the predicted box and the fused box to the union area. Th is the threshold set to determine whether the predicted box and the fused box are the same target.
[0051] If no matching box is found in the box fusion list F, an empty new set is added to the end of the box cluster list L and the box fusion list F, and the predicted box B in the predicted box list B is added to the end of the box cluster list L and the box fusion list F. i Add to the collection respectively;
[0052] If a matching box is found in the box fusion list F, the predicted box B i Add to the set of corresponding positions in the box cluster list L, and recalculate the coordinates and confidence of the corresponding fused box in the box fusion list F according to all the predicted boxes of the set in the box cluster list L;
[0053] The box fusion list F obtained after the traversal is the inference result of the target in the final image.
[0054] The confidence score of the fused frame described in S4-3-4 is the average confidence of all frames that make up the confidence score. The coordinates of the fused frame are the weighted sum of the frame coordinates, where the weight is the confidence score of the corresponding frame. The specific calculation formula is as follows:
[0055]
[0056] In formula (2)-(4):
[0057] C represents the predicted box B i The confidence of the fused box in the corresponding box fusion list F;
[0058] n represents the predicted box B i The number of all predicted boxes in the corresponding box fusion list L;
[0059] C i Represents the confidence of the i-th prediction box in the set;
[0060] Represents the horizontal coordinate of the corner point of the kth prediction box of the i-th prediction box in the set, k = 1, 2, 3, 4;
[0061] Represents the ordinate of the corner point of the kth prediction box of the i-th prediction box in the set, k = 1, 2, 3, 4;
[0062] x k Represents the horizontal coordinate of the corner point of the k-th prediction box of the fusion box, k = 1, 2, 3, 4;
[0063] y k Represents the ordinate of the corner point of the k-th predicted box of the fused box, k = 1, 2, 3, 4.
[0064] Compared with the prior art, the present invention has the following beneficial effects:
[0065] 1. In the process of network training, the present invention fully combines the diversity training augmentation strategy to obtain more diverse training samples and improve the generalization ability of the model; a rotating box progressive prediction head is proposed. For a single model, a multi-stage prediction branch is cascaded after the region proposal network, and then combined with a multi-model integrated prediction framework composed of multi-scale image input and decision-level fusion, the model prediction accuracy is fully improved. Experimental results show that this method can improve mAP by more than 10% compared with a single model prediction, and can fully tap the potential of existing detection methods for remote sensing target detection and recognition.
[0066] 2. In order to reduce the influence of false alarm sources on the decision-making accuracy of a single model in complex remote sensing scenes, the present invention makes full use of the complementary advantages of Transformer and CNN structures, selects the two as the backbone feature extraction networks, and performs multi-scale scaling on the images input to each model in the reasoning stage. The detection and recognition results are fused at the decision level from both aspects of the model and between models to improve the target extraction accuracy in complex scenes.
[0067] 3. In order to alleviate the recognition difficulties caused by large differences within the target class and small differences between classes, the present invention cascades multi-stage prediction branches after the region proposal network for the single model before fusion. The cascaded prediction heads are trained and inferred under different IoU thresholds, so that the network can predict difficult samples from coarse to fine, thereby significantly improving the prediction quality of the single model. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 is a flow chart of the present invention;
[0069] Figure 2 It is a schematic diagram of the multi-model integrated prediction framework;
[0070] Figure 3 It is a visualization diagram of the detection and recognition results, wherein: the first row is the true value, and the second row is the prediction result of the present invention. DETAILED DESCRIPTION
[0071] The technical solution of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0072] A remote sensing image target detection and recognition method based on multi-model integration and progressive prediction, the method comprising the following steps:
[0073] S1: Divide the data set into a training set and a validation set in a ratio of 4:1, input the training set into the network for training; perform online augmentation of the training set through the training augmentation strategy to obtain the augmented input data;
[0074] The online augmentation described in S1 includes the following steps:
[0075] S1-1: Use the Mosaic augmentation strategy to perform online augmentation on the training set: randomly crop the four pictures used for training, and splice the cropped pictures into one picture as training data for training; Mosaic data augmentation greatly enriches the background of the image through random cropping, which in turn increases the number of samples and batch training number, and accelerates the convergence speed of the model.
[0076] S1-2: Use the Copy-paste augmentation strategy to perform online augmentation on the training set: randomly copy the target to the background to generate an augmented image; this method can enrich the combination of background and target in the data. The present invention sets the probability of performing augmentation to 0.25.
[0077] S1-3: Use the Mixup augmentation strategy to augment the training set online: randomly extract two images and add them in a certain ratio to improve the generalization ability of the model and the robustness to adversarial samples. The present invention sets the probability of performing augmentation to 0.25.
[0078] S2: Input the augmented input data into the backbone feature extraction network composed of Swin-Transformer Blocks to obtain feature maps of different levels, ensuring the effective capture of information at multiple scales. Input the feature maps of different levels into the basic detection and recognition network composed of the rotating box multi-stage progressive prediction head to perform target positioning and classification training, and obtain the trained model A;
[0079] The acquisition of feature maps of different levels described in S2 includes the following steps:
[0080] S2-1-1: Each image is divided into multiple non-overlapping batches through batches, and the channel data of each pixel is linearly transformed; input an RGB image of size H×W×3, firstly the batch partitioning module is used to divide the input RGB image into non-overlapping batches, and its features are set as the concatenation of the original pixel RGB values. Taking the size of 4×4 as an example, the feature dimension of each batch is 4×4×3=48, that is, the image is divided into Batch
[0081] S2-1-2: Apply a linear transformation layer to the original value feature and project it to any dimension (denoted as C, C is a hyperparameter, the setting of the C value is related to its scale, the present invention adopts Swin-Large, C = 196), and then apply two Swin Transformer Blocks with multi-head self-attention calculation on all batches; Swin Transformer Blocks consists of two series-connected Swin Transformer Blocks, the first Swin Transformer Block is based on a multi-head self-attention module (W-MSA) with a conventional window configuration, and the second Swin Transformer Block is based on a multi-head self-attention module (SW-MSA) with an offset window, and Swin Transformer Blocks are used to maintain the number of batches Output feature map.
[0082] The Swin Transformer Block described in S2-1-2 includes a multi-head self-attention module based on a regular window / offset window and a double-layer multi-layer perception module thereafter, a layer normalization is applied before the multi-head self-attention module based on a regular window / offset window and the double-layer multi-layer perception module, and a residual connection auxiliary training is applied after the multi-head self-attention module based on a regular window / offset window and the double-layer multi-layer perception module;
[0083] In the conventional window-based multi-head self-attention module, the M×M batches are merged into a set of non-overlapping windows, and then each batch is expanded into a one-dimensional vector x p , as the input of Transformer, self-attention calculation is performed in different windows;
[0084] In the multi-head self-attention module based on the offset window, the window position is moved on the basis of the window division of the multi-head self-attention module based on the regular window in the previous layer Swin Transformer Block to calculate the self-attention. The calculation process is as follows:
[0085]
[0086] In formula (1):
[0087] Represents the output features of the multi-head self-attention module based on regular window / offset window of the lth Swin Transformer Block;
[0088] z l Represents the output features of the double-layer multi-layer perception module of the lth Swin Transformer Block;
[0089] W-MSA means multi-head self-attention using a multi-head self-attention module with regular windows;
[0090] SW-MSA stands for Multi-Head Self-Attention using Multi-Head Self-Attention Module with offset window;
[0091] MLP stands for a two-layer multi-layer perception module operation;
[0092] LN represents the layer normalization operation.
[0093] S2-1-3: Reduce the number of batches through batch merging layers to produce hierarchical feature map representations; the first batch merging layer concatenates the features of each group of 2×2 neighborhood blocks to obtain The feature map is taken from the 4C-dimensional concatenated features, and a 1×1 linear convolution layer is applied to reduce the batch size by a factor of 2×2=4 (2× resolution downsampling), and the output dimension is set to 2C. The Swin Transformer block is then used to transform the features, and the resolution is kept at get feature map.
[0094] S2-1-4: Take the feature map obtained in S2-1-3 as input, repeat the operation of S2-1-3 twice, and get different ( and ) feature maps to obtain feature map representations at different levels.
[0095] The target positioning and classification training described in S2 includes the following steps:
[0096] S2-2-1: Initial Candidate Box Generation: In the first stage, the model first uses a conventional rotation box region generation network to generate a series of target candidate boxes, which will serve as input for subsequent stages.
[0097] S2-2-2: Multi-stage training of cascaded prediction heads: Cascade three stages of detectors for prediction, and each stage of the prediction head corresponds to an independent detector, which are trained at different IoU thresholds. The first detector is trained at a lower IoU threshold (such as 0.5), while subsequent detectors are trained at higher IoU thresholds (such as 0.6, 0.7). As the number of stages increases, the IoU threshold gradually increases, allowing the detectors at each stage to focus on targets that are more difficult to predict.
[0098] S2-2-3: Step-by-step optimization. The detector in each cascade stage will use the output candidate box of the previous stage for further bounding box regression and classification. Through step-by-step optimization, each detector can output more accurate coordinates and categories, gradually improving the quality of the prediction results.
[0099] S3: Input the augmented input data into the feature extraction network with ReResNet as the backbone to obtain a high-level feature map, which ensures the effective representation of targets at different angles and improves the model's ability to correctly identify rotated objects. The high-level feature map is input into the basic detection and recognition network composed of the multi-stage progressive prediction head of the rotating box to perform target positioning and classification training to obtain the trained model B;
[0100] The ReResNet backbone feature extraction network described in S3 reimplements all layers of the ResNet backbone network based on the rotational equivariant network of e2cnn. Inputting the image into the rotational equivariant network can obtain the rotational equivariant feature map Different from conventional feature maps, this feature map has N orientation channels, corresponding to N discrete rotation transformations.
[0101] S4: Use Model A and Model B to build a multi-model integrated prediction framework and output target detection and recognition results.
[0102] S4-1: Upsample and downsample the validation set images of model A and model B in the inference phase to obtain multi-scale input images with scale factors of 0.5, 1.0, and 1.5 to improve the reliability of inference;
[0103] S4-2: Use model A and model B to infer the validation set respectively and obtain their respective inference results, including the category, confidence, and coordinates of the four corner points of the detection box;
[0104] S4-2-1: Use model A to infer the multi-scale input image, and obtain the multi-scale input image inference results with scale factors of 0.5, 1.0 and 1.5 respectively, and reverse sample the inference results to restore them to the original size inference results, and finally obtain image inference results A-1, A-2 and A-3;
[0105] S4-2-2: Use model B to infer the multi-scale input image, and obtain the multi-scale input image inference results with scale scaling factors of 0.5, 1.0 and 1.5 respectively, and reverse sample the inference results to restore them to the original size inference results, and finally obtain image inference results B-1, B-2 and B-3;
[0106] S4-3: Use WBF (Weighted boxes fusion) to perform decision-level fusion on the reasoning results, adjust the coordinates and confidence scores based on the reasoning results of the two models, and obtain the final target recognition result, realizing the complementary advantages of different models.
[0107] S4-3-1: Create three empty lists:
[0108] Prediction box list B is used to place the prediction boxes to be processed;
[0109] The box cluster list L is used to place the set of target prediction boxes;
[0110] The box fusion list F is used to place the fusion boxes processed based on the predicted boxes.
[0111] That is, each position in L contains a set of all prediction boxes of the same target, while each position in F contains only one fused prediction box;
[0112] S4-3-2: Put the multi-scale input image inference results A-1, A-2, A-3, B-1, B-2 and B-3 of different models obtained in S4-2-1 and S4-2-2 into the prediction box list B;
[0113] S4-3-3: Sort the prediction box list B in descending order of confidence score C;
[0114] S4-3-4: Traverse the prediction box list B in order, and for each prediction box B i Find the matching box in the box fusion list F, and define the match as the intersection-over-union ratio IoU>Th between the predicted box and the fused box, where IoU is the ratio of the intersection area of the predicted box and the fused box to the union area, and Th is the threshold set for determining whether the predicted box and the fused box are the same target, which is Th=0.6 in the present invention;
[0115] If no matching box is found in the box fusion list F, an empty new set is added to the end of the box cluster list L and the box fusion list F, and the predicted box B in the predicted box list B is added to the end of the box cluster list L and the box fusion list F. i Add to the collection respectively;
[0116] If a matching box is found in the box fusion list F, the predicted box B i Add to the set of corresponding positions in the box cluster list L, and recalculate the coordinates and confidence of the corresponding fused box in the box fusion list F according to all the predicted boxes of the set in the box cluster list L;
[0117] The box fusion list F obtained after the traversal is the inference result of the target in the final image.
[0118] The confidence score of the fused frame described in S4-3-4 is the average confidence of all frames that make up the confidence score. The coordinates of the fused frame are the weighted sum of the frame coordinates, where the weight is the confidence score of the corresponding frame. The specific calculation formula is as follows:
[0119]
[0120] In formula (2)-(4):
[0121] C represents the predicted box B i The confidence of the fused box in the corresponding box fusion list F;
[0122] n represents the predicted box B i The number of all predicted boxes in the corresponding box fusion list L;
[0123] C i Represents the confidence of the i-th prediction box in the set;
[0124] Represents the horizontal coordinate of the corner point of the k-th prediction box of the i-th prediction box in the set, k = 1, 2, 3, 4;
[0125] Represents the ordinate of the corner point of the kth prediction box of the i-th prediction box in the set, k = 1, 2, 3, 4;
[0126] x k Represents the horizontal coordinate of the corner point of the k-th prediction box of the fusion box, k = 1, 2, 3, 4;
[0127] y k Represents the ordinate of the corner point of the k-th predicted box of the fused box, k = 1, 2, 3, 4.
[0128] Figure 3 Qualitative visualization results and quantitative evaluation results are given, which can prove the superiority of the present invention in remote sensing image target detection and recognition.
[0129] Table 1 Experimental evaluation results of detection and recognition of specific aircraft targets on the FAIR1M dataset
[0130]
[0131]
[0132] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations within the meaning and range of equivalents of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.
[0133] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.
Claims
1. A remote sensing image target detection and recognition method based on multi-model integration and progressive prediction, characterized by: The method comprises the following steps: S1: Divide the data set into a training set and a validation set in proportion, input the training set into the network for training; perform online augmentation of the training set through a training augmentation strategy to obtain augmented input data; S2: Input the augmented input data into the backbone feature extraction network composed of Swin-Transformer Blocks to obtain feature maps of different levels, and input the feature maps of different levels into the basic detection and recognition network composed of the rotating box multi-stage progressive prediction head to perform target positioning and classification training to obtain the trained model A; The acquisition of feature maps of different levels described in S2 includes the following steps: S2-1-1: Divide each image into multiple non-overlapping batches, and perform linear transformation on the channel data of each pixel; S2-1-2: Apply a linear transformation layer to the original value features, project them to an arbitrary dimension, and then apply two Swin Transformer Blocks with multi-head self-attention calculations on all batches; S2-1-3: Reduce the number of batches through batch merging layers to produce a hierarchical feature map representation; S2-1-4: Using the feature map obtained in S2-1-3 as input, repeat the operation of S2-1-3 twice to obtain different feature maps respectively, thereby obtaining feature map representations at different levels; The target positioning and classification training described in S2 includes the following steps: S2-2-1: Initial candidate box generation: In the first stage, the model first uses a conventional rotation box region generation network to generate a series of target candidate boxes, which will serve as input for subsequent stages; S2-2-2: Multi-stage training of cascaded prediction heads: The detectors of three stages are cascaded for prediction. Each stage of the prediction head corresponds to an independent detector. These detectors are trained under different IoU thresholds. As the number of stages increases, the IoU threshold gradually increases. S2-2-3: Step-by-step optimization: The detector in each cascade stage uses the output candidate box of the previous stage for further bounding box regression and classification; S3: Input the augmented input data into the feature extraction network with ReResNet as the backbone to obtain a high-level feature map, and input the high-level feature map into the basic detection and recognition network composed of the rotating box multi-stage progressive prediction head to perform target positioning and classification training to obtain the trained model B; S4: Use model A and model B to build a multi-model integrated prediction framework and output target detection and recognition results; The S4 comprises the following steps: S4-1: Upsample and downsample the validation set images of model A and model B in the inference phase to obtain multi-scale input images with scaling factors of 0.5, 1.0, and 1.5; S4-2: Use model A and model B to infer the validation set and obtain their respective inference results; S4-3: Use WBF to perform decision-level fusion on the inference results, adjust the coordinates and confidence scores based on the inference results of the two models, and obtain the final target recognition result, realizing the complementary advantages of different models.
2. The method for remote sensing image target detection and recognition based on multi-model integration and progressive prediction according to claim 1, characterized in that: The online augmentation described in S1 includes the following steps: S1-1: Use the Mosaic augmentation strategy to augment the training set online: randomly crop the images used for training, and splice the cropped images into one image as training data for training; S1-2: Use the Copy-paste augmentation strategy to augment the training set online: randomly copy the target to the background to generate an augmented image; S1-3: Use the Mixup augmentation strategy to augment the training set online: randomly select two images and add them in proportion.
3. The method for remote sensing image target detection and recognition based on multi-model integration and progressive prediction according to claim 2, characterized in that: The Swin Transformer Block described in S2-1-2 includes a multi-head self-attention module based on a regular window / offset window and a double-layer multi-layer perception module thereafter, a layer normalization is applied before the multi-head self-attention module based on a regular window / offset window and the double-layer multi-layer perception module, and a residual connection auxiliary training is applied after the multi-head self-attention module based on a regular window / offset window and the double-layer multi-layer perception module; In the conventional window-based multi-head self-attention module, The batches are merged into a set of non-overlapping windows, and then each batch is expanded into a one-dimensional vector , as the input of Transformer, self-attention calculation is performed in different windows; In the multi-head self-attention module based on the offset window, the window position is moved on the basis of the window division of the multi-head self-attention module based on the regular window in the previous layer Swin Transformer Block to calculate the self-attention. The calculation process is as follows: (1) In formula (1): Indicates Output features of a Swin Transformer Block based on a regular window / offset window multi-head self-attention module; Indicates The output features of the double-layer multi-layer perception module of the Swin Transformer Block; Represents multi-head self-attention using a multi-head self-attention module with regular windows; Represents the multi-head self-attention module using offset window; Represents the operation of a two-layer multi-layer perception module; Represents the layer normalization operation.
4. The method for remote sensing image target detection and recognition based on multi-model integration and progressive prediction according to claim 1, characterized in that: The ReResNet backbone feature extraction network described in S3 reimplements all layers in the ResNet backbone network based on the rotational equivariant network of e2cnn.
5. The method for remote sensing image target detection and recognition based on multi-model integration and progressive prediction according to claim 4, characterized in that: The S4-2 comprises the following steps: S4-2-1: Use model A to infer the multi-scale input image, and obtain the multi-scale input image inference results with scale factors of 0.5, 1.0 and 1.5 respectively, and reverse sample the inference results to restore them to the original size inference results, and finally obtain image inference results A-1, A-2 and A-3; S4-2-2: Use model B to infer the multi-scale input image, and obtain the multi-scale input image inference results with scale scaling factors of 0.5, 1.0 and 1.5 respectively, and reverse sample the inference results to restore them to the original size inference results, and finally obtain image inference results B-1, B-2 and B-3.
6. The method for remote sensing image target detection and recognition based on multi-model integration and progressive prediction according to claim 5, characterized in that: The S4-3 comprises the following steps: S4-3-1: Create three empty lists: Prediction box list B is used to place the prediction boxes to be processed; The box cluster list L is used to place the set of target prediction boxes; The box fusion list F is used to place the fusion boxes processed based on the predicted boxes; That is, each position in L contains a set of all prediction boxes of the same target, while each position in F contains only one fused prediction box; S4-3-2: Put the multi-scale input image inference results A-1, A-2, A-3, B-1, B-2 and B-3 of different models obtained in S4-2-1 and S4-2-2 into the prediction box list B; S4-3-3: Score the prediction box list B by confidence C Sort in descending order; S4-3-4: Traverse the prediction box list B in order, and for each prediction box B i Find the matching box in the box fusion list F, and define the match as the intersection over union ratio (IoU) > Th between the predicted box and the fused box. IoU is the ratio of the intersection area of the predicted box and the fused box to the union area. Th is the threshold set to determine whether the predicted box and the fused box are the same target. If no matching box is found in the box fusion list F, an empty new set is added to the end of the box cluster list L and the box fusion list F, and the predicted box in the predicted box list B is added to the end of the box cluster list L and the box fusion list F. B i Add to the collection separately; If a matching box is found in the box fusion list F, the predicted box B i Add to the set of corresponding positions in the box cluster list L, and recalculate the coordinates and confidence of the corresponding fused box in the box fusion list F according to all the predicted boxes of the set in the box cluster list L; The box fusion list F obtained after the traversal is the inference result of the target in the final image.
7. The method for remote sensing image target detection and recognition based on multi-model integration and progressive prediction according to claim 6, characterized in that: The confidence score of the fused frame described in S4-3-4 is the average confidence of all frames that make up the confidence score. The coordinates of the fused frame are the weighted sum of the frame coordinates, where the weight is the confidence score of the corresponding frame. The specific calculation formula is as follows: (2) (3) (4) In formulas (2)-(4): Represents the prediction box B i The confidence of the fused box in the corresponding box fusion list F; Represents the prediction box B i The number of all predicted boxes in the corresponding box fusion list L; Indicates the first The confidence of the predicted box; Indicates the first The prediction box The horizontal coordinates of the corner points of the prediction box, ; Indicates the first The prediction box The ordinate of the corner point of the prediction box, ; Indicates the first The horizontal coordinates of the corner points of the prediction box, ; Indicates the first The ordinate of the corner point of the prediction box, .
Citation Information
Patent Citations
Multi-scale feature fusion remote sensing image segmentation method, device, equipment and memory
CN113688813A
Directed target detection method based on semi-supervised learning
CN116452794A