Remote Sensing Image Target Detection Method Based on DMSE-YOLO Network
By introducing the C2f_MCA module, MSPConv and MLBlock, LSKA attention mechanism and lightweight MSPOBB detection head, the problems of poor detection of rotating objects and unbalanced detection of multi-scale objects in remote sensing images are solved, improving detection accuracy and reducing computing resource consumption.
Patent Information
- Application Number
- CN202510474469.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-16
AI Technical Summary
In remote sensing images, there are problems such as poor detection effect of rotating objects, uneven detection accuracy of multi-scale object, and the difficulty of traditional bounding boxes to accurately surround rotating or tilting targets.
The C2f_MCA module was introduced to extract directional features and enhance position features. MSPConv was used to improve multi-scale feature extraction capabilities, combined with MLBlock and LSKA attention mechanism, improved neck network structure to enhance feature fusion, and used a lightweight MSPOBB detection head.
It improves the robustness and accuracy of remote sensing image object detection, reduces computing resource consumption, and is suitable for devices with limited computing resources.
Smart Images

Figure CN119992276B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and relates to object detection, remote sensing and aerial image analysis, image classification, etc., and specifically presents a method for object detection in remote sensing images based on the DMSE-YOLO network. Background Art
[0002] As an important means of obtaining surface information, remote sensing technology has been widely applied in fields such as resource investigation, environmental monitoring, urban planning, disaster assessment, and military reconnaissance. With the continuous improvement of the imaging capabilities of earth observation satellites, remote sensing images have made remarkable progress in terms of spatial resolution, temporal resolution, and spectral resolution, and can quickly and comprehensively provide rich information about the earth's surface. These advantages make object detection in remote sensing images a key task in many application scenarios, and its goal is to automatically identify and locate various ground objects in the image, such as airplanes, ships, vehicles, and buildings.
[0003] In recent years, the rapid development of deep learning technology has greatly promoted the research progress of object detection in remote sensing images. Especially the single-stage detection algorithms represented by YOLO (You Only Look Once), with their end-to-end training mode and excellent detection speed, have received extensive attention in practical applications. The YOLO series of models have high detection accuracy while ensuring real-time performance, and have become one of the mainstream methods in object detection in remote sensing images.
[0004] Although the existing methods have improved the detection performance to a certain extent, there are still many challenges in remote sensing images, which seriously restrict the further improvement of the detection effect. First, the scale differences of different types of objects in remote sensing images are significant. Large-scale objects such as football fields and athletic fields usually occupy a large area of the image, while small-scale objects such as vehicles and ships only occupy a very small area, resulting in uneven detection accuracy when the model processes objects of different sizes. In addition, even for objects of the same category, their sizes may also show large differences, increasing the detection difficulty of the model. Second, the objects in remote sensing images have the characteristics of diverse directions and irregular arrangements. Traditional horizontal bounding boxes (Horizontal Bounding Box, HBB) are difficult to accurately enclose rotated or tilted objects, resulting in positioning offsets and recognition errors. Although using oriented bounding boxes (Oriented Bounding Box, OBB) can well enclose rotated objects, how to accurately identify rotated objects is still an urgent problem to be solved. Summary of the Invention
[0005] In view of the above problems existing in the prior art, the present invention proposes a remote sensing image target detection method based on the DMSE-YOLO network. The main content of this method is as follows: (1) To solve the problem of poor detection effect of rotating objects, the C2f_MCA module is proposed based on the C2f module. This module introduces the VSSBlock module of the VMamba network to scan the graph in four directions to obtain direction features, and on this basis, the Coordinate attention (CA) attention mechanism is introduced to obtain accurate position features. (2) To improve the multi-scale feature extraction ability of the model, the present invention draws on the idea of fasternet and proposes an efficient multi-scale convolution MSPConv. (3) To solve the problem of poor detection effect in multi-scale scenarios, the MSPConv and the LSKA attention mechanism are combined to form the MLBlock, and the MLBlock and the C2f module are used to form the C2f_ML, thereby further improving the multi-scale detection ability of the model. (4) To better perform feature fusion, a more refined feature fusion method CRC is introduced. In addition, the neck network structure is improved to increase the fusion path between the shallow feature map and the deep feature map to supplement detailed information. The improved neck network is called the feature fusion enhanced neck network. (5) To further reduce the computational amount and the number of parameters of the model, the OBB detection head of the model is improved using MSPConv to form MSPOBB, and the DMSE-YOLO model using MSPOBB is called L-DMSE-YOLO.
[0006] The technical solution adopted by the present invention to solve its technical problems includes the following steps:
[0007] S1. Preprocessing of the data set:
[0008] The present invention conducts training based on the DOTAV1.5 data set. Since the resolution of the original image is relatively high, it is necessary to preprocess the image through the processing script provided by the official for cropping operations and then send it into the model for training.
[0009] S2. Configuration of the training environment:
[0010] In terms of parameter configuration, it is necessary to reasonably adjust key parameters such as the input image size, the number of model iteration rounds, and the number of target categories in combination with the computer video memory capacity and the graphics card performance. At the same time, it is necessary to ensure the driver compatibility between the development environment and the hardware device.
[0011] S3. Improve the YOLOv8 model structure to obtain the DMSE-YOLO network model:
[0012] The present invention improves the C2f module of YOLOv8, uses the VSSBlock as the feature extraction module to replace the original Bottleneck; then adds a CA attention mechanism at the end of the C2f module for position weighting to finally obtain the output, and this module is called C2f_MCA. C2f_MCA uses the VSSBlock to replace the Bottleneck and utilizes the CA attention mechanism, improving the robustness and accuracy of the model in identifying objects with different orientations. Preferably, the last two C2f modules in the backbone part of YOLOv8 are replaced with the C2f_MCA module.
[0013] The present invention proposes a convolutional MSPConv more suitable for multi-scale feature extraction. It divides the input into three branches, performs convolutions of different scales on the first two branches respectively to obtain multi-scale features, and does not perform convolution operation on the third branch. Then, a convolutional layer, MSPConv, and LSKA attention mechanism are used to form the MLBlock module, and the MLBlock module is used as the feature extraction module in C2f to form C2f_ML. A new feature fusion method CRC is introduced in the neck network of YOLOv8 to replace the original concat module, which can use channel weighting to fuse the feature maps more precisely. Then, the structure of the neck network is improved by adding a cross-scale connection to supplement the detailed information from the shallow feature maps. This neck network is called the feature fusion enhanced neck network.
[0014] Use MSPConv to improve the regression branch and classification branch in the OBB detection head, and replace the ordinary convolution in them with MSPConv to achieve a lightweight multi-scale detection head.
[0015] S4. Train the network model:
[0016] Based on the improved DMSE-YOLO model, deploy the model to the configured deep learning framework environment, and import the optimized hyperparameter configuration file into the DMSE-YOLO model. Correctly configure the storage paths of the training set and validation set through the data loading module to ensure the normal reading of image data. After the model training stage, the model needs to be quantitatively evaluated through an independent test set to obtain key performance indicators such as mAP and recall rate.
[0017] S5. Analyze the experimental results:
[0018] The DMSE-YOLO model will generate corresponding weight files after training. Then, import these trained weight files, the images to be detected, and the corresponding labels. After running the program, the detected data and images are obtained; finally, it will be compared whether the recognition effect and detection accuracy meet the expected requirements.
[0019] Advantages of the present invention:
[0020] The present invention discloses a remote sensing image target detection method based on the DMSE-YOLO network: In this research, (1) to improve the robustness and accuracy of the model in identifying objects in different orientations, C2f_MCA is proposed. C2f_MCA can extract the direction features of objects and enhance the attention to the object positions. (2) To enhance the multi-scale detection ability of the model, a new multi-scale convolution MSPConv is proposed. Compared with the standard convolution, it can obtain multi-scale features while reducing the computational cost. Then, MSPConv and LSKA are used to form the MLBlock module, and based on this, C2f_ML is proposed. Using C2f_ML instead of C2f can extract rich multi-scale features and global information. (3) To better perform feature fusion, the neck network is improved. First, a cross-scale connection is added to the structure to ensure that there is a feature map from the shallow layer every time feature combination occurs in the neck network, optimizing the fusion process of multi-scale semantic information. Secondly, a channel-weighted feature fusion method CRC is introduced. The weight assignment for channels can more finely select the channel information required for model prediction and better achieve feature fusion of different channels. The improved neck network is called the neck network with enhanced feature fusion. (4) To reduce the computational cost and the number of parameters of the model while minimizing the loss of detection performance, lightweight MSPConv is used to reconstruct the OBB detection head, forming the lightweight detection head MSPOBB, which can effectively reduce the consumption of computing resources. Description of the drawings
[0021] Figure 1 : Overall model structure diagram.
[0022] Figure 2 : Structure diagram of the C2f_MCA module.
[0023] Figure 3 : Schematic diagram of the MSPConv module.
[0024] Figure 4 : Structure diagram of the C2f_ML module.
[0025] Figure 5 : Structure diagram of the MSPOBB module.
[0026] Figure 6 : Detection effect diagram of the DOTA dataset.
[0027] Figure 7 : Comparison diagram of the detection effects of DMSE-YOLO. Detailed implementation manners
[0028] The present invention will be further described below in conjunction with the accompanying drawings and specific cases. It should be noted, however, that the present invention is not limited to the following embodiments. Those skilled in the art should recognize that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the protection scope of the present invention should not be limited to the content disclosed in the embodiments.
[0029] 1. Data acquisition:
[0030] First, the acquisition of the remote sensing dataset is carried out. The invention uses the DOTAv1.5 dataset. This dataset is an improvement and extension of the DOTAv1.0 dataset. It uses the same images as DOTAv1.0, but also annotates extremely small targets (less than 10 pixels). In addition, a new category "container crane" is added. The images come from different sensors and platforms, including satellite images and UAV images, with diverse imaging conditions and scenarios. The dataset contains a total of 2,806 remote sensing images, including 16 types of targets, including airplanes, ships, storage tanks, baseball fields, tennis courts, basketball courts, athletic fields, ports, bridges, large vehicles, small vehicles, helicopters, roundabouts, football fields, swimming pools, and container cranes, with a total of 403,318 instances.
[0031] 2. Image preprocessing
[0032] As a benchmark dataset for aerial image object detection, the original image resolution of the DOTAv1.5 dataset can reach the order of 20,000×20,000 pixels. Due to hardware computing power limitations and the characteristics of small size and high density of the targets in the images, there are significant difficulties in directly inputting them into model training. For this reason, an image tiling strategy is adopted to cut the large-size image into sub-images of 1024×1024 pixels and set an overlapping area of 200 pixels between adjacent sub-images. This sliding window segmentation strategy can effectively retain the integrity of cross-block targets, thereby reducing the probability of small target missed detection. This preprocessing method improves the model training efficiency while significantly improving the detection accuracy of dense small targets.
[0033] 3. Configuration of YOLOv8 model parameters
[0034] After processing the dataset, next, create the DOTA.yaml file for the dataset, and then write the paths of train and val as well as all the class information in this dataset under the DOTA.yaml file. Then, modify parameters such as the number of training times and batch-size in train.py according to the requirements of the invention. The environment of this invention is: cuda11.8, deep learning framework pytorch1.12.1, Intel core i9-13900ks CPU, 128G memory, GPU is NVIDIA GeForce RTX4090, and the video memory is 24G.
[0035] Step 4: Improve the YOLOv8 model structure to obtain the DMSE-YOLO network model. The structure of the DMSE-YOLO network model is as follows:
[0036] First is the backbone part. The input passes through a total of ten layers in sequence: convolution, convolution, C2f, convolution, C2f, convolution, C2f_MCA, convolution, C2f_MCA, and SPPF. Then is the neck part. The output of the backbone part passes through a total of twelve layers in the neck in sequence: upsampling, CRC, C2f_ML, upsampling, CRC, C2f_ML, convolution, CRC, C2f_ML, convolution, CRC, C2f_ML. There are many feature map fusions at different stages in the neck part. Among them, the same as the original YOLOv8 are: the C2f layer of the fifth layer of the backbone part is used as an additional input to the CRC layer of the fifth layer of the neck, the C2f_MCA layer of the seventh layer of the backbone part is used as an additional input to the CRC layer of the second layer of the neck, the SPPF layer of the tenth layer of the backbone part is used as an additional input to the CRC layer of the eleventh layer of the neck, and the C2f_ML layer of the third layer of the neck is used as an additional input to the CRC layer of the eighth layer of the neck. In addition, this model adds an extra cross-scale connection from the C2f_MCA layer of the seventh layer of the backbone part to the CRC layer of the eighth layer of the neck, thus constituting the feature fusion enhanced neck of this model. Finally, the sixth, ninth, and twelfth layers of the feature fusion enhanced neck are respectively input into the OBB detection head to form the head of DMSE-YOLO.
[0037] Compared with the original YOLOv8 model structure, the improvement points of the DMSE-YOLO network model of this invention are as follows:
[0038] (1) In order to enhance the detection ability of rotated objects, the C2f_MCA module is proposed to extract the direction and position features of the image. Its structure diagram is as Figure 2As shown below. First, the VSSBlock module of VMamba is introduced to extract features from the input. It can scan the input in four different directions, thereby integrating feature information from different directions. Compared with traditional convolutional operations that only have local perception, VSSBlock aggregates information globally for each pixel point during the scanning process. By selectively focusing on features in different directions, it can capture long-range dependencies.
[0039] To further enhance the feature information extracted by VSSBlock, CA is introduced. It can aggregate the input features into two independent direction-aware feature maps using global pooling operations in the horizontal and vertical directions, encode them into an attention map containing spatial coordinate information, and finally multiply it by the input feature map to enhance the position features.
[0040] In the YOLOv8 network, the first few layers are responsible for extracting low-level features such as edges and textures, while the latter layers process more abstract and high-level semantic information. In this invention, C2f_MCA is only applied to the sixth and eighth layers of YOLOv8 for the following reasons:
[0041] ① This can make the attention mechanism focus on more complex and globally contextual features. The CA attention mechanism can enhance the network's attention to important features, especially in the processing of these high-level features, improving the model's sensitivity to object detection. If the attention mechanism is introduced in the first few layers, it may cause the network to focus on local information prematurely and interfere with the extraction of basic features.
[0042] ② Replacing all four C2f modules with C2f_MCA may lead to overfitting of the network. The attention mechanism can improve the model's performance on the training set, but if too many attention modules are used, the model may focus too much on the details in the training data, affecting the generalization performance. By only replacing the latter two modules, a certain structural diversity is retained in the model, reducing the risk of overfitting, thereby improving the performance on different datasets.
[0043] ③ Multi-scale feature fusion in YOLOv8 is an important mechanism for detecting small targets. Keeping the first two C2f modules allows low-level features to pass through relatively "cleanly", while the C2f_MCA modules in the latter two layers introduce stronger context awareness for the fused deep features. This combination may be more beneficial for the model to balance global and local information when dealing with targets of different scales.
[0044] (2) To enhance the model's ability to learn multi-scale features, multi-scale partial convolution (MSPConv) was proposed. As proposed in Fasternet, there is often a high degree of similarity between different numbers of channels in the feature map. Only a part of the channels need to be convolved, and the rest do not perform any operations, which can effectively reduce the computational cost while achieving an effect similar to that of ordinary convolution. The structural diagram of MSPConv is as shown in Figure 3 . MSPConv is divided into three branches. The first two branches use 1 / 4 of the number of channels respectively to perform feature extraction at different scales through 3×3 convolution. The third branch does not perform any operations on the remaining 1 / 2 of the number of channels, and can still obtain rich multi-scale information while reducing the computational cost. Finally, the features output by the three branches are concatenated, and the concatenated feature map is used for feature fusion through 1×1 convolution, so as to obtain a more rich feature representation. In addition, the 3×3 convolution used in MSPConv is a depthwise separable convolution. Each convolution kernel only convolves a single channel, which not only improves the operation efficiency of the model but also reduces the extraction of redundant features.
[0045] In addition, large kernel attention LSKA was introduced. Using MSPConv and LSKA, MLBlock was proposed, and Bottleneck in C2f was replaced with MLBlock to form C2f_ML to obtain multi-scale features and global information. Its structure is as shown in Figure 4 . First, Conv is used for channel compression and local feature extraction, and then multi-scale convolution is introduced through MSPConv, enabling it to perform feature extraction under different receptive fields, so as to obtain richer feature information. In addition, with the help of LSKA, a large receptive field and global information are obtained, and LSKA is used to process the output of MSPConv to help the model perform more refined selection on these multi-scale features, improving the model's performance in complex scenarios. Specifically, the multi-scale convolution of MSPConv provides rich multi-scale features. As an attention mechanism, LSKA can further enhance the expressive ability of these features, and can focus on more important features and suppress noise, improving the expressiveness of features and the robustness of the model. Therefore, by integrating MLBlock into C2f, the generalization ability of the model can be improved, and the overall performance of the model in the object detection task can be improved.
[0046] (3) To better perform feature fusion, I proposed a feature fusion enhanced neck network. As shown in Figure 1As shown in []. First, in order to better fuse features at different levels, the feature fusion method of BIFPN is borrowed, and a cross-scale connection is added to ensure that there are feature maps from the shallow layer every time feature combination occurs in the neck network, optimizing the fusion process of multi-scale semantic information. Second, in order to better utilize the features of different channels, Channel ReweightConcat (CRC) is introduced. When fusing features, CRC assigns a weight to each channel of the input feature map, which avoids some important channels being ignored and suppresses channels containing noise at the same time. The weight assignment can more finely select the channel information required for model prediction and better achieve the feature fusion of different channels. The cross-scale connection and CRC optimize the feature fusion process of different levels and different channels, providing rich and effective information for subsequent multi-scale feature extraction.
[0047] (4) This invention uses YOLOv8n-OBB as the benchmark. By analyzing the computational complexity of each module of the YOLOv8n-OBB model, it is found that the computational complexity of the detection head accounts for more than 40% of the total computational complexity. Therefore, the lightweighting of the detection head is the goal of this paper. YOLOv8 has two detection heads, OBB and Detect. The Detect detector consists of two branches, the regression (Cls.) branch and the classification (Bbox.) branch. The OBB detection head has one more angle branch than the Detect detection head to predict the rotation angle of the anchor box. In YOLOv8n, the computational complexity of the Detect detection head is 2.99 GFLOPS, and the computational complexity of the OBB detection head is 3.24 GFLOPS, only 0.25 GFLOPS more than the Detect detection head. This is because there is only one piece of information to be predicted, i.e., the rotation angle, in the angle branch of OBB. Therefore, the number of channels it requires is much smaller than those of the regression branch and the classification branch. So, although there is one more branch, the increase in computational complexity is very small. Due to the above reasons, there is no need to lightweight the angle branch.
[0048] To lightweight the OBB detection head, MSPConv is used to reconstruct the OBB detection head, which is called MSPOBB, as Figure 5As shown in the figure. From top to bottom are the classification branch, the regression branch, and the angle branch. The structure of each branch is roughly the same: the first two convolutions are used for feature extraction, and the third convolution is used to change the number of channels and output the information required by the branch at the end. For the classification branch and the regression branch, MSPConv is used for feature extraction. The angle branch uses Conv for feature extraction. Only half of the channels are operated on in MSPConv, which can effectively reduce redundant calculations and memory access, achieving the purpose of a lightweight detection head. In addition, through MSPConv, MSPOBB can capture features at different scales, thereby improving the model's multi-scale detection ability. The model adopting MSPOBB is called L-DMSE-YOLO.
[0049] Step 5. Train with the modified model:
[0050] In the present invention, the dataset is divided into a training set, a validation set, and a test set in a ratio of 6:2:2. The number of training rounds is set to 400, and 8 pictures are input for each training. During the training process, the training process is observed in real time through wandb. After the training is completed, the trained weights are saved. The following describes the effects achieved by the present invention in combination with the drawings and data. In order to further test the effects of each module, I conducted ablation experiments on the DOTAv1.5 dataset. The baseline model YOLOv8n-obb uses the obb detection head provided by the official ultralytics. The experimental results are shown in Table 1. The ablation experiment proves the effectiveness of each model. The DMSE-YOLO model proposed in the present invention reduces the number of parameters by 10% and the computational amount by 8% compared with the original model, while the value of mAP increases by 5.45 percentage points. On the basis of DMSE-YOLO, although using the MSPOBB detection head will cause the mAP to drop by 0.2 percentage points, the number of parameters and the computational amount are greatly reduced again, which enables the model to be applied to some devices with severely limited computing resources.
[0051] Table 1 shows the ablation experiment results of DMSE-YOLO
[0052]
[0053] To demonstrate the effects achieved by the invention, in combination with the attached Figure 6 explanation, in the attachment Figure 6 The first row is the original picture, and the second row is the detection result of the DMSE-YOLO model. From the first column and the third column, it can be seen the model's detection ability for objects of different scales in the picture. From the second column and the fourth column, it can be seen the model's detection ability for different rotated objects.
[0054] In addition, to verify the effectiveness of this method, the present invention compared this model with mainstream models on the DOTA v1.5 dataset and the DIOR-R dataset. The results of the comparative experiments are shown in Table 2 and Table 3 respectively. It can be seen from the comparative experiments that this method can achieve good detection results in different remote sensing datasets.
[0055] Table 2 is the comparative experiment on the DOTA v1.5 dataset
[0056] Method mAP(%) YOLOv8obb 64.45 FCOSR-S 66.37 ReDet 66.86 FR-O 62.00 Mask R-CNN 62.67 HTC 63.40 YOLOv10-obb 65.1 YOLOv11-obb 66.05 DMSE-YOLO(ours) 69.90 L-DMSE-YOLO(ours) 69.76
[0057] Table 3 is the comparative experiment on the DIOR-R dataset
[0058] Method mAP(%) RetinaNet-O 57.55 Faster RCNN-O 59.54 Gliding Vertex 60.06 RoI Transformer 63.87 LSKNet-S 65.9 Oriented-DETR(Swin-T) 74.26 MAE+MTP 74.54 YOLOv8-obb 75.14 DMSE-YOLO(ours) 77.6 L-DMSE-YOLO(ours) 77.2
[0059] To verify the effectiveness of this method, in another publicly available remote sensing dataset, DIOR-R, the detection effect diagrams of the DMSE-YOLO model and the original model were compared, as shown in the attached Figure 7 description. The upper row in the first column is the detection diagram of the original model, and the second row is the detection diagram of the improved model. As shown in the Figure 7 first column, the original model misidentified the iron frame in the upper left corner as an airplane, while the improved model did not have a misdetection. In the second column, the original model misidentified the wooden path as a port, while the improved model did not have a misdetection; in the third column, the original model had a missed detection and did not identify the vehicle and the bridge, while the improved model could correctly identify them; in the fourth column, the original model did not detect a chimney, while the improved model could correctly detect it. Through the comparison of these diagrams, it can be seen that the detection performance of the DMSE-YOLO model in remote sensing images is better than that of the original YOLOv8 model.
Claims
1. A remote sensing image target detection method based on the DMSE-YOLO network, characterized in that, It includes the following steps: S1. Preprocessing of the dataset; S2. Configuring the training environment; S3. Improving the YOLOv8 model structure to obtain the DMSE-YOLO network model; Firstly, the C2f_MCA module is proposed. The C2f_MCA module uses the VSSBlock to replace the Bottleneck and utilizes the CA attention mechanism; meanwhile, the last two C2f modules in the backbone part of YOLOv8 are replaced with the C2f_MCA module; Secondly, the multi-scale partial convolution MSPConv is proposed. It divides the input into three branches, performs convolutions of different scales on the first two branches respectively to obtain multi-scale features, and does not perform convolution operation on the third branch; then, a convolutional layer, MSPConv, and LSKA attention mechanism are used to form the MLBlock module, and the MLBlock module is used as the feature extraction module in C2f to form C2f_ML; Then, a feature fusion enhanced neck network is proposed, that is, a feature fusion method CRC is introduced in the YOLOv8 neck network to replace the original concat module, and a cross-scale connection is added to supplement the detailed information from the shallow feature map; S4. Training the network model; S5. Analyzing the experimental results.
2. The remote sensing image target detection method based on the DMSE-YOLO network according to claim 1, characterized in that, The acquisition of the dataset is specifically as follows: Firstly, the remote sensing dataset is acquired, and the DOTAv1.5 dataset is used; Secondly, a new category "container crane" is added.
3. The remote sensing image target detection method based on the DMSE-YOLO network according to claim 1, characterized in that, The preprocessing of the dataset is specifically as follows: An image tiling strategy is adopted to cut the large-size image into sub-images of 1024×1024 pixels, and an overlapping area of 200 pixels is set between adjacent sub-images.
4. The remote sensing image target detection method based on the DMSE-YOLO network according to claim 1, characterized in that, The structure of the DMSE-YOLO network model is as follows: Firstly, it is the backbone part. The input passes through a convolutional layer, a convolutional layer, C2f, a convolutional layer, C2f, a convolutional layer, C2f_MCA, a convolutional layer, C2f_MCA, and SPPF in sequence, a total of ten layers; Then, it is the neck part. The output of the backbone part passes through upsampling, CRC, C2f_ML, upsampling, CRC, C2f_ML, a convolutional layer, CRC, C2f_ML, a convolutional layer, CRC, C2f_ML in sequence in the neck, a total of twelve layers. There are many feature map fusions at different stages in the neck part. Among them, the same as the original YOLOv8 are: the C2f layer of the fifth layer of the backbone part is used as an additional input to the CRC layer of the fifth layer of the neck, the C2f_MCA layer of the seventh layer of the backbone part is used as an additional input to the CRC layer of the second layer of the neck, the SPPF layer of the tenth layer of the backbone part is used as an additional input to the CRC layer of the eleventh layer of the neck, and the C2f_ML layer of the third layer of the neck is used as an additional input to the CRC layer of the eighth layer of the neck; Finally, a cross-scale connection is additionally added to this model, from the C2f_MCA layer of the seventh layer of the backbone part to the CRC layer of the eighth layer of the neck, thus forming the feature fusion enhanced neck of this model; finally, the sixth layer, the ninth layer, and the twelfth layer of the feature fusion enhanced neck are respectively input into the OBB detection head to form the head of DMSE-YOLO.
5. The remote sensing image target detection method based on the DMSE-YOLO network according to claim 1 or 4, characterized in that The structure of the multi-scale partial convolution MSPConv is specifically as follows: The multi-scale partial convolution (MSPConv) is divided into three branches. The first two branches use 1 / 4 of the number of channels to perform feature extraction at different scales through 3×3 convolutions, and the third branch leaves the remaining 1 / 2 of the channels without any operation. Finally, the features output by the three branches are concatenated, and the concatenated feature map is used for feature fusion through 1×1 convolution. The 3×3 convolution used in MSPConv is a depthwise separable convolution.
6. The remote sensing image target detection method based on the DMSE-YOLO network according to claim 1 or 4, characterized in that The MLBlock module is proposed using the multi-scale partial convolution (MSPConv) and the large kernel attention (LSKA), and the Bottleneck in C2f is replaced with the MLBlock module to form C2f_ML for obtaining multi-scale features and global information. The structure of C2f_ML is as follows: First, the convolution (Conv) is used for channel compression and local feature extraction. Then, the multi-scale partial convolution (MSPConv) is introduced to enable feature extraction under different receptive fields. Finally, the large kernel attention (LSKA) is used to obtain a large receptive field and global information, and the output of the multi-scale partial convolution (MSPConv) is processed using the large kernel attention (LSKA).
7. The remote sensing image target detection method based on the DMSE-YOLO network according to claim 1 or 4, characterized in that, The multi-scale partial convolution (MSPConv) is used to reconstruct the OBB detection head, called MSPOBB. Its structure is as follows: from top to bottom are the classification branch, the regression branch, and the angle branch. The structure of each branch is the same: the first two convolutions are used for feature extraction, and the third convolution is used to change the number of channels and output the information required at the end of the branch. For the classification branch and the regression branch, the multi-scale partial convolution (MSPConv) is used for feature extraction. For the angle branch, the Conv is used for feature extraction.
Citation Information
Patent Citations
High-resolution remote sensing image semantic segmentation method based on multi-scale depth supervision
CN119559403A
Method for classifying images using novel classes
GB2617440A