Remote sensing image target detection method based on DMSE-YOLO network

By introducing VSSBlock, CA attention mechanism, MSPConv, LSKA attention mechanism, cross-scale connection and CRC feature fusion methods in the DMSE-YOLO network, the problems of rotation and multi-scale object detection in remote sensing images are solved, and higher detection accuracy and lower computing resource consumption are achieved.

CN119992276AActive Publication Date: 2025-05-13HANGZHOU DIANZI UNIV

Patent Information

Application Number
CN202510474469.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The existing remote sensing image object detection method has the problem of unbalanced detection accuracy when processing rotating or tilting targets, and it is difficult to effectively extract multi-scale features, resulting in poor detection results.

Method used

A remote sensing image object detection method based on DMSE-YOLO network is proposed. By introducing VSSBlock and Coordinate attention (CA) attention mechanisms in the C2f module, the direction and position features are extracted; multi-scale convolution MSPConv and LSKA attention mechanisms are used to form MLBlock, which improves the multi-scale feature extraction capability; and cross-scale connection and Channel ReweightConcat (CRC) feature fusion method are introduced in the neck network, and the OBB detection head is improved to a lightweight MSPOBB.

Benefits of technology

It improves the model's detection robustness and accuracy of rotating objects, improves multi-scale detection capabilities, enhances feature fusion effect, and improves detection accuracy while reducing the amount of calculation and parameter.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992276A_ABST
    Figure CN119992276A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image target detection method based on a DMSE-YOLO network. The method comprises the following steps: S1, preprocessing a data set; s2, configuring a training environment; s3, the structure of the YOLOv8 model is improved, and a DMSE-YOLO network model is obtained; s4, training a network model; and S5, analyzing an experiment result. In order to improve the robustness and accuracy of the model for identifying objects in different orientations, C2fMCA is provided; in order to improve the multi-scale detection capability of the model, multi-scale convolution MSPConv is proposed, and compared with standard convolution, the multi-scale convolution MSPConv can obtain multi-scale features while reducing the calculation amount; an MLBlock module is formed by using the MSPConv and the LSKA; in order to better perform feature fusion, a cross-scale connection is added to the structure of the neck network, then a channel weighted feature fusion mode CRC is introduced, channel information required by model prediction can be selected more finely through weight distribution of channels, and feature fusion of different channels is better realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision technology, and relates to target detection, remote sensing and aerial image analysis, image classification, etc., and is specifically manifested in a remote sensing image target detection method based on a DMSE-YOLO network. Background Art

[0002] As an important means of obtaining surface information, remote sensing technology has been widely used in resource surveys, environmental monitoring, urban planning, disaster assessment, military reconnaissance and other fields. With the continuous improvement of Earth observation satellite imaging capabilities, remote sensing images have made significant progress in spatial resolution, temporal resolution and spectral resolution, and can quickly and comprehensively provide rich information on the Earth's surface. These advantages make remote sensing image target detection a key task in many application scenarios. Its goal is to automatically identify and locate various types of objects in the image, such as aircraft, ships, vehicles and buildings.

[0003] In recent years, the rapid development of deep learning technology has greatly promoted the research progress of remote sensing image target detection. In particular, single-stage detection algorithms represented by YOLO (You Only Look Once) have received widespread attention in practical applications due to their end-to-end training mode and excellent detection speed. The YOLO series of models have high detection accuracy while ensuring real-time performance, and have become one of the mainstream methods in remote sensing image target detection. Although existing methods have improved the detection performance to a certain extent, there are still many challenges in remote sensing images, which seriously restrict the further improvement of detection effects. First, the scale differences of objects of different categories in remote sensing images are significant. Large-scale targets such as football fields and athletic fields usually occupy a large image area, while small-scale targets such as vehicles and ships only occupy a small area, resulting in uneven detection accuracy when the model handles targets of different sizes. In addition, even objects of the same category may have large differences in size, which increases the difficulty of model detection. Secondly, the targets in remote sensing images have the characteristics of directional diversity and irregular arrangement. The traditional horizontal bounding box (Horizontal Bounding Box, HBB) is difficult to accurately surround the rotated or tilted targets, resulting in positioning offset and recognition errors. Although the use of the oriented bounding box (OBB) can well surround the rotating target, how to accurately identify the rotating object is still an urgent problem to be solved. Summary of the invention

[0004] In view of the above problems existing in the prior art, the present invention proposes a remote sensing image target detection method based on the DMSE-YOLO network. The main contents of the method are: (1) In order to solve the problem of poor detection effect of rotating objects, the C2f_MCA module is proposed on the basis of the C2f module. The module introduces the VSSBlock module of the VMamba network to scan the graphics in four directions to obtain directional features, and on this basis, the Coordinate attention (CA) attention mechanism is introduced to obtain accurate position features. (2) In order to improve the multi-scale feature extraction capability of the model, the present invention draws on the idea of ​​fasternet and proposes an efficient multi-scale convolution MSPConv. (3) In order to solve the problem of poor detection effect in multi-scale scenes, MLBlock is formed by combining MSPConv and LSKA attention mechanism, and C2f_ML is formed by using MLBlock and C2f module, thereby further improving the multi-scale detection capability of the model. (4) In order to better perform feature fusion, a more sophisticated feature fusion method CRC is introduced. In addition, the neck network structure is improved, the fusion path of shallow feature maps and deep feature maps is increased, and detailed information is supplemented. The improved neck network is called the feature fusion enhanced neck network. (5) In order to further reduce the amount of calculation and parameters of the model, the OBB detection head of the model is improved using MSPConv to form MSPOBB. The DMSE-YOLO model using MSPOBB is called L-DMSE-YOLO.

[0005] The technical solution adopted by the present invention to solve the technical problem includes the following steps: S1. Preprocessing of the dataset: The present invention carries out training based on the DOTAV1.5 dataset. Since the original image has a high resolution, the image needs to be cropped in advance using the official processing script and then sent to the model for training.

[0006] S2. Configure the training environment: In terms of parameter configuration, it is necessary to combine the computer's video memory capacity and graphics card performance to reasonably adjust key parameters such as input image size, model iteration rounds, and the number of target categories. At the same time, it is necessary to ensure the driver compatibility between the development environment and the hardware device.

[0007] S3. Improve the YOLOv8 model structure to obtain the DMSE-YOLO network model: The present invention improves the C2f module of YOLOv8, uses VSSBlock as a feature extraction module to replace the original Bottleneck; then adds a CA attention mechanism at the end of the C2f module to perform position weighting and finally obtain the output, and this module is called C2f_MCA. C2f_MCA uses VSSBlock to replace Bottleneck, and utilizes the CA attention mechanism to improve the robustness and accuracy of the model in identifying objects in different orientations. Preferably, the C2f_MCA module is used to replace the last two C2f modules of the YOLOv8 backbone.

[0008] The present invention proposes a convolutional MSPConv that is more suitable for multi-scale feature extraction. It divides the input into three branches, performs convolutions of different scales on the first two branches to obtain multi-scale features, and does not perform convolution on the third branch. Then, the MLBlock module is constructed using convolution, MSPConv and LSKA attention mechanisms, and the MLBlock module is used as the feature extraction module in C2f to form C2f_ML. A new feature fusion method CRC is introduced in the YOLOv8 neck network to replace the original concat module, which can use channel weighting to more finely fuse feature maps. Then, the structure of the neck network is improved, and a cross-scale connection is added to supplement the detail information from the shallow feature map. This neck network is called a feature fusion enhanced neck network.

[0009] MSPConv is used to improve the regression branch and classification branch in the OBB detection head, and MSPConv is used to replace the ordinary convolution to achieve a lightweight multi-scale detection head.

[0010] S4. Training network model: Based on the improved DMSE-YOLO model, deploy the model to the configured deep learning framework environment, and import the optimized hyperparameter configuration file into the DMSE-YOLO model. Correctly configure the storage path of the training set and the validation set through the data loading module to ensure that the image data is read normally. After the model training phase is completed, the model needs to be quantitatively evaluated through an independent test set to obtain key performance indicators such as mAP and recall rate.

[0011] S5. Analyze the experimental results: After training, the DMSE-YOLO model will generate a corresponding weight file, and then import these trained weight files, the images to be detected and the corresponding labels. After running the program, the detected data and images will be obtained. Finally, the recognition effect and detection accuracy will be compared to see if they meet the expected requirements.

[0012] Beneficial effects of the present invention: The present invention discloses a remote sensing image target detection method based on a DMSE-YOLO network: (1) In order to improve the robustness and accuracy of the model in identifying objects in different orientations, C2f_MCA is proposed. C2f_MCA can extract the directional features of the object and enhance the focus on the position of the object. (2) In order to improve the multi-scale detection capability of the model, a new multi-scale convolution MSPConv is proposed. Compared with the standard convolution, it can obtain multi-scale features while reducing the amount of calculation. Then, the MLBlock module is constructed using MSPConv and LSKA, and C2f_ML is proposed on this basis. Using C2f_ML instead of C2f can extract rich multi-scale features and global information. (3) In order to better perform feature fusion, the neck network is improved. First, a cross-scale connection is added to the structure to ensure that each time the features are combined in the neck network, there are feature maps from the shallow layer, which optimizes the fusion process of multi-scale semantic information. Secondly, a channel-weighted feature fusion method CRC is introduced. The weight allocation of the channel can more finely select the channel information required for model prediction and better realize the feature fusion of different channels. The improved neck network is called the feature fusion enhanced neck network. (4) In order to reduce the amount of computation and parameters of the model without losing the detection performance as much as possible, the lightweight MSPConv is used to reconstruct the OBB detection head to form a lightweight detection head MSPOBB, which can effectively reduce the consumption of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 : Overall model structure diagram.

[0014] Figure 2 :C2f_MCA module structure diagram.

[0015] Figure 3 : Schematic diagram of MSPConv module.

[0016] Figure 4 :C2f_ML module structure diagram.

[0017] Figure 5 : MSPOBB module structure diagram.

[0018] Figure 6 :DOTA dataset detection effect diagram.

[0019] Figure 7 : DMSE-YOLO detection effect comparison chart. DETAILED DESCRIPTION

[0020] The present invention is further described below in conjunction with the accompanying drawings and specific examples, but it should be noted that the present invention is not limited to the following implementation examples. Those skilled in the art should recognize that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the protection scope of the present invention should not be limited to the contents disclosed in the implementation examples.

[0021] 1. Data acquisition: The first is to obtain remote sensing data sets, and the invention uses the DOTAv1.5 data set. This data set is an improvement and extension of the DOTAv1.0 data set. It uses the same images as DOTAv1.0, but also annotates very small targets (less than 10 pixels). In addition, a new category "container crane" has been added. The images come from different sensors and platforms, including satellite images and drone images, with diverse imaging conditions and scenes. The data set contains a total of 2,806 remote sensing images, including 16 types of targets, including aircraft, ships, storage tanks, baseball fields, tennis courts, basketball courts, athletic fields, ports, bridges, large vehicles, small vehicles, helicopters, roundabouts, football fields, swimming pools, container cranes, a total of 403,318 instances.

[0022] 2. Image Preprocessing The DOTAv1.5 dataset is a benchmark dataset for aerial image target detection. Its original image resolution can reach 20,000×20,000 pixels. Due to the limitation of hardware computing power and the small size and high density of targets in the image, it is significantly difficult to directly input them into model training. For this reason, it is necessary to adopt an image segmentation strategy to cut large-size images into 1024×1024 pixel sub-images and set an overlapping area of ​​200 pixels between adjacent sub-images. This sliding window segmentation strategy can effectively retain the integrity of cross-block targets, thereby reducing the probability of missed detection of small targets. This preprocessing method significantly improves the detection accuracy of dense small targets while improving the efficiency of model training.

[0023] 3. Configuration of YOLOv8 model parameters After processing the data set, the DOTA.yaml file of the data set is created next, and then the paths of train and val and all the category information in this data set are written in the DOTA.yaml file, and then the parameters such as the number of training times and batch-size under tain.py are modified according to the requirements of the invention. The environment of the present invention is: cuda11.8, deep learning framework pytorch1.12.1, Intel core i9-13900ks CPU, 128G memory, GPU is NVIDIA GeForce RTX4090, and the video memory is 24G.

[0024] Step 4: Improve the YOLOv8 model structure to obtain a DMSE-YOLO network model. The structure of the DMSE-YOLO network model is as follows: First is the backbone part, the input goes through ten layers in sequence: convolution, convolution, C2f, convolution, C2f, convolution, C2f_MCA, convolution, C2f_MCA, and SPPF; then comes the neck part, the output of the backbone part goes through twelve layers in the neck: upsampling, CRC, C2f_ML, upsampling, CRC, C2f_ML, convolution, CRC, C2f_ML, convolution, CRC, and C2f_ML. There are many feature map fusions at different stages in the neck part, among which the same as the original YOLOv8 are: the fifth C2f layer of the backbone part serves as an additional input to the fifth CRC layer of the neck, the seventh C2f_MCA layer of the backbone part serves as an additional input to the second CRC layer of the neck, the tenth SPPF layer of the backbone part serves as an additional input to the eleventh CRC layer of the neck, and the third C2f_ML layer of the neck serves as an additional input to the eighth CRC layer of the neck. In addition, this model adds an extra cross-scale connection from the seventh C2f_MCA layer of the trunk to the eighth CRC layer of the neck, thus forming the feature fusion enhanced neck of this model. Finally, the sixth, ninth, and twelfth layers of the feature fusion enhanced neck are respectively input into the OBB detection head to form the head of DMSE-YOLO.

[0025] Compared with the original YOLOv8 model structure, the DMSE-YOLO network model of the present invention has the following improvements: (1) In order to enhance the detection capability of rotating objects, the C2f_MCA module is proposed to extract the direction and position features of the image. Its structure is shown in the figure Figure 2 As shown in the figure. First, the VSSBlock module of VMamba is introduced to extract features from the input. It can scan the input in four different directions to integrate feature information from different directions. Compared with the traditional convolution operation that only perceives locally, VSSBlock aggregates information on a global scale for each pixel during the scanning process, and captures long-distance dependencies by selectively focusing on features in different directions.

[0026] In order to further enhance the feature information extracted by VSSBlock, CA is introduced, which can aggregate the input features into two independent direction-aware feature maps using global pooling operations in both horizontal and vertical directions, and encode them into attention maps containing spatial coordinate information, which are finally multiplied by the input feature map to enhance the position features.

[0027] In the YOLOv8 network, the first few layers are responsible for extracting low-level features, such as edges, textures, etc., while the latter layers process more abstract and high-level semantic information. In this invention, C2f_MCA is only applied to the sixth and eighth layers of YOLOv8 for the following reasons: ① This allows the attention mechanism to focus on more complex and more global context features. The CA attention mechanism can enhance the network's attention to important features, especially in these high-level feature processing, and improve the model's sensitivity to target detection. If the attention mechanism is also introduced in the first few layers, the network may focus on local information too early, interfering with the extraction of basic features.

[0028] ②Replacing all four C2f modules with C2f_MCA may cause the network to overfit. The attention mechanism can improve the performance of the model on the training set, but if too many attention modules are used, the model may pay too much attention to the details in the training data, affecting the generalization performance. By replacing only the last two modules, a certain structural diversity is retained in the model, reducing the risk of overfitting, thereby improving the performance on different datasets.

[0029] ③ Multi-scale feature fusion in YOLOv8 is an important mechanism for detecting small objects. Retaining the first two C2f modules allows low-level features to pass more "cleanly", while the C2f_MCA modules in the last two layers introduce stronger context-awareness to the fused deep features. This combination may be more conducive to the model taking into account both global and local information when dealing with objects of different scales.

[0030] (2) In order to enhance the model's ability to learn multi-scale features, a multi-scale partial convolution (MSPConv) was proposed. In Fasternet, it is proposed that there is often a high degree of similarity between the different numbers of channels in the feature map. Only a part of the channels needs to be convolved, and the rest do not need to be operated. This can effectively reduce the amount of calculation while achieving an effect similar to that of ordinary convolution. The structure diagram of MSPConv is shown in Figure 3 As shown in the figure, MSPConv is divided into three branches. The first two branches use 1 / 4 of the number of channels to extract features of different scales through 3×3 convolution. The third branch does not perform any operation on the remaining 1 / 2 of the number of channels, which can reduce the amount of calculation while still obtaining rich multi-scale information. Finally, the features output by the three branches are spliced, and the spliced ​​feature maps are fused using 1×1 convolution to obtain richer feature representations. In addition, the 3×3 convolution used in MSPConv is a depth-separable convolution. Each convolution kernel only performs convolution on a single channel, which not only improves the computational efficiency of the model but also reduces the extraction of redundant features.

[0031] In addition, the large kernel attention LSKA is introduced, and MLBlock is proposed using MSPConv and LSKA. MLBlock is used to replace Bottleneck in C2f to form C2f_ML to obtain multi-scale features and global information. Its structure is as follows Figure 4 As shown in the figure. Conv is first used to perform channel compression and local feature extraction, and then multi-scale convolution is introduced through MSPConv, so that it can extract features under different receptive fields, thereby obtaining richer feature information. In addition, LSKA is used to obtain a large receptive field and global information, and LSKA is used to process the output of MSPConv, which helps the model to perform more refined selection on these multi-scale features and improves the expressiveness of the model in complex scenes. Specifically, the multi-scale convolution of MSPConv provides rich multi-scale features. LSKA as an attention mechanism can further enhance the expressiveness of these features, and can focus on more important features and suppress noise, thereby improving the expressiveness of features and the robustness of the model. Therefore, by integrating MLBlock into C2f, the generalization ability of the model can be improved, and the overall performance of the model in target detection tasks can be improved.

[0032] (3) In order to better perform feature fusion, I proposed a feature fusion enhanced neck network. Figure 1 As shown in . First, in order to better fuse features at different levels, the feature fusion method of bifpn is used for reference, and a cross-scale connection is added to ensure that every time the features are combined in the neck network, there are feature maps from the shallow layer, which optimizes the fusion process of multi-scale semantic information. Secondly, in order to better utilize the features of different channels, Channel ReweightConcat (CRC) is introduced. When fusing features, CRC assigns a weight to each channel of the input feature map, which avoids the neglect of some important channels and suppresses the channels containing noise. Weight allocation can more finely select the channel information required for model prediction and better realize the feature fusion of different channels. Cross-scale connection and CRC optimize the feature fusion process of different channels at different levels, providing rich and effective information for subsequent multi-scale feature extraction.

[0033] (4) The present invention uses YOLOv8n-OBB as a benchmark. By analyzing the computational complexity of each module of the YOLOv8n-OBB model, it is found that the computational complexity of the detection head accounts for more than 40% of the total computational complexity. Therefore, this paper aims to make the detection head lightweight. YOLOv8 has two detection heads, OBB and Detect. The Detect detector consists of two branches, the regression (Cls.) branch and the classification (Bbox.) branch. The OBB detection head has an additional angle branch compared to the detect detection head, which is used to predict the rotation angle of the anchor box. In YOLOv8n, the computational complexity of the Detect detection head is 2.99GFLOPS, and the computational complexity of the OBB detection head is 3.24GFLOPS, which is only 0.25GFLOPS more than the Detect detection head. This is because the only information that needs to be predicted in the angle branch of the OBB is the rotation angle, so the number of channels it requires is much smaller than that of the regression branch and the classification branch. Therefore, although there is one more branch, the computational complexity increases very little. Due to the above reasons, there is no need to lightweight the angle branch.

[0034] In order to make the OBB detection head lightweight, the OBB detection head is reconstructed using MSPConv, called MSPOBB, as shown in Figure 5 As shown in the figure. From top to bottom, there are classification branch, regression branch and angle branch. The structure of each branch is roughly the same: the first two convolutions are used for feature extraction, the third convolution is used to change the number of channels, and output the information needed by the branch at the end. For both the classification branch and the regression branch, MSPConv is used for feature extraction. The angle branch uses Conv for feature extraction. In MSPConv, only 1 / 2 of the number of channels is operated, which can effectively reduce redundant calculations and memory accesses, and achieve the purpose of lightweight detection head. In addition, through MSPConv, MSPOBB can capture features of different scales, thereby improving the model's multi-scale detection capabilities. The model using MSPOBB is called L-DMSE-YOLO.

[0035] Step 5. Train with the modified model: In the present invention, the data set is divided into a training set, a validation set and a test set according to a ratio of 6:2:2. The number of training rounds is set to 400. Eight pictures are input for each training. During the training process, the training process is observed in real time through wandb. After the training is completed, the trained weights are saved. The following is a description of the effects achieved by this invention in conjunction with the accompanying drawings and data. In order to further test the effects of each module, I performed an ablation test on the DOTAv1.5 data set. The benchmark model YOLOv8n-obb uses the obb detection head officially provided by ultralytics. The experimental results are shown in Table 1. The effectiveness of each model is proved by the ablation experiment. The DMSE-YOLO model proposed in the present invention reduces the number of parameters by 10% and the amount of calculation by 8% compared to the original model, while the mAP value increases by 5.45 percentage points. On the basis of DMSE-YOLO, the use of the MSPOBB detection head will reduce the mAP by 0.2 percentage points, but the number of parameters and the amount of calculation have been greatly reduced, which allows the model to be applied to some devices with severely limited computing resources.

[0036] Table 1 shows the DMSE-YOLO ablation experiment results In order to demonstrate the effect achieved by the invention, Figure 6 Description, attached Figure 6 The first row is the original image, and the second row is the detection result of the DMSE-YOLO model. From the first and third columns, we can see the model's detection ability when the image contains objects of different sizes. From the second and fourth columns, we can see the model's detection ability for objects with different rotations.

[0037] In addition, in order to verify the effectiveness of this method, the present invention compares this model with the mainstream model on the DOTAv1.5 dataset and the DIOR-R dataset. The results of the comparative test are shown in Table 2 and Table 3 respectively. Through the comparative experiment, it can be seen that this method can achieve good detection results in different remote sensing datasets.

[0038] Table 2 shows the comparison test of DOTAv1.5 dataset Method mAP(%) YOLOv8obb 64.45 FCOSR-S 66.37 ReDet 66.86 FR-O 62.00 Mask R-CNN 62.67 HTC 63.40 YOLOv10-obb 65.1 YOLOv11-obb 66.05 DMSE-YOLO(ours) 69.90 L-DMSE-YOLO(ours) 69.76 Table 3 shows the comparison test of DIOR-R dataset Method mAP(%) RetinaNet-O 57.55 Faster RCNN-O 59.54 Gliding Vertex 60.06 RoI Transformer 63.87 LSKNet-S 65.9 Oriented-DETR(Swin-T) 74.26 MAE+MTP 74.54 YOLOv8-obb 75.14 DMSE-YOLO(ours) 77.6 L-DMSE-YOLO(ours) 77.2 In order to verify the effectiveness of this method, the detection effect diagram of the DMSE-YOLO model is compared with the original model in another public remote sensing dataset DIOR-R. Figure 7 The first line above is the original model detection diagram, and the second line is the improved model detection diagram. Figure 7As shown in the first column, the original model mistakenly identified the iron frame in the upper left corner as an airplane, while the improved model did not make any false detections. In the second column, the original model mistakenly identified the wooden path as a port, while the improved model did not make any false detections; in the third column, the original model missed detections and did not recognize vehicles and bridges, while the improved model was able to correctly identify them; in the fourth column, the original model did not detect a chimney, while the improved model was able to detect it correctly. By comparing these pictures, it can be seen that the detection performance of the DMSE-YOLO model in remote sensing images is better than that of the original YOLOv8 model.

Claims

1. A remote sensing image target detection method based on DMSE-YOLO network, characterized in that: The steps include: S1. Preprocessing of dataset; S2. Configure the training environment; S3. Improve the YOLOv8 model structure to obtain the DMSE-YOLO network model; First, the C2f_MCA module is proposed. The C2f_MCA module uses VSSBlock instead of Bottleneck and utilizes the CA attention mechanism. At the same time, the C2f_MCA module is used to replace the last two C2f modules in the YOLOv8 backbone. Secondly, we propose a multi-scale partial convolution (MSPConv), which divides the input into three branches. We perform convolutions of different scales on the first two branches to obtain multi-scale features, and do not perform convolution on the third branch. We then use convolution, MSPConv, and LSKA attention mechanisms to form an MLBlock module, and use the MLBlock module as a feature extraction module in C2f to form C2f_ML. Then, a feature fusion enhanced neck network is proposed, that is, the feature fusion method CRC is introduced into the YOLOv8 neck network to replace the original concat module, and a cross-scale connection is added to supplement the detailed information from the shallow feature map; S4, training network model; S5. Analyze the experimental results.

2. The remote sensing image target detection method based on the DMSE-YOLO network according to claim 1, characterized in that: The data set acquisition is as follows: First, we acquire remote sensing data sets, using the DOTAv1.5 data set; The second is the addition of a new category "container crane".

3. The remote sensing image target detection method based on the DMSE-YOLO network according to claim 1, characterized in that: The preprocessing of the data set is as follows: The image partitioning strategy is adopted to cut the large-size image into sub-images of 1024×1024 pixels, and an overlapping area of ​​200 pixels is set between adjacent sub-images.

4. The remote sensing image target detection method based on the DMSE-YOLO network according to claim 1, characterized in that: The structure of the DMSE-YOLO network model is as follows: First is the backbone part, where the input goes through ten layers in sequence: convolution, convolution, C2f, convolution, C2f, convolution, C2f_MCA, convolution, C2f_MCA, and SPPF; Then comes the neck part. The output of the trunk part goes through upsampling, CRC, C2f_ML, upsampling, CRC, C2f_ML, convolution, CRC, C2f_ML, convolution, CRC, C2f_ML, a total of twelve layers in the neck. There are many feature map fusions at different stages in the neck part. Among them, the same as the original YOLOv8 are: the fifth C2f layer of the trunk part is used as an additional input of the fifth CRC layer of the neck, the seventh C2f_MCA layer of the trunk part is used as an additional input of the second CRC layer of the neck, the tenth SPPF layer of the trunk part is used as an additional input of the eleventh CRC layer of the neck, and the third C2f_ML layer of the neck is used as an additional input of the eighth CRC layer of the neck. Finally, this model adds an additional cross-scale connection from the seventh C2f_MCA layer of the backbone to the eighth CRC layer of the neck, thus forming the feature fusion enhanced neck of this model; finally, the sixth, ninth, and twelfth layers of the feature fusion enhanced neck are respectively input into the OBB detection head to form the head of DMSE-YOLO.

5. The remote sensing image target detection method based on the DMSE-YOLO network according to claim 1 or 4, characterized in that: The structure of multi-scale partial convolution MSPConv is as follows: Multi-scale partial convolution MSPConv is divided into three branches. The first two branches use 1 / 4 of the number of channels to extract features of different scales through 3×3 convolution, and the third branch does not perform any operation on the remaining 1 / 2 of the number of channels. Finally, the features output by the three branches are concatenated, and the concatenated feature maps are fused using 1×1 convolution. The 3×3 convolution used in MSPConv is a depth-wise separable convolution.

6. The remote sensing image target detection method based on the DMSE-YOLO network according to claim 1 or 4, characterized in that: The MLBlock module is proposed by using multi-scale partial convolution MSPConv and large kernel attention LSKA, and the MLBlock module is used to replace the Bottleneck in C2f to form C2f_ML to obtain multi-scale features and global information. The structure of C2f_ML is as follows: First, convolution Conv is used to perform channel compression and local feature extraction; then multi-scale convolution is introduced through multi-scale partial convolution MSPConv, so that it can extract features under different receptive fields; finally, with the help of large kernel attention LSKA, a large receptive field and global information are obtained, and the output of multi-scale partial convolution MSPConv is processed using large kernel attention LSKA.

7. The remote sensing image target detection method based on the DMSE-YOLO network according to claim 1 or 4, characterized in that: The OBB detection head is reconstructed using multi-scale partial convolution MSPConv, called MSPOBB, and its structure is as follows: from top to bottom, there are classification branch, regression branch and angle branch, and the structure of each branch is the same: the first two convolutions are used for feature extraction, and the third convolution is used to change the number of channels and output the information needed by the branch at the end; for both the classification branch and the regression branch, multi-scale partial convolution MSPConv is used for feature extraction; for the angle branch, Conv is used for feature extraction.

Citation Information

Patent Citations

  • Multi-scale feature fusion triple branch network method for multi-organ segmentation

    CN119478404A

  • High-resolution remote sensing image semantic segmentation method based on multi-scale depth supervision

    CN119559403A

  • Prostate MRI image segmentation method based on Mamba-Unet

    CN119579627A

  • Unmanned aerial vehicle aerial photography target detection method based on improved YOLOv8

    CN119832456A

  • Method for classifying images using novel classes

    GB2617440A

Cited By

  • Field-shielded broccoli ball detection method based on improved YOLOv11

    CN121884330A