Optical remote sensing image small target detection method based on improved YOLOv5
By adding a small object detection layer in the YOLOv5 model, introducing feature enhancement module MFFM and fusion attention mechanism module CASA, and using MPDIoU loss function, the problem of low detection performance of small object in remote sensing images is solved, and higher detection accuracy and lower error detection rate are achieved.
Patent Information
- Application Number
- CN202510276753.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-11
AI Technical Summary
The existing small object detection methods have problems with low detection performance and serious missed detection in remote sensing images, especially in remote sensing images with complex backgrounds and high spatial resolution, which are difficult to effectively extract features and distinguish small objects.
Based on YOLOv5, a small object detection layer is added, a feature enhancement module MFFM and a fusion attention mechanism module CASA are introduced, and an MPDIoU loss function is used to form a FMCM-YOLO model to improve the detection performance of small objects in remote sensing images.
The detection performance of small targets in remote sensing images is improved, the probability of missed detection and false detection is reduced, and the detection accuracy and generalization ability of the model are improved.
Smart Images

Figure CN120298658A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a small target detection method for optical remote sensing images based on improved YOLOv5, belonging to the field of image processing. Background Art
[0002] Remote sensing generally refers to all non-contact remote detections. Generally, it refers to the technology of observing objects on the earth's surface using remote sensing detection sensors on artificial earth satellites, space shuttles or other aircraft. Classifying remote sensing technology according to the spectral range of electromagnetic waves can be divided into visible light remote sensing, infrared remote sensing, microwave remote sensing, etc. When the working band of the remote sensing detection sensor is in the visible light band, the image collected from the earth's surface is called an optical remote sensing image. The optical remote sensing image contains rich detail information of the object, can intuitively present physical information such as the texture structure and position of the object, and is conducive to people's observation and analysis of the object. The present invention mainly focuses on optical remote sensing images, hereinafter simply referred to as remote sensing images.
[0003] With the continuous development and progress of remote sensing technology, the information provided by remote sensing images is becoming increasingly rich, which also puts forward higher requirements for remote sensing image processing technology, especially for the detection of small targets in remote sensing images. Small targets refer to target objects that are relatively small, have low contrast and lack details in images or videos. Since the information contained in small targets in remote sensing images is limited and they are easily affected by the background, the detection and classification accuracy of small targets in remote sensing images is usually low.
[0004] Early small target detection algorithms were mainly based on manually designed features, combined with traditional classifiers such as support vector machines to detect and classify targets. It mainly includes three steps: First, locate the target, extract candidate boxes, and traverse the graph by means of a sliding window; secondly, based on methods of features such as color, shape, and texture information, extract basic features; finally, use a classifier to classify small targets in the image to complete the detection and classification of the target. The extraction of candidate boxes requires generating a large number of sliding windows, which will bring additional calculations, and the manually designed features mainly rely on shallow features such as color and shape for extraction. The feature expression ability is weak and the robustness is low, which is not suitable for complex and variable environments and cannot meet the actual application requirements.
[0005] With the rapid development of deep learning, small object detection methods have also made further breakthroughs. Currently, small object detection methods based on deep learning are divided into two-stage detection methods and one-stage detection methods. Two-stage detection methods first generate candidate boxes based on the positions of objects, and then locate and classify the candidate boxes. Representative algorithms include Fast-CNN, Faster-CNN, Mask-CNN, etc. One-stage detection methods directly generate denser candidate boxes and directly classify the targets when generating candidate boxes, making the detection more efficient. Representative algorithms include SDD (Single Shot MultiBox Detector) and the YOLO (you only look once) series. Among them, YOLOv5 in the YOLO series is widely used in small object detection, aiming to obtain higher detection performance and faster detection speed.
[0006] However, the above methods are mainly designed for small object detection in natural images. However, when directly applied to small objects in remote sensing images, the detection performance will be greatly reduced. This is because there are certain differences between remote sensing images and natural images. Remote sensing images contain various complex terrains and landforms and have a high spatial resolution, resulting in more complex background information in remote sensing images, including more target categories and specific details. In addition, the position where remote sensing images are taken is far from the ground. Therefore, compared with traditional natural images, the target pixels in remote sensing images are smaller, the proportion of small objects is larger, and higher requirements are placed on the detection performance of the model. Summary of the Invention
[0007] The purpose of the present invention is to provide a small object detection method for optical remote sensing images based on improved YOLOv5, aiming to solve the technical problems of low detection performance and serious false detection and missed detection problems existing in existing small object detection methods when facing difficulties in feature extraction, foreground and background confusion, and large target scale changes in remote sensing images.
[0008] To achieve the above purpose, the technical solution adopted by the present invention is: a small object detection method for optical remote sensing images based on improved YOLOv5, including the following steps:
[0009] Step1: Select an optical remote sensing image dataset and divide it into a training set and a validation set;
[0010] Step2: Taking YOLOv5 as the baseline algorithm, add a small object detection layer and corresponding detection heads, and transform the model into a four-head detection model to improve the detection performance of small objects;
[0011] Step3: Introduce a feature enhancement module MFFM into the Neck network of YOLOv5 to improve the ability to extract features;
[0012] Step 4: Introduce the fusion attention mechanism module CASA into the neck network Neck of YOLOv5 to improve the sensitivity to small targets;
[0013] Step 5: Introduce the MPDIoU loss function to obtain the improved detection model FMCM-YOLO;
[0014] Step 6: Use the training set to train the FMCM-YOLO model to obtain the training weights, and use the validation set to validate the model, ultimately achieving the detection of small targets in optical remote sensing images.
[0015] The Step 2 is specifically as follows:
[0016] YOLOv5 is divided into three parts: Backbone, Neck and Head. The Backbone part is composed of several CBS modules, C3 modules and SPPF modules; the Neck part is composed of several CBS modules, C3 modules, Upsample modules and Concat modules; the Head part is composed of three Detect modules. The CBS module is used to downsample the feature map and extract the feature information; the C3 module improves the model's ability to extract features by increasing the number of model layers and increasing the network's receptive field; the SPPF module performs pooling operations at different scales through pyramid pooling, and splices and outputs the pooling results of each scale; the Concat module is used for splicing feature maps; the Upsample module upsamples the feature map through nearest neighbor interpolation; the Detect module is used to finally detect small targets from the feature map;
[0017] Specifically, the first CBS module after the input layer of the YOLOv5 model is defined as layer 0, the second CBS module is layer 1, and so on, and the last C3 module is layer 23;
[0018] The added small target detection layer is to sequentially connect C3, CBS, Upsample, Concat, C3, CBS and Concat modules between the 16th and 17th layers of the original YOLOv5 model, for a total of 7 modules;
[0019] After inserting the small target detection layer, the number of layers of all modules is redefined. The number of layers of modules before the 16th layer of the original YOLOv5 model remains unchanged. After inserting the small target detection layer, the original 17th layer becomes the 24th layer, and the subsequent modules are similar. The last C3 module is the 30th layer.
[0020] Concatenate the feature maps of the 2nd and 19th layers; concatenate the feature maps of the 18th and 22nd layers; add the Detect module after the 23rd layer to complete the introduction of the entire small target detection layer;
[0021] The small object detection layer undergoes two CBS modules based on the original image for 4x downsampling, introducing a feature map of 160×160, where each 1×1 pixel point only contains the information of 4×4 pixel points in the original image.
[0022] Specifically, Step 3 is as follows:
[0023] The feature enhancement module MFFM adopts a multi-branch convolution structure to extract various detailed semantic information. Each branch adopts the form of cascading standard convolution and dilated convolution, and a parallel structure is formed between multiple branches to fully extract the features of small objects. Specifically:
[0024] Each branch first passes through a standard two-dimensional convolution with a convolution kernel of 1×1 to reduce the number of channels and reduce the computational complexity;
[0025] The first branch passes through a dilated convolution with a convolution kernel of 3×3 and a dilation rate of 3;
[0026] The second branch sequentially passes through two dilated convolutions with a convolution kernel size of 3×3 and dilation rates of 2 and 5 respectively;
[0027] The third branch sequentially passes through two dilated convolutions with a convolution kernel size of 3×3 and dilation rates of 3 and 4 respectively;
[0028] The fourth branch is a residual structure that only passes through a 1×1 convolution to reduce the number of channels and retain the shallow texture information of small objects;
[0029] Specifically, by reducing the number of channels to reduce the computational complexity, it is ensured that the feature enhancement module MFFM does not increase the computational amount and has relatively excellent feature extraction ability. By increasing dilated convolutions with different dilation rates, the receptive field is enhanced, and the detection ability of small objects is improved. The input feature map X obtains the enhanced feature map Y after passing through the feature enhancement module MFFM, and the calculation formula is:
[0030]
[0031] In the formula, y n represents the output feature map of the nth branch of the feature enhancement module MFFM, where n ∈ (1, 2, 3, 4); represents the standard two-dimensional convolution with a convolution kernel size of 1×1; represents the dilated convolution with a dilation rate of m and a convolution kernel size of 3×3, where m ∈ (2, 3, 4, 5); concat represents the splicing operation of feature maps; represents the bitwise addition operation of feature maps.
[0032] The feature enhancement module MFFM can fuse multi-level features, enabling the model to learn richer context features, while maintaining the computational cost and increasing the receptive field, thereby improving the model's ability to represent small targets in remote sensing images.
[0033] Specifically, Step4 is as follows:
[0034] The fusion attention mechanism module CASA is divided into two branches;
[0035] In the first branch, the original input feature map of CASA first passes through two CBS modules connected in series to obtain the input feature map of the channel attention mechanism module; after inputting the input feature map into the channel attention mechanism module, the channel attention mechanism weight is obtained. Subsequently, the input feature map and the channel attention mechanism weight are subjected to a matrix multiplication operation to obtain the output feature map of the channel attention mechanism module; the output feature map is used as the input of the spatial attention mechanism module to obtain the spatial attention mechanism module weight; subsequently, the output feature map and the spatial attention mechanism module weight are subjected to a matrix multiplication operation to obtain the output feature map of the spatial attention mechanism module, which is used as the output of the first branch;
[0036] The second branch is a residual structure that directly outputs the original input feature map;
[0037] The output feature maps of the two branches are subjected to a feature map splicing operation to obtain the CASA output feature map.
[0038] The present invention introduces the attention mechanism module CASA to imitate the human visual and cognitive system, focusing on the parts of small targets when the model processes the input feature map, so as to solve the problems of foreground and background confusion and serious omission of small targets in remote sensing images, and improve the performance and generalization ability of the model for small target detection in remote sensing images.
[0039] Specifically, Step5 is as follows:
[0040] The original loss function of YOLOv5 is CIoU, which is specifically defined as:
[0041]
[0042]
[0043] In the formula, L CIoU represents the loss function CIoU; IoU represents the intersection over union between the ground truth box and the predicted box; ρ(b gt , b prd ) is the Euclidean distance between the center point coordinates of the ground truth box and the predicted box; b gt and b prdThey are the center point coordinates of the ground truth box and the predicted box respectively; c represents the diagonal length of the smallest bounding rectangle that can enclose the ground truth box and the predicted box; α represents the balance parameter; v represents the parameter for measuring the aspect ratio consistency; w gt and h gt represent the width and height of the ground truth box; w prd and h prd represent the width and height of the predicted box.
[0044] The original loss function CIoU of YOLOv5 only considers the relative value of the aspect ratio of the predicted box, rather than the absolute value. When the ground truth box and the predicted box have the same aspect ratio but different specific values of width and height, the loss function CIoU will lose its effectiveness, which will greatly limit the convergence speed and prediction accuracy of the model.
[0045] The present invention introduces the MPDIoU loss function to replace the original loss function. The MPDIoU loss function is defined as:
[0046]
[0047] In the formula, L MPDIoU represents the loss function MPDIoU; represent the coordinate values of the upper left and lower right corners of the ground truth box and the predicted box respectively; gt represents the preset ground truth box; prd represents the candidate box generated by the model; d1 and d2 represent the distances between the upper left corners and the lower right corners of the ground truth box and the predicted box respectively; w and h are the width and height of the input feature map respectively.
[0048] The MPDIoU loss function achieves better performance of the loss function by minimizing the distances between the upper left corners and the lower right corners of the ground truth box and the predicted box.
[0049] The beneficial effects of the present invention are as follows: The present invention adds a small target detection layer and a corresponding detection head, transforms the model into a four-head detection model, and improves the detection performance of the model for small targets; a feature enhancement module MFFM is introduced into the neck network to improve the feature extraction ability of the model; some C3 modules in the neck network are replaced with a fusion attention mechanism module CASA to focus more attention on small targets and improve the attention degree of the model to small targets; finally, the MPDIoU loss function is introduced, and the improved model FMCM-YOLO is finally obtained. In summary, the improvements proposed by the present invention can improve the detection performance of the model for small targets in optical remote sensing images and reduce the probability of missed detection and false detection. Description of the Drawings
[0050] Figure 1 is the flowchart of the steps of the present invention;
[0051] Figure 2This is the structural diagram of the FMCM-YOLO model of the present invention;
[0052] Figure 3 This is the structural diagram of the MFFM module of the present invention;
[0053] Figure 4 This is the structural diagram of the CASA module of the present invention. Detailed implementation manner
[0054] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation manners.
[0055] Example 1: An optical remote sensing image small target detection method based on improved YOLOv5. For the step flow chart, see Figure 1 For the specific structural diagram of the FMCM-YOLO model, see Figure 2 The technical solutions adopted mainly include the following steps:
[0056] Step1: Select an optical remote sensing image dataset and divide it into a training set and a validation set. Specifically:
[0057] Prepare the experimental configurations required for the present invention: The operating system uses Ubuntu22.04.4, the NVDIARTX4090D graphics card, the selected deep learning computing framework is PyTorch, the CUDA11.8 deep learning platform is used to manage the tool libraries required for running the code, and the PyCharm software is used to run the model code and configure the virtual environment and tool libraries required by YOLOv5.
[0058] In order to train the improved model of the present invention and verify the effectiveness and generality of the model, the USOD dataset and the AI-TOD dataset are selected to train and verify the model. The USOD dataset is a remote sensing image small target dataset constructed based on UNICORN2008. This dataset contains a total of 3000 images, all with a size of 640×640, including 43378 vehicle instances. The proportion of objects smaller than 16×16 is 96.3%, and the proportion of objects smaller than 32×32 is 99.9%, meeting the requirements for the remote sensing image small target detection method. The AI-TOD dataset is a dataset for detecting tiny objects in aerial images, including 28036 remote sensing pictures, with a total of 700621 object instances, divided into 8 object categories in total. The proportion of objects smaller than 32×32 is 97.9%. In addition, the two datasets are divided into a training set and a validation set according to a ratio of 7:3.
[0059] Step2: Taking YOLOv5 as the baseline algorithm, add a small target detection layer and the corresponding detection head, and transform the model into a four-head detection model to improve the detection performance of small targets.
[0060] Specifically, define the first CBS module after the input layer of the YOLOv5 model as layer 0, the second CBS module as layer 1, and so on, and the last C3 module as layer 23;
[0061] The added small target detection layer is formed by sequentially connecting C3, CBS, Upsample, Concat, C3, CBS, and Concat modules in series between the 16th and 17th layers of the original YOLOv5 model, for a total of 7 modules;
[0062] After inserting the small target detection layer, redefine the layer numbers of all modules. The layer numbers of the modules before the 16th layer of the original YOLOv5 model remain unchanged. After inserting the small target detection layer, the original 17th layer becomes the 24th layer, and the subsequent modules follow this pattern. The last C3 module is the 30th layer;
[0063] Concatenate the feature maps of layer 2 and layer 19; concatenate the feature maps of layer 18 and layer 22; add a Detect module after layer 23 to complete the introduction of the entire small target detection layer. See the Figure 2 dotted box part.
[0064] The small target detection layer undergoes two CBS modules on the basis of the original image for 4x downsampling, introducing a 160×160 feature map. Among them, each 1×1 pixel point only contains the information of 4×4 pixel points in the original image, which is beneficial for the model to detect small targets.
[0065] Step3: Introduce a feature enhancement module MFFM in the Neck network of YOLOv5 to improve the ability to extract features.
[0066] The feature enhancement module MFFM adopts a multi-branch convolution structure. Each branch adopts the form of a standard convolution and a dilated convolution in series, and a parallel structure is formed between the multi-branches. Specifically:
[0067] Each branch first passes through a standard two-dimensional convolution with a convolution kernel of 1×1;
[0068] The first branch passes through a dilated convolution with a convolution kernel of 3×3 and a dilation rate of 3;
[0069] The second branch sequentially passes through two dilated convolutions with a convolution kernel size of 3×3 and dilation rates of 2 and 5 respectively;
[0070] The third branch sequentially passes through two dilated convolutions with a convolution kernel size of 3×3 and dilation rates of 3 and 4 respectively;
[0071] The fourth branch is a residual structure and only passes through a 1×1 convolution; See the specific MFFM module structure in Figure 3 .
[0072] The input feature map X is enhanced to obtain the enhanced feature map Y after passing through the feature enhancement module MFFM. The calculation formula is as follows:
[0073]
[0074] In the formula, y n represents the output feature map of the nth branch of the feature enhancement module MFFM, where n ∈ (1, 2, 3, 4); represents a standard two-dimensional convolution with a convolution kernel size of 1×1; represents a dilated convolution with a dilation rate of m and a convolution kernel size of 3×3, where m ∈ (2, 3, 4, 5); concat represents the concatenation operation of feature maps; represents the element-wise addition operation of feature maps.
[0075] Step4: Introduce the fusion attention mechanism module CASA into the Neck network of YOLOv5 to improve the sensitivity to small targets.
[0076] Specifically, the fusion attention mechanism module CASA is divided into two branches;
[0077] The first branch first passes the original input feature map of CASA through two CBS modules connected in series to obtain the input feature map of the channel attention mechanism module; after inputting the input feature map into the channel attention mechanism module, the channel attention mechanism weight is obtained. Subsequently, the input feature map and the channel attention mechanism weight are multiplied matrix-wise to obtain the output feature map of the channel attention mechanism module; the output feature map is used as the input of the spatial attention mechanism module to obtain the spatial attention mechanism module weight; subsequently, the output feature map and the spatial attention mechanism module weight are multiplied matrix-wise to obtain the output feature map of the spatial attention mechanism module, which is used as the output of the first branch;
[0078] The second branch is a residual structure that directly outputs the original input feature map;
[0079] The output feature maps of the two branches are concatenated to obtain the CASA output feature map; for the specific CASA structure diagram, see Figure 4 . The calculation process of the fusion attention mechanism module CASA can be expressed as:
[0080] F CA = f CBS (f CBS (F))
[0081]
[0082] In the formula, the input feature map of the CASA module is denoted as F, and the output feature map is denoted as Z; f CBS denotes that the feature map passes through CBS and outputs; F CA , and w CA respectively denote the input feature map, output feature map and weight of the channel attention mechanism module; F SA and w SA respectively denote the input feature map and weight of the spatial attention mechanism module; · denotes matrix multiplication operation on the feature map; denotes the bitwise addition operation on the feature map.
[0083] Step5: Introduce the MPDIoU loss function to obtain the improved detection model FMCM-YOLO.
[0084] Specifically, the original loss function of YOLOv5 is CIoU, which is defined as:
[0085]
[0086] In the formula, L CIoU denotes the loss function CIoU; IoU denotes the intersection over union between the ground truth box and the predicted box; ρ(b gt ,b prd ) is the Euclidean distance between the center point coordinates of the ground truth box and the predicted box; b gt and b prd are the center point coordinates of the ground truth box and the predicted box respectively; c represents the diagonal length of the smallest bounding rectangle that can enclose the ground truth box and the predicted box; α represents the balance parameter; v represents the parameter for measuring the aspect ratio consistency; w gt and h gt denote the width and height of the ground truth box; w prd and h prd denote the width and height of the predicted box.
[0087] The original loss function CIoU of YOLOv5 only considers the relative value of the aspect ratio of the predicted box, rather than the absolute value. When the ground truth box and the predicted box have the same aspect ratio but different specific values of width and height, the loss function CIoU will lose effectiveness, which will greatly limit the convergence speed and prediction accuracy of the model.
[0088] The present invention introduces the MPDIoU loss function to replace the original loss function. The MPDIoU loss function is defined as:
[0089]
[0090] In the formula, L MPDIoU denotes the loss function MPDIoU; They respectively represent the coordinate values of the upper left and lower right corners of the ground truth box and the predicted box; gt represents the preset ground truth box; prd represents the candidate box generated by the model; d1 and d2 respectively represent the distances between the upper left corners and the lower right corners of the ground truth box and the predicted box; w and h are respectively the width and height of the input feature map.
[0091] The MPDIoU loss function achieves better loss function performance by minimizing the distances between the upper left corners and the lower right corners of the ground truth box and the predicted box.
[0092] Step6: Use the training set to train the FMCM-YOLO model to obtain the training weights, and use the validation set to validate the model, and finally achieve the detection of small targets in optical remote sensing images.
[0093] Specifically, after building the FMCM-YOLO model, the present invention uses the small target dataset USOD of remote sensing images to train and validate the superiority of the model. The SDG optimizer is used in the training, and the hyperparameter settings are shown in Table 1.
[0094] Table 1 Hyperparameter settings for model training
[0095]
[0096] To verify the superiority of the proposed FMCM-YOLO model of the present invention, under the conditions of the same experimental environment and the same network parameters, it is compared with other mainstream object detection algorithms based on two datasets respectively, and the results are shown in Tables 2 and 3.
[0097] Table 2 Detection effects of different algorithms on the USOD dataset
[0098]
[0099]
[0100] The performance evaluation metrics adopted by the USOD dataset are precision (P), recall (R), and mean average precision (mAP). Among them, mAP50 represents the mean average precision of all target categories when the intersection over union (IoU) threshold is 0.5; mAP50:95 represents the mean average precision at 10 different IoU thresholds from 0.5 to 0.95 with a step of 0.05. The experimental results in Table 2 show that the detection performance of the present invention is better than other mainstream detection algorithms in the USOD dataset. Compared with the baseline algorithm YOLOv5m, the performance metrics P, R, mAp50, and mAP50:95 are improved by 3.8%, 2.6%, 2.8%, and 2.2% respectively.
[0101] Table 3 Detection effects of different algorithms on the AI-TOD dataset
[0102]
[0103] In addition to mAP50 and mAP50:95, the AI-TOD dataset also uses mAPvt, mAPt, and mAps as evaluation metrics. Among them, mAPvt represents the mAP of objects with a size smaller than 8×8; mAPt represents the mAP of objects with a size between 8×8 and 16×16; mAPs represents the mAP of objects with a size between 16×16 and 32×32. It can be seen from the results in Table 3 that the detection performance of the present invention in the AI-TOD dataset is better than that of other mainstream detection algorithms. Compared with the baseline algorithm YOLOv5m, the performance indicators mAp50, mAP50:95, mAPvt, mAPt, and mAPs have been improved by 5.9%, 5%, 2.1%, 6.5%, and 5.1% respectively.
[0104] In summary, the algorithm proposed by the present invention can effectively improve the detection performance of small targets in optical remote sensing images, is more superior than other mainstream algorithms, and has good versatility and practicality.
[0105] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. A small target detection method for optical remote sensing images based on improved YOLOv5, characterized by: Step 1: Select the optical remote sensing image dataset and divide it into training set and validation set; Step 2: Using YOLOv5 as the baseline algorithm, add a small target detection layer and the corresponding detection head, and transform the model into a four-head detection model to improve the detection performance of small targets; Step 3: Introduce the feature enhancement module MFFM into the neck network Neck of YOLOv5 to improve the ability of feature extraction; Step 4: Introduce the fusion attention mechanism module CASA into the neck network Neck of YOLOv5 to improve the sensitivity to small targets; Step 5: Introduce the MPDIoU loss function to obtain the improved detection model FMCM-YOLO; Step 6: Use the training set to train the FMCM-YOLO model to obtain the training weights, and use the validation set to validate the model, ultimately achieving the detection of small targets in optical remote sensing images.
2. The optical remote sensing image small target detection method based on improved YOLOv5 according to claim 1, wherein The Step 2 is specifically as follows: Define the first CBS module after the input layer of the YOLOv5 model as layer 0, the second CBS module as layer 1, and so on, and the last C3 module as layer 23; The added small target detection layer is to sequentially connect C3, CBS, Upsample, Concat, C3, CBS and Concat modules between the 16th and 17th layers of the original YOLOv5 model, for a total of 7 modules; After inserting the small target detection layer, the number of layers of all modules is redefined. The number of layers of modules before the 16th layer of the original YOLOv5 model remains unchanged. After inserting the small target detection layer, the original 17th layer becomes the 24th layer, and the subsequent modules are similar. The last C3 module is the 30th layer. Concatenate the feature maps of the 2nd and 19th layers; concatenate the feature maps of the 18th and 22nd layers; add the Detect module after the 23rd layer to complete the introduction of the entire small target detection layer; The small target detection layer passes through the CBS module twice on the basis of the original image, performs 4-fold downsampling, and introduces a 160×160 feature map, in which each 1×1 pixel only contains the information of 4×4 pixels in the original image.
3. A small target detection method for optical remote sensing images based on improved YOLOv5 according to claim 1, characterized in that, The Step 3 is specifically as follows: The feature enhancement module MFFM adopts a multi-branch convolution structure, each branch adopts the form of standard convolution and hole convolution in series, and multiple branches form a parallel structure, specifically: Each branch first passes through a standard two-dimensional convolution with a convolution kernel of 1×1; The first branch passes through a 3×3 dilated convolution with a dilation rate of 3; The second branch passes through two dilated convolutions in sequence, with a kernel size of 3×3 and dilation rates of 2 and 5 respectively; The third branch passes through two dilated convolutions with kernel size of 3×3 and dilation rates of 3 and 4 respectively; The fourth branch is a residual structure, which only passes 1×1 convolution; The input feature map X is enhanced by the feature enhancement module MFFM to obtain the enhanced feature map Y. The calculation formula is: where y n represents the output feature map of the n-th branch of the feature enhancement module MFFM, where n ∈ (1, 2, 3, 4); represents a standard two-dimensional convolution with a convolution kernel size of 1×1; represents a dilated convolution with a dilation rate of m and a convolution kernel size of 3×3, where m ∈ (2, 3, 4, 5); concat represents the concatenation operation of feature maps; represents the bitwise addition operation of feature maps.
4. A small target detection method for optical remote sensing images based on improved YOLOv5 according to claim 1, characterized in that, The Step 4 is specifically as follows: The fusion attention mechanism module CASA is divided into two branches; In the first branch, the original input feature map of CASA first passes through two consecutively connected CBS modules to obtain the input feature map of the channel attention mechanism module. After the input feature map is input into the channel attention mechanism module, the channel attention mechanism weights are obtained. Subsequently, the input feature map and the channel attention mechanism weights are subjected to a matrix multiplication operation to obtain the output feature map of the channel attention mechanism module. The output feature map is used as the input of the spatial attention mechanism module to obtain the spatial attention mechanism module weights; Subsequently, the output feature map and the spatial attention mechanism module weights are subjected to a matrix multiplication operation to obtain the output feature map of the spatial attention mechanism module, which is used as the output of the first branch; The second branch is a residual structure that directly outputs the original input feature map; The output feature maps of the two branches are subjected to a feature map concatenation operation to obtain the CASA output feature map.