Synthetic aperture radar image small target detection method based on deep learning
By introducing an improved SHViT backbone network and efficient multi-scale attention EMA mechanism in the YOLOv8 model, the problem of low detection accuracy of small objects in SAR images is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411835415.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art has low accuracy in synthetic aperture radar (SAR) images, resulting in low detection rate and high error detection rate.
Using an improved method based on the YOLOv8 model, feature extraction and detection performance is enhanced by introducing an improved SHViT backbone network and an efficient multi-scale attention EMA mechanism.
It significantly improves the detection accuracy and robustness of small targets in SAR images, reduces the false detection rate and missed detection rate, and improves the detection performance.
Smart Images

Figure CN119942172A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and in particular relates to a method for detecting small targets in synthetic aperture radar images based on deep learning. Background Art
[0002] Synthetic Aperture Radar (SAR) is an all-day, all-weather sensor that transmits and receives radar signals multiple times to synthesize echo signals at different times to obtain high-resolution ground imaging. SAR radar technology is widely used in geological exploration, environmental monitoring, agricultural monitoring, disaster assessment, and military reconnaissance. In these applications, it is crucial to be able to accurately detect targets in images.
[0003] However, the detection of small targets in SAR images faces great challenges. In the field of computer vision, small targets are usually defined as those whose square root of the ratio of their bounding box area to the image area is less than 0.03. Such targets usually have low scattering intensity and occupy few pixels in the image, which can be easily submerged by background noise. In addition, due to the limitation of SAR image resolution, the imaging of small targets may appear blurred, thus affecting the accuracy of detection results.
[0004] In recent years, object detection methods based on deep learning and neural networks have made significant progress, including two-stage detection models such as R-CNN, Fast R-CNN, and Faster R-CNN. These models extract candidate regions through selective search, and then extract and classify these regions. Although the detection accuracy is very high, the cost is that the detection speed is extremely slow. In addition, there are some one-stage detection models with convolutional neural networks (CNN) as the backbone network, such as YOLO (You Only LookOnce) and SSD (Single Shot Multibox Detector). These models improve the detection speed through local receptive fields and shared weights.
[0005] Although these deep learning methods have shown good performance and robustness in target detection, they still have shortcomings in the detection accuracy of small targets, resulting in low detection rate and high false detection rate of small targets. Summary of the invention
[0006] Purpose of the invention: The present invention aims to solve the defects of low accuracy in target detection in synthetic aperture radar (SAR) images in the prior art, especially the problem of insufficient detection of small targets. To this end, the present invention provides a new improved SAR image target detection method based on the YOLOv8 model. The method adopts an innovative backbone network and feature extraction neck network, thereby significantly improving the detection performance of the model for various types of targets in SAR images. This improvement effectively enhances the accuracy and robustness of target detection in complex environments and is suitable for a variety of practical application scenarios.
[0007] In order to achieve the above technical objectives, the present invention provides a method for detecting small targets in synthetic aperture radar images based on deep learning, comprising the following steps:
[0008] Step 1: Establish an experimental target detection dataset;
[0009] Step 2, using the original YOLOv8 model as the baseline model, improve the model and generate a new YOLOv8 model;
[0010] Step 3, using the established data set to train the new YOLOv8 model to obtain a SAR image target detection prediction model;
[0011] Step 4: Input the images of the test set into the SAR image target detection prediction model one by one, and output the detection results to evaluate the detection performance of the improved model.
[0012] Step 1 includes the following steps:
[0013] Step 1.1, collect multiple synthetic aperture radar SAR images of multiple categories, and use the YOLO format to annotate the targets in each image to establish an experimental target detection dataset; YOLO uses bounding boxes and category labels to annotate the targets in the image. In this way, each image will have a corresponding label file containing the category information and location information of the target. These annotation information enables the semantic information of the image to be correctly read;
[0014] Step 1.2, randomly divide the experimental target detection dataset into training set, validation set and test set in a ratio of 6:2:2; the training set is used to train the model, the validation set is used to adjust the hyperparameters and model selection, and the test set is used to finally evaluate the performance of the model;
[0015] Step 1.3, perform data augmentation on the training set obtained in step 1.2 to obtain enhanced data.
[0016] In step 1.1, YOLO uses bounding boxes and category labels to mark the objects in the image. Each image has a corresponding label file that contains the category information and location information of the object.
[0017] In step 1.3, the data augmentation operations include translation, scaling, flipping, random cropping, brightness and saturation adjustment, random fusion, and mosaic enhancement. These operations increase the generalization ability and robustness of the model by introducing more changes during the training process.
[0018] Step 2 includes:
[0019] Step 2.1, build a SAR image target detection model based on YOLOv8. The model includes an input end, a backbone network, a neck network, and a detection head.
[0020] Step 2.1 includes the following steps:
[0021] Step 2.1.1, the input end divides the training set obtained in step 1.3 according to the set batch size, and then outputs the divided training set to the backbone network;
[0022] Step 2.1.2, in the backbone network, an improved SHViT (Single Head Vision Transformer) is introduced as a new backbone network for feature extraction; SHViT is a visual model based on the Transformer architecture that captures long-distance dependencies in images through a self-attention mechanism. The improved SHViT includes four stages stage1 to stage4 corresponding to the residual neural network ResNet and the sliding window Transformer (in the backbone of the network, four feature maps of different sizes are generally output to the neck network behind, and the structure between the output of the previous size feature map and the next size feature map is called a stage), stage1 to stage2 are called shallow stages, and stage3 to stage 4 are called deep stages. The single head self-attention mechanism is omitted in the shallow stage, and the calculation formula of the shallow stage is:
[0023] F i =residual(FFN(residual(BN(DWconv(F i-1 )))))i=1or 2
[0024] Among them, FFN is a fully connected layer, residual is a residual block connection, BN is batch normalization, DWconv is a depth-separable convolution, F i It is the feature map output by the i-th stage;
[0025] In the shallow stage, the feature map F from the previous stage i-1 Introduced into the SHViT block (the basic component of the model is block, i.e. SHViT block), the feature map F i-1 After a depth-wise separable convolution with a kernel of 3×3, the model is then passed through a batch normalization layer to stabilize the training process and a residual connection is used to enhance the direct transfer of gradients. Finally, the feature maps are integrated through a fully connected layer and passed to the next stage.
[0026] In the deep stage, the convolution part is the same as the shallow stage, and the single-head self-attention calculation formula is:
[0027]
[0028] Among them, concat is the concatenation layer, split is the splitting layer, Attention is the attention calculation, W Q ,W K ,W V is the projection weight, d qk is the dimension of Q vector and K vector, F att ,F res Respectively represent the features for attention calculation and the remaining features; represents the features calculated by self-attention; T represents transposition; softmax represents the normalized exponential function;
[0029] In the deep stage, the processing flow of the feature map from the previous stage includes: first, the feature map passes through a 3×3 depth-separable convolution and uses residual connections to enhance the transmission of features; then, the feature map undergoes layer normalization to separate the features of a specific part of the channel rC of the feature map and perform attention calculation, while the feature map of the remaining channel (1-r)C remains unchanged to maintain the original feature representation;
[0030] After the attention calculation is completed, the partial channel feature maps separated for attention processing are concatenated with the remaining channel feature maps that remain unchanged through a concat operation. Subsequently, the concatenated feature maps pass through a projection layer and finally the output F is obtained. i , the output F i It is then passed to the next stage of the network;
[0031] The SHViT blocks at different stages (i.e., stage1 to stage2 or stage2 to stage3 or stage3 to stage4) are connected by downsampling layers. In the downsampling layers, the feature maps pass through three convolutional layers in sequence, and the SE attention mechanism is introduced. The formula for SE attention calculation is:
[0032]
[0033] Where x is the input, act is the activation function, the activation function used here is the ReLU function, GAP is the global average pooling, scale is the scaling operation, sigmoid is the S-type activation function, and SE(x) is the final output of the SE attention mechanism;
[0034] Finally, it passes through the fast pyramid pooling SPPF layer, which is composed of the convolution layer and the maximum pooling layer. After feature extraction, the backbone network outputs feature maps of different sizes to the neck network.
[0035] Step 2.1.3, the neck network includes upsampling, downsampling, splicing and C2f modules, wherein upsampling and downsampling are to change the size of the feature map, splicing is to splice the feature maps of the same size according to the channel to achieve feature fusion, and the C2f module further extracts and enhances the fused features. Finally, feature maps of different sizes are output to the detection head;
[0036] The C2f module in the neck network is improved by introducing an efficient multi-scale attention EMA mechanism to obtain the C2f_EMA module. The C2f_EMA module uses a parallel convolution kernel structure to establish local cross-channel interactions by increasing the width and feature grouping. It uses a cross-spatial aggregation method to fuse the feature maps output by the two branches in parallel in the X and Y spatial dimensions to obtain feature aggregation information and output it to the detection head.
[0037] Step 2.1.4, the detection head has two branches, each of which contains two convolution layers with a convolution kernel size of 3, a convolution layer with a convolution kernel size of 1, and loss function calculation; among them, two types of convolution layers are used for information extraction, and the loss function calculation is divided into two items: bounding box regression loss function and classification loss function; an additional detection head for shallow feature maps is introduced in the detection head, and the backbone network and neck network are adaptively modified: the feature map calculated by the backbone network in the shallow stage stage1 is also used as input and passed to the neck network for feature fusion; in the neck network, an additional upsampling process is added to realize the fusion of shallow feature maps, and the fused feature maps are directly input into the newly added detection head for detection. This change significantly improves the detection performance of small targets and reduces the false detection rate and missed detection rate of small targets.
[0038] Step 3 includes the following steps:
[0039] Step 3.1: Input the training set obtained in step 1 into the new YOLOv8 model for training. The hyperparameters set during training include learning rounds and batch size. Generally, the learning rounds are 300 and the batch size is 32. The optimizer uses the stochastic gradient descent strategy with momentum, and the default input size is 384×384.
[0040] Step 3.2, save the trained model and weight files for verification and testing of model detection performance; and record the model parameter quantity and calculation amount.
[0041] Step 4 includes the following steps:
[0042] Step 4.1, load the model and weights obtained in step 3, and evaluate the model performance on the test set. The evaluation indicators include accuracy, recall rate and average detection precision.
[0043] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the described method.
[0044] The present invention also provides a storage medium storing a computer program or instruction. When the computer program or instruction is run on a computer, the steps of the method described are executed.
[0045] Beneficial effects: The deep learning-based SAR image small target detection method of the present invention has excellent detection performance, and improves the detection index of all types of targets without increasing the computing burden of the equipment. In particular, in the detection of small targets, the false detection rate and missed detection rate are significantly reduced, and fast and accurate detection in complex scenes is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.
[0047] Figure 1 It is the overall structure flow chart of the present invention.
[0048] Figure 2 It is a complete structural diagram of the model proposed in the present invention.
[0049] Figure 3 It is a structural diagram of the SHViT block in the improved backbone network of the present invention.
[0050] Figure 4 It is a diagram of the SPPF layer structure used in the present invention.
[0051] Figure 5This is the complete structural diagram of the improved C2f_EMA in the present invention. DETAILED DESCRIPTION
[0052] like Figure 1 As shown, an embodiment of the present invention provides a method for detecting small targets in synthetic aperture radar images based on deep learning, comprising the following steps:
[0053] Step 1: Collect a diverse SAR image dataset, including target types such as ships, oil tanks, aircraft, and bridges. Label the targets in the dataset in the format required by the YOLO model, and divide the dataset into training set, validation set, and test set in a ratio of 6:2:2. Preprocess the training set to improve data quality and reduce interference factors in training, thereby helping the model learn and generalize better and improve prediction accuracy.
[0054] Step 2: Optimize and improve the model structure. The improved model structure is shown in the figure below: Figure 2 The specific measures are as follows:
[0055] In the backbone network of the model, a modified SHViT network architecture is used. At the beginning of the network (stem part), stride convolution is applied, which helps to extract local information of the image more effectively. In the SHViT block part, the depth-separable convolution is combined with the single-head self-attention mechanism. This composite strategy can effectively capture the global and local dependencies in the feature map without significantly increasing the computational and memory burden.
[0056] In order to further improve the performance and computational efficiency of the model, the present invention divides the network structure into four stages corresponding to ResNet and Swin Transformer. Considering that the self-attention mechanism is not necessary in the shallow stage of the network and may lead to a waste of computing resources, the single-head self-attention mechanism is omitted in the shallow stage to improve the computational efficiency of the model. Two different SHViT block structures are shown as follows: Figure 3 shown.
[0057] The calculation formula for the shallow stage is:
[0058] F i =residual(FFN(residual(BN(DWconv(F i-1 )))))i=1or 2
[0059] Among them, FFN is a fully connected layer, residual is a residual block connection, BN is batch normalization, DWconv is a depth-separable convolution, Fi It is the feature map output by the i-th stage.
[0060] In the shallow stage of the network, the feature map F from the previous stage i-1 Introduced into the SHViT block, these feature maps first undergo a depthwise separable convolution with a kernel of 3×3, then pass through a batch normalization layer to stabilize the training process, and use residual connections to enhance the direct transfer of gradients. Finally, the feature maps are integrated through a fully connected layer and passed to the next stage. This feature processing flow reduces the computational burden while improving training speed and model accuracy.
[0061] In the deep stage, the convolution part is the same as the shallow stage, and the single-head self-attention calculation formula is as follows:
[0062]
[0063] Among them, concat is the concatenation layer, split is the splitting layer, Attention is the attention calculation, W Q ,W K ,W V is the projection weight, d qk is the dimension of Q vector and K vector, F att ,F res Respectively represent the features for attention calculation and the remaining features, F i It is the feature map output by the i-th stage.
[0064] In the deep stage of the network, the processing flow of the feature map from the previous stage is as follows: First, the feature map passes through a 3×3 depth-separable convolution and uses residual connections to enhance feature transmission. Then, the feature map undergoes layer normalization processing to separate the features of a specific part of the channel (denoted as rC) of the feature map and perform attention calculation. The feature map of the remaining channel (denoted as (1-r)C) remains unchanged and maintains its original feature representation.
[0065] After the attention calculation is completed, the partial channel feature maps separated for attention processing are concatenated with the remaining channel feature maps that remain unchanged through a concat operation. Subsequently, the concatenated feature maps pass through a projection layer to ensure the consistency of feature dimensions, and finally the output F is obtained. i , which is then passed to the next stage of the network.
[0066] The present invention uses an efficient downsampling layer to connect SHViT blocks at different stages. In the downsampling layer, the feature map passes through three convolutional layers in sequence to achieve spatial dimension reduction of features and information extraction. In this process, the SE (Squeeze-and-Excitation) attention mechanism is introduced. The SE attention mechanism globally averages the pooled feature map through the Squeeze operation to capture the dependencies between channels, and recalibrates the channel responses through the Excitation operation. Such a mechanism effectively improves the expressiveness of important features while suppressing unimportant features. There is no information loss during the sampling process and the representation ability of the feature map is improved. The formula for SE attention calculation is as follows:
[0067]
[0068] Among them: x is the input, act is the activation function, here it is the ReLU function, GAP is the global average pooling, and scale is the scaling operation.
[0069] Finally, it passes through the SPPF (Spatial Pyramid Pooling-Fast) layer. The SPPF layer is composed of a convolutional layer and a maximum pooling layer. Its structure is as follows: Figure 4 shown.
[0070] The above content together constitutes the backbone network part of the improved network, which realizes feature extraction and obtains multi-scale feature information, thereby improving the accuracy and robustness of the model.
[0071] In the neck network, a detection head for small target detection is first added. To adapt to it, other structures of the network are also modified accordingly, including the introduction of upsampling layers, convolutional layers, and concat layers. This small target detection head mainly receives features from the shallow layers of the network. Since these features undergo fewer downsampling times, they retain more spatial details and low-level semantic information, which helps to achieve more accurate target positioning. Although this design increases the number of parameters and computation of the model, it can capture the features of small targets more effectively than traditional structures.
[0072] The C2f structure in the neck network was improved, and the efficient multi-scale attention EMA (Efficient Multi-Scale Attention) mechanism was introduced to obtain a C2f_EMA module with better performance. The structure diagram is shown in the figure Figure 5As shown in the figure. The efficient multi-scale attention EMA mechanism uses the parallel convolution kernel structure to establish local cross-channel interactions by increasing the width and feature grouping. It uses the cross-spatial aggregation method to fuse the feature maps output by the two branches in parallel in the two spatial dimensions of X and Y to obtain rich feature aggregation information. The fusion of C2f and EMA attention improves the expressiveness of features while maintaining computational efficiency.
[0073] After being processed by the backbone network and the neck network, four feature maps with different sizes and numbers of channels are finally generated. These feature maps carry rich semantic information and cover different levels of features from coarse to fine. These four feature maps are input into the detection head for further calculation. The detection head predicts the target location, category and corresponding confidence in the image by using these multi-scale features.
[0074] Step 3: After getting the trained model, input the test set into the model for prediction, and evaluate the model performance through the prediction results. The evaluation indicators include: mean average precision (mAP), precision, and recall. Their calculation formulas are as follows:
[0075]
[0076] Where TP, FP, and FN are respectively the number of truly positive samples that are correctly predicted as positive samples by the model, the number of actually negative samples that are incorrectly predicted as positive samples by the model, and the number of actually positive samples that are incorrectly predicted as negative samples by the model; n is the number of predicted target types; P i is the accuracy under the i-th recall rate, ΔR i It is the difference between the i-th recall rate and the i+1-th recall rate. The calculation of AP (Average Precision) requires the use of interpolation method.
[0077] After multiple experiments, compared with the basic model, the mAP increased by 8.3%, and the mAP in the specific category of aircraft increased by 18.7%. At the same time, the precision and recall increased by 2.5% and 10% respectively. The results fully demonstrate that the present invention is superior to the traditional YOLOv8 target detection model in detection performance. Compared with the prior art, the present invention achieves better detection performance, especially in small target detection; at the same time, it still maintains a low model parameter amount and calculation amount to ensure the efficiency and convenience of model deployment.
[0078] The present invention provides a method for detecting small targets in synthetic aperture radar images based on deep learning. There are many methods and ways to implement the technical solution. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the protection scope of the present invention. All components not specified in this embodiment can be implemented by existing technologies.
Claims
1. A method for detecting small targets in synthetic aperture radar images based on deep learning, characterized in that: The following steps are involved: Step 1: Establish an experimental target detection dataset; Step 2, using the original YOLOv8 model as the baseline model, improve the model and generate a new YOLOv8 model; Step 3, using the established data set to train the new YOLOv8 model to obtain a SAR image target detection prediction model; Step 4: Input the images of the test set into the SAR image target detection prediction model one by one, and output the detection results to evaluate the detection performance of the improved model.
2. The method according to claim 1, characterized in that Step 1 includes the following steps: Step 1.1, collect synthetic aperture radar SAR images and use the YOLO format to annotate the targets in each image to establish an experimental target detection dataset; Step 1.2, randomly divide the experimental target detection dataset into training set, validation set and test set according to the proportion; the training set is used to train the model, the validation set is used to adjust the hyperparameters and model selection, and the test set is used to finally evaluate the performance of the model; Step 1.3, perform data augmentation on the training set obtained in step 1.2 to obtain enhanced data.
3. The method according to claim 2, characterized in that In step 1.1, YOLO uses bounding boxes and category labels to mark the objects in the image. Each image has a corresponding label file that contains the category information and location information of the object.
4. The method according to claim 3, characterized in that In step 1.3, the data enhancement operations include translation, scaling, flipping, random cropping, brightness and saturation adjustment, random fusion and mosaic enhancement.
5. The method according to claim 4, characterized in that Step 2 includes: Step 2.1, build a SAR image target detection model based on YOLOv8. The model includes an input end, a backbone network, a neck network, and a detection head.
6. The method according to claim 5, characterized in that Step 2.1 includes the following steps: Step 2.1.1, the input end divides the training set obtained in step 1.3 according to the set batch size, and then outputs the divided training set to the backbone network; Step 2.1.2, in the backbone network, an improved SHViT is introduced as a new backbone network for feature extraction; the improved SHViT includes four stages stage1 to stage4 corresponding to the residual neural network ResNet and the sliding window Transformer, stage1 to stage 2 are called shallow stages, stage3 to stage 4 are called deep stages, and the single-head self-attention mechanism is omitted in the shallow stage. The calculation formula of the shallow stage is: F i =FFN(residual(BN(DWconv(F i-1 ))))i=1or 2 Among them, FFN is a fully connected layer, residual is a residual block connection, BN is batch normalization, DWconv is a depth-separable convolution, F i It is the feature map output by the i-th stage; In the shallow stage, the feature map F from the previous stage i-1 Introduced into the SHViT block, the feature map F i-1 After a depth-wise separable convolution with a kernel of 3×3, the model is then passed through a batch normalization layer to stabilize the training process and a residual connection is used to enhance the direct transfer of gradients. Finally, the feature maps are integrated through a fully connected layer and passed to the next stage. In the deep stage, the convolution part is the same as the shallow stage, and the single-head self-attention calculation formula is: Among them, concat is the concatenation layer, split is the splitting layer, Attention is the attention calculation, W Q ,W K ,W V is the projection weight, d qk is the dimension of Q vector and K vector, F att ,F res Respectively represent the features for attention calculation and the remaining features; Represents the features calculated by self-attention; T stands for transpose; Softmax represents the normalized exponential function; In the deep stage, the processing flow of the feature map from the previous stage includes: first, the feature map passes through a 3×3 depth-separable convolution and uses residual connections to enhance the transmission of features; then, the feature map undergoes layer normalization to separate the features of a specific part of the channel rC of the feature map and perform attention calculation, while the feature map of the remaining channel (1-r)C remains unchanged to maintain the original feature representation; After the attention calculation is completed, the partial channel feature maps separated for attention processing are concatenated with the remaining channel feature maps that remain unchanged through a concat operation. Subsequently, the concatenated feature maps pass through a projection layer and finally the output F is obtained. i , the output F i It is then passed to the next stage of the network; The downsampling layer is used to connect the SHViT blocks at different stages. In the downsampling layer, the feature map passes through three convolutional layers in sequence, and the SE attention mechanism is introduced. The formula for SE attention calculation is: Where x is the input, act is the activation function, the activation function used here is the ReLU function, GAP is the global average pooling, scale is the scaling operation, sigmoid is the S-type activation function, and SE(x) is the final output of the SE attention mechanism; Finally, it passes through the fast pyramid pooling SPPF layer, which is composed of the convolution layer and the maximum pooling layer. After feature extraction, the backbone network outputs feature maps of different sizes to the neck network. Step 2.1.3, the neck network includes upsampling, downsampling, splicing and C2f modules, wherein upsampling and downsampling are to change the size of the feature map, splicing is to splice the feature maps of the same size according to the channel to achieve feature fusion, and the C2f module further extracts and enhances the fused features. Finally, feature maps of different sizes are output to the detection head; The C2f module in the neck network is improved by introducing an efficient multi-scale attention EMA mechanism to obtain the C2f_EMA module. The C2f_EMA module uses a parallel convolution kernel structure to establish local cross-channel interactions by increasing the width and feature grouping. It uses a cross-spatial aggregation method to fuse the feature maps output by the two branches in parallel in the X and Y spatial dimensions to obtain feature aggregation information and output it to the detection head. In step 2.1.4, the detection head has two branches, each of which contains two convolutional layers with a convolution kernel size of 3, a convolutional layer with a convolution kernel size of 1, and a loss function calculation; wherein, two types of convolutional layers are used for information extraction, and the loss function calculation is divided into two items: a bounding box regression loss function and a classification loss function; an additional detection head for shallow feature maps is introduced into the detection head, and the backbone network and the neck network are adaptively modified: the feature map calculated by the backbone network in the shallow stage stage1 is also used as input and passed to the neck network for feature fusion; in the neck network, an additional upsampling process is added to realize the fusion of shallow feature maps, and the fused feature maps are directly input into the newly added detection head for detection.
7. The method according to claim 6, characterized in that Step 3 includes the following steps: Step 3.1: Input the training set obtained in step 1 into the new YOLOv8 model for training. The hyperparameters set during training include the learning round and batch size. The optimizer uses the stochastic gradient descent strategy with momentum. Step 3.2, save the trained model and weight files for verification and testing of model detection performance; and record the model parameter quantity and calculation amount.
8. The method according to claim 7, characterized in that Step 4 includes the following steps: Step 4.1, load the model and weights obtained in step 3, and evaluate the model performance on the test set. The evaluation indicators include accuracy, recall rate and average detection precision.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.
10. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 8 are executed.
Citation Information
Cited By
Separable multi-scale feature fusion method for image small target detection
CN120747689A
Underwater target detection method based on dual supervision guidance
CN121033646A
Network model and method for identifying dry and wet pores of poultry and chicken, and storage medium
CN121214071A