Lightweight small target detection method based on deep learning
By building the improved yolo11s model and introducing lightweight network design and loss functions, the existing small object detection methods are solved, and the high-precision and low-latency small object detection capabilities are achieved on edge devices, which are suitable for intelligent traffic monitoring.
Patent Information
- Application Number
- CN202510452405.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
AI Technical Summary
Due to the huge parameter scale, the existing small-objective detection methods have large computing volume, making it difficult to achieve efficient detection on mobile devices, and in the face of complex interference, the scenario generalization capability is insufficient.
By adopting a lightweight small object detection method based on deep learning, by building an improved yolo11s model, adding lightweight selective channel downsampling module SCDown, feature compression module SEA, lightweight cross-feature fusion module CFF and upsampling module DySample, the model structure is optimized, the parameter amount is reduced, and the Powerful-IoU loss function is introduced to improve detection accuracy.
It realizes high-precision and low-latency small target recognition capabilities on edge computing devices, and is suitable for intelligent traffic monitoring scenarios, improves traffic management efficiency and resource allocation optimization, and ensures traffic safety.
Smart Images

Figure CN119992075A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a lightweight small target detection method based on deep learning. Background Art
[0002] Small target detection technology has gradually become a core technology pillar in the field of intelligent traffic monitoring due to its high sensitivity, multi-scale adaptability and real-time processing capabilities. The accurate feature extraction and anti-interference recognition capabilities of tiny targets are the key bottlenecks in achieving high-precision detection.
[0003] The current mainstream methods mainly rely on multi-branch network architectures with large parameter scales, such as complex feature pyramid systems built by stacking densely connected modules. Such models improve detection accuracy through multi-level feature fusion, but information compatibility issues between different levels may still lead to small target feature matching deviations. Some schemes use parameter-intensive attention enhancement modules. Although they can suppress background interference through global context modeling, their huge amount of computation seriously restricts their application in mobile devices. In addition, although the super-resolution scheme based on parameter-intensive generative adversarial networks can enhance the details of the input image, the inference delay problem caused by the bloated model significantly reduces its practicality. Faced with complex interference such as sudden changes in illumination and dynamic blur, the contradiction between the high dependence of such large-parameter models on computing power and their scene generalization capabilities has become increasingly prominent. Summary of the invention
[0004] Purpose of the invention: The technical problem to be solved by the present invention is to provide a lightweight small target detection method based on deep learning in view of the shortcomings of the prior art, comprising the following steps:
[0005] Step 1, data preprocessing: Use a script to process the public dataset visdrone2019 to obtain the training set, validation set, and test set;
[0006] Step 2: Build an improved yolo11s model and input the training set for training;
[0007] Step 3: Use the trained yolo11s model to detect small targets.
[0008] Step 1 includes: after the script runs, first parse the command line parameters, obtain the path of the public dataset visdrone2019, and convert the public dataset visdrone2019 into a Path object;
[0009] Then, the script will process the three subdirectories in the public dataset visdrone2019 in turn: the training set VisDrone2019-DET-train, the validation set VisDrone2019-DET-val, and the test set VisDrone2019-DET-test-dev; the training set is used for model training, the validation set is used to adjust the model effect, and the test set is used for the final evaluation of the model;
[0010] For each subdirectory, the script creates a folder named labels to store the converted YOLO annotation files; labels means labels;
[0011] Next, the script will traverse the annotations folder in the subdirectory, read the .txt annotation files in it, and parse the annotation information line by line. For each line X1, the script will check whether the annotated category is an ignored area. If so, the line X1 will be skipped. The ignored area refers to the category number 0; otherwise, the script will convert the bounding box format (x1, y1, x2, y2) of the dataset visdrone2019 to the YOLO format (x_center, y_center, w, h), and normalize it to the range of [0, 1], and subtract 1 from the category number to adapt to the YOLO indexing method, where x1, y1 represent the horizontal and vertical coordinates of the vertex in the lower left corner of the image, respectively, and x2, y2 represent the horizontal and vertical coordinates of the vertex in the upper right corner, respectively; x_center, y_center represent the horizontal and vertical coordinates of the center of the image, respectively; w represents the width of the image, and h represents the height of the image;
[0012] The converted annotation information will be formatted into the format of the YOLO annotation file and written into the corresponding labels folder;
[0013] Finally, when all the annotation files in the subdirectories have been processed, the script ends and writes the converted YOLO annotation files to the labels folder in each subdirectory.
[0014] Step 2 includes: Optimizing on the basic yolo11s model:
[0015] A lightweight selective channel downsampling module SCDown is added to the backbone network of the YOLO11s model to optimize the downsampling process of YOLOv8s. At the same time, a feature compression module SEA (Squeeze and Excitation Attention) is added to reduce the number of parameters.
[0016] A lightweight cross-feature fusion module CFF is added to the neck network of the yolo11s model to reduce redundant features and enhance the focus on important features; at the same time, an upsampling module DySample and a lightweight selective channel downsampling module SCDown are added;
[0017] Add Powerful-IoU (penalized intersection over union) loss function during model training , used to optimize the regression of the detection box.
[0018] In step 2, the Powerful-IoU loss function The formula is: ,
[0019] Among them, IoU is the intersection over union ratio of the predicted box and the true box, a and b are hyperparameters that will be automatically adjusted during the learning process, and ϵ is the preset minimum value to avoid division by zero.
[0020] Step 3 includes: the trained yolo11s model receives an RGB image of size B×L×3, normalizes the image pixel values to the range of 0 to 1, where B represents the width of the image and L represents the height of the image;
[0021] The normalized image is input to the first convolution layer of the yolo11s model, and a 3×3 convolution kernel is used to extract low-level features and output a 64-channel feature map F1:
[0022] ,
[0023] Among them, W1 is the convolution kernel weight, b1 is the bias term, σ is the activation function, I normalized is the normalized input image.
[0024] Step 3 also includes: performing the first downsampling of the feature map through the SCDown module to reduce the spatial dimension while maintaining the channel information. The SCDown module implements the downsampling operation by combining point-by-point convolution and depth-wise separable convolution:
[0025] Point-by-point convolution: Use a 1×1 convolution kernel to reduce the number of channels of the input feature map. The calculation formula is:
[0026] ,
[0027] Among them, W pw is the point-by-point convolution kernel weight, b pw is the bias term, It is the feature map after point-by-point convolution;
[0028] Depthwise separable convolution: Use a 3×3 convolution kernel for spatial downsampling with a step size of 2. The calculation formula is:
[0029] ,
[0030] Among them, W dw is the depth-wise separable convolution kernel weight, b dw is the bias term, It is the feature map after downsampling;
[0031] Feature map after downsampling The input is sent to the C3k2 module of the yolo11s model, and the feature expression capability is enhanced through convolution and residual connection. The following two convolution operations are performed inside the C3k2 module:
[0032] First convolution: use convolution kernel to extract features , the calculation formula is:
[0033] ,
[0034] Among them, W conv1 is the weight of the first convolution, b conv1 is the bias term;
[0035] Second convolution: get enhanced features :
[0036] ,
[0037] Among them, W conv2 is the weight of the second convolution, b conv2 is the bias term;
[0038] Add the results of the two convolutions to get feature F3:
[0039] ,
[0040] Repeat the downsampling and feature enhancement operations on F3, gradually increasing the number of channels and reducing the size of the feature map, including:
[0041] Second downsampling: Use the SCDown module to increase the number of channels of feature F3 from 128 to 256, and perform spatial downsampling to output feature map F4 with a size of W / 4×H / 4×256;
[0042] The third downsampling: Use the SCDown module to increase the number of channels of feature map F4 from 256 to 512, and perform spatial downsampling to output feature map F shallow , size is W / 8×H / 8×512;
[0043] The fourth downsampling: Use the SCDown module to downsample the feature map F shallowThe number of channels is kept at 512, and spatial downsampling is performed to output a feature map F6 with a size of W / 16×H / 16×512.
[0044] Step 3 also includes: applying the SPPF module on the deep feature map F6, fusing multi-scale features, enhancing the model's adaptability to objects of different sizes, and outputting the feature map F7:
[0045] ,
[0046] Concat is a connection operation, F pool1 、F pool2 、F pool3 It is the feature map after pooling at different scales, generated by the following operations:
[0047] Use a 5×5 pooling window to perform maximum pooling on F6 with a step size of 1 to obtain F pool1 , use a 9×9 pooling window to perform maximum pooling on F6 with a step size of 1, and get F pool2 , use a 13×13 pooling window to perform maximum pooling on F6 with a step size of 1, and get F pool3 ;
[0048] The 512-channel feature map F7 output by the SPPF module is used as the input of the feature compression module SEA, which performs the following operations:
[0049] Perform global average pooling on the feature map F7 to obtain the channel feature vector v:
[0050] ,
[0051] Among them, the feature map F7 is a three-dimensional feature map with a size of H×W×C, where H is the height, W is the width, and C is the number of channels; i and j are the indexes in the height and width directions of the feature map, respectively, and the symbol : means taking the values of all channels; It means summing all spatial positions of feature map F7;
[0052] The ReLU activation function is used to reduce the feature dimension to 32 and obtain the feature map :
[0053] ,
[0054] in, and denote the first weight and the first bias respectively;
[0055] Restore the dimension to 512 and get the feature map :
[0056] ,
[0057] in, and denote the second weight and the second bias respectively;
[0058] Normalized by sigmoid function, the channel attention weight vector is obtained :
[0059] ,
[0060] in, is the sigmoid function;
[0061] The channel attention weight vector Multiply element-by-element with the original feature map F7 to obtain the output feature map F of the feature compression module SEA sea :
[0062] ,
[0063] Among them, ⊙ represents element-by-element multiplication;
[0064] Feature map F output by the backbone network sea As input to the neck network;
[0065] The upsampling module DySample is used to upsample the feature map, and specifically performs the following operations:
[0066] Offset calculation: for the input feature map F sea Perform convolution operation to obtain the offset feature map O:
[0067] ,
[0068] in, is the convolution kernel weight, is the bias term;
[0069] Generate a grid: Generate a sampling grid based on the offset feature map O and the preset initial position:
[0070] ,
[0071] Among them, grid is the generated two-dimensional sampling grid, is the preset initial position, O(i,j) is the value of the offset feature map at position (i,j);
[0072] According to the position of the sampling grid, the weight of the bilinear interpolation is calculated:
[0073] ,
[0074] Among them, weight(i,j,m,n) represents the weight between the target position (i,j) and the pixel position (m,n) in the original feature map, and m and n are the height index and width index of the original feature map respectively;
[0075] Bilinear interpolation: First initialize a feature map F up , and then use the generated sampling grid to input feature map F sea Perform bilinear interpolation to obtain the upsampled feature map F up :
[0076] ,
[0077] Among them, F up (i,j) represents the feature map F up The pixel value at position (i, j) in the image is calculated using this formula to get F up The pixel value at the corresponding position in is filled into the empty feature map F at the beginning. up The upsampling operation is completed at the corresponding position in ;
[0078] The CFF module receives the upsampled feature map F up and the shallower feature map F in the backbone network shallow as input.
[0079] Step 3 also includes: reorganizing the input feature map and adjusting the number of channels through convolution operations, specifically including:
[0080] For the feature map F up Perform convolution operation and adjust the number of channels and F shallow same;
[0081] Using the feature map after feature reorganization, calculate the weight of each feature map: perform global average pooling on each feature map to obtain the channel feature vector:
[0082] ,
[0083] ,
[0084] in, and Represents F up The channel eigenvector and F shallow The channel feature vector of ;
[0085] Use the fully connected layer to transform the feature vector and generate the initial weight value:
[0086] ,
[0087] ,
[0088] in, and Represents F up The initial weight value and F shallow The initial weight value of is the weight matrix of the fully connected layer, is the bias term;
[0089] Assign weights to different feature maps according to the proportion of initial weight values:
[0090] ,
[0091] ,
[0092] Among them, α and β are F up The weight and F shallow The weight of
[0093] Get the fused feature map F fuse :
[0094] ,
[0095] The fused feature map F fuse Input to the C3k2 module, and further enhance the feature expression ability through convolution and residual connection to obtain the feature map F ;
[0096] Repeat the upsampling, feature fusion and feature enhancement processes to process feature maps of different scales respectively, and finally generate a multi-scale fused feature map :
[0097] ,
[0098] in, is the output of F4 in the backbone network through the neck network, F shallow Output through the neck network;
[0099] Multi-scale fusion feature map output by the neck network As the input of the detection head;
[0100] For each scale feature map , i is 1, 2, 3, and the feature map is processed by convolution operation Adjust the number of channels:
[0101] ,
[0102] Among them, W det is the convolution kernel weight, b det is the bias term, F det_i is the feature map after processing;
[0103] Detection box parameter prediction: For each processed feature map F det_i , predict the detection box parameters at each location, including the coordinates, category probability and confidence of the bounding box:
[0104] ,
[0105] Among them, B represents the predicted detection box set, including the bounding box coordinates (x, y, w, h), category probability pc and confidence c; Predict represents the prediction function of the original yolo11s model.
[0106] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the described method.
[0107] The present invention also provides a storage medium storing a computer program or instruction. When the computer program or instruction is run on a computer, the steps of the method described are executed.
[0108] By integrating model compression technology with multi-scale perception architecture, the present invention realizes high-precision, low-latency small target recognition capabilities (detection size is usually less than 32×32 pixels) on edge computing devices, which is particularly suitable for intelligent traffic monitoring scenarios. This technology breaks through the limitations of traditional detection models with large parameters and strong hardware dependence, and can monitor traffic flow, vehicle speed and road conditions in real time. By optimizing the model structure and introducing lightweight network design, the present invention achieves efficient computing performance on edge devices while maintaining high detection accuracy, providing an ideal solution for traffic monitoring scenarios with high real-time requirements and limited deployment environments. This capability not only improves traffic management efficiency, but also optimizes resource allocation, ensures traffic safety, and promotes the development of green transportation.
[0109] The present invention has the following beneficial effects: Efficient feature extraction and fusion: Through multiple convolution and downsampling operations in the backbone network, the model can extract feature maps of different scales. These feature maps contain rich information from local details to global semantics, providing a solid foundation for subsequent target detection.
[0110] Excellence in Small Object Detection: Improvements in the neck network effectively fuse feature maps of different scales, and enhance feature relevance through weight calculation and weighted summation of features, which excels in small object detection tasks in particular.
[0111] High computational efficiency: By combining point-by-point convolution and depth-wise separable convolution, feature map downsampling is performed efficiently, reducing the consumption of computing resources while maintaining the integrity of feature information.
[0112] Small model size: The overall model design focuses on lightweight, which reduces the number of parameters and computational complexity while ensuring detection accuracy, making it suitable for deployment in resource-constrained environments.
[0113] Good adaptability: Through multi-scale feature extraction and feature fusion, the model can adapt to targets of different scales and shapes, improving its adaptability to various scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0114] Figure 1 Flowchart of the model algorithm.
[0115] Figure 2 This is a schematic diagram of the SCDown module processing flow.
[0116] Figure 3 Schematic diagram of the SEA module processing flow.
[0117] Figure 4 This is a schematic diagram of the DySample module processing flow.
[0118] Figure 5 This is a schematic diagram of the CFF module processing flow.
[0119] Figure 6 This is the target detection result image at the road intersection in the example.
[0120] Figure 7 This is the target detection result image on the road in the example. DETAILED DESCRIPTION
[0121] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.
[0122] The embodiment of the present invention provides a lightweight small target detection method based on deep learning, comprising the following steps:
[0123] Step 1, data preprocessing: Use a script to perform simple processing on the public dataset visdrone2019 so that it can be used as a training dataset for this technical model.
[0124] After the script runs, it first parses the command line parameters, obtains the dataset path, and converts it into a Path object for subsequent operations.
[0125] The script then processes the three subdirectories in the dataset in turn: VisDrone2019-DET-train (training set), VisDrone2019-DET-val (validation set), and VisDrone2019-DET-test-dev (test set).
[0126] For each subdirectory, the script creates a folder named labels to store the converted YOLO annotation files.
[0127] Next, the script will traverse the annotations folder in the subdirectory, read the .txt annotation files in it, and parse the annotation information line by line.
[0128] For each row, the script checks if the labeled category is "ignore region" (category number 0), and if so, skips the row. Otherwise, the script converts the bounding box format of VisDrone (x_min, y_min, w, h) to YOLO format (x_center, y_center, w, h), normalizes it to the range [0, 1], and subtracts 1 from the category number to adapt to YOLO's indexing method.
[0129] The converted annotation information will be formatted into the format of the YOLO annotation file and written to the corresponding labels folder.
[0130] Finally, when all the annotation files in the subdirectories are processed, the script ends and the converted YOLO annotation files are written to the labels folder in each subdirectory. These files can be used directly for training or testing the YOLO model.
[0131] Step 2: Build the model;
[0132] First, build the basic yolo11s model and optimize it on this basis:
[0133] The SCDown module is introduced into the backbone network to optimize the downsampling process of YOLOv8s. At the same time, a more lightweight feature compression module SEA is introduced to reduce the number of parameters and pay more attention to global feature information.
[0134] A lightweight cross-feature fusion module CFF, which is more suitable for small target detection, is introduced into the neck network. By reducing redundant features and enhancing the focus on important features, it can improve the accuracy and efficiency of detection in complex image tasks. At the same time, a lighter upsampling module DySample and a lighter downsampling module SCDown are added, which not only improves efficiency but also reduces the consumption of computing resources.
[0135] The loss function of the model introduces the Piou loss function, which is more suitable for small targets. The Powerable-IoU (PIoU) loss function optimizes the detection frame regression of the head, and combines the adaptive penalty factor and gradient adjustment function to improve the adaptability and stability of the model to different targets and reduce false detections. These four modules work together to enable the system to achieve high-precision target detection while maintaining lightweight and high efficiency.
[0136] Step 3: Use the constructed model to detect small targets;
[0137] The model receives an RGB image of size W×H×3 and normalizes its pixel values to the range of 0~1.
[0138] The normalized image is input to the first convolution layer, and a 3×3 convolution kernel is used to extract low-level features, and a 64-channel feature map F1 is output. The formula is:
[0139] ,
[0140] Among them, W1 is the convolution kernel weight, b1 is the bias term, σ is the activation function, I normalized is the normalized input image;
[0141] The feature map is downsampled through SCDown to reduce the spatial dimension while maintaining the channel information. The SCDown module implements the downsampling operation by combining point-by-point convolution and depth-wise separable convolution;
[0142] Point-by-point convolution: Use a 1×1 convolution kernel to reduce the number of channels of the input feature map. The calculation formula is:
[0143] ,
[0144] Among them, W pw is the point-by-point convolution kernel weight, b pw is the bias term.
[0145] Depthwise separable convolution: Use a 3×3 convolution kernel for spatial downsampling with a step size of 2. The calculation formula is:
[0146] ,
[0147] Among them, Wdw is the depth-wise separable convolution kernel weight, b dw is the bias term.
[0148] The downsampled feature map F dw The input is sent to the C3k2 module, and the feature expression capability is enhanced through convolution and residual connection. Two convolution operations are performed inside the C3k2 module:
[0149] First convolution: Use the convolution kernel to extract features. The calculation formula is:
[0150] ,
[0151] Among them, W conv1 is the weight of the first convolution, b conv1 is the bias term;
[0152] Second convolution: further enhance the features, the calculation formula is:
[0153] ,
[0154] Among them, W conv2 is the weight of the second convolution, b conv2 is the bias term;
[0155] Add the results of the two convolutions to enhance the feature expression capability:
[0156] ,
[0157] Repeat the downsampling and feature enhancement operations on F3, gradually increase the number of channels and reduce the size of the feature map to obtain a deeper feature representation:
[0158] Second downsampling: Use the SCDown module to increase the number of channels of F3 from 128 to 256, and perform spatial downsampling to output the feature map F4 with a size of W / 4×H / 4×256.
[0159] The third downsampling: Use the SCDown module to increase the number of channels of F4 from 256 to 512, and perform spatial downsampling to output the feature map F5 with a size of W / 8×H / 8×512.
[0160] Fourth downsampling: Use the SCDown module to keep the number of channels of F5 to 512, and perform spatial downsampling to output the feature map F6 with a size of W / 16×H / 16×512.
[0161] Each downsampling operation uses the SCDown module, and its specific process is the same as above, that is, first adjust the number of channels through point-by-point convolution, and then use depth-wise separable convolution for spatial downsampling;
[0162] Apply the SPPF module on the deep feature map F6 to fuse multi-scale features, enhance the model's adaptability to objects of different sizes, and output the feature map F7. The SPPF module is a native module of YOLOv8s. Its function is to extract features through pooling operations of different scales and splice these features with the original feature map to enrich the feature information:
[0163] ,
[0164] Among them, F pool1 、F pool2 、F pool3 They are the feature maps after pooling at different scales. These pooled feature maps are generated by the following operations:
[0165] Use a 5×5 pooling window to perform maximum pooling on F6 with a step size of 1 to obtain F pool1 , use a 9×9 pooling window to perform maximum pooling on F6 with a step size of 1, and get F pool2 , use a 13×13 pooling window to perform maximum pooling on F6 with a step size of 1, and get F pool3 ;
[0166] The 512-channel feature map F7 output by the SPPF module is used as the input of the feature compression module SEA. The following is the specific workflow of the feature compression module SEA:
[0167] Perform global average pooling on F7 to obtain the channel feature vector v, the formula is:
[0168] ,
[0169] Where F7 is a three-dimensional feature map with a shape of H×W×C, where H is the height, W is the width, and C is the number of channels.
[0170] i and j are the indices in the height and width directions of the feature map, respectively. The symbol: means taking the values of all channels.
[0171] It means to sum all spatial positions of feature map F7 (i.e. all combinations of i and j). For each channel c, the sum of the values of the channel at all spatial positions is calculated.
[0172] Reduce the feature dimension to 32:
[0173] ,
[0174] Among them, ReLU represents the activation function, and represent the first weight and the first bias respectively.
[0175] Restore the dimension to 512:
[0176] ,
[0177] in, and Represented as the second weight and the second bias respectively.
[0178] Normalized by sigmoid function, the channel attention weight vector is obtained :
[0179] ,
[0180] in, is the sigmoid function, The final calculated value of each element will be between 0 and 1.
[0181] The channel attention weight vector Multiply element-by-element with the original feature map F7 to obtain the output feature map F of the feature compression module SEA sea :
[0182] ,
[0183] Among them, ⊙ represents element-by-element multiplication, F sea Rich in global feature information and with fewer parameters.
[0184] The feature map F output by the backbone network is rich in global feature information sea As input to the neck network.
[0185] The upsampling module DySample is used to upsample the feature map to improve the resolution of the feature map for subsequent feature fusion and target detection.
[0186] Offset calculation: for the input feature map F sea Perform convolution operation to obtain the offset feature map O:
[0187] ,
[0188] in, is the convolution kernel weight, is the bias term.
[0189] Generate a grid: Generate a sampling grid based on the offset feature map O and the preset initial position:
[0190] ,
[0191] Among them, init_pos is the preset initial position, and O(i,j) is the value of the offset feature map at position (i,j).
[0192] According to the position of the sampling grid, the weight of the bilinear interpolation is calculated. The weight calculation formula is:
[0193] ,
[0194] Among them, (m,n) is the position in the original feature map, and (i,j) is the position in the upsampled feature map.
[0195] Bilinear interpolation: Using the generated sampling grid, the input feature map F sea Perform bilinear interpolation to obtain the upsampled feature map F up :
[0196] ,
[0197] in, is the weight of bilinear interpolation, which is determined by the position of the sampling grid.
[0198] The CFF module receives the upsampled feature map F up and the shallower feature map F in the backbone network shallow As input;
[0199] The input feature map is reorganized and the number of channels is adjusted through convolution operations, including:
[0200] For the feature map F up Perform convolution operation and adjust the number of channels and F shallow same;
[0201] Using the feature map after feature reorganization, calculate the weight of each feature map: perform global average pooling on each feature map to obtain the channel feature vector:
[0202] ,
[0203] ,
[0204] Among them, v gloabl_up and v gloabl_shallow Represents F up and F shallow The channel feature vector of ;
[0205] Use the fully connected layer to transform the feature vector and generate the initial weight value:
[0206] ,
[0207] ,
[0208] in, and Represents F up The initial weight value and F shallow The initial weight value of is the weight matrix of the fully connected layer, is the bias term;
[0209] Assign weights to different feature maps according to the proportion of initial weight values:
[0210] ,
[0211] ,
[0212] Among them, α and β are F up The weight and F shallow The weight of
[0213] Get the fused feature map F fuse :
[0214] ,
[0215] The fused feature map F fuse Input to the C3k2 module, and further enhance the feature expression ability through convolution and residual connection to obtain the feature map F ;
[0216] Repeat the upsampling, feature fusion and feature enhancement processes to process feature maps of different scales respectively, and finally generate a multi-scale fused feature map :
[0217] ,
[0218] in, is the output of F4 in the backbone network through the neck network, F shallow Output through the neck network
[0219] Multi-scale fusion feature map output by the neck network As the input of the detection head;
[0220] For each scale feature map , i is 1, 2, 3, and the feature map is processed by convolution operation Adjust the number of channels:
[0221] ,
[0222] Among them, Wdet is the convolution kernel weight, b det is the bias term, F det_i is the feature map after processing.
[0223] Detection box parameter prediction: For each processed feature map F det_i , predict the detection box parameters at each location, including the coordinates, category probability and confidence of the bounding box:
[0224] ,
[0225] Among them, B represents the predicted detection box set, including the bounding box coordinates (x, y, w, h), category probability pc and confidence c; Predict represents the prediction function of the original yolo11s model.
[0226] The Powerful-IoU (PIoU) loss function is also introduced during the training process to optimize the regression of the detection box, improve the adaptability and stability of the model to different targets, and reduce false detections.
[0227] PIoU loss function The formula is:
[0228] ,
[0229] Among them, IoU is the intersection over union ratio of the predicted box and the true box, a and b are hyperparameters that will be automatically adjusted during the learning process, and ϵ is the preset minimum value to avoid division by zero.
[0230] Step 4: Use the trained target detection model to detect road traffic conditions.
[0231] Figure 6 , Figure 7 , is the final prediction result display, Figure 6 It can be seen that in the case of many small targets on the road in the image, objects including pedestrians, cars, and motorcycles can be accurately identified. Figure 7 It can be seen that even when blocked by trees, the cars on both sides of the road can still be accurately identified.
[0232] In the process of detecting objects in images, it is necessary not only to correctly predict the category of the object in the image, but also to predict the specific location of the object. The result evaluation can be divided into two parts, one is the prediction index of the prediction box, and the other is the classification prediction index.
[0233] 1. Intersection over Union (IoU)
[0234] IoU is used to measure the overlap between the predicted bounding box and the true bounding box. It is an important part of calculating the average precision mAP. The calculation formula is:
[0235]
[0236] Where Z1 represents the union area of the predicted box and the true box, and Z2 represents the intersection area of the predicted box and the true box;
[0237] The higher the IoU value, the greater the overlap between the predicted box and the true box, and the higher the regional accuracy of the object. Usually, the IoU threshold is set to 0.5, that is, when IoU ≥ 0.5, the prediction is considered to be a positive example.
[0238] 2. Confusion Matrix:
[0239] Used to show the relationship between the model prediction results and the true label, usually divided into four categories: true positive, false positive, true negative and false negative.
[0240] Through the confusion matrix, the indicator precision Precision can be calculated. Precision measures the proportion of samples predicted by the model as positive samples that are actually positive samples. The calculation formula is:
[0241] ,
[0242] Among them, TP represents the number of positive samples predicted correctly, and FP represents the number of positive samples predicted incorrectly.
[0243] With the prediction indicators of the prediction box and the indicators of classification prediction, the two can be combined into the evaluation indicators of the model:
[0244] mAP@0.5 means that when the IoU threshold is 0.5, the average precision of each category is calculated. The higher the mAP@0.5 value, the better the overall detection performance of the model.
[0245] Indicators such as IoU, confusion matrix, precision and mAP@0.5 together constitute a comprehensive evaluation system for the performance of the target detection model. In order to evaluate the importance of the improved module in this example and its contribution to the detection algorithm. Table 1 shows the performance comparison between the yolo11s detection algorithm and the target detection model in this example.
[0246] Table 1
[0247] Model mAP@0.5(%) Param(M) Size(MB) YOLO11s 33.3 9.458 18.29 Improve yolo11s 34.3 5.921 11.65
[0248] In the table, mAP@0.5 represents the average accuracy of each category when the IoU threshold is 0.5, Param represents the parameter quantity, M represents million, and the unit of parameter quantity is one million; Size represents the model size, and MB represents megabytes, which is the unit of model size. From the experimental results in Table 1 above, it can be seen that the results of the improved model are higher than those of the existing technical model. While greatly compressing the model parameters and model size, it can also improve the model detection accuracy.
[0249] The present invention provides a lightweight small target detection method based on deep learning. There are many methods and ways to implement the technical solution. The above is only a preferred implementation of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention. All components not specified in this embodiment can be implemented by existing technologies.
Claims
1. A lightweight small target detection method based on deep learning, characterized in that: The following steps are involved: Step 1, data preprocessing: Use a script to process the public dataset visdrone2019 to obtain the training set, validation set, and test set; Step 2: Build an improved yolo11s model and input the training set for training; Step 3: Use the trained yolo11s model to detect small targets.
2. The method according to claim 1, characterized in that: Step 1 includes: after the script runs, first parse the command line parameters, obtain the path of the public dataset visdrone2019, and convert the public dataset visdrone2019 into a Path object; Then, the script will process the three subdirectories in the public dataset visdrone2019 in turn: the training set VisDrone2019-DET-train, the validation set VisDrone2019-DET-val, and the test set VisDrone2019-DET-test-dev; the training set is used for model training, the validation set is used to adjust the model effect, and the test set is used for the final evaluation of the model; For each subdirectory, the script creates a folder named labels to store the converted YOLO annotation files; labels means labels; Next, the script will traverse the annotations folder in the subdirectory, read the .txt annotation files in it, and parse the annotation information line by line. For each line X1, the script will check whether the annotated category is an ignored area. If so, the line X1 will be skipped. The ignored area refers to the category number 0; otherwise, the script will convert the bounding box format (x1, y1, x2, y2) of the dataset visdrone2019 to the YOLO format (x_center, y_center, w, h), and normalize it to the range of [0, 1], and subtract 1 from the category number to adapt to the YOLO indexing method, where x1, y1 represent the horizontal and vertical coordinates of the vertex in the lower left corner of the image, respectively, and x2, y2 represent the horizontal and vertical coordinates of the vertex in the upper right corner, respectively; x_center, y_center represent the horizontal and vertical coordinates of the center of the image, respectively; w represents the width of the image, and h represents the height of the image; The converted annotation information will be formatted into the format of the YOLO annotation file and written into the corresponding labels folder; Finally, when all the annotation files in the subdirectories have been processed, the script ends and writes the converted YOLO annotation files to the labels folder in each subdirectory.
3. The method according to claim 2, characterized in that Step 2 includes: Optimizing on the basic yolo11s model: A lightweight selective channel downsampling module SCDown is added to the backbone network of the YOLO11s model to optimize the downsampling process of YOLOv8s, and a feature compression module SEA is added to reduce the number of parameters. A lightweight cross-feature fusion module CFF is added to the neck network of the yolo11s model to reduce redundant features and enhance the focus on important features; at the same time, an upsampling module DySample and a lightweight selective channel downsampling module SCDown are added; Add Powerful-IoU loss function during model training , used to optimize the regression of the detection box.
4. The method according to claim 3, characterized in that: In step 2, the Powerful-IoU loss function The formula is: , Among them, IoU is the intersection over union ratio of the predicted box and the true box, a and b are hyperparameters, and ϵ is the preset minimum value to avoid division by zero.
5. The method according to claim 4, characterized in that Step 3 includes: the trained yolo11s model receives an RGB image of size B×L×3, normalizes the image pixel values to the range of 0 to 1, where B represents the width of the image and L represents the height of the image; The normalized image is input to the first convolution layer of the yolo11s model, and a 3×3 convolution kernel is used to extract low-level features and output a 64-channel feature map F1: , Among them, W1 is the convolution kernel weight, b1 is the bias term, σ is the activation function, I normalized is the normalized input image.
6. The method according to claim 5, characterized in that Step 3 also includes: performing the first downsampling of the feature map through the SCDown module to reduce the spatial dimension while maintaining the channel information. The SCDown module implements the downsampling operation by combining point-by-point convolution and depth-wise separable convolution: Point-by-point convolution: Use a 1×1 convolution kernel to reduce the number of channels of the input feature map. The calculation formula is: , Among them, W pw is the point-by-point convolution kernel weight, b pw is the bias term, It is the feature map after point-by-point convolution; Depthwise separable convolution: Use a 3×3 convolution kernel for spatial downsampling with a step size of 2. The calculation formula is: , Among them, W dw is the depth-wise separable convolution kernel weight, b dw is the bias term, It is the feature map after downsampling; Feature map after downsampling The input is sent to the C3k2 module of the yolo11s model, and the feature expression capability is enhanced through convolution and residual connection. The following two convolution operations are performed inside the C3k2 module: First convolution: use convolution kernel to extract features , the calculation formula is: , Among them, W conv1 is the weight of the first convolution, b conv1 is the bias term; Second convolution: get enhanced features : , Among them, W conv2 is the weight of the second convolution, b conv2 is the bias term; Add the results of the two convolutions to get feature F3: , Repeat the downsampling and feature enhancement operations on F3, gradually increasing the number of channels and reducing the size of the feature map, including: Second downsampling: Use the SCDown module to increase the number of channels of feature F3 from 128 to 256, and perform spatial downsampling to output feature map F4 with a size of W / 4×H / 4×256; The third downsampling: Use the SCDown module to increase the number of channels of feature map F4 from 256 to 512, and perform spatial downsampling to output feature map F shallow , size is W / 8×H / 8×512; The fourth downsampling: Use the SCDown module to downsample the feature map F shallow The number of channels is kept at 512, and spatial downsampling is performed to output a feature map F6 with a size of W / 16×H / 16×512.
7. The method according to claim 6, characterized in that Step 3 also includes: applying the SPPF module on the deep feature map F6, fusing multi-scale features, enhancing the model's adaptability to objects of different sizes, and outputting the feature map F7: , Concat is a connection operation, F pool1 、F pool2 、F pool3 It is the feature map after pooling at different scales, generated by the following operations: Use a 5×5 pooling window to perform maximum pooling on F6 with a step size of 1 to obtain F pool1 , use a 9×9 pooling window to perform maximum pooling on F6 with a step size of 1, and get F pool2 , use a 13×13 pooling window to perform maximum pooling on F6 with a step size of 1, and get F pool3 ; The 512-channel feature map F7 output by the SPPF module is used as the input of the feature compression module SEA, which performs the following operations: Perform global average pooling on the feature map F7 to obtain the channel feature vector v: , Among them, the feature map F7 is a three-dimensional feature map with a size of H×W×C, where H is the height, W is the width, and C is the number of channels; i and j are the indexes in the height and width directions of the feature map, respectively, and the symbol : means taking the values of all channels; It means summing all spatial positions of feature map F7; The ReLU activation function is used to reduce the feature dimension to 32 and obtain the feature map : , in, and denote the first weight and the first bias respectively; Restore the dimension to 512 and get the feature map : , in, and denote the second weight and the second bias respectively; Normalized by sigmoid function, the channel attention weight vector is obtained : , in, is the sigmoid function; The channel attention weight vector Multiply element-by-element with the original feature map F7 to obtain the output feature map F of the feature compression module SEA sea : , Among them, ⊙ represents element-by-element multiplication; Feature map F output by the backbone network sea As input to the neck network; The upsampling module DySample is used to upsample the feature map, and specifically performs the following operations: Offset calculation: for the input feature map F sea Perform convolution operation to obtain the offset feature map O: , in, is the convolution kernel weight, is the bias term; Generate a grid: Generate a sampling grid based on the offset feature map O and the preset initial position: , Among them, grid is the generated two-dimensional sampling grid, is the preset initial position, O(i,j) is the value of the offset feature map at position (i,j); According to the position of the sampling grid, the weight of the bilinear interpolation is calculated: , Among them, weight(i,j,m,n) represents the weight between the target position (i,j) and the pixel position (m,n) in the original feature map, and m and n are the height index and width index of the original feature map respectively; Bilinear interpolation: First initialize a feature map F up , and then use the generated sampling grid to input feature map F sea Perform bilinear interpolation to obtain the upsampled feature map F up : , where F up (i,j) represents the feature map F up The pixel value at position (i,j) in ; The CFF module receives the upsampled feature map F up and the shallower feature map F in the backbone network shallow as input.
8. The method according to claim 7, characterized in that Step 3 also includes: reorganizing the input feature map and adjusting the number of channels through convolution operations, specifically including: For the feature map F up Perform convolution operation and adjust the number of channels and F shallow same; Using the feature map after feature reorganization, calculate the weight of each feature map: perform global average pooling on each feature map to obtain the channel feature vector: , , in, and Represents F up The channel eigenvector and F shallow The channel feature vector of ; Use the fully connected layer to transform the feature vector and generate the initial weight value: , , in, and Represents F up The initial weight value and F shallow The initial weight value of is the weight matrix of the fully connected layer, is the bias term; Assign weights to different feature maps according to the proportion of initial weight values: , , Among them, α and β are F up The weight and F shallow The weight of Get the fused feature map F fuse : , The fused feature map F fuse Input to the C3k2 module, and further enhance the feature expression ability through convolution and residual connection to obtain the feature map F ; Repeat the upsampling, feature fusion and feature enhancement processes to process feature maps of different scales respectively, and finally generate a multi-scale fused feature map : , in, is the output of F4 in the backbone network through the neck network, F shallow Output through the neck network; Multi-scale fusion feature map output by the neck network As the input of the detection head; For each scale feature map , i is 1, 2, 3, and the feature map is processed by convolution operation Adjust the number of channels: , Among them, W det is the convolution kernel weight, b det is the bias term, F det_i is the feature map after processing; Detection box parameter prediction: For each processed feature map F det_i , predict the detection box parameters at each location, including the coordinates, category probability and confidence of the bounding box: , Among them, B represents the predicted detection box set, including the bounding box coordinates (x, y, w, h), category probability pc and confidence c; Predict represents the prediction function of the original yolo11s model.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.
10. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 8 are executed.
Citation Information
Cited By
Multi-modal ship target individual identification method and system
CN120472250A