Real-time lightweight small target detection method

By improving the structure of the YOLOv8s model, the problem of difficult balance between accuracy and lightweight design of remote sensing image object detection algorithm in the prior art is solved, and efficient, fast and suitable target detection for complex traffic environments is achieved.

CN119942091APending Publication Date: 2025-05-06NANJING UNIV OF INFORMATION SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510389441.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing remote sensing image object detection algorithms are difficult to balance between accuracy and lightweight designs, making it difficult to deploy on hardware-limited mobile or edge devices, especially in dense scenarios and complex traffic environments.

Method used

By improving the structure of the YOLOv8s model, including using the DWR_C2F module in the backbone network, the C_BiFPN module in the neck network, and the Detect_LSCD detection head in the head network, the model is optimized to improve detection accuracy and reduce computational complexity.

Benefits of technology

It achieves the realization that while maintaining high detection accuracy, the calculation complexity and number of parameters are greatly reduced, the running speed and adaptability of the model are improved, and the targets in small targets and complex scenarios can be better detected.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942091A_ABST
    Figure CN119942091A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time lightweight small target detection method, and belongs to the technical field of target detection, and the method comprises the steps: firstly obtaining a traffic road condition image; then a basic YOLOv8s model is constructed and optimized, specifically, a DWRC2F module is constructed in a backbone network, a CBiFPN module is constructed in a neck network, and a DetectLSCD detection head is constructed in a head network; performing iterative training on the optimized YOLOv8s model by adopting a training set to obtain a target detection model; and finally, detecting an urban traffic target by using the trained target detection model. According to the method, the detection precision is remarkably improved, and particularly, the method is excellent in detection of small targets and complex scenes; the calculation complexity of the improved YOLOv8s model is greatly reduced, and the number of parameters is remarkably reduced, so that the operation speed and efficiency of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a target detection method, in particular to a real-time lightweight small target detection method, which belongs to the technical field of computer vision and deep learning. Background Art

[0002] With the acceleration of urbanization and the rapid development of the transportation industry, the number and types of targets on urban roads are increasing, leading to aggravated traffic congestion and increasingly serious safety issues. In intelligent transportation systems, accurate and efficient detection of targets in dense remote sensing images is becoming increasingly important for applications such as traffic monitoring, autonomous driving, and parking.

[0003] Traditional remote sensing image target detection algorithms mainly rely on manual feature extraction and pattern matching, which have limitations such as low detection accuracy, slow processing speed, high computational cost, and insufficient adaptability to dense scenes. Although the latest advances in deep learning have significantly improved the target detection performance of remote sensing images, existing remote sensing image detection algorithms often have difficulty balancing accuracy and lightweight design, and are difficult to deploy on mobile or edge devices with limited hardware. Summary of the invention

[0004] Purpose of the invention: In view of the above problems, the purpose of the present invention is to provide a real-time lightweight small target detection method, which improves the accuracy, speed and adaptability of urban traffic target detection by improving the existing deep learning algorithm, so that it can better meet the needs of intelligent transportation systems.

[0005] Technical solution: A real-time lightweight small target detection method of the present invention comprises the following steps: Step 1: Obtain traffic condition images and preprocess them, use the preprocessed images to construct a data set, and divide them into a training set and a test set in proportion; Step 2: Build the basic YOLOv8s model, including the backbone network, neck network, and head network; Step 3: Optimize the YOLOv8s model structure, specifically: Construct the DWR_C2F module in the backbone network to replace the original C2F module; construct the C_BiFPN module in the neck network to replace the original feature pyramid network; construct the Detect_LSCD detection head in the head network to replace the original detection head; Step 4, use the training set to iteratively train the optimized YOLOv8s model to obtain the target detection model; Step 5: Use the trained target detection model to detect urban traffic targets.

[0006] Furthermore, the working process of the backbone network in the optimized YOLOv8s model includes: The image input to the optimized YOLOv8s model is denoted as I. Image I first passes through a 3×3 convolutional layer to extract the initial features and obtain a new feature map F1. The formula is: , In the formula, Conv3×3 represents a standard 3×3 convolutional layer; Then the feature map F1 is extracted through the DWR_C2F module: First, a standard 3×3 convolutional layer is used to extract the initial features of the feature map F1 to generate the feature map F conv3×3 , expressed as: , Then, three dilated convolutional layers are used, with different dilation rates set to capture features of different scales and generate feature maps. , expressed as: , Where r represents the expansion rate, which can be 1, 3, or 5; D-r3x3DConv represents the convolution operation with different expansion rates according to the r value, thus obtaining , and ; Next, the feature maps output by the three dilated convolutional layers , and Concatenate on the channel dimension to get the feature map , and then adjust the number of channels through a 1×1 convolution layer to obtain the feature map , expressed as: , , In the formula, Concatenate represents the concatenation operation, and Conv1x1 represents the convolution operation with a convolution kernel size of 1; The feature map obtained After adding to image I, we get feature map F DWR , as the output item of DWR_C2F module; The feature map processed by the DWR_C2F module is extracted through three residual blocks, and the feature representation is gradually enhanced to obtain the feature map F2, which is expressed as: , Where ResBlock represents the residual operation; In the process of feature extraction, downsampling operations are performed periodically, and convolutional layers or maximum pooling layers with a stride of 2 are used to reduce the resolution of the feature map while increasing the number of channels to obtain the feature map F3, which is expressed as: , Where, DownSample represents the downsampling operation; Then, through the downsampling operation, three feature maps of different scales are generated, which are denoted as F P3 、F P4 and F P5 , respectively expressed as: , , .

[0007] Furthermore, the working process of the C_BiFPN module in the optimized YOLOv8s model includes: The C_BiFPN module receives three feature maps of different scales F from the backbone network. P3 、F P4 and F P5 As input, a top-down feature pass is first performed: For the feature map F P5 Perform an upsampling operation to increase its resolution to the same level as the feature map F P4 Same as feature map F P4 Add element by element to generate feature map F P4_top_down , expressed as: , In the formula, Upsample indicates the upsampling operation through the CARAFE module; For the feature map F P4_top_down Upsample to make its resolution the same as the feature map F P3 Same as feature map F P3 Add element by element to generate feature map F P3_top_down , expressed as: , Then perform bottom-up feature transfer: For the feature map F P3_top_down Perform downsampling to make its resolution consistent with the feature map F P4_top_down Same as feature map F P4_top_down Add element by element to generate feature map F P4_bottom_up , expressed as: , Where, Downsample represents the downsampling operation; For the feature map F P4_bottom_up Downsample to make its resolution the same as the feature map F P5 Same as feature map FP5 Add element by element to generate feature map F P5_bottom_up , expressed as: , For the feature map F P4_bottom_up Upsample to make its feature rate consistent with the feature map F P3_top_down Same as feature map F P3_top_down Add element by element to generate feature map F P3_bottom_up , expressed as: , Because the feature map F P5 In the top-down feature transfer path, it is the top-level feature map, so there is the following relationship: , For each level, the feature maps from different paths are concatenated and then Convolution reduces the number of channels and fuses features. Finally, the fused features are weighted summed through learnable weights to generate a feature map. , expressed as: , In the formula, α n and It is the weight parameter obtained through training and learning. n represents P3, P4, and P5, which are used to represent different feature levels. Conv represents the convolution layer, including convolution operation, group normalization, and activation function.

[0008] Furthermore, the working process of the Detect_LSCD detection head in the optimized YOLOv8s model includes: The feature map input to the Detect_LSCD detection head is , the feature map First, feature extraction is performed through the convolution layer to generate the feature map F conv , the formula is: , Where Conv represents the convolution layer, including convolution operation, group normalization and activation function; Detect_LSCD introduces separable convolution and shared convolution, where separable convolution decomposes the standard convolution into depth convolution and point-by-point convolution. First, each input channel is convolved separately through depth convolution to extract spatial features and generate feature map F. dw , the feature map F conv The shape of [C in ,H,W], the depth convolution uses C in 3×3 convolution kernel, then the feature map F dw The shape is still [Cin ,H,W], feature map F dw It is expressed as: , Where DepthwiseConv represents the deep convolution operation; C in , H, W represent the number of channels, height, and width of the feature map, respectively; Then, a point-by-point convolution operation is performed: by using a 1×1 convolution kernel to perform convolution on the channel dimension, the features of different channels are integrated, and the number of output channels of the point-by-point convolution is recorded as C out , then the generated feature map F sep The shape is [C out ,H,W], feature map F sep Expressed as: , In the formula, PointwiseConv represents the point-by-point convolution operation; Then the feature map F sep Perform a shared convolution operation: Use the shared convolution kernel to perform a convolution operation on the feature map to generate a feature map F share , expressed as: , In the formula, SharedConv represents the shared convolution operation, and the weights are shared between different detection tasks or feature maps; Finally, bounding box regression and category probability generation are performed, and the generated bounding box position information and the corresponding category probability are spliced ​​together to generate the final detection result.

[0009] Furthermore, the upsampling operation performed by the CARAFE module includes: The characteristic graph F input into the CARAFE module P3 、F P4 、F P5 Unified as F input , the feature map F input The features are reorganized through a convolution layer and a transposed convolution layer to obtain the reorganized feature map F reorg , expressed as: , Where ConvTranspose represents the deconvolution operation; Then, using the reorganized feature map F reorg Generate upsampling convolution kernel K and transform the feature map F reorg The convolution kernel is used to upsample and obtain the feature map Upsample (F input ), expressed as: .

[0010] Furthermore, bounding box regression predicts the bounding box position and size of the target through convolutional layers and learnable scaling parameters. The calculation formula is: , In the formula, Conv box is a convolutional layer for bounding box regression, and scale is a learnable scaling parameter; Category probability generation predicts the probability of the category to which the target belongs through the convolution layer and the Sigmoid activation function. The calculation formula is: , In the formula, Conv cls represents the convolutional layer for class probability generation, and σ represents the Sigmoid activation function.

[0011] Furthermore, the loss function of the optimized YOLOv8s model is the PIoU loss function, and the calculation formula is: , In the formula, IoU represents the ratio of the intersection area to the union area, and P represents the adaptive penalty factor.

[0012] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: 1. Improve detection accuracy: The present invention significantly improves detection accuracy, especially in the detection of small targets and complex scenes; 2. Reduce computational complexity: The computational complexity of the present invention is greatly reduced, and the number of parameters is significantly reduced, thereby improving the running speed and efficiency of the model; 3. Enhanced real-time performance: The present invention can achieve rapid target detection while maintaining high detection accuracy, meeting real-time requirements, and is suitable for application scenarios with high real-time requirements such as urban traffic monitoring and autonomous driving; 4. Strong adaptability: The present invention has good adaptability to targets of different scales, can effectively detect dense small targets and obscured targets, and is suitable for complex remote sensing image scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a flowchart of a real-time lightweight small target detection method; Figure 2 It is a schematic diagram of the DWR_C2F structure; Figure 3 Schematic diagram of separable convolution in Detect_LSCD; Figure 4 This is the target detection result diagram of the road intersection in the example; Figure 5This is the target detection result image on the road in the example. DETAILED DESCRIPTION

[0014] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0015] The flowchart of a real-time lightweight small target detection method described in this embodiment is as follows: Figure 1 As shown, the method comprises the following steps: Step 1: Obtain traffic condition images and preprocess them, use the preprocessed images to build a data set, and divide them into a training set and a test set in proportion.

[0016] In one example, a traffic image is obtained, and the image contains multiple targets in urban traffic, such as vehicles, pedestrians, etc. The size of the input image may not be consistent with the input size required by the target detection model, so the image size needs to be adjusted. For example, if the target detection model requires the input image size to be 640×640, and the actual image size is 1280×720, the width of the original image needs to be scaled to 640 pixels, and the height needs to be scaled to 360 pixels accordingly. Methods such as bilinear interpolation can be used to adjust the image size to maintain the clarity of the image. Pixels are filled around the scaled image to make its size 640×640, and the filled pixel value is fixed to 114. The image pixel values ​​are then mapped from the range of 0-255 to the range of 0-1 for normalization.

[0017] Specifically, for each pixel value p, the new pixel value p′ obtained after mapping is calculated as: , the pixel value can be linearly mapped from the range of 0-255 to the range of 0-1. Normalization helps to accelerate the convergence of the network and improve the numerical stability of the target detection model.

[0018] The preprocessed images are used to construct a data set and divided into a training set and a test set in a ratio of 6:4. The training set is used to train and optimize the YOLOv8s model, and the test set is used to verify and test the target detection model obtained after training.

[0019] Step 2: Build the basic YOLOv8s model, including the backbone network, neck network, and head network.

[0020] Step 3: Optimize the YOLOv8s model structure, specifically: Construct the DWR_C2F module in the backbone network to replace the original C2F module; construct the C_BiFPN module in the neck network to replace the original feature pyramid network; construct the Detect_LSCD detection head in the head network to replace the original detection head.

[0021] The YOLOv8s model is a single-stage target detection model. It has been widely used in the field of target detection for its high efficiency and accuracy. It directly performs convolution and pooling operations on the image, generates candidate boxes, classifies objects, and regresses bounding boxes, thereby seeking a balance between detection speed and accuracy. The basic YOLOv8s model includes a backbone network, a neck network, and a head network. In this example, the YOLOv8s model is improved and optimized. The backbone network is used to extract the initial features of the image. After replacing the C2F module with the DWR_C2F module in the backbone network, the receptive field is enlarged by dilated convolution, and multi-scale features are extracted to provide rich context information for subsequent modules, especially for the detection of small targets. It can enhance the feature extraction ability. The C_BiFPN module is used to replace the original feature pyramid network as the new neck network of the model, which is responsible for the fusion and enhancement of multi-scale features. The complementarity of features is enhanced through bidirectional connections and learnable weighting mechanisms. The Detect_LSCD detection head is used to replace the original detection head, which is responsible for the final detection task. The number of parameters is reduced by sharing convolution kernels, and the detection speed is improved while maintaining the detection performance.

[0022] Furthermore, the working process of the backbone network in the optimized YOLOv8s model includes: The image input to the optimized YOLOv8s model is denoted as I. Image I first passes through a 3×3 convolutional layer to extract the initial features and obtain a new feature map F1. The formula is: , In the formula, Conv3×3 represents a standard 3×3 convolutional layer; Then the feature map F1 is extracted by the DWR_C2F module, as shown in Figure 2 As shown: First, a standard 3×3 convolutional layer is used to extract the initial features of the feature map F1 to generate the feature map F conv3×3 , expressed as: , Then, three dilated convolutional layers are used, with different dilation rates set to capture features of different scales and generate feature maps. , expressed as: , In the formula, r represents the expansion rate, which can be 1, 3, or 5. D-r3x3DConv represents the convolution operation with different expansion rates according to the r value. When r is 1, Figure 2 In the figure, it is represented as 3x3DConv, thus we get , and ; Next, the feature maps output by the three dilated convolutional layers , and Concatenate on the channel dimension to get the feature map , and then adjust the number of channels through a 1×1 convolution layer to obtain the feature map , expressed as: , , In the formula, Concatenate represents the concatenation operation, and Conv1x1 represents the convolution operation with a convolution kernel size of 1; The feature map obtained After adding to image I, we get feature map F DWR , as the output item of DWR_C2F module; The feature map processed by the DWR_C2F module is extracted through three residual blocks, and the feature representation is gradually enhanced to obtain the feature map F2, which is expressed as: , Where ResBlock represents the residual operation; In the process of feature extraction, downsampling operations are performed periodically, and convolutional layers or maximum pooling layers with a stride of 2 are used to reduce the resolution of the feature map while increasing the number of channels to obtain the feature map F3, which is expressed as: , Where, DownSample represents the downsampling operation; Then, through the downsampling operation, three feature maps of different scales are generated, which are denoted as F P3 、F P4 and F P5 , respectively expressed as: , , .

[0023] In this example, the dilated residual module (DWR) is improved to construct the DWR_C2F module. The residual connection has been proven to be an effective network structure in traditional deep learning and convolutional neural networks, which helps to learn more complex features and solve problems such as gradient disappearance and gradient explosion. The present invention increases the receptive field by inserting holes in the convolution kernel, thereby allowing the capture of a wider range of contextual information. Specifically, a multi-branch feature extraction strategy is adopted to capture multi-scale contextual information through convolution operations with different expansion rates. This design not only enhances the network's ability to extract detailed features, but is also particularly suitable for detecting small targets in remote sensing images. In addition, the feature representation capability is further enhanced through the combination of skip connections and attention weights, so that the model improves the accuracy and robustness of detection while maintaining computational efficiency.

[0024] Furthermore, the working process of the C_BiFPN module in the optimized YOLOv8s model includes: The C_BiFPN module receives three feature maps of different scales F from the backbone network. P3 、F P4 and F P5 As input, these feature maps have different resolutions and numbers of channels. The larger the number, the lower the resolution of the feature map but the richer the semantic information. First, the top-down feature transfer is performed: For the feature map F P5 Perform an upsampling operation to increase its resolution to the same level as the feature map F P4 Same as feature map F P4 Add element by element to generate feature map F P4_top_down , expressed as: , In the formula, Upsample indicates the upsampling operation through the CARAFE module; For the feature map F P4_top_down Upsample to make its resolution the same as the feature map F P3 Same as feature map F P3 Add element by element to generate feature map F P3_top_down , expressed as: , Then perform bottom-up feature transfer: For the feature map F P3_top_down Perform downsampling to make its resolution consistent with the feature map F P4_top_down Same as feature map F P4_top_down Add element by element to generate feature map F P4_bottom_up , expressed as: , Where, Downsample represents the downsampling operation; For the feature map F P4_bottom_up Downsample to make its resolution the same as the feature map F P5 Same as feature map F P5 Add element by element to generate feature map F P5_bottom_up , expressed as: , For the feature map F P4_bottom_up Upsample to make its feature rate consistent with the feature map F P3_top_down Same as feature map F P3_top_down Add element by element to generate feature map F P3_bottom_up , expressed as: , Because the feature map F P5 In the top-down feature transfer path, the top-level feature map is not fused with any higher-level feature maps, so the following relationship exists: , For each level, the feature maps from different paths are concatenated, and then the number of channels is reduced and the features are fused through 1×1 convolution. Finally, the fused features are weighted summed through learnable weights to generate a feature map. , expressed as: , In the formula, α n and It is the weight parameter obtained through training and learning. n represents P3, P4, and P5, which are used to represent different feature levels. Conv represents the convolution layer, including convolution operation, group normalization, and activation function.

[0025] This example improves the bidirectional feature pyramid network (BiFPN) and constructs the C_BiFPN module. The traditional feature pyramid network and path aggregation network have some limitations in feature fusion, especially when dealing with multi-scale features, features are lost or insufficiently fused, which affects the recognition and positioning performance of remote sensing images with small, blurred or complex backgrounds. In this example, a new feature map is created in the C_BiFPN module by using bidirectional connections to fuse multiple feature layers. Specifically, the complementary information between features is enhanced through top-down and bottom-up propagation paths, and a learnable weighting mechanism is used to automatically assign different fusion weights according to the importance of the features. This design not only improves the overall feature integration and performance, but also enables the model to improve the accuracy and robustness of detection while maintaining computational efficiency.

[0026] Furthermore, the working process of the Detect_LSCD detection head in the optimized YOLOv8s model includes: The feature map input to the Detect_LSCD detection head is , the feature map First, feature extraction is performed through the convolution layer to generate the feature map F conv , the formula is: , Where Conv represents the convolution layer, including convolution operation, group normalization and activation function; Detect_LSCD introduces separable convolution and shared convolution, where separable convolution decomposes the standard convolution into depth convolution and point-by-point convolution, such as Figure 3 As shown in Figure 2, each input channel is first convolved separately through deep convolution to extract spatial features and generate feature maps F. dw , the feature map F conv The shape of [C in ,H,W], the depth convolution uses C in 3×3 convolution kernel, then the feature map F dw The shape is still [C in ,H,W], feature map F dw It is expressed as: , Where DepthwiseConv represents the deep convolution operation; C in , H, W represent the number of channels, height, and width of the feature map, respectively; Then, a point-by-point convolution operation is performed: by using a 1×1 convolution kernel to perform convolution on the channel dimension, the features of different channels are integrated, and the number of output channels of the point-by-point convolution is recorded as C out , then the generated feature map F sep The shape is [C out ,H,W], feature map F sep It is expressed as: , In the formula, PointwiseConv represents the point-by-point convolution operation; Then the feature map F sep Perform a shared convolution operation: Use the shared convolution kernel to perform a convolution operation on the feature map to generate a feature map F share , expressed as: , In the formula, SharedConv represents the shared convolution operation, and the weights are shared between different detection tasks or feature maps; Finally, bounding box regression and category probability generation are performed, and the generated bounding box position information and the corresponding category probability are spliced ​​together to generate the final detection result.

[0027] Traditional object detection models usually adopt decoupled classification and regression heads, using independent convolutional layers for each task. While this design allows specialized feature extraction, it introduces a lot of parameter redundancy and computational overhead, which is particularly challenging for remote sensing applications that require real-time processing. In this example, the Detect_LSCD detection head is constructed by improving the lightweight shared convolutional detection head (LSCD), which significantly reduces the scale and computational cost of the object detection model by sharing convolution kernels and reducing the number of parameters. Specifically, the Detect_LSCD detection head introduces parameter sharing between detection tasks while maintaining detection accuracy. Multiple detection tasks share common convolutional layers for initial feature processing, significantly reducing parameter redundancy. Only the last few layers are dedicated to classification and regression, which allows task-specific feature refinement while maintaining efficiency. In addition, depthwise separable convolutions are used to further reduce computational complexity while maintaining strong detection performance.

[0028] Furthermore, the upsampling operation performed by the CARAFE module includes: The characteristic graph F input into the CARAFE module P3 、F P4 、F P5 Unified as F input , the feature map F input The features are reorganized through a convolution layer and a transposed convolution layer to obtain the reorganized feature map F reorg , expressed as: , Where ConvTranspose represents the deconvolution operation; Then, using the input feature map F input Generate the upsampling convolution kernel K, expressed as: , In the formula, Conv 1x1 Represents a convolutional layer with a convolution kernel size of 1, which is used to compress the size of the feature map. 3x3 represents a convolution layer with a convolution kernel size of 3, which is used to enlarge the size of the feature map. PixelShuffle represents a pixel shuffle operation, which is used to change the resolution of the feature map. Softmax is an activation function used to normalize the convolution kernel. Through the above operations, we can capture F input The features in the feature map F reorg , thereby realizing the interactive fusion of features; The feature map Freorg The convolution kernel is used to upsample and obtain the feature map Upsample (F input ), expressed as: .

[0029] Furthermore, bounding box regression predicts the bounding box position and size of the target through convolutional layers and learnable scaling parameters. The calculation formula is: , In the formula, Conv box is a convolutional layer for bounding box regression, and scale is a learnable scaling parameter; Category probability generation predicts the probability of the category to which the target belongs through the convolution layer and the Sigmoid activation function. The calculation formula is: , In the formula, Conv cls represents the convolutional layer for class probability generation, and σ represents the Sigmoid activation function.

[0030] The fused bounding box and class probability are combined into the final detection result. The output result includes the detection box, class probability and class label.

[0031] Step 4: Use the training set to iteratively train the optimized YOLOv8s model to obtain the target detection model.

[0032] Furthermore, the loss function of the optimized YOLOv8s model is the PIoU loss function, and the calculation formula is: , In the formula, IoU represents the ratio of the intersection area to the union area, and P represents the adaptive penalty factor.

[0033] In this example, the PIoU loss function is used to optimize the detection box regression of the Detect_LSCD head. During the detection box regression process, the PIoU loss function is used to optimize the position and size of the detection box. The PIoU loss function combines the target size adaptive penalty factor and the gradient adjustment function based on the quality of the anchor box, solving the problem that the traditional loss function is not strongly correlated with IoU, and is especially suitable for target detection of different scales and shapes. By introducing the PIoU loss function, the target detection model in this example can adjust the detection box more accurately and reduce false detections and missed detections.

[0034] Step 5: Use the trained target detection model to detect urban traffic targets.

[0035] Figures 4 to 5The test result diagram shows the final prediction results. It can be seen that the targets in the image, including pedestrians, cars, and motorcycles, are accurately identified.

[0036] In the process of detecting objects in images, it is necessary not only to correctly predict the category of the object, but also to predict the specific location of the object. The result evaluation can be divided into two parts, one is the prediction index of the prediction box, and the other is the classification prediction index.

[0037] 1. Intersection over Union (IoU) IoU is used to measure the overlap between the predicted bounding box and the true bounding box. It is an important part of calculating the average precision mAP. The calculation formula is:

[0038] The higher the IoU value, the greater the overlap between the predicted box and the true box, and the higher the regional accuracy of the object. Usually, the IoU threshold is set to 0.5, that is, when IoU ≥ 0.5, the prediction is considered to be a positive example.

[0039] 2. Confusion Matrix: Used to show the relationship between the model prediction results and the true label, usually divided into four categories: true positive, false positive, true negative and false negative.

[0040] Through the confusion matrix, the indicator precision Precision can be calculated. Precision measures the proportion of samples predicted by the model as positive samples that are actually positive samples. The calculation formula is: , Among them, TP represents the number of positive samples predicted correctly, and FP represents the number of positive samples predicted incorrectly.

[0041] With the prediction indicators of the prediction box and the indicators of classification prediction, the two can be combined into the evaluation indicators of the model: mAP@0.5 means that when the IoU threshold is 0.5, the average precision of each category is calculated. The higher the mAP@0.5 value, the better the overall detection performance of the model.

[0042] Indicators such as IoU, confusion matrix, precision, and mAP@0.5 together constitute a comprehensive evaluation system for the performance of the target detection model. In order to evaluate the importance of the improved modules in this example and their contribution to the detection algorithm, the following Table 1 shows the impact of different modules and their combinations on the detection performance and lightweight performance of the YOLOv8 detection algorithm, where RTL-Net represents the target detection model in this example.

[0043] Table 1

[0044] In Table 1, Param is the parameter quantity, M stands for million, and the unit of parameter quantity is one million; Size is the model size, and MB stands for megabytes, which is the unit of model size. From the experimental results in Table 1 above, it can be seen that RTL-Net has the best results in all aspects, with mAP higher than the original model and the parameter quantity and model size significantly lower than the original model, which fully demonstrates that the optimized model structure in the present invention has achieved a great degree of lightweight while improving accuracy.

Claims

1. A real-time lightweight small target detection method, characterized in that: The steps include: Step 1: Obtain traffic condition images and preprocess them, use the preprocessed images to construct a data set, and divide them into a training set and a test set in proportion; Step 2: Build the basic YOLOv8s model, including the backbone network, neck network, and head network; Step 3: Optimize the YOLOv8s model structure, specifically: Construct the DWR_C2F module in the backbone network to replace the original C2F module; construct the C_BiFPN module in the neck network to replace the original feature pyramid network; construct the Detect_LSCD detection head in the head network to replace the original detection head; Step 4, use the training set to iteratively train the optimized YOLOv8s model to obtain the target detection model; Step 5: Use the trained target detection model to detect urban traffic targets.

2. A real-time lightweight small target detection method according to claim 1, characterized in that: The working process of the backbone network in the optimized YOLOv8s model includes: The image input to the optimized YOLOv8s model is denoted as I. Image I first passes through a The convolution layer extracts the initial features and obtains a new feature map F1. The formula is: , In the formula, Represents a standard 3×3 convolutional layer; Then the feature map F1 is extracted through the DWR_C2F module: First, a standard 3×3 convolutional layer is used to extract the initial features of the feature map F1 to generate the feature map F conv3×3 , expressed as: , Then, three dilated convolutional layers are used, with different dilation rates set to capture features of different scales and generate feature maps. , expressed as: , Where r represents the expansion rate, which can be 1, 3, or 5; D-r3x3DConv represents the convolution operation with different expansion rates according to the r value, thus obtaining , and ; Next, the feature maps output by the three dilated convolutional layers , and Concatenate on the channel dimension to get the feature map , and then adjust the number of channels through a 1×1 convolution layer to obtain the feature map , expressed as: , , In the formula, Concatenate represents the concatenation operation, and Conv1x1 represents the convolution operation with a convolution kernel size of 1; The feature map obtained After adding to image I, we get feature map F DWR , as the output item of DWR_C2F module; The feature map processed by the DWR_C2F module is extracted through three residual blocks, and the feature representation is gradually enhanced to obtain the feature map F2, which is expressed as: , Where ResBlock represents the residual operation; In the process of feature extraction, downsampling operations are performed periodically, and convolutional layers or maximum pooling layers with a stride of 2 are used to reduce the resolution of the feature map while increasing the number of channels to obtain the feature map F3, which is expressed as: , Where, DownSample represents the downsampling operation; Then, through the downsampling operation, three feature maps of different scales are generated, which are denoted as F P3 、F P4 and F P5 , respectively expressed as: , , 。 3. A real-time lightweight small target detection method according to claim 2, characterized in that: The working process of the C_BiFPN module in the optimized YOLOv8s model includes: The C_BiFPN module receives three feature maps of different scales F from the backbone network. P3 、F P4 and F P5 As input, a top-down feature pass is first performed: For the feature map F P5 Perform an upsampling operation to increase its resolution to the same level as the feature map F P4 Same as the feature map F P4 Add element by element to generate feature map F P4_top_down , expressed as: , In the formula, Upsample indicates the upsampling operation through the CARAFE module; For the feature map F P4_top_down Upsample to make its resolution the same as the feature map F P3 Same as the feature map F P3 Add element by element to generate feature map F P3_top_down , expressed as: , Then perform bottom-up feature transfer: For the feature map F P3_top_down Perform downsampling to make its resolution consistent with the feature map F P4_top_down Same as the feature map F P4_top_down Add element by element to generate feature map F P4_bottom_up , expressed as: , Where, Downsample represents the downsampling operation; For the feature map F P4_bottom_up Downsample to make its resolution the same as the feature map F P5 Same as the feature map F P5 Add element by element to generate feature map F P5_bottom_up , expressed as: , For the feature map F P4_bottom_up Upsample to make its feature rate consistent with the feature map F P3_top_down Same as the feature map F P3_top_down Add element by element to generate feature map F P3_bottom_up , expressed as: , Because the feature map F P5 In the top-down feature transfer path, it is the top-level feature map, so there is the following relationship: , For each level, the feature maps from different paths are concatenated and then Convolution reduces the number of channels and fuses features. Finally, the fused features are weighted summed through learnable weights to generate a feature map. , expressed as: , In the formula, α n and β n It is the weight parameter obtained through training and learning. n represents P3, P4, and P5, which are used to represent different feature levels. Conv represents the convolution layer, including convolution operation, group normalization, and activation function.

4. A real-time lightweight small target detection method according to claim 3, characterized in that: The working process of the Detect_LSCD detection head in the optimized YOLOv8s model includes: The feature map input to the Detect_LSCD detection head is , the feature map First, feature extraction is performed through the convolution layer to generate the feature map F conv , the formula is: , Where Conv represents the convolution layer, including convolution operation, group normalization and activation function; Detect_LSCD introduces separable convolution and shared convolution, where separable convolution decomposes the standard convolution into depth convolution and point-by-point convolution. First, each input channel is convolved separately through depth convolution to extract spatial features and generate feature map F. dw , the feature map F conv The shape of [C in ,H,W], the depth convolution uses C in 3×3 convolution kernel, then the feature map F dw The shape is still [C in ,H,W], feature map F dw It is expressed as: , Where DepthwiseConv represents the deep convolution operation; C in , H, W represent the number of channels, height, and width of the feature map, respectively; Then perform point-by-point convolution: convolution is performed on the channel dimension using a 1×1 convolution kernel to integrate the features of different channels, and the number of output channels of the point-by-point convolution is recorded as C out , then the generated feature map F sep The shape is [C out ,H,W], feature map F sep Expressed as: , In the formula, PointwiseConv represents the point-by-point convolution operation; Then the feature map F sep Perform a shared convolution operation: Use the shared convolution kernel to perform a convolution operation on the feature map to generate a feature map F share , expressed as: , In the formula, SharedConv represents the shared convolution operation, and the weights are shared between different detection tasks or feature maps; Finally, bounding box regression and category probability generation are performed, and the generated bounding box position information and the corresponding category probability are spliced ​​together to generate the final detection result.

5. A real-time lightweight small target detection method according to claim 4, characterized in that: The upsampling operation through the CARAFE module includes: The characteristic graph F input into the CARAFE module P3 、F P4 、F P5 Unified as F input , the feature map F input The features are reorganized through a convolution layer and a transposed convolution layer to obtain the reorganized feature map F reorg , expressed as: , Where ConvTranspose represents the deconvolution operation; Then, using the reorganized feature map F reorg Generate upsampling convolution kernel K and transform the feature map F reorg The convolution kernel is used to upsample and obtain the feature map Upsample (F input ), expressed as: 。 6. A real-time lightweight small target detection method according to claim 5, characterized in that: Bounding box regression predicts the bounding box position and size of the target through convolutional layers and learnable scaling parameters. The calculation formula is: , In the formula, Conv box is a convolutional layer for bounding box regression, and scale is a learnable scaling parameter; Category probability generation predicts the probability of the category to which the target belongs through the convolution layer and the Sigmoid activation function. The calculation formula is: , In the formula, Conv cls represents the convolutional layer for class probability generation, and σ represents the Sigmoid activation function.

7. A real-time lightweight small target detection method according to any one of claims 1 to 6, characterized in that: The loss function of the optimized YOLOv8s model is the PIoU loss function, and the calculation formula is: , In the formula, IoU represents the ratio of the intersection area to the union area, and P represents the adaptive penalty factor.

Citation Information

Cited By

  • Satellite target detection method and system based on multi-scale feature fusion enhancement

    CN122336588A

  • A Satellite Target Detection Method and System Based on Multi-Scale Feature Fusion Enhancement

    CN122336588B