Target detection method based on EcoDetect-YOLOv2
By adding small object detection layer P2 on the basis of YOLOv8s and introducing an efficient multi-scale attention mechanism EMA, combined with Dysample upsampling technology and GhostConv module, the existing object detection methods are solved, and the problem of small object detection is difficult and computational complexity is high in the perspective of the surveillance camera, achieving efficient and robust object detection performance.
Patent Information
- Application Number
- CN202510541044.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing target detection methods face problems such as difficult detection of small targets, complex background interference, target object occlusion and diversity of morphological changes from the perspective of the surveillance camera, resulting in a lot of room for improvement in detection accuracy and real-timeness.
The object detection method based on EcoDetect-YOLOv2, by adding a small object detection layer P2 on the basis of YOLOv8s, and introducing an efficient multi-scale attention mechanism EMA, using Dysample upsampling technology and GhostConv module, the calculation complexity and inference time of the model are optimized.
It significantly improves the detection ability of small targets, enhances the robustness of the model in complex backgrounds and noise environments and the generalization performance of cross-scale targets, and reduces the computational complexity and inference time of the model.
Smart Images

Figure CN120070873A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an object detection method based on EcoDetect - YOLOv2. Background Art
[0002] The research on object detection technology has a long history. Before the rise of deep learning, this field mainly relied on detection frameworks based on image processing and traditional machine learning. Early methods mainly relied on manually extracting geometric shapes, edge features, and texture information, and used classifiers for object recognition. However, such methods generally have problems such as limited feature expression ability, high computational complexity, low detection accuracy, and insufficient generalization ability. In recent years, the breakthrough progress of deep learning technology has greatly improved the performance of computer vision tasks [9]. Among them, the convolutional neural network (CNN), with its end - to - end feature extraction ability, has shown excellent performance in object detection and image classification tasks. The multi - layer structure of CNN can hierarchically learn the high - level semantic information of images, getting rid of the shackles of traditional manually designed features and promoting the development of object detection towards intelligence and efficiency.
[0003] Existing object detection methods are mainly divided into two categories: two - stage detection (Two - Stage) and one - stage detection (One - Stage). Two - stage methods first generate candidate regions and then use classification and regression modules to optimize object localization. Among them, Faster R - CNN based on the Region Proposal Network (RPN) is a representative model of this type of method. Nowakowski and Pamula conducted electronic waste detection based on Faster R - CNN and achieved an average accuracy of over 90%. However, although this method performs excellently in high - precision object detection tasks, it has a high computational cost, a slow inference speed, and limited small - object detection ability. In contrast, one - stage detection methods, such as SSD (Single Shot MultiBox Detector) and YOLO (You Only Look Once) series, improve the detection speed by directly regressing object categories and bounding boxes, but the detection accuracy is often lower than that of two - stage methods. However, current garbage detection models are mainly trained based on relatively simple scenarios and still face many challenges under the perspective of surveillance cameras, such as the difficulty of small - object detection, complex background interference, object occlusion, and diverse morphological changes. This leads to a large room for improvement in the detection accuracy and real - time performance of existing algorithms. In addition, most advanced detectors have a high computational complexity and are difficult to meet the application requirements of real - time monitoring.
[0004] Therefore, in view of the above problems, an object detection method based on EcoDetect - YOLOv2 is provided. Summary of the Invention
[0005] The object of the present invention is to provide an object detection method based on EcoDetect-YOLOv2 to overcome the existing defects, which reduces the computational complexity and inference time while ensuring the detection accuracy, and improves the detection accuracy and efficiency of small target garbage.
[0006] The technical solution to achieve the above object is as follows: An object detection method based on EcoDetect-YOLOv2, comprising: Step S1, select a multi-target garbage exposure detection data set from the perspective of a monitoring camera, and randomly divide the data set into a training set and a test set according to a ratio of 8:2; Step S2, preprocess the multi-target garbage exposure detection data set; Step S3, construct an object detection model EcoDetect-YOLOv2; Step S4, train the object detection model EcoDetect-YOLOv2 with the training set, and use the trained object detection model EcoDetect-YOLOv2 to detect the test set, and output the detection result.
[0007] Preferably, in the step S1, the multi-target garbage exposure detection data set includes nine types of garbage detection target categories, namely paper garbage, plastic garbage, snake skin bags, packaging garbage, stone waste, sand waste, cardboard boxes, foam garbage and metal waste.
[0008] Preferably, in the step S2, the preprocessing includes: Adopt the Mosaic (a technique for enhancing the diversity of training data) data augmentation method to enrich the training data by randomly cropping and splicing different images; In terms of loss calculation, adopt the TaskAlignedAssigner (dynamic assignment strategy) in task-aligned object detection as the dynamic object assignment strategy to optimize the object matching mechanism. Among them, task-aligned object detection uses the following formula for object assignment: ; In the formula, is the predicted score corresponding to the labeled category, is the intersection over union between the predicted box and the ground truth box, and are both weight parameters; The loss function is defined as follows: ; In the formula, is the sample index, is the predicted probability of the model for the sample, is the true probability of the sample, is the total number of samples.
[0009] Preferably, in the step S3, constructing the target detection model EcoDetect-YOLOv2 includes: Select YOLOv8s as the base model and add the small target detection layer P2 on the basis of YOLOv8s; Introduce the multi-scale attention mechanism EMA before the small target detection layer P2; In the Neck (neck) network, use Dysample (ultra-lightweight dynamic upsampling operator) upsampling to replace the original nearest neighbor upsampling method to generate the feature map; In the Neck network, use GhostConv (improved convolutional layer design) to replace Conv (traditional convolutional layer design) in the Neck network, and based on GhostConv, further propose the GhostResBottleneck structure that combines GhostConv and residual connection; And adopt the One-Shot aggregation strategy to design the cross-stage partial network module ResGhostCSP (a neural network design that combines ResNet, Ghost module, and CSP structure, where ResNet is a residual network, Ghost module is a lightweight neural network module, and CSP structure is a network structure design) to replace C2f in the Neck network; Furthermore, complete the construction of the target detection model EcoDetect-YOLOv2; Among them, residual connections are introduced in both the GhostResBottleneck and ResGhostCSP structures.
[0010] Preferably, in the step S3, to fully fuse the multi-scale features of channels and space, introduce the multi-scale attention mechanism EMA. Given a feature map, its input tensor is defined as follows: ; In the formula, is the number of channels, and are the spatial dimensions of the input feature map, that is, the height and width of the feature map. The multi-scale attention mechanism EMA first divides into groups of sub-features, that is: ; The multi-scale attention mechanism EMA adopts a multi-scale feature extraction strategy, using two 1×1 branches and a 3×3 convolutional kernel for parallel operation to construct three feature extraction paths; Among them, the 1×1 branch uses global average pooling, while the 3×3 branch extracts features through multiple paths to capture the dependencies between channels and reduce the computational complexity; In addition, this mechanism encodes the features and performs feature fusion along the height direction. Without reducing the number of channels, it shares the 1×1 convolution operation, divides the output into two vectors, and simultaneously uses the non-linear Sigmoid function to establish cross-channel interactions between the branches; The 3×3 branch interacts with the features through convolution operations to expand the feature space and retain the spatial structure information; Subsequently, two-dimensional global average pooling is used to encode the spatial information of the three branches, that is: ; In the formula, is the 1×1 convolutional kernel, is the 3×3 convolutional kernel; Among them, the two-dimensional average pooling The calculation formula is as follows: ; In the formula, is the value at the position in the input feature map.
[0011] Preferably, in step S3, in the Neck network, Dysample upsampling is used instead of the original nearest neighbor upsampling method to generate the feature map. Among them, In the DySample design, the sampling point generator is one of the key components. Let the size of the input feature map X be , and the size of the sampling set is , where the first two dimensions represent and coordinates. Using the grid_sample function, the input feature map is resampled with the coordinates provided by the sampling set . This process is based on the bilinear interpolation method to achieve feature mapping and generate a new feature map with a size of , and the specific definition is as follows: ; Let the upsampling scale factor be , then the size of the input feature map is , to achieve upsampling, a linear transformation layer is first used, with an input channel number of , and an output channel number of , thereby generating an offset with a size of ; Subsequently, according to the pixel rearrangement algorithm, the offset is rearranged to a size of ; Finally, the sampling set is obtained by superimposing the offset and the original sampling grid . The calculation process is as follows: ; ; Finally, the upsampled feature map generated by the sampling set and the grid_sample function, whose dimension is .
[0012] Preferably, in step S3, residual connections are introduced in both the GhostResBottleneck and ResGhostCSP structures, and their calculation formula is as follows: ; In the formula, is the input of the residual block, is the output of the residual block, is the non - linear transformation of the input feature, including but not limited to convolution and activation function operations.
[0013] The beneficial effects of the present invention are as follows: Based on YOLOv8s, by adding a P2 small - target detection layer and introducing an efficient multi - scale attention mechanism EMA, the detection ability of small targets is significantly improved, and at the same time, the robustness of the model in complex background and noise environments and the generalization performance of cross - scale targets are enhanced; in terms of feature fusion, the Dysample upsampling technique is used to replace the traditional nearest - neighbor upsampling, optimizing the information fusion process and further improving the discrimination ability of overlapping targets; in addition, by introducing GhostConv and the ResGhostCSP module designed based on the one - shot aggregation strategy, the present invention effectively reduces the model calculation complexity and inference time while ensuring the detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a flowchart of an object detection method based on EcoDetect - YOLOv2 of the present invention; Figure 2It is a distribution diagram of the number of images and the number of instances corresponding to each category in the multi-object garbage exposure detection dataset of the present invention; Figure 3 It is a schematic diagram of the GhostResBottleneck and ResGhostCSP structures in the present invention.
[0015] Figure 4 It is the confusion matrix of the baseline model YOLOv8s under the mAP0.5 metric in the present invention; Figure 5 It is the confusion matrix of the EcoDetect-YOLOv2 model under the mAP0.5 metric in the present invention. Detailed implementation manners
[0016] Next, the technical solution of the present invention will be clearly and completely described in conjunction with the accompanying drawings. In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0017] Next, the present invention will be further described in conjunction with the accompanying drawings.
[0018] As Figure 1 shown, a target detection method based on EcoDetect-YOLOv2 includes: Step S1, select the multi-object garbage exposure detection dataset from the perspective of the monitoring camera, and randomly divide the dataset into a training set and a test set according to the ratio of 8:2.
[0019] In the embodiment, the multi-object garbage exposure detection dataset includes nine types of garbage detection target categories, namely paper trash (Paper_trash), plastic trash (Plastic_trash), snakeskin bag (Snakeskin_bag), packed trash (Packed_trash), stone waste (Stone_waste), sand waste (Sand_waste), carton (Carton), foam trash (Foam_trash), and metal waste (Metal_waste). The number of images (ImageCount) and the number of instances (InstanceCount) corresponding to each category are as Figure 2As shown. During the experiment, the dataset was randomly divided into a training set and a test set at a ratio of 8:2 for model training and testing.
[0020] Step S2: Preprocess the multi-object garbage exposure detection dataset.
[0021] In the embodiment, the preprocessing includes: The Mosaic data augmentation method is used to enrich the training data by randomly cropping and splicing different images, improving the generalization ability of the model for different scenarios and targets. However, Mosaic data augmentation may lead to overfitting problems during training. To avoid this impact, in the present invention, Mosaic is turned off in the last 10 Epochs of model training, enabling the model to finally converge on the original dataset without cropping, thereby alleviating the potential drawbacks brought by Mosaic data augmentation.
[0022] In terms of loss calculation, TaskAlignedAssigner in task-aligned object detection is adopted as the dynamic object assignment strategy to optimize the object matching mechanism. Among them, task-aligned object detection uses the following formula for object assignment: ; In the formula, is the predicted score corresponding to the annotation category, is the intersection over union between the predicted bounding box and the ground truth bounding box, and are both weight parameters; The loss function is defined as follows: ; In the formula, is the sample index, is the predicted probability of the model for the sample, is the true probability of the sample, is the total number of samples. This task-aligned strategy can dynamically adjust object assignment, improve the matching accuracy of detection bounding boxes, and thus enhance the detection ability of the model.
[0023] Step S3: Build the object detection model EcoDetect-YOLOv2.
[0024] In the embodiment, building the object detection model EcoDetect-YOLOv2 includes: Select YOLOv8s as the base model and add the small object detection layer P2 on the basis of YOLOv8s; Introduce the multi-scale attention mechanism EMA before the small object detection layer P2; In the Neck network, Dysample upsampling is adopted to replace the original nearest neighbor upsampling method to generate feature maps; In the Neck network, GhostConv is used to replace Conv in the Neck network, and based on GhostConv, a GhostResBottleneck structure combining GhostConv and residual connection is further proposed; And the One-Shot aggregation strategy is adopted to design the cross-stage partial network module ResGhostCSP to replace C2f in the Neck network; Furthermore, the construction of the object detection model EcoDetect-YOLOv2 is completed; Among them, residual connections are introduced in both the GhostResBottleneck and ResGhostCSP structures.
[0025] In the embodiment, the original backbone network structure of Yolov8 performs the feature extraction process through top-down hierarchical downsampling; in the Backbone layer of YOLOv8, as the network depth increases, the representation ability of the obtained high-level semantic information is enhanced, but the feature information of the obtained samples is significantly reduced; when performing small object detection, since the small object samples occupy a small size in the image, when passing through the Neck layer, it is difficult for the model to learn the feature information of the small object samples, and in the three detectors of the Head layer, smaller objects cannot be detected, resulting in the situation of missed detection of small object samples in the detection results; in the original YOLOv8 model's network structure, there are three detectors. When the input image size is 640*640, after multiple downsamplings of the Backbone layer and feature fusion of the Neck layer for detection, detection feature maps of 20x20, 40x40, and 80x80 can be obtained respectively, which can detect objects with target sizes larger than 32x32, 16x16, and 8x8; if there are samples with sizes smaller than 8x8 in the input image, then it is very difficult for the YOLOv8 model to detect these samples, resulting in missed detection; and from the perspective of a surveillance camera, the problem of garbage exposure involves a large number of small objects, especially paper scraps and other garbage, whose features often occupy only a very small number of pixels; in this case, relying solely on the P3 detection head may lead to the omission of a large number of small objects; therefore, in order to improve the small object detection ability, a P2 (160×160) detection head is additionally introduced in the network in the present invention to enhance the recognition ability of fine-grained objects.
[0026] In the embodiment, EMA (Efficient Multi-scale Attention) is a multi-scale attention module for cross-space learning. This module abandons the traditional convolutional dimensionality reduction method and instead selects some channels and reshapes them into batch channels to achieve dimensionality reduction. At the same time, it uses a cross-space learning strategy for multi-scale feature extraction. In fields such as object detection and semantic segmentation, convolutional operations have been widely used due to their excellent feature learning ability. However, as the depth of the convolutional network increases, its memory occupancy and computational complexity also increase significantly. In addition, the translational invariance property of convolution leads to feature localization, restricting the effective modeling of global information. Therefore, reasonably introducing an attention mechanism in deep networks helps improve feature discrimination ability.
[0027] To ensure the efficiency of garbage recognition from the perspective of the monitoring camera and fully integrate multi-scale features of channels and space, a multi-scale attention mechanism EMA is introduced. Given a feature map, its input tensor is defined as follows: ; In the formula, is the number of channels, and are the spatial dimensions of the input feature map, that is, the height and width of the feature map. The multi-scale attention mechanism EMA first divides into groups of sub-features, that is: ; The multi-scale attention mechanism EMA adopts a multi-scale feature extraction strategy and uses two 1×1 branches and a 3×3 convolutional kernel for parallel operation to construct three feature extraction paths; Among them, the 1×1 branch uses global average pooling, while the 3×3 branch extracts features through multiple paths to capture the dependencies between channels and reduce computational complexity; In addition, this mechanism encodes the features and performs feature fusion along the height direction. Without reducing the number of channels, it shares the 1×1 convolution operation and divides the output into two vectors. At the same time, it uses the non-linear Sigmoid function to establish cross-channel interaction between the branches; The 3×3 branch interacts with the features through convolution operations to expand the feature space and retain spatial structure information; Subsequently, two-dimensional global average pooling is used to encode the spatial information of the three branches, that is: ; In the formula, is the 1×1 convolutional kernel, is the 3×3 convolutional kernel; Among them, the two-dimensional average pooling The calculation formula is as follows: ; In the formula, is the value at position in the input feature map.
[0028] To further improve the calculation efficiency, the multi-scale attention mechanism EMA combines the Softmax function with average pooling operations for linear transformation, and finally outputs a set of spatial attention weights to achieve feature enhancement. In summary, the multi-scale attention mechanism EMA ensures the accuracy and calculation efficiency of garbage recognition from the perspective of surveillance cameras by effectively modeling the multi-scale feature relationships between channels and spaces. Its core advantage lies in efficiently learning feature semantics and fully integrating multi-scale feature information, providing an innovative solution for garbage recognition tasks in complex scenarios.
[0029] In the embodiment, in the Neck network, Dysample upsampling is used instead of the original nearest neighbor upsampling method to generate the feature map. Among them, DySample improves the utilization efficiency of computing resources by avoiding dynamic convolution operations and using a point-based sampling method for upsampling; in the YOLOv8 model, the traditional UpSample method usually requires a large amount of computing resources and parameters, which limits the lightweight deployment of the model from the perspective of surveillance cameras and thus affects its detection performance for small target garbage; in actual surveillance applications, small target garbage such as paper trash is often small in size and vulnerable to pixel distortion, resulting in the loss of fine details and posing challenges to feature learning; to solve this problem, the present invention introduces a lightweight and efficient dynamic upsampling method - DySample as an alternative to UpSample; DySample can improve the recognition ability for small target garbage under the condition of low surveillance image quality; its core idea is to combine a point-based sampling strategy with a learning sampling method to achieve efficient upsampling; this method can not only reduce the consumption of computing resources, but also effectively improve the image resolution without increasing the computational burden, thereby enhancing the overall performance and calculation efficiency of the model; In the DySample design, the sampling point generator is one of the key components. Let the size of the input feature map X be , and the size of the sampling set be , where the first two dimensions represent and coordinates. Using the grid_sample function, resample the input feature map with the coordinates provided by the sampling set . This process is based on the bilinear interpolation method to achieve feature mapping and generate a size of The new feature map is specifically defined as follows: ; Let the upsampling ratio factor be , then the size of the input feature map is . To achieve upsampling, first a linear transformation layer is used, whose input channel number is , and the output channel number is , thereby generating an offset with a size of ; Subsequently, according to the pixel rearrangement algorithm, the offset is rearranged to a size of ; Finally, the sampling set is obtained by superimposing the offset and the original sampling grid . Its calculation process is as follows: ; ; Finally, the upsampled feature map generated by the sampling set and the grid_sample function, whose dimension is .
[0030] In the embodiment, the computational cost of GhostConv is about 50% of that of the standard convolution (Conv), but its feature learning ability can be comparable to that of the standard convolution. Based on GhostConv, the present invention further proposes a GhostResBottleneck structure that combines GhostConv with a residual connection (Residual Connection); in addition, to enhance the feature learning ability of the CNN, the present invention combines a generalized deep learning optimization strategy and adopts a One-Shot aggregation strategy to design an efficient and hardware-friendly cross-stage partial (CSP) network module - ResGhostCSP, which can effectively reduce the computational complexity and inference time while maintaining high accuracy.
[0031] To ensure effective feature extraction and improve the stability of the network, residual connections are introduced into both the GhostResBottleneck and ResGhostCSP structures. Among them, the structures of GhostResBottleneck and ResGhostCSP are as Figure 3As shown in the figure, residual connections can effectively alleviate problems that may occur during deep network training, such as overfitting, gradient vanishing, and gradient explosion. The calculation formula is as follows: ; In the formula, is the input of the residual block, is the output of the residual block, It is a nonlinear transformation of input features, including but not limited to convolution and activation function operations.
[0032] Compared with the traditional feature concatenation (Concat) operation, the introduction of residual connection can more effectively alleviate the gradient problem in deep network training and significantly improve the stability of the model.
[0033] GhostResBottleneck and ResGhostCSP combine the advantages of GhostConv's low-cost feature map generation and the efficient feature transfer capability of residual connections. While reducing the consumption of computing resources, they further improve the model's feature extraction capabilities and provide strong support for lightweight and efficient network design.
[0034] Step S4: train the target detection model EcoDetect-YOLOv2 using the training set, detect the test set using the trained target detection model EcoDetect-YOLOv2, and output the detection result.
[0035] In order to intuitively demonstrate the advancedness of the proposed EcoDetect-YOLOv2, the present invention compares it with the baseline model YOLOv8s, and the experimental results are shown in Table 1.
[0036] Table 1 The proposed EcoDetect-YOLOv2 was verified using the test set. The experimental results are shown in Table 1. The precision, recall, mAP0.5 and mAP0.5:0.95 of the proposed EcoDetect-YOLOv2 increased by 1.0%, 4.6%, 4.8% and 3.1% respectively, while the computational complexity was slightly increased, and the number of parameters decreased by 19.3%.
[0037] Figure 4 , Figure 5The confusion matrix of the baseline model YOLOv8s and EcoDetect-YOLOv2 proposed in the present invention under the mAP0.5 metric is shown, fully demonstrating that the classification accuracy of EcoDetect-YOLOv2 has been significantly improved in almost all categories. Notably, the classification accuracy of the paper trash class, which are extremely small targets that YOLOv8s often misses, has been increased by 13% in EcoDetect-YOLOv2 proposed in the present invention. The classification accuracies of the plastic_trash and packed_trash classes, which are small targets, have been increased by 8% and 6% respectively, showing a significant improvement in the small target detection ability of EcoDetect-YOLOv2. In addition, the classification accuracies of the sand_waste and metal_waste, which are relatively large targets, have been increased by 16% and 14% respectively, indicating that while improving the small target detection ability, the model has also enhanced the detection performance of other targets of different sizes. The Snakeskin_bag and stone_waste have maintained their original good accuracy. The accuracy of Carton has slightly decreased. Generally speaking, the EcoDetect-YOLOv2 model introduced in the present invention has a greater improvement compared to the baseline YOLOv8s, achieving more superior prediction and generalization capabilities.
[0038] In addition, to further analyze the garbage exposure detection ability of EcoDetect-YOLOv2 from the perspective of surveillance cameras, the present invention selects multiple images in the test set for visual comparison. The experimental results further show that YOLOv8s has significant detection defects in multiple complex scenarios.
[0039] When facing various small target detection scenarios, the baseline model YOLOv8s shows a relatively serious problem of missed detections. Specifically, due to the small size of the paper_trash targets and their extremely low proportion in the image, YOLOv8s can hardly detect all such targets. In contrast, although there are still a small number of missed detections in EcoDetect-YOLOv2 proposed in the present invention, it has been able to effectively detect the small target paper_trash in the image. Not only limited to paper_trash, for other targets placed far away or with small sizes, YOLOv8s also shows a high missed detection rate. For example, the carbon placed at a distance cannot be correctly detected by YOLOv8s, while EcoDetect-YOLOv2 successfully completes the detection; similarly, the stone_waste cannot be recognized by YOLOv8s due to its small size, while EcoDetect-YOLOv2 overcomes this problem.
[0040] In addition, YOLOv8s also has a relatively serious misdetection problem. This model misclassifies the white parking line as paper_trash and misidentifies the foundation as plastic_trash, while EcoDetect-YOLOv2 demonstrates better detection robustness in such scenarios.
[0041] In low-light environments, the detection performance of YOLOv8s further deteriorates. In night scenes, YOLOv8s fails to detect the clearly present plastic_trash, while EcoDetect-YOLOv2 can still successfully detect the target despite having a relatively low confidence level. These experimental results indicate that EcoDetect-YOLOv2 can effectively adapt to different target sizes and characteristic variations, significantly improving the accuracy and robustness of object detection.
[0042] To systematically evaluate the advantages of the EMA attention mechanism, this invention is based on the YOLOv8s model and introduces various mainstream attention mechanisms such as CA (Coordinate Attention), SE (Squeeze-and-Excitation), NonlocalBlockND, CBAM (Convolutional Block Attention Module), and SCSA (Spatial-Channel SynergicAttention), and conducts comparative experiments to explore the performance of each mechanism in this task.
[0043] Among them, the CA attention mechanism realizes efficient encoding of feature space information by introducing position sensitivity and direction perception ability, thereby improving the accuracy of the model's target localization. The SE attention mechanism dynamically adjusts the weights according to the importance of each channel to compensate for the information loss caused by uneven channel weight distribution during the feature extraction process. NonlocalBlockND is based on non-local operations and can calculate the correlation between any two positions in the feature map, thus effectively capturing long-range dependence information and enhancing the model's global context understanding ability. However, due to its high computational complexity, the computational cost is relatively large when processing high-resolution feature maps. CBAM combines channel attention and spatial attention, calculates the channel weights through global average pooling and max pooling, and optimizes the information distribution of the feature channels; at the same time, its spatial attention module effectively highlights the features of the key regions by compressing the channel dimension. Although CBAM has certain advantages in improving the model's detection performance, its computational efficiency is slightly insufficient compared to lightweight attention mechanisms such as EMA. SCSA overcomes the limitations of traditional attention mechanisms in local and global feature modeling through spatial-channel collaborative modeling and multi-semantic information fusion.
[0044] Table 2 As shown in Table 2, after introducing the EMA attention mechanism based on YOLOv8s, the overall detection performance of the model reaches the best. Compared with the original YOLOv8s, EMA reaches 68.9%, 53.5%, 57.3% and 33.3% in terms of precision (P), recall (R), mAP0.5 and mAP0.5:0.95 respectively, and both mAP0.5 and mAP0.5:0.95 achieve the best performance. Although the precision and recall of EMA do not reach the highest values separately, it achieves the best balance between the two, making the model have better robustness in the detection task. In addition, the introduction of EMA does not significantly increase FLOPs and the parameter scale, further verifying that this mechanism can effectively improve the object detection performance of the model while maintaining light weight.
[0045] To verify whether the improved modules in EcoDetect-YOLOv2 achieve the expected effects, the present invention conducts systematic ablation experiments on each component, as shown in Table 3. Among them, Model 0 is the baseline YOLOv8s, and Models 1 to 15 represent different combinations of improvement schemes to evaluate the impact of each module on the detection performance.
[0046] Table 3 The experimental results in Table 3 reveal the following key findings: (1) Model 1 shows that adding the P2 small object detection layer effectively enhances the detection ability of the model in the garbage exposure detection task from the perspective of a surveillance camera mainly for small object detection, increasing the precision (P), recall (R), mAP0.5 and mAP0.5:0.95 by 2.2%, 1.9% and 0.4% respectively. Although the precision (P) decreases slightly, this phenomenon conforms to the precision-recall trade-off in the object detection task, that is, in the process of improving the recall and detection coverage, there may be a certain number of false detections. However, from the overall detection performance, this trade-off is acceptable and helps to improve the robustness of the model.
[0047] (2) Model 2 shows that the addition of the EMA module significantly enhances the detection ability. Specifically, through the cross-space learning mechanism, the EMA module reshapes some channel dimensions into the batch dimension and processes them in groups, captures features in the multi-scale space by combining 3×3 convolutions, and realizes the joint modeling of channel and spatial information. Its parallel sub-network structure performs multi-scale convolution processing on the grouped sub-features, suppresses background interference and enhances target response through the dynamic modulation mechanism, while avoiding the dimension reduction problem in the traditional attention mechanism. In addition, EMA adopts a cross-space feature aggregation strategy, and through lightweight decomposition and depthwise separable convolution design, it strengthens the feature expression of small targets while reducing the computational complexity. The experimental results show that the introduction of the EMA module increases the precision (P), recall (R), mAP0.5, and mAP0.5:0.95 by 0.7%, 1.4%, 2.2%, and 1.5% respectively, verifying its efficient feature representation ability in complex scenarios.
[0048] (3) Model 3 shows that after the introduction of DySample, the recall (R), mAP0.5, and mAP0.5:0.95 are all significantly improved (increased by 1.0%, 0.8%, and 0.7% respectively), indicating that through the dynamic upsampling mechanism and lightweight decomposition design, it enhances the multi-scale feature expression ability while only increasing a very small number of parameters, thus strengthening the capture effect of small targets in complex scenarios. Although the precision (P) slightly decreases from 68.2% to 67.8%, this may be due to the fact that when DySample improves the feature sensitivity, the response threshold for edge-blurred regions or background noise decreases. However, from the overall performance, DySample effectively improves the detection robustness while maintaining the efficiency of the model through the cross-scale feature aggregation strategy and dynamic sampling point segmentation technology.
[0049] (4) The experimental results of Model 4 show that ResGhostCSP proposed based on GhostConv in the present invention maintains the target detection performance while significantly reducing the computational complexity of the model, thus achieving the dual optimization of lightweight and detection accuracy.
[0050] (5) Models 5 to 14 represent the combined optimization experiments of different improvement modules. Although the introduction of multiple improvement components may slightly increase the computational complexity and the number of parameters, overall, the target detection accuracy is improved. In addition, the introduction of ResGhostCSP can effectively reduce the computational cost of the model on the premise of ensuring the detection performance, further verifying the effectiveness of multi-module collaborative optimization.
[0051] (6) Model 15, as the model proposed in this invention, performs excellently in this task. Compared with the baseline YOLOv8s, this model has improved by 0.5%, 10.9%, 9.0% and 6.1% respectively in terms of precision (P), recall (R), mAP0.5 and mAP0.5:0.95. In addition, the number of parameters of this model has decreased by 19.3%, and the computational complexity has only increased slightly, fully demonstrating that EcoDetect-YOLOv2 can effectively improve the garbage detection ability under the perspective of surveillance cameras while maintaining light weight. This achievement provides an efficient and reliable solution for intelligent environmental monitoring and digital governance.
[0052] The present invention selects object detection models of the same scale as YOLOv8s, namely YOLOv3-tiny, YOLOv5s, YOLOv8s, YOLO11s models, as well as the m and l series models corresponding to the YOLOv5, v8, 11 series, and compares them with the proposed EcoDetect-YOLOv2. In addition, several advanced detection models are also adopted, such as Waste-YOLO, YOLOv5s-OCDS, EcoDetect-YOLO and GCC-YOLO. Under the same experimental environment and parameters, the results generated by model training are shown in Table 4.
[0053] Table 4 As shown in Table 4, compared with the existing object detection models, EcoDetect-YOLOv2 proposed in the present invention has achieved the highest recall (R), mAP0.5 and mAP0.5:0.95 while the computational complexity and the number of parameters have been significantly reduced. Although the precision (P) is not the best, its overall detection performance still remains leading.
[0054] Specifically, compared with YOLOv3-tiny, YOLOv5s, YOLOv8s, and YOLO11s of the same scale, EcoDetect-YOLOv2 demonstrates superior detection performance. This may be because other models are more vulnerable to background interference in complex scenarios. At the same time, a larger downsampling ratio and a larger-sized detection head limit their ability to capture small targets. Compared with YOLOv8n, although it has the smallest scale and complexity, its weak detection ability is insufficient to achieve effective garbage detection. Compared with the m and l series models corresponding to the larger-scale YOLOv5, v8, and 11 series, although the model performance improves with the increase in model size and complexity, the increase amplitude is limited and far less than that of EcoDetect-YOLOv2 proposed in this invention. It is worth noting that the overall detection effect of YOLO11l is similar to that of YOLO11m, and even far lower than YOLO11m in terms of precision (P), indicating that excessive model complexity may lead to overfitting. Compared with a variety of current state-of-the-art models, EcoDetect-YOLOv2 also achieves the highest performance in terms of recall rate, mAP0.5, and mAP0.5:0.95. Although the precision rate is not the highest, it demonstrates better robustness. In summary, EcoDetect-YOLOv2 proposed in this invention achieves the best comprehensive detection performance while maintaining a low computational complexity, fully reflecting its efficiency and robustness.
[0055] In summary, in response to the problem of accurate identification and positioning of multi-scale and multi-type garbage in complex backgrounds from the perspective of surveillance cameras, this invention proposes a lightweight and efficient detection model - EcoDetect-YOLOv2. Based on YOLOv8s, this model significantly improves the detection ability of small targets by adding a P2 small target detection layer and introducing an efficient multi-scale attention mechanism (EMA), while enhancing the robustness of the model in complex backgrounds and noise environments and the generalization performance of cross-scale targets. In terms of feature fusion, the Dysample upsampling technique is used to replace the traditional nearest neighbor upsampling, optimizing the information fusion process and further improving the ability to distinguish overlapping targets. In addition, by introducing GhostConv and the ResGhostCSP module designed based on the one-shot aggregation strategy, this invention effectively reduces the model's computational complexity and inference time while ensuring the detection accuracy.
[0056] The experimental results show that when the number of model parameters decreases by 19.3%, compared with the baseline model YOLOv8s, EcoDetect-YOLOv2 improves the precision, recall, mAP0.5 and mAP0.5:0.95 metrics by 1.0%, 4.6%, 4.8% and 3.1% respectively, verifying its ability to achieve real-time multi-object detection of garbage from the perspective of surveillance cameras. Generally speaking, EcoDetect-YOLOv2 provides strong technical support for the automation of urban waste management and the development of the digital economy, and also proposes an effective solution to the problem of small object detection in complex scenarios. Future research will further focus on improving the adaptability of the model in larger-scale datasets and more complex scenarios, as well as the integration and optimization of real-time systems, to promote the in-depth development of the intelligent garbage management process.
[0057] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A target detection method based on EcoDetect-YOLOv2, characterized in that: include: Step S1, select a multi-target garbage exposure detection dataset from the perspective of a surveillance camera, and randomly divide the dataset into a training set and a test set in a ratio of 8:2; Step S2, preprocessing the multi-target garbage exposure detection data set; Step S3, constructing the target detection model EcoDetect-YOLOv2; Step S4: train the target detection model EcoDetect-YOLOv2 using the training set, detect the test set using the trained target detection model EcoDetect-YOLOv2, and output the detection result.
2. The target detection method based on EcoDetect-YOLOv2 according to claim 1, characterized in that: In step S1, the multi-target garbage exposure detection data set includes nine types of garbage detection target categories, namely paper garbage, plastic garbage, snakeskin bags, packaging garbage, stone waste, sand waste, cardboard boxes, foam garbage and metal waste.
3. The target detection method based on EcoDetect-YOLOv2 according to claim 1, characterized in that: In step S2, the preprocessing includes: Mosaic data augmentation method is used to enrich training data by randomly cropping and splicing different images; In terms of loss calculation, TaskAlignedAssigner in task-aligned target detection is used as a dynamic target assignment strategy to optimize the target matching mechanism. Task-aligned target detection uses the following formula for target assignment: ; In the formula, is the prediction score corresponding to the labeled category, is the intersection-over-union ratio between the predicted box and the true box, and All are weight parameters; The loss function is defined as follows: ; In the formula, is the sample index, is the model’s predicted probability for the sample, is the true probability of the sample, is the total number of samples.
4. The target detection method based on EcoDetect-YOLOv2 according to claim 1, characterized in that: In step S3, a target detection model EcoDetect-YOLOv2 is constructed, including: Select YOLOv8s as the basic model, and add a small target detection layer P2 on the basis of YOLOv8s; Introduce the multi-scale attention mechanism EMA before the small object detection layer P2; In the Neck network, Dysample upsampling is used to generate feature maps instead of the original nearest neighbor upsampling method; In the Neck network, GhostConv is used to replace Conv in the Neck network, and based on GhostConv, a GhostResBottleneck structure combining GhostConv and residual connection is further proposed; And adopt the One-Shot aggregation strategy to design the cross-stage partial network module ResGhostCSP to replace C2f in the Neck network; Then complete the construction of the target detection model EcoDetect-YOLOv2; Among them, residual connections are introduced in both GhostResBottleneck and ResGhostCSP structures.
5. The target detection method based on EcoDetect-YOLOv2 according to claim 4, characterized in that: In step S3, in order to fully integrate the multi-scale features of channels and spaces, a multi-scale attention mechanism EMA is introduced. Given a given feature map, its input tensor is defined as follows: ; In the formula, is the number of channels, and are the spatial dimensions of the input feature map, i.e., the height and width of the feature map. The multi-scale attention mechanism EMA first converts Divide into Group characteristics, namely: ; The multi-scale attention mechanism EMA adopts a multi-scale feature extraction strategy, using two 1×1 branches and a 3×3 convolution kernel to operate in parallel to construct three feature extraction paths; The 1×1 branch uses global average pooling, while the 3×3 branch extracts features through multiple paths to capture the dependencies between channels and reduce computational complexity. In addition, the mechanism encodes the features and performs feature fusion along the height direction, shares the 1×1 convolution operation without reducing the number of channels, and divides the output into two vectors, while using the nonlinear Sigmoid function to establish cross-channel interactions between branches; The 3×3 branch interacts with features through convolution operations, expands the feature space, and retains spatial structure information; Then, two-dimensional global average pooling is used to encode the spatial information of the three branches, namely: ; In the formula, is a 1×1 convolution kernel, is a 3×3 convolution kernel; Among them, two-dimensional average pooling The calculation formula is as follows: ; In the formula, is the position in the input feature map The value on .
6. The target detection method based on EcoDetect-YOLOv2 according to claim 4, characterized in that: In step S3, in the Neck network, Dysample upsampling is used to replace the original nearest neighbor upsampling method to generate a feature map, wherein: In the DySample design, the sampling point generator is one of the key components. Suppose the size of the input feature map X is , sampling set The size is , where the first two dimensions represent and Coordinates, using the grid_sample function, with the sampling set The provided coordinates are input to the feature map Resampling is performed based on the bilinear interpolation method to achieve feature mapping and generate a size of New feature map , which is defined as follows: ; Assume the upsampling factor is , then the input feature map The size is , in order to achieve upsampling, a linear transformation layer is first used, and the number of input channels is , the number of output channels is , thus generating a size of The offset ; Then, according to the pixel rearrangement algorithm, the offset Rearrange to size ; Finally, the sample set By offset With the original sampling grid The calculation process is as follows: ; ; Finally, by sampling the set And the upsampled feature map generated by the grid_sample function , whose dimensions are .
7. The target detection method based on EcoDetect-YOLOv2 according to claim 4, characterized in that: In step S3, residual connections are introduced into both GhostResBottleneck and ResGhostCSP structures, and the calculation formula is as follows: ; In the formula, is the input of the residual block, is the output of the residual block, It is a nonlinear transformation of input features, including but not limited to convolution and activation function operations.
Citation Information
Patent Citations
Multi-target garbage detection method based on improved YOLOv5 model
CN116452950A
Damaged traffic sign detection method and device
CN116580378A
Deep learning-based urine formed component detection method
CN117292375A
Marine litter target detection method based on yov8-CA-AFPN model
CN117523380A
Real-time household garbage detection system applied to monitoring camera
CN118736480A