An outdoor garbage image detection method and system
Patent Information
- Application Number
- CN202410424762.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-10
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-04-10
AI Technical Summary
因此,现有的评价指标不适合户外垃圾检测场景,导致现有的户外垃圾目标检测技术普遍存在检测准确率低的问题
[0029]本发明的有益效果:本发明首先引入改进的综合评价指标GD-F1,在目标检测模型测试时,计算综合评价指标GD-F1,通过综合评价指标GD-F1调整目标模型在训练阶段的分类损失,迭代训练,得到更适合户外垃圾目标检测的目标检测模型,提高检测准确率;其次,本发明在目标检测模型中加入检测层和多支路空洞卷积自适应融合模块,提高对小目标的检测准确率。
Smart Images

Figure CN118314324B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of target detection technology, and in particular relates to an image detection method for outdoor garbage. Background Technology
[0002] Deep learning models possess powerful abilities to learn the inherent functional patterns of sample datasets and analyze abstract features. In recent years, they have assisted people in making decisions in many fields and provided solutions to many complex recognition and classification problems. Deep learning has achieved excellent results in fields such as bioinformatics, image recognition, speech recognition, autonomous vehicles, artistic creation, emotion recognition, natural language processing, and banking. Furthermore, with the continuous efforts of researchers, the efficiency of deep learning models has been continuously improved. In recent years, with the advent of convolutional neural networks, image recognition technology based on deep neural networks has reached new heights.
[0003] Jin Guang et al. proposed a method for detecting floating objects in river images from UAV aerial photography in patent CN115115934A, based on an improved version of YOLOv5. This patented technical solution improves YOLOv5, primarily in its attention mechanism.
[0004] Most existing datasets for garbage detection directly use classification datasets or select a few common types of garbage, and the scenarios are mostly indoors or with simple backgrounds. Many scholars both domestically and internationally have applied deep learning to garbage detection and classification research. In outdoor garbage detection, unlike conventional object detection, the main focus is on detecting the presence and location of garbage, regardless of the type of garbage or whether it's a single piece of garbage (i.e., a non-dense target) or a pile of garbage (i.e., a dense target). For example, when sanitation trucks or garbage sweepers are cleaning and detecting garbage, they don't care whether it's a single piece or a pile, because as long as garbage is detected, the entire pile will be swept away. Therefore, existing evaluation metrics for conventional object detection, such as Map, Recall, and Precision, are not suitable for outdoor garbage detection. Existing evaluation metrics mean that detecting a single piece of garbage counts as recalling that garbage, and within a pile of garbage, the number of pieces detected counts as recalling that garbage, while undetected garbage is not considered recalled. Outdoor garbage detection, however, should consider the entire pile of garbage (i.e., a dense target) as detected if any garbage is found within it. Therefore, existing evaluation indicators are not suitable for outdoor waste detection scenarios, resulting in low detection accuracy in existing outdoor waste target detection technologies. Summary of the Invention
[0005] This invention aims to improve the accuracy of outdoor litter target detection, and proposes an image detection method and system for outdoor litter.
[0006] The first aspect of this invention proposes an image detection method for outdoor litter, the method comprising:
[0007] Acquire an image of outdoor litter to be detected, input the outdoor litter image into a trained target prediction model, and obtain the detection result of the outdoor litter image;
[0008] The training process of the target prediction model includes:
[0009] S1: Obtain a set of outdoor garbage images, perform data preprocessing on the outdoor garbage image set, and obtain preprocessed training samples and test samples;
[0010] S2: Using the training samples, train the target prediction model and calculate the classification loss of the target prediction model;
[0011] S3: Test the target prediction model using test samples, and calculate the comprehensive evaluation index GD-F1 based on the test results;
[0012] S4: Adjust the parameters of the classification loss function according to the comprehensive evaluation index GD-F1;
[0013] S5: Iterate the training until the loss function converges, then stop training to obtain the trained target prediction model.
[0014] Furthermore, the specific formula for calculating the comprehensive evaluation index GD-F1 is as follows:
[0015]
[0016]
[0017]
[0018] Wherein, GD-F1 represents the comprehensive evaluation index of outdoor litter target detection, GD_Precision represents the accuracy of outdoor litter target detection, GD_Recall represents the recall of outdoor litter target detection, TP represents the number of positive samples correctly predicted by the model, FP represents the number of positive samples incorrectly predicted by the model, FN represents the number of negative samples incorrectly predicted by the model, and n represents the number of undetected targets in a dense target group when at least one target is detected.
[0019] Furthermore, the classification loss function is specifically as follows:
[0020]
[0021] Among them, L conf Let represent the classification loss function, z represent the classification score predicted by the object detection model, α, β, γ, θ, and σ represent different hyperparameters, y = 1 indicates that the sample predicted by the object detection model is a positive sample, y = 0 indicates that the sample predicted by the object detection model is a negative sample, m = 0 indicates that the sample predicted by the object detection model is a non-dense object box, m = 1 indicates that the sample predicted by the object detection model is a dense object box, and q represents the intersection-union ratio between the predicted box and the ground truth box predicted by the regression branch of the object detection model.
[0022] Furthermore, in step S2, the target prediction model is an improved YOLOv8 network, which adds a detection layer and a multi-branch dilated convolution adaptive fusion module to the original YOLOv8 network.
[0023] A second aspect of the present invention provides an image detection system for outdoor litter, the device comprising:
[0024] Data acquisition module: This module is equipped with a camera for taking photos or videos. It also includes a communication module for transmitting photos or videos.
[0025] Object detection module: This module includes a pre-trained object detection model for object detection in real-time or offline images and videos;
[0026] Output display module: This module is used to output and display the detection results of the target detection module;
[0027] Processor module: Used to run software programs and modules stored in memory, and to perform various functions and data processing of the computer;
[0028] Storage module: This module includes a program storage area and a data storage area, used to store data, software programs, or modules.
[0029] The beneficial effects of this invention are as follows: First, this invention introduces an improved comprehensive evaluation index, GD-F1. During the testing of the target detection model, the comprehensive evaluation index GD-F1 is calculated. The classification loss of the target model during the training phase is adjusted by the comprehensive evaluation index GD-F1, and iterative training is conducted to obtain a target detection model that is more suitable for outdoor garbage target detection, thereby improving the detection accuracy. Second, this invention adds a detection layer and a multi-branch dilated convolution adaptive fusion module to the target detection model to improve the detection accuracy of small targets. Attached Figure Description
[0030] Figure 1 This is a flowchart of the steps in Embodiment 1 of the present invention;
[0031] Figure 2 This is a schematic diagram of the structure of the improved YOLOv8 model in Embodiment 1 of the present invention;
[0032] Figure 3 This is a schematic diagram of the structure of the multi-branch dilated convolution adaptive fusion module in Embodiment 1 of the present invention;
[0033] Figure 4 This is a schematic diagram of the adaptive fusion of the multi-branch dilated convolution adaptive fusion module in Embodiment 1 of the present invention;
[0034] Figure 5 This is a comparison chart of the test results of the embodiments of the present invention and the test results of existing target test models;
[0035] Figure 6 This is a comparison table of the results obtained from the ablation experiments of each module in Embodiment 1 of the present invention. Detailed Implementation
[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0037] Example 1
[0038] Embodiment 1 of this invention proposes an image detection method for outdoor waste, referring to... Figure 1 As shown, the method includes:
[0039] Acquire an image of outdoor litter to be detected, input the outdoor litter image into a trained target prediction model, and obtain the detection result of the outdoor litter image;
[0040] The training process of the target prediction model includes:
[0041] S1: Obtain a set of outdoor garbage images, perform data preprocessing on the outdoor garbage image set, and obtain preprocessed training samples and test samples;
[0042] S2: Using the training samples, train the target prediction model and calculate the classification loss of the target prediction model;
[0043] S3: Test the target prediction model using test samples, and calculate the comprehensive evaluation index GD-F1 based on the test results;
[0044] S4: Adjust the parameters of the classification loss function according to the comprehensive evaluation index GD-F1;
[0045] S5: Iterate the training until the loss function converges, then stop training to obtain the trained target prediction model.
[0046] Specifically, in step S1, the outdoor garbage image set is collected in an outdoor setting, so there is a certain degree of difficulty in detecting samples, and some negative samples are easily misdetected as positive samples.
[0047] Furthermore, in step S1, the specific process of data preprocessing includes:
[0048] S101: Collect a set of images of outdoor litter and use the labelimg tool to label the images as litter.
[0049] S102: Perform data augmentation on the labeled image set. Data augmentation methods include: horizontal flipping, cropping, mosaicking, stitching, rotation, translation, adding noise, and blurring / sharpening.
[0050] Specifically, in step S2, the target prediction model is trained using the training samples. The training process mainly involves training and optimizing the parameters (such as bias and weight) of the target prediction model and calculating the classification loss of the target prediction model.
[0051] In conventional target detection, the definitions and formulas of each indicator are shown in Table 1:
[0052] True value = 1 TP FN True value = 0 FP TN
[0053] Table 1 Confusion Matrix
[0054] Recall is the proportion of instances correctly identified as positive (true positives) by a target prediction model out of all actual positive instances. In the context of object detection, recall represents the proportion of correctly detected targets out of all actual targets, specifically expressed as:
[0055]
[0056] Where TP (True Positive) represents the number of positive samples correctly predicted by the model, i.e., true positives; FP (False Positive) represents the number of positive samples incorrectly predicted by the model, i.e., false positives; and FN (False Negative) represents the number of negative samples incorrectly predicted by the model, i.e., false negatives. Accuracy refers to the percentage of correctly detected targets. It is usually defined as the ratio of the number of correctly detected targets to the total number of detected targets, specifically expressed as:
[0057]
[0058] Wherein, TP (True Positive) represents the number of positive samples correctly predicted by the model, i.e., true positives; FP (False Positive) represents the number of positive samples incorrectly predicted by the model, i.e., false positives; and FN (False Negative) represents the number of negative samples incorrectly predicted by the model, i.e., false negatives.
[0059] This invention improves upon existing evaluation indicators to obtain the comprehensive evaluation index GD-F1, which is more suitable for evaluating the effectiveness of outdoor waste detection.
[0060] Furthermore, the specific formula for calculating the comprehensive evaluation index GD-F1 is as follows:
[0061]
[0062]
[0063]
[0064] Wherein, GD-F1 represents the comprehensive evaluation index of outdoor litter target detection, GD_Precision represents the accuracy of outdoor litter target detection, GD_Recall represents the recall of outdoor litter target detection, TP represents the number of positive samples correctly predicted by the model, FP represents the number of positive samples incorrectly predicted by the model, FN represents the number of negative samples incorrectly predicted by the model, and n represents the number of undetected targets in a dense target group when at least one target is detected.
[0065] This invention proposes an evaluation index for outdoor litter detection. The first step is to determine how many undetected targets exist within a dense cluster of targets (e.g., a pile of litter) when at least one target is detected. The method for determining this specific number is as follows:
[0066] (1) Mark dense targets (e.g., a pile of garbage) and non-dense targets (e.g., a single piece of garbage or scattered pieces of garbage): In the labeled file, find whether each target box intersects with other target boxes. This can be done by calculating whether the area of intersection between target boxes is greater than 0. If it is greater than 0, it means that there are other target boxes intersecting with that target. Then, the target box is regarded as dense garbage. For txt format, simply add the number 1 to the end of each line of the txt file to indicate that the target box is dense garbage; otherwise, add the number 0.
[0067] (2) Find the images and txt files that have been missed: Compare the multiple detected txt files with the multiple labeled txt files, that is, match each predicted txt file with the labeled txt and image. Since the four coordinate values of the txt label format are a ratio calculated based on the size of the image, the image also needs to be matched. If a target box is not detected in a certain txt file, the detected txt file and the corresponding image are extracted.
[0068] (3) Find specific target boxes that belong to dense targets but have not been detected: Each box in the labeled txt file is sequentially compared with all the boxes in the corresponding detected txt file to calculate the Intersection over Union (IoU). If the maximum IoU value is less than the threshold, the box is considered not detected. Determine whether these undetected boxes belong to dense garbage, that is, determine whether the last number of each line in the txt file is 1.
[0069] (4) Find out how many undetected boxes belong to dense targets. Since they are in a pile of garbage and at least one of the garbage has been detected, they can be considered as detected and recalled. After finding the undetected boxes that belong to dense garbage, find the target boxes that intersect with these boxes and determine whether these target boxes have been recalled. If they have been recalled, then the undetected boxes that belong to dense garbage can also be considered as recalled. If these target boxes have not been recalled, then continue to find the boxes that intersect with these target boxes, that is, the indirectly intersecting target boxes. In order to avoid finding duplicate target boxes, a set can be used to store the target boxes that have been found. In this way, the search can be continued in a loop until a box that directly or indirectly intersects with the undetected boxes that belong to dense garbage is found to be detected, or no target box different from the set can be found.
[0070] By following the four steps above, we can find out how many undetected targets are in a dense array of targets when at least one target is detected. After calculating this number (assuming it is n), we can obtain the evaluation index for outdoor waste detection.
[0071] Furthermore, the classification loss function of the target detection model is specifically as follows:
[0072]
[0073] Among them, L confLet represent the classification loss function, z represent the classification score predicted by the object detection model, and α, β, γ, θ, and σ represent different hyperparameters. To reduce the need for hyperparameter tuning during model training, we can directly set σ = 1 and θ = 2, because σ, as the exponent of q, is used to make the model more biased towards regressing high-quality bounding boxes, and θ, as the exponent of (qz), is used to measure the consistency between the regression branch and the classification branch. y = 1 indicates that the sample predicted by the object detection model is a positive sample, y = 0 indicates that the sample predicted by the object detection model is a negative sample, m = 0 indicates that the sample predicted by the object detection model is a non-dense bounding box, m = 1 indicates that the sample predicted by the object detection model is a dense bounding box, q represents the intersection-union ratio between the predicted bounding box and the ground truth bounding box predicted by the regression branch of the object detection model, hyperparameter α is used to reduce the weight contribution of dense samples, hyperparameter β is used to adjust the overall weight contribution of positive and negative samples to the classification loss, and hyperparameter γ, as the exponent of z, is used to represent the difficulty of classifying negative samples.
[0074] Add q to the positive samples σ This means that in classification, the model is biased towards samples with better bounding box regression quality, thereby accelerating the model's convergence speed. (qz) θ The main goal is to improve the consistency between classification and regression. q is the Intersection over Union (IoU) value obtained from the regression branch, and z is the classification score. If q and z are inconsistent, meaning one value is larger and the other smaller, then (qz) θ A larger value indicates that the model will be more biased towards these inconsistent samples, resulting in more consistent classification and regression performance during the inference phase, thus improving the model's detection performance and avoiding a situation where one performs well while the other performs poorly in classification and bounding box regression. γ Its main purpose is to enhance the model's attention to difficult negative samples, that is, to balance the difficulty of detection samples, thereby improving the model's classification accuracy for negative samples. For example, if the sample is negative and the predicted classification score z is small, then after z... γ After that, the classification loss for the negative sample will be very small, and the model will not pay much attention to this type of sample. Conversely, if the predicted classification score z is large, then after z... γ Subsequently, the classification loss of this negative sample becomes significantly larger compared to those easily classified negative samples, causing the model to favor the difficult-to-classify negative samples and improve the classification accuracy of the samples. α is mainly designed for the proposed evaluation index for outdoor waste detection. In this index, since the entire dense target group is considered detected as long as some targets are detected, the weight contribution of dense samples to the classification loss can be appropriately reduced, thus making the model favor non-dense samples and improving the overall detection performance.
[0075] Step S4 specifically involves adjusting the hyperparameters of the classification loss function based on the calculated comprehensive evaluation index GD_F1, followed by iterative training and model testing. Specifically, the hyperparameters α, β, and γ of the classification loss function are adjusted: hyperparameter α is used to reduce the weight contribution of dense samples, and its value can range from 0 to 1; hyperparameter β is used to adjust the overall weight contribution of positive and negative samples, and its value can also range from 0 to 1; hyperparameter γ, as the exponent of z, represents the difficulty of classifying negative samples, and its value can be 1, 2, or 3, thereby making z... γ The magnitudes of these differences vary. When adjusting hyperparameters, two hyperparameters can be fixed, for example, α and β. Both hyperparameters α and β can be set to 0.2, or one can be set to 0.1 and the other to 0.2. This allows the hyperparameter γ to be set to 1, 2, and 3 respectively. Iterative training can then be performed to observe the convergence effect of the classification loss. If the convergence effect is poor, the hyperparameters can be adjusted further.
[0076] Figure 2 This is a schematic diagram of the structure of the improved YOLOv8 model in Embodiment 1 of the present invention. Figure 2 In CSPLayer_2Conv, ConvModule is a convolutional module consisting of convolution, normalization, and activation functions. k=3 indicates a kernel size of 3, s=2 indicates a stride of 2, and p=1 indicates the padding size. CSPLayer_2Conv is composed of various ConvModules, where add=True indicates using a residual structure, add=False indicates not using a residual structure, n=3×d indicates the output channels are 3 times the number of input channels, Concat indicates concatenating two feature maps along the channel dimension, Upsample indicates upsampling, SPPF is feature pyramid pooling, and the AFM module is... Figure 3 The multi-branch dilated convolution adaptive fusion module, where Conv represents convolution.
[0077] Reference Figure 2 As shown, since there are some small targets in outdoor garbage detection, shallow detection layers are more likely to detect small targets. Therefore, a small target detection layer is added on the basis of the original YOLOv8. At the same time, in order to enhance feature extraction and enable the network to learn contextual information better, a multi-branch dilated convolution adaptive fusion module is added to the feature map of each layer in the neck of the network. This module can comprehensively utilize feature maps of different receptive fields to improve the detection effect.
[0078] Furthermore, in step S2, the target prediction model is an improved YOLOv8 network, which adds a detection layer and a multi-branch dilated convolution adaptive fusion module (AFM) to the original YOLOv8 network.
[0079] Reference Figure 3 As shown, Conv represents convolution, Rate refers to the dilation rate of the dilated convolution, ReLU represents the activation function, and upperlayer represents the feature map of the previous layer. The multi-branch dilated convolution adaptive fusion module utilizes the generated spatial weight map, similar to the weight map in spatial attention. It uses an attention-like mechanism to guide feature maps using dilated convolutions with different dilation rates. Different dilation rates result in different receptive fields for the feature maps, thus fusing features from different receptive fields. The attention mechanism can retain more effective information, achieving a better feature map fusion effect, hence the term adaptive fusion. It mainly performs adaptive fusion of feature maps from several different receptive fields, adaptively learning the fusion spatial weight coefficients for each feature map from different receptive fields. Each feature map has a corresponding spatial weight, and these weight maps guide how to fuse feature maps from different receptive fields together, retaining useful information and filtering out useless information. By adaptively fusing feature maps with different receptive fields, the model can extract features from targets of different scales effectively. This is because the feature map after adaptive feature fusion combines target features from different receptive fields. Regions with smaller receptive fields are mainly used to extract features from small targets, while regions with larger receptive fields are mainly used to extract features from large targets. By fusing feature maps with different receptive fields, the resulting feature map can simultaneously capture target features from different receptive fields, thus enabling the model to perform well in detecting targets of different scales.
[0080] Adaptive fusion is performed using a multi-branch dilated convolution adaptive fusion module, and the specific process is as follows:
[0081] S201: Perform a 1×1 convolution operation on the feature map of each layer in the neck of the network to reduce the number of channels, obtain a dimensionality-reduced feature map, and reduce the computational cost of the model.
[0082] S202: The dimensionality-reduced feature map is passed through four branches of receptive field to obtain corresponding feature information. The receptive field of the feature map is increased by using 3×3 dilated convolution with dilation rates of 1, 3, 5 and 7 respectively.
[0083] S203: Adaptive feature fusion and 1×1 convolution operation are performed on the features of different receptive fields to obtain feature maps. During the learning process, the network adaptively selects the best receptive field fusion method according to different scale targets, and can adaptively focus on important features at different scales.
[0084] S204: It effectively alleviates the problem of loss value explosion and vanishing during network gradient propagation by using residuals, speeds up training, and fuses the original features and multi-receptive field fusion features to obtain the final generated feature map.
[0085] The final generated feature map is specifically represented as follows:
[0086] y i,j =a i,j ×A i,j +b i,j ×B i,j +c i,j ×C i,j +d i,j ×D i,j
[0087] Among them, y i,j It refers to the value of the feature map at position (i,j) obtained after adaptive fusion, a i,j A is the weight parameter at position (i,j) of the first feature map after softmax. i,j It is the value of the first feature map at position (i,j). i,j B is the weight parameter at position (i,j) of the second feature map after softmax. i,j c is the value of the second feature map at position (i,j). i,j C is the weight parameter at position (i,j) of the third feature map after softmax. i,j These are the values of the three feature maps at position (i,j). d i,j D is the weight parameter at position (i,j) of the fourth feature map after softmax. i,j It is the value of the fourth feature map at position (i,j).
[0088] Specifically, refer to Figure 4 As shown, the specific process of feature fusion performed by the Adaptive Fusion Module (AFM) is as follows: Figure 4The feature maps from left to right are named A, B, C, and D. A 1×1 convolution is performed on each of these four feature maps to obtain four weight maps, each with a shape of (1, h, w), where h represents the height and w represents the width. Therefore, for the same feature map, the weights on different channels are the same, meaning each channel has equal importance. However, the weight coefficients are different for different feature maps, meaning that different feature maps are weighted and fused to obtain four weight maps. After obtaining these four weight maps, softmax processing is needed to ensure that the weight parameters are all between 0 and 1, and that the sum of the weight coefficients at the same spatial location in the four feature maps is 1. The formula is as follows:
[0089]
[0090] Among them, a i,j λ represents the weight parameters at position (i,j) of the first feature map after softmax. ai,j λ represents the weight parameter of the first feature map (i.e., feature map A) at position (i,j). bi,j λ represents the weight parameter of the second feature map (i.e., feature map B) at position (i,j). ci,j λ represents the weight parameter of the third feature map (i.e., feature map C) at position (i,j). di,j This represents the weight parameters of the fourth feature map (i.e., feature map D) at position (i,j).
[0091] Once the weight parameters of the four feature maps are obtained, the final feature map can be generated using these adaptive weight parameters, as shown in the formula:
[0092] y i,j =a i,j ×A i,j +b i,j ×B i,j +c i,j ×C i,j +d i,j ×D i,j
[0093] Among them, y i,j a represents the value of the feature map at position (i,j) obtained after adaptive fusion. i,j Let A represent the weight parameters at position (i,j) of the first feature map (i.e., feature map A) after softmax. i,j b represents the value of the first feature map (i.e., feature map A) at position (i,j). i,j B represents the weight parameters at position (i,j) of the second feature map (i.e., feature map B) after softmax processing. i,jc represents the value of the second feature map (i.e., feature map B) at position (i,j). i,j C represents the weight parameters at position (i,j) of the third feature map (i.e., feature map C) after softmax. i,j d represents the value of the third feature map (i.e., feature map C) at position (i,j). i,j Let D denote the weight parameters at position (i,j) of the fourth feature map (i.e., feature map D) after softmax. i,j This represents the value of the fourth feature map (i.e., feature map D) at position (i,j).
[0094] The entire adaptive fusion module, as Figure 4 In this context, Conv represents convolution, Softmax is an activation function that maps the value of each element to between 0 and 1. A, B, C, and D are the four feature maps input to the adaptive fusion module, a, b, c, and d are the spatial weight maps of the adaptive fusion corresponding to each feature map, and Y is the final feature map obtained.
[0095] Example 2
[0096] Embodiment 2 of the present invention proposes an image detection system for outdoor waste, the system comprising:
[0097] Data acquisition module: This module is equipped with a camera for taking photos or videos. This module also includes a communication module for transmitting photos or videos.
[0098] Object detection module: This module includes a pre-trained object detection model for object detection in real-time or offline images and videos;
[0099] Output display module: This module is used to output and display the detection results of the target detection module;
[0100] Processor module: Used to run software programs and modules stored in memory, and to perform various functions and data processing of the computer;
[0101] Storage module: This module includes a program storage area and a data storage area, used to store data, software programs, or modules.
[0102] The target detection model of this invention (the improved model) and the existing target testing model (the unimproved model) were tested using test samples, and a comparison chart of the results was obtained. (Refer to...) Figure 5As shown, the target detection model of this invention significantly improves the detection performance of outdoor litter (such as bottle caps, straws, and newspapers) compared to the existing target testing model (the model before improvement). The existing target testing model (the model before improvement) can detect piles of bottle caps, but it cannot detect individual bottle caps, resulting in missed detections. The target detection model of this invention (the improved model) can detect individual bottle caps, although the confidence level is not high, it still exceeds 0.5. Before the model improvement, straws and newspapers could not be detected, but after the improvement, they can be detected.
[0103] Ablation experiments for each module Figure 6 As shown, Parameters represents the number of parameters in the model.
[0104] The beneficial effects of this invention are as follows: First, this invention introduces an improved comprehensive evaluation index, GD-F1. During the testing of the target detection model, the comprehensive evaluation index GD-F1 is calculated. The classification loss of the target model during the training phase is adjusted through the comprehensive evaluation index GD-F1, and iterative training is conducted to obtain a target detection model that is more suitable for outdoor garbage target detection, thereby improving the detection accuracy. Second, this invention adds a detection layer and a multi-branch dilated convolution adaptive fusion module to the target detection model to improve the detection accuracy for small targets.
[0105] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include ROM, RAM, disk, or optical disk, etc.
[0106] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for image detection of outdoor waste, characterized in that, include: Acquire an image of outdoor litter to be detected, input the outdoor litter image into a trained target prediction model, and obtain the detection result of the outdoor litter image; The training process of the target prediction model includes: S1: Obtain a set of outdoor garbage images, perform data preprocessing on the outdoor garbage image set, and obtain preprocessed training samples and test samples; S2: Using the training samples, train the target prediction model and calculate the classification loss of the target prediction model; S3: The target prediction model is tested using test samples, and the comprehensive evaluation index GD-F1 is calculated based on the test results. The specific formula for calculating the comprehensive evaluation index GD-F1 is as follows: , , , Among them, GD-F1 represents the comprehensive evaluation index for outdoor litter target detection. This indicates the accuracy rate of outdoor litter target detection. TP represents the recall rate of outdoor litter target detection, FP represents the number of positive samples correctly predicted by the model, FN represents the number of negative samples incorrectly predicted by the model, and n represents the number of undetected targets in a dense target pool when at least one target is detected. S4: Adjust the parameters of the classification loss function according to the comprehensive evaluation index GD-F1, wherein the classification loss function is specifically: ; in, Let z represent the classification loss function, and z represent the classification score predicted by the object detection model. , , , ,and y=1 indicates that the sample predicted by the object detection model is a positive sample, y=0 indicates that the sample predicted by the object detection model is a negative sample, m=0 indicates that the sample predicted by the object detection model is a non-dense object box, m=1 indicates that the sample predicted by the object detection model is a dense object box, and q indicates the cross-union ratio between the predicted box predicted by the regression branch of the object detection model and the ground truth box. S5: Iterate the training until the loss function converges, then stop training to obtain the trained target prediction model.
2. The image detection method for outdoor waste according to claim 1, characterized in that, In step S1, the specific process of data preprocessing includes: S101: Collect a set of images of outdoor litter and use the labelimg tool to label the images with litter information; S102: Perform data augmentation on the labeled image set. Data augmentation methods include: horizontal flipping, cropping, mosaicking, stitching, rotation, translation, adding noise, and blurring / sharpening.
3. The image detection method for outdoor waste according to claim 1, characterized in that, In step S2, the target prediction model is an improved YOLOv8 network, which adds a detection layer and a multi-branch dilated convolution adaptive fusion module to the original YOLOv8 network.
4. An image detection system for outdoor waste, the system being used to implement the image detection method for outdoor waste as described in claim 1, characterized in that, include: Data acquisition module: This module is equipped with a camera for taking photos or videos. It also includes a communication module for transmitting photos or videos. Object detection module: This module includes a pre-trained object detection model for object detection in real-time or offline images and videos; Output display module: This module is used to output and display the detection results of the target detection module; Processor module: Used to run software programs and modules stored in memory, and to perform various functions and data processing of the computer; Storage module: This module includes a program storage area and a data storage area, used to store data, software programs, or modules.
Citation Information
Patent Citations
Deep learning method for remote sensing extraction of typical rural roads
CN117173557A
Garbage detection method in complex scene based on improved YOLOv8 model
CN117710771A