An Unmanned Aerial Vehicle Image Surface Solid Waste Detection Model and Detection Method
Through the drone image surface solid waste detection model, multiple feature extraction, asymmetric decomposition and semantic flow alignment processing are used to solve the problem of limited feature extraction performance in garbage detection, and the precise detection of complex backgrounds and diverse garbage is achieved, which is suitable for urban-level full-domain applications.
Patent Information
- Application Number
- CN202210973707.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-15
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-08-15
AI Technical Summary
In the garbage detection, the existing technology has problems such as irregular changes in the edges of solid waste stacking points, different scales, numerous types, diverse forms, sparse distribution, and complex backgrounds, resulting in limited feature extraction performance and lack of large-scale application capabilities, especially the applicability in different regions.
The drone image surface solid waste detection model is adopted, and through multiple feature extraction, asymmetric decomposition, channel attention processing and semantic flow alignment processing, combined with hollow convolution group and deconvolution module, feature extraction and information transmission are optimized to achieve more accurate garbage detection.
It realizes accurate detection of solid waste with complex backgrounds, different scales, diverse forms, sparse and other types. It is suitable for urban-level full-domain inspection, with a wider scope of application and higher detection accuracy.
Smart Images

Figure CN115330720B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of garbage detection, and in particular relates to an unmanned aerial vehicle (UAV) image-based surface solid waste detection model and a detection method. Background Art
[0002] With the continuous increase in the global population and the acceleration of the urbanization process, the generation of solid waste is increasing day by day, and its management and disposal are a global challenge. Driven by illegal interests, behaviors such as unauthorized dumping or stacking of construction waste and domestic waste, and unauthorized establishment of waste disposal sites to receive construction waste occur from time to time, making supervision difficult. Among them, a few supervision blind spots form solid waste storage sites, causing a huge impact on the environment.
[0003] The detection of garbage dumps is mainly the responsibility of urban management (housing construction), ecological environment, water conservancy, agriculture and rural areas at all levels. The traditional detection methods are based on public reports or regular on-site inspections, both of which are manual methods. The traditional detection methods require a large amount of manpower and material resources, and the investigation is usually lagging. Moreover, due to the diverse distribution environments of garbage, there are inspection blind spots in complex environmental areas that staff cannot reach, easily forming a long-term accumulation situation. Therefore, seeking a non-formal garbage stacking detection method with high precision, high efficiency, low cost and informatization is an urgent problem to be solved by urban management and other departments for rectification and investigation.
[0004] For this reason, scholars at home and abroad have proposed to apply remote sensing technology to garbage stacking points. Relevant research mainly includes the analysis of the suitability of garbage dump site selection, the monitoring of garbage dump capacity, the settlement of garbage dumps, the identification and recycling of garbage types, and the monitoring of environmental impacts. However, most of these studies are based on the known locations of garbage dumps. However, the hidden illegal garbage stacking points are not treated, and relatively speaking, they cause greater damage to the environment. Therefore, it is necessary to seek a fast and efficient way to detect the locations of potential garbage stacking points.
[0005] In response, some experts and scholars have conducted regional prediction research from the perspective of temperature characteristics. Gill processed Landsat imagery of Kuwait from 1985 to 1994, using thermal remote sensing to measure the surface temperature (LST). He calculated LST within landfills and, by combining multi-temporal LST contours and overlay analysis, delineated the most likely dumping areas within landfills. Temperature-based detection methods have low sensitivity, and small landfills do not significantly alter surface temperature, making them only suitable for detecting large landfills. Other researchers have considered feature extraction. Angelino et al. developed a multi-feature detection algorithm based on four key features: type, state, location, and activity, using optical satellite imagery of the provinces of Naples and Caserta. This algorithm automated the detection of illegal dumping and tracked the evolution of identified landfills. Hema Begur proposed an edge-based intelligent mobile service system for illegal dumping detection and monitoring in San José. Wang Chen et al. (2016) integrated empirical knowledge of image texture, hue, shape, and location distribution to construct a remote sensing recognition expert knowledge base for landfill information extraction. This type of traditional extraction algorithm mainly focuses on designing a suitable feature set for the target for expression and automatic identification. However, manually designed features require a lot of preliminary work to analyze and compare various features. Although the extraction effect is good in some samples, due to the complex and diverse structure and texture of garbage dump sites, it cannot be applied on a large scale to full-domain automated extraction.
[0006] In recent years, artificial intelligence, and in particular deep learning, has gradually emerged as a new machine learning model whose concept originates from the study of artificial neural networks. The term "deep" is used in contrast to shallow learning methods such as support vector machines (SVMs), boosting, and maximum entropy methods (Yin Baocai et al., 2015). Compared with conventional methods, deep learning possesses more powerful feature learning and expression capabilities. It can automatically and multi-layeredly extract abstract features of complex objects, thereby mining the deep feature information implicit in image data, and can be applied to a variety of complex classification scenarios. Currently, it has achieved a series of breakthrough research results in image classification, object detection, semantic segmentation, and face recognition.
[0007] Deep learning has been applied at all levels of remote sensing data analysis: from traditional image preprocessing, object classification and recognition, extracting task features, and understanding high-level semantic information such as image scenes. Specifically, it has found widespread application in urban land use classification, road extraction, building extraction, cloud detection, and change detection. Numerous experimental results demonstrate the excellent performance of deep learning-based algorithms in remote sensing data analysis.
[0008] Applying deep learning to the field of garbage detection, Youme proposed an automatic solution for detecting hidden landfills in the Saint-Louis region of Senegal, West Africa, using drone imagery. The solution selected the Single Shot Detector (SSD) as the detection algorithm and used VGG16 in Convolutional Neural Networks (CNNs) as the base model. During the training process, prior boxes with different scales or aspect ratios were set for each unit to reduce the training difficulty, and finally, convolution was directly used to extract the detection results from different feature maps. Yu proposed a street garbage detection algorithm based on Sift feature image registration and region-based convolutional neural network (R-CNN). First, a CNN network trained on a sample set composed of garbage and non-garbage images was established on real-time street images. Second, for each clean street image, Sift features were obtained and image registration was performed with the real-time street image to obtain a difference image, reducing the detection range. Finally, the real-time image was input into the CNN to obtain the detection results. Cui Wenjing (2019) studied marine floating garbage using video surveillance and constructed a convolutional neural network model based on semantic segmentation. Four training schemes were designed to extract orange strip-shaped substances (possibly formed by a mixture of foam and suspended sediment produced by ship sewage discharge) and wood chip-like garbage in the pictures. Wu Tong et al. (2020) used convolutional neural network technology to propose a sample updated-Retina Net (SU-Retina Net) framework for detecting informal garbage dumps in high-resolution remote sensing images, analyzed the influence of different parameters and network structures on the model detection effect, and improved the detection efficiency of informal garbage dumps.
[0009] However, although many scholars have proposed ideas for applying deep learning to the field of garbage detection, the research on solid waste detection using deep learning is still in its infancy, and the above-mentioned related solutions still have the following problems:
[0010] (1) The edges of solid waste dumps change irregularly, with different scales, numerous types, diverse shapes, sparse distributions, and complex backgrounds. No model extraction module has been designed specifically for the characteristics of solid waste, and the extraction performance for various types of garbage dumps is limited.
[0011] (2) Most are limited to distinguishing and identifying garbage types in close-range views, lacking large-scale applications. Many literatures often only study the garbage extraction effect in different regions of the same image. There are doubts about whether the network trained on a specific image can maintain a comparable extraction ability for different regions. Summary of the Invention
[0012] The object of the present invention is to provide a surface solid waste detection model and detection method for UAV images in view of the above problems.
[0013] To achieve the above object, the present invention adopts the following technical solutions:
[0014] A surface solid waste detection method for UAV images, characterized in that the method includes:
[0015] S1. Perform multiple feature extractions on the input image to obtain multiple layers of features in sequence.
[0016] S2. Perform asymmetric decomposition on the deepest layer feature in step S1 and use its output as the restored deepest layer feature.
[0017] S3. Perform channel attention processing on the previous layer feature, perform semantic flow alignment processing on the previous layer feature and the restored next layer feature, and then superimpose the output of the channel attention processing and the output of the semantic flow alignment processing to output the restored previous layer feature.
[0018] Perform the above processing layer by layer on each layer of features until the first layer of features is restored.
[0019] S4. Perform deconvolution processing on the restored first layer of features to restore the features to the input size.
[0020] S5. Output a prediction result based on the features restored in step S4.
[0021] This method performs multiple feature extractions on the input image, performs asymmetric decomposition processing on the deepest layer, and uses it as the restored deepest layer feature. At the same time, starting from the restored deepest layer feature, channel attention and semantic flow alignment processing are used for each subsequent layer of features to restore them layer by layer, which can enhance the feature extraction ability of the model, achieve more effective information extraction and transmission, and effectively optimize the problem of edge semantic misalignment caused in the resolution restoration process.
[0022] In the above surface solid waste detection method for UAV images, the spatial resolution of the input image is 0.038 m, the scale is 256 * 256, and the channel dimension is 3. In step S1, four feature extractions are performed on the input image, and the size of the finally obtained deepest layer feature map is 16 * 16, and the channel dimension is 256.
[0023] In the above surface solid waste detection method for UAV images, step S2 specifically includes:
[0024] S21. Process the deepest layer feature using the first layer of dilated convolutional layer.
[0025] S22. Process the output of the first layer of dilated convolutional layer using the second layer of dilated convolutional layer.
[0026] S23. Use regularization and the ReLU function to perform arithmetic processing on the output after the processing in step S22, and output the asymmetric decomposition atrous convolution operation results with different dilation rates;
[0027] S24. Superimpose the deepest layer features and the convolution outputs with different dilation rates in S23 and use them as the restored deepest layer features.
[0028] In the above-mentioned method for detecting surface solid waste in UAV images, each atrous convolution layer includes a plurality of atrous convolutions arranged in parallel.
[0029] In the above-mentioned method for detecting surface solid waste in UAV images, each atrous convolution layer includes four atrous convolutions arranged in parallel, and the dilation rates of each atrous convolution are 3, 5, 7, and 9 respectively.
[0030] In the above-mentioned method for detecting surface solid waste in UAV images, in step S3, the method for performing channel attention processing on the features of the previous layer includes:
[0031] S31A. Compress the features of the previous layer into 1*1*C, where C represents the channel dimension;
[0032] S32A. Perform convolution and activation function processing on the compressed features obtained in step S31A to obtain the weights of each channel;
[0033] S33A. Multiply the result of step S32A by the features of the previous layer described in step S31A and then output it to achieve weighted adjustment of the channel features and complete the recalibration of the remote low-dimensional features in the channel dimension.
[0034] In the above-mentioned method for detecting surface solid waste in UAV images, in step S3, the method for performing semantic flow alignment processing on the features of the previous layer and the restored features of the next layer includes:
[0035] S31B. Perform 1*1 convolution on the features of the previous layer and the restored features of the next layer respectively;
[0036] S32B. Upsample the restored features of the next layer to obtain an upsampled feature map;
[0037] S33B. Superimpose the upsampled feature map and the features of the previous layer in the channel dimension;
[0038] S34B. Perform convolution operation on the superimposed result in step S33B to obtain the semantic offset field between the upsampled feature map and the features of the previous layer;
[0039] S35B. Guide the upsampling of the restored features of the next layer according to the semantic offset field to obtain the output of the accurately decoded semantic feature map.
[0040] In the above-mentioned method for detecting surface solid waste in UAV images, in step S5, the Sigmoid function is used to output pixel-by-pixel prediction results based on the features restored in step S4; and the pixels with probability values greater than the probability threshold are used as the target recognition results.
[0041] In step S1, the four times of feature extraction performed on the input image are as follows:
[0042] For the first time, the input image of 3*256*256 is downsampled and convolved to obtain the first-layer features with a size of 64*128*128.
[0043] For the second time, the first-layer features of 64*128*128 are downsampled and convolved to obtain the second-layer features with a size of 64*64*64.
[0044] For the third time, the second-layer features of 64*64*64 are downsampled and convolved to obtain the third-layer features with a size of 128*32*32.
[0045] For the fourth time, the third-layer features of 128*32*32 are downsampled and convolved to obtain the fourth-layer features with a size of 256*16*16.
[0046] The first value of the above feature size is the channel dimension of the corresponding feature, and the second and third values are the sizes of the corresponding feature maps.
[0047] A model for detecting surface solid waste in UAV images includes a feature extraction module, an atrous convolution group, a high-dimensional feature precise semantic alignment module, and a transposed convolution module.
[0048] The said feature extraction module is used to perform multiple feature extractions on the input image.
[0049] The said atrous convolution group is used to perform asymmetric decomposition on the deepest layer of features extracted by the feature extraction module and output it as the restored deepest layer of features.
[0050] The said high-dimensional feature precise semantic alignment module is used to perform attention processing and semantic alignment processing based on the restored next layer of features and the previous layer of features extracted by the feature extraction module to restore the previous layer of features layer by layer until the first layer of features is restored.
[0051] The said transposed convolution module is used to perform transposed convolution operations on the restored first layer of features to restore the features to the input size for subsequent result prediction.
[0052] In the above-mentioned model for detecting surface solid waste in UAV images, the said feature extraction module includes four layers of feature extraction layers, which are used to perform four times of feature extraction on the input image with a size of 3*256*256, and finally the size of the deepest layer of features obtained is 256*16*16.
[0053] The described dilated convolution group includes a first - layer dilated convolution layer, a second - layer dilated convolution layer, and a regularization + relu function operation layer. The first - layer dilated convolution layer is used to process the deepest - layer features. The second - layer dilated convolution layer is used to process the output of the first - layer dilated convolution layer. The regularization + relu function operation layer is used to perform operation processing on the output processed by the second - layer dilated convolution layer, superimpose the deepest - layer features and the convolution outputs with different dilation rates after passing through the activation function, and finally output the operation result of the asymmetric - decomposition dilated convolution group, which is used as the restored deepest - layer features;
[0054] And each layer of the dilated convolution layer includes four dilated convolutions in parallel, and the dilation rates of each dilated convolution are 3, 5, 7, and 9 respectively. The scale of the deepest - layer feature map in this scheme is 16 * 16. Selecting the dilation rates of 3, 5, 7, and 9 can achieve the required multi - scale feature learning. This structure can provide a larger receptive field to aggregate multi - scale context information without using a pooling layer. The original input features are superimposed with the convolution outputs of different dilation rates, making the module tend to learn the differences between the input features and the ground truth during the learning process, preventing the model from degrading during the learning process and further approaching the identity mapping.
[0055] The described high - dimensional feature precise semantic alignment module includes an attention module and a semantic flow alignment module. The attention module is used to perform:
[0056] S31A. Compress the features of the previous layer into 1 * 1 * C;
[0057] S32A. Obtain the weights of each channel by performing convolution and activation function processing on the compressed features obtained in step S31A;
[0058] S33A. Multiply the result of step S32A with the features of the previous layer described in step S31A and output;
[0059] The described semantic flow alignment module is used to perform:
[0060] S31B. Perform 1 * 1 convolutions on the features of the previous layer and the restored features of the next layer respectively;
[0061] S32B. Upsample the restored features of the next layer to obtain an upsampled feature map;
[0062] S33B. Superimpose the upsampled feature map and the features of the previous layer in the channel dimension;
[0063] S34B. Perform convolution operation on the superimposed result in step S33B to obtain the semantic offset field between the upsampled feature map and the features of the previous layer;
[0064] S35B. Upsample the restored next-layer features according to the semantic offset field to obtain an output of an accurately decoded semantic feature map;
[0065] The high-dimensional feature accurate semantic alignment module superimposes the output of the attention module and the output of the semantic flow alignment module and outputs to obtain the restored upper-layer features. This structure can obtain an accurately decoded semantic feature map and achieve accurate semantic fusion of high-dimensional information and low-dimensional information.
[0066] The advantages of the present invention are as follows:
[0067] 1. This solution realizes more effective information extraction and transmission based on the asymmetric decomposition dilated convolution group and the high-dimensional feature accurate semantic alignment module, and can optimize the edge semantic misalignment problem caused in the resolution restoration process, so as to achieve more accurate and complete feature extraction;
[0068] 2. This solution can achieve relatively accurate detection and recognition of solid wastes with complex backgrounds, different scales, diverse forms, and different sparsities;
[0069] 3. The solid waste recognition network structure SWE proposed in this solution can be used for city-level global detection, can be applied to various solid waste types and complex distribution scenarios, and has a wider application range and higher accuracy advantage compared with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 is the model structure block diagram of the surface solid waste detection model of the unmanned aerial vehicle image of the present invention;
[0071] Figure 2 is an example diagram of the asymmetric decomposition of a 3×3 dilated convolution with a parameter dilation rate of 2;
[0072] Figure 3 is the structure diagram of the dilated convolution group in the present invention;
[0073] Figure 4 is the structure diagram of the high-dimensional feature accurate semantic alignment module of the present invention;
[0074] Figure 5 is the model structure of the surface solid waste detection model of the unmanned aerial vehicle image of the present invention and the flowchart of its feature extraction method;
[0075] Figure 6 is a comparison example diagram of multi-class solid waste detection;
[0076] Figure 7 is an example diagram of the solid waste detection result in a local area of Jiaxing City. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0077] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0078] This solution provides a model for detecting surface solid waste in UAV remote sensing images, and realizes the accurate detection of surface solid waste based on this model.
[0079] As Figure 1 shown, the UAV image surface solid waste detection model includes a feature extraction module 1, an atrous convolution group 2, a high-dimensional feature precise semantic alignment module 3, and a deconvolution module 4;
[0080] The feature extraction module 1 is used to perform multiple feature extractions on the input image; in this embodiment, the feature extraction module 1 includes four layers of feature extraction layers, which are used to perform four feature extractions on the input image with a size of 3*256*256 and finally obtain the deepest layer feature with a size of 256*16*16;
[0081] The atrous convolution group 2 is used to perform asymmetric decomposition convolution on the deepest layer feature extracted by the feature extraction module 1 and output it as the restored deepest layer feature, that is, the fourth layer feature here.
[0082] Figure 2 Shown is an example of the asymmetric decomposition of a 3*3 atrous convolution with a parameter dilation rate of 2. Replace the 3×3 convolution in the convolution group with the combination of a 3×1 convolution (three dark blocks in the horizontal direction) and a 1×3 convolution (three dark blocks in the vertical direction) through asymmetric decomposition. They have the same receptive field, and the replaced combination can reduce the number of parameters by 33% and speed up the model running speed. In addition, the asymmetric decomposition also enables the model to learn the target to extend features in the horizontal and vertical directions, which helps to locate the target edge.
[0083] The high-dimensional feature precise semantic alignment module 3 is used to perform attention processing and semantic alignment processing based on the restored next layer feature and the upper layer feature extracted by the feature extraction module 1 to restore the upper layer feature layer by layer until the first layer feature is restored. That is, based on the restored fourth layer feature and the third layer feature extracted by the feature extraction module 1, perform attention processing and semantic alignment processing to obtain the restored third layer feature. Based on the restored third layer feature and the second layer feature extracted by the feature extraction module 1, perform attention processing and semantic alignment processing to obtain the restored second layer feature. Based on the restored second layer feature and the first layer feature extracted by the feature extraction module 1, perform attention processing and semantic alignment processing to obtain the restored first layer feature. By using the attention mechanism to optimize the remote low-dimensional information for the semantic misalignment problem that is prone to occur when fusing high-dimensional information and low-dimensional information in semantic segmentation, and at the same time improving the matching accuracy of high-dimensional information in the resolution restoration process by adopting a semantic flow alignment structure, it is possible to avoid the adverse impact on feature extraction caused by the disorder of the solid waste edge.
[0084] The deconvolution module 4 is used to perform deconvolution operations on the restored first-layer features to restore the features to the input size for subsequent result prediction, that is, to restore the restored first-layer features to 2*256*256. Subsequently, the Sigmoid function module outputs pixel-by-pixel prediction results based on the restored 2*256*256 feature map.
[0085] Specifically, as Figure 3 shown, the dilated convolution group 2 of this embodiment includes a first-layer dilated convolution layer and a second-layer dilated convolution layer. After the two convolution groups, regularization and relu function operations are uniformly used to output as the deepest restored features. Figure 3 In represents convolution, represents upsampling, represents dimension stacking, represents numerical stacking, represents offset field scheduling.
[0086] Dilated convolution uses a jumping receptive field to replace the conventional receptive field. The parameter dilation rate defines the spacing of each value when the convolution kernel processes data. This module is composed of four parameter dilation rates of [3, 5, 7, 9]. Different combinations of dilation rates enable each intermediate pixel of the image to encode semantic features from different scales. Since the pixel sampling of the convolution layer with a large dilation rate is relatively sparse, the combination of dilation rates enables each pixel to obtain more dense information within the neighborhood interval range, which can ensure the feature proportion and spatial correlation.
[0087] Here, the dilated convolution group 2 is used to extract multi-scale features at the deepest layer of the network. The dilated convolution group 2 has a total of 5 branches, one of which is the original input feature, and the remaining branches are asymmetrically decomposed dilated convolutions, composed of four parameter dilation rates of [3, 5, 7, 9]. The operation results of the 5 branches are summed and output. Such a structure can enhance the model's ability to extract spatial context information of the input image, enable the model to be applicable to the extraction of multi-modal and multi-scale target features, can provide a larger receptive field to aggregate multi-scale context information without using a pooling layer, and at the same time make the module tend to learn the difference between the input feature and the ground truth during the learning process, so that the model will not degenerate during the learning process and further approximate the identity mapping. And adding regularization and relu function operations can prevent the model from overfitting and enhance non-linearity.
[0088] Specifically, as Figure 4As shown in the figure, the high-dimensional feature precise semantic alignment module 3 includes an attention module (the dotted part in the figure) and a semantic flow alignment module. For the attention module, the Squeeze-Excitation method is adopted. First, the feature is compressed into 1*1*C. The compressed feature has a global receptive field and can represent the main information of each channel. The weights of each channel are obtained through convolution and multiplied by the input feature to achieve weighted adjustment of the channel features, complete the recalibration of the remote low-dimensional features in the channel dimension, and are superimposed on the input feature, enabling the attention module to learn the difference between the existing feature and the true value feature, and realizing the integration and optimization of multi-channel information. Regarding the semantic flow alignment module, the low-dimensional feature and the high-dimensional feature are each subjected to a 1*1 convolution for feature channel information interaction. The high-dimensional feature is upsampled to obtain an upsampled feature map, and then the two features are superimposed in the dimension, and a convolution operation is performed to obtain the semantic offset field between the upsampled feature map and the low-dimensional feature. Finally, the high-dimensional feature map is guided by the semantic offset field to be upsampled to obtain an accurately decoded semantic feature map. Finally, the sum of the features optimized by the two modules is output to achieve the precise semantic fusion of high-dimensional information and low-dimensional information.
[0089] Finally, the high-dimensional feature precise semantic alignment module 3 superimposes the output of the attention module and the output of the semantic flow alignment module and outputs to obtain the restored upper-layer feature.
[0090] Furthermore, as Figure 5 shown, the method for detecting surface solid waste in UAV images in this solution is as follows. In the figure, IFM represents the high-dimensional feature precise semantic alignment module 3. The size of the input image is 256*256, and the number of channel dimensions is 3:
[0091] S1. The feature extraction module 1 performs four times of feature extraction on the input image to obtain multiple layers of features in sequence. The four times of feature extraction are respectively:
[0092] For the first time, the input image of 3*256*256 is subjected to downsampling convolution to obtain the first-layer feature with a size of 64*128*128;
[0093] For the second time, the first-layer feature of 64*128*128 is subjected to downsampling convolution to obtain the second-layer feature with a size of 64*64*64;
[0094] For the third time, the second-layer feature of 64*64*64 is subjected to downsampling convolution to obtain the third-layer feature with a size of 128*32*32;
[0095] For the fourth time, the third-layer feature of 128*32*32 is subjected to downsampling convolution to obtain the fourth-layer feature with a size of 256*16*16;
[0096] The first value in each of the above dimensions is the channel dimension of the corresponding feature, and the second and third values are the sizes of the corresponding feature maps, that is, the width and length.
[0097] S21. Process the deepest layer features using the first dilated convolutional layer;
[0098] S22. Process the output of the first dilated convolutional layer using the second dilated convolutional layer;
[0099] S23. Perform arithmetic operations on the output processed in step S22 using regularization and the relu function, and output the asymmetric decomposition dilated convolutional operation results with different dilation rates;
[0100] S24. Superimpose the deepest layer features and the convolutional outputs with different dilation rates in S23 and use them as the restored deepest layer features.
[0101] The attention module of the high-dimensional feature precise semantic alignment module 3 performs the following steps:
[0102] S31A. Compress the features of the previous layer into 1*1*C, where C represents the channel dimension;
[0103] S32A. Perform convolution and activation function processing on the compressed features obtained in step S31A to obtain the weights of each channel;
[0104] S33A. Multiply the result of step S32A by the features of the previous layer in step S31A and output it to achieve weighted adjustment of the channel features and complete the recalibration of the remote low-dimensional features in the channel dimension.
[0105] The semantic flow alignment module of the high-dimensional feature precise semantic alignment module 3 performs the following steps:
[0106] S31B. Perform 1*1 convolution on the features of the previous layer and the restored features of the next layer respectively;
[0107] S32B. Upsample the restored features of the next layer to obtain an upsampled feature map;
[0108] S33B. Superimpose the upsampled feature map and the features of the previous layer in the channel dimension;
[0109] S34B. Perform convolution operation on the superimposed result in step S33B to obtain the semantic offset field between the upsampled feature map and the features of the previous layer;
[0110] S35B. Guide the upsampling of the restored features of the next layer according to the semantic offset field to obtain the output of the precisely decoded semantic feature map.
[0111] Then, step S36 is executed: S36. The high-dimensional feature precise semantic alignment module 3 superimposes the output of the attention module and the output of the semantic flow alignment module and outputs to obtain the restored feature of the upper layer.
[0112] The above steps S31A - S36 are processed layer by layer for each layer of features until the first layer of features is restored. At this time, the size of the restored first layer of features is 64 * 128 * 128. There is no corresponding low-dimensional feature map for the 128 scale, so the image size is restored to 256 using the transposed convolution method later.
[0113] S4. The transposed convolution module 4 performs transposed convolution processing on the restored first layer of features to restore the features to the input size of 3 * 256 * 256.
[0114] S5. Use the Sigmoid function to output the pixel-by-pixel prediction result based on the features restored in step S4; and use the pixels with probability values greater than the probability threshold as the target recognition result. The probability threshold can be set to 0.5.
[0115] The solid waste recognition network structure proposed in this solution strengthens the feature extraction ability of the model by adopting a dilated convolution group for asymmetric decomposition at the deepest layer of the network. At the same time, a high-dimensional feature precise semantic alignment module is provided. This module performs a series of processes on high-dimensional features and low-dimensional features, and guides the upsampling of the high-dimensional feature map according to the semantic offset field to obtain an accurately decoded semantic feature map. Finally, it can achieve the precise semantic fusion of high-dimensional information and low-dimensional information. Through the combination of the aforementioned modules, more effective information extraction and transmission are realized, and the problem of edge semantic misalignment caused during the resolution restoration process is optimized, enabling the model to maintain high detection performance and accurate detection effects in solid waste detection scenarios with complex backgrounds, different scales, diverse shapes, and sparse differences. Therefore, it can be used for urban-level global detection, and the model is applicable to various solid waste types and complex distribution scenarios.
[0116] Through experimental comparison, it is found that this solution has a relatively obvious accuracy advantage compared with other existing semantic segmentation methods. The following is the experimental comparison process and comparison results:
[0117] The experimental data is the local UAV image data of Nanhu District, Jiaxing City, Zhejiang Province, China collected in June 2021. The spatial resolution is 0.038m. The original images contain two scenes with 194374×301992 pixels and 146917×160405 pixels respectively, and the total area is about 100 square kilometers. Artificial annotation is performed on the garbage targets larger than 0.5 square meters, and 256×256 pixel image patches are cropped centered on the targets. The images without targets are not used as the dataset. No data augmentation is performed. After distribution, the training set, validation set, and test set are 410, 90, and 137 respectively.
[0118] The network of this solution is implemented on the open-source deep learning framework Pytorch 1.2.0. All experiments are trained for 100 batches. The initial learning rate is set to 0.0001, the optimizer is Adam, and the loss function is BCEloss. The network is trained on NVIDIA RTX2080Ti, and the batch size is usually set to 8. For large networks, the maximum runnable value is taken.
[0119] The comparison networks in this experiment include ENet, UNet, NestedUnet, DeepLabv3+, DaNet, PSPNet, SegFormer, ERFNet, DDRNet, and the model SWE proposed in this solution. The F1 value, overall accuracy, recall rate, and precision are selected as the numerical indicators for performance analysis and comparison. Four quantitative indicators, namely Precision, Recall, F1-score, and overall accuracy (OA), are selected to evaluate the comparative methods in the invention.
[0120]
[0121]
[0122]
[0123]
[0124] In the above formulas, pixels are used as the evaluation unit, f i is the predicted value, y i is the true value. The symbols in the binary classification confusion matrix are as follows: TP: The target is correctly predicted, actually true and predicted true; TN: The target is missed, actually true and predicted false; FP: The target is misdetected, actually false and predicted true; FN: The background is correctly detected, actually false and predicted false. Precision represents the proportion of correct predictions in the extraction results; Recall represents the proportion of all targets extracted; The F1-score is the harmonic mean of Precision and Recall, and evaluates the algorithm performance from a comprehensive perspective; OA represents the ratio between all correctly predicted pixels and the total pixels.
[0125] The accuracy evaluation results of ten different models are shown in Table 1:
[0126] Table 1 Model Accuracy Evaluation Results
[0127]
[0128]
[0129] Here, they are arranged in ascending order of F1 value, and the highest value of each index is bolded. As can be seen from the table, the lowest F1 value in the comparison method is SegFormer, only 74.587%, the highest of NestedUNet is 86.56%, the F1 values of other models are around 85%, while the F1 value of this solution is as high as 89.467%, which is 2.9% higher than the highest NestedUNet method. In addition, it also has the highest total accuracy and recall rate among all methods, and is only 0.5 percentage points lower than the ENet model in terms of precision. Generally speaking, the model proposed in this paper performs the best in the quantitative analysis of indicators.
[0130] In addition, Figure 6 are some typical images tested in the dataset, and SWE is the image obtained by this solution:
[0131] The target in Region 1 is construction waste, and the contrast between the background and the target is relatively obvious. There are a large number of missed extractions in the extraction results of SegFormer and PSPNet in the comparison models, and other models can extract the target well, with only a small amount of error at the edge.
[0132] Region 2 is mainly household waste. The upper left corner is connected to other garbage, so it is marked. The area of the lower left corner is less than 0.5 square meters, so it is not marked. The models SegFormer, Deeplabv3+, and RAST can extract the targets in the two small regions, while other models all have missed extractions.
[0133] Region 3 includes white industrial waste and a small amount of other sundries. It is easy to mis-extract by confusing the background bare land with sundries. In the comparison models, DDRNet and ERFNet have too much over-extraction, and UNet and NestedUNet also have mis-extractions of bare land. PSPNet performs better in this region, but there are still a small number of missed extractions. Although this solution mis-extracts the blue box in the upper left corner as a target and has a small amount of mis-extraction at the edge, the overall extraction result is relatively accurate and little affected by the background.
[0134] Region 4 has a green space background and there are two scattered garbage piles. DDRNet, ENet, PSPNet, and NestedUNet do not recognize the lower targets, and other models also have errors at the edge. Only this solution extracts a relatively complete and accurate image.
[0135] In summary, it can be seen that the SWE proposed in this solution has the best extraction ability for various types of solid waste.
[0136] Generally speaking, the variety, scale, texture, and morphology of garbage are diverse, making it difficult to extract the target well. The complexity of the background further increases the test of the model's learning ability. The model proposed in this solution can adapt to various complex environments and has a relatively complete recognition ability for different types and distributions of targets, showing the best performance among multiple models. In addition, applying this solution to the UAV images of a local area in Jiaxing City for solid waste detection, the result examples are as shown in Figure 7 shown. The actual detection result map uses different colors to identify different wastes, and the identification results are close to the actual distribution of solid waste. The detection results show that this solution can be applied in a large-scale space.
[0137] The specific embodiments described in this article are only illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar methods to replace them, but they will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A method for detecting surface solid waste in UAV images, characterized in that, The method includes: S1. Perform multiple feature extractions on the input image to sequentially obtain multiple layers of features; S2. Asymmetrically decompose the deepest layer feature in step S1 and output it as the restored deepest layer feature; S3. Perform channel attention processing on the previous layer feature, perform semantic flow alignment processing on the previous layer feature and the restored next layer feature, and then superimpose the output of the channel attention processing and the output of the semantic flow alignment processing to output the restored previous layer feature; Perform the above processing layer by layer for each layer of features until the first layer of features is restored; S4. Perform deconvolution processing on the restored first layer of features to restore the features to the input size; S5. Output the prediction result based on the features restored in step S4.
2. The method for detecting surface solid waste in drone images according to claim 1, wherein, The spatial resolution of the input image is 0.038m, the scale is 256*256, and the number of channels is 3. In step S1, the input image is subjected to four feature extractions, and the size of the finally obtained deepest layer feature map is 16*16, and the number of channels is 256.
3. The method for detecting surface solid waste in UAV images according to claim 1 or 2, characterized in that, Step S2 specifically includes: S21. Process the deepest layer feature using the first layer of dilated convolutional layer; S22. Process the output of the first layer of dilated convolutional layer using the second layer of dilated convolutional layer; S23. Perform arithmetic processing on the output processed in step S22 using regularization and relu functions to output the asymmetric decomposition dilated convolutional operation results with different dilation rates; S24. Superimpose the deepest layer feature and the convolutional outputs with different dilation rates in S23 and use it as the restored deepest layer feature.
4. The method for detecting surface solid waste in UAV images according to claim 3, wherein Each layer of dilated convolutional layer includes a plurality of dilated convolutions arranged in parallel.
5. The method for detecting surface solid waste in UAV images according to claim 4, characterized in that, Each layer of dilated convolutional layer includes four dilated convolutions arranged in parallel, and the dilation rates of each dilated convolution are 3, 5, 7, and 9 respectively.
6. The method for detecting surface solid waste from UAV images according to claim 1 or 2, characterized in that, In step S3, the method for performing channel attention processing on the previous layer feature includes: S31A. Compress the previous layer feature to 1*1*C, where C represents the channel dimension; S32A. Perform convolution and activation function processing on the compressed feature obtained in step S31A to obtain the weights of each dimension; S33A. Multiply the result of step S32A by the previous layer feature described in step S31A and then output.
7. The method for detecting surface solid waste from UAV images according to claim 6, wherein, In step S3, the method for performing semantic flow alignment processing on the previous layer feature and the restored next layer feature includes: S31B. Perform 1*1 convolution on the previous layer feature and the restored next layer feature respectively; S32B. Upsample the restored next layer feature to obtain an upsampled feature map; S33B. Superimpose the upsampled feature map and the previous layer feature in the channel dimension; S34B. Perform convolution operation on the superimposed result in step S33B to obtain the semantic offset field between the upsampled feature map and the previous layer feature; S35B. Guide the upsampling of the restored next layer feature according to the semantic offset field to obtain the output of the accurately decoded semantic feature map.
8. The method for detecting surface solid waste in UAV images according to claim 1, wherein In step S5, use the Sigmoid function to output the per-pixel prediction result based on the features restored in step S4; and use the pixels with probability values greater than the probability threshold as the target recognition result; In step S1, the four feature extractions performed on the input image are respectively: First, perform downsampling convolution on the input image of 3*256*256 to obtain the first layer of features with a size of 64*128*128; Second, perform downsampling convolution on the first layer of features of 64*128*128 to obtain the second layer of features with a size of 64*64*64; Third, perform downsampling convolution on the second layer of features of 64*64*64 to obtain the third layer of features with a size of 128*32*32; Fourth, perform downsampling convolution on the third layer of features of 128*32*32 to obtain the fourth layer of features with a size of 256*16*16.
Citation Information
Patent Citations
Real-time streetscape image semantic segmentation method based on staged feature semantic alignment
CN113011429A
Polarimetric SAR image classification method based on a channel attention depth network
CN113240040A