Automatic identification method for remote sensing image of solid waste landfill based on deep learning

By constructing the GSFF-VIT model, the problems of inaccurate identification results, scarce data sets and relying on manual feature extraction in automatic identification of remote sensing images in solid waste landfills are solved, and high-accuracy and high-efficiency landfill recognition are achieved.

CN119942327AActive Publication Date: 2025-05-06GUANGZHOU UNIVERSITY

Patent Information

Application Number
CN202510010736.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

The prior art has problems such as inaccuracy of identification results, scarcity of data sets, limitations of identification methods and relying on manual feature extraction in automatic identification of remote sensing images in solid waste landfills.

Method used

Using a deep learning-based method, the GSFF-VIT model is built, and the remote sensing image data of solid waste landfills is trained and tested by integrating multiple deep learning models to achieve automatic recognition. By improving the VIT model, this model enhances the integration of semantic feature information and spatial coordinates in picture patches, and improves the accuracy of recognition.

Benefits of technology

It improves the accuracy and efficiency of solid waste landfill identification, realizes the transition from manual feature extraction to data-driven end-to-end feature extraction, which is suitable for landfill identification in multiple regions, filling the gap in existing landfill remote sensing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942327A_ABST
    Figure CN119942327A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic identification method for a remote sensing image of a solid waste landfill based on deep learning, and relates to computer vision and collection of remote sensing image data of the solid waste landfill. Fusing the remote sensing image data of the solid waste landfill into a data set RSD46-WHU, and constructing a new data set RSSCD47-GU; constructing a GSF-VIT model improved based on a VIT model, training and testing the data set RSSCD47-GU through the GSF-VIT model and a plurality of deep learning models to obtain a confusion matrix, and calculating an identification index of each model according to the confusion matrix; and performing comparative analysis on each model according to the confusion matrix and the identification index, and judging the effectiveness of the GSFF-VIT model. According to the method, the problems of difficult feature extraction and low efficiency in the landfill identification process are solved, meanwhile, the model can be quickly suitable for landfill identification of other regions, and the accuracy of solid waste landfill identification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer vision, and more specifically, to an automatic recognition method of remote sensing images of solid waste landfills based on deep learning. Background Art

[0002] With the increase in the amount of solid waste generated, existing landfills cannot meet the current landfill demand. Many regions are facing the dilemma of landfill saturation, which has directly led to the emergence of many small-scale "illegal landfills". These privately-built illegal landfills are hidden in their locations, and their construction has not been approved by government departments. Most of them lack necessary environmental protection measures and effective risk control measures, resulting in a sharp increase in safety hazards and environmental pollution risks, and increased management difficulties. Therefore, automatic and accurate identification of landfills is crucial for effective supervision and management of solid waste landfills. Rapid and accurate identification of solid waste landfills helps to monitor the status and location of solid waste landfills in a timely manner, detect abnormal events as early as possible, reduce the occurrence of accidents or reduce the harm caused by accidents.

[0003] Due to the scattered nature of solid waste landfills, the sporadic and hidden nature of illegal landfills, there are still challenges in identifying landfills. At present, the identification methods of landfills can be summarized into the following three categories:

[0004] The first is the field survey method, which obtains the accurate location and size of the landfill through manual on-site visits. This method is time-consuming and requires a lot of human resources. The coverage is very limited, and most of them can only identify a small number of legal landfills, but cannot identify those hidden illegal landfills.

[0005] The second method is to combine visual interpretation with mathematical statistical analysis, which aims to explore the spatial distribution of landfills based on the visual interpretation of remote sensing images and the basic characteristics of landfills, such as vegetation cover, spectral characteristics, surface temperature, etc. This method is suitable for identifying large landfills, but since illegal landfills are usually small, it is difficult to effectively identify them based on these remote sensing indicators alone. Therefore, some studies have introduced other variables, such as population density and traffic convenience, to identify illegal landfills. However, these studies can only obtain probability maps of landfill distribution, and the identification mainly relies on the external features around the landfill, rather than the characteristics of the landfill itself, and fails to achieve accurate identification of landfills.

[0006] The second method is to use traditional machine learning methods, where domain experts extract image features and then input them into the machine learning model for model training and testing. This method has achieved certain results in improving recognition efficiency. However, the feature extraction process still relies on manual operation, and the quality of feature extraction directly determines the performance of the model. In addition, this method is also prone to misjudging solid waste landfills in satellite images as large-scale engineering foundation pit excavation areas, which shows that there is still room for improvement in feature differentiation and model accuracy.

[0007] In general, when carrying out landfill identification work, we must not only consider the accuracy of the identification results, but also fully consider the resource consumption during the identification process. With the increase in the number of landfills, especially the increase in the proportion of illegal landfills, the management of landfills has gradually become more difficult. If the efficiency of identification is ignored, the identification process becomes too cumbersome, time-consuming and resource-intensive, which will affect the subsequent management activities. Moreover, although existing research has provided several methods for landfill identification, these methods still require manual intervention to a large extent, resulting in a time-consuming and inefficient identification process, which is still far from the goal of achieving rapid and accurate identification. Therefore, the analysis of existing landfill identification methods and related research shows that the relevant technologies still have the following technical problems that need to be solved:

[0008] 1. Remote sensing datasets for landfills are still relatively scarce. They are usually created independently by some scientific research projects or cooperative institutions, and the data acquisition channels are mostly specific cooperation or research applications, with poor openness;

[0009] 2. Existing identification methods are limited to specific areas and are based on the external features around the landfill rather than focusing on the characteristics of the landfill itself. The results only provide a probability map of the landfill distribution, which limits the accuracy and application of the identification results.

[0010] 3. Existing identification methods rely too much on manual experience and judgment in landfill feature extraction, which affects the accuracy and efficiency of identification. Summary of the invention

[0011] The technical problem to be solved by the present invention is to address the deficiencies in the prior art and provide an automatic recognition method for remote sensing images of solid waste landfills based on deep learning, thereby improving the accuracy of solid waste landfill recognition.

[0012] The present invention provides a method for automatically identifying remote sensing images of solid waste landfills based on deep learning, the method comprising the following steps:

[0013] Step 1: Collect remote sensing image data of solid waste landfill;

[0014] Step 2: Integrate the solid waste landfill remote sensing image data into the dataset RSD46-WHU to construct a new dataset RSSCD47-GU;

[0015] Step 3: construct a GSFF-VIT model improved based on the VIT model, train and test the data set RSSCD47-GU through the GSFF-VIT model and multiple deep learning models to obtain a confusion matrix, and calculate the recognition index of each model according to the confusion matrix;

[0016] Step 4: Compare and analyze the models according to the confusion matrix and recognition index to determine the effectiveness of the GSFF-VIT model.

[0017] Preferably, the steps of constructing the GSFF-VIT model are:

[0018] Step 1: Take the image in the dataset RSSCD47-GU as the input image x, and divide the input image x into P small blocks, each of which has a size of Then it is mapped to the D-dimensional feature space to obtain the block feature image x′; where H and W are the height and width of the input image x respectively;

[0019] Step 2: Add a category token to the beginning of the block feature image x′ to obtain a new block feature image x″;

[0020] Step 3: Add position embedding for each small block and category token in the new block feature image x″;

[0021] Step 4: Upsample the position embedding of the image token of the new block feature image x″ to match the new feature shape;

[0022] Step 5: Concatenate the position embedding of the image token of the sampled new block feature image x″ with the category token of the new block feature image x″ to obtain a tensor;

[0023] Step 6: Construct a bottle neck structure to fuse the information of the tensor;

[0024] Step 7: Add the information-fused tensor to the new block feature image x″ to obtain the block feature image x″′;

[0025] Step 8: Perform an exit operation on the block feature image x″′ to obtain an image feature x″″;

[0026] Step 9: Process the image feature x″″ through L layers of Transformer blocks to obtain the output of L Transformer blocks;

[0027] Step 10: Perform LayerNorm normalization on the output of the Transformer block of the last layer to obtain the normalized image feature x final ;

[0028] Step 11: Extract the normalized image feature x final The features of the class token x class , according to the feature x class Classify the input image x.

[0029] Preferably, the information of the tensor is fused, specifically:

[0030] The dimension of the tensor is transformed, and the features of the transformed tensor are fused and the dimension is amplified using two linear layers, and then the dimension is transformed back to the original dimension.

[0031] Preferably, the recognition indicators include Top-1 Accuracy, Top-5 Accuracy, Mean Precision and Mean Recall; wherein,

[0032] Top-1 Accuracy indicates the proportion of samples whose predicted category with the highest probability is consistent with the true category to the total samples.

[0033] Top-5 Accuracy indicates the proportion of samples in the top five most probable categories predicted by the model that contain the true category to the total number of samples;

[0034] Mean Precision represents the average precision;

[0035] Mean Recall represents the average recall rate.

[0036] Preferably, the Top-1 Accuracy is calculated by the following formula:

[0037]

[0038] Among them, TP i Indicates that the model correctly predicts the sample whose true category is i as category i; TN i Indicates that the model correctly predicts samples whose true category is not i as non-i category; FP i Indicates that the model mistakenly predicts samples whose true category is not i as category i; FN i It means that the model mistakenly predicts samples with true category i as non-i category; the category is the collection scene of remote sensing image data of solid waste landfill.

[0039] Preferably, the Mean Precision is calculated by the following formula:

[0040]

[0041] Here, k is the total number of categories.

[0042] Preferably, the Mean Recall is calculated by the following formula:

[0043]

[0044] Preferably, the multiple deep learning models include convolutional neural network models and Transformer architecture models, specifically: MobileNetV2, ResNet18, Vgg16, ResNet34, ResNet50, Swin Transformer, VIT.

[0045] Preferably, step four is specifically as follows: visualizing the recognition indicators of all models in graphs, analyzing the visualized graphs to determine whether the GSFF-VIT model is superior to other models in processing remote sensing image data of solid waste landfills; by analyzing the confusion matrix, finding categories in each model whose errors exceed a set threshold as target categories, extracting the confidence of the target category from the prediction results of each model for the data set RSSCD47-GU, and determining the effectiveness of the GSFF-VIT model by visually analyzing the confidence.

[0046] Preferably, a model with the best confidence is selected from the convolutional neural network model as a comparison model, and multiple categories other than the target category are selected from all categories as comparison categories. The comparison model and the GSFF-VIT model are visualized and analyzed based on the confidence of the comparison categories to further judge the effectiveness of the GSFF-VIT model.

[0047] Beneficial Effects

[0048] The advantages of the present invention are:

[0049] 1. The RSD46-WHU dataset was improved and a multi-regional solid waste landfill remote sensing image dataset was created, filling the gap in existing landfill remote sensing data.

[0050] 2. It realizes the transformation from manual feature extraction to data-driven end-to-end feature extraction, which solves the problem of difficult and inefficient feature extraction in the landfill identification process. At the same time, the model can be quickly applied to landfill identification in other regions, improving the efficiency of identification and application.

[0051] 3. A deep learning model for automatic identification of large-scale solid waste landfills was established. The Transformer architecture was introduced into the landfill identification task, and the Vision Transformer model was optimized to enable it to focus more effectively on the landfill’s own characteristics, thereby improving the accuracy of solid waste landfill identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a schematic flow chart of the automatic identification method of remote sensing images of solid waste landfills of the present invention;

[0053] FIG. 2 (a)-(b) are solid waste landfill remote sensing image data collection process and image display diagrams of the present invention;

[0054] Figure 3 A schematic diagram of an example of the RSSCD47-GU data set of the present invention;

[0055] Figure 4 Schematic diagram of the architecture of the GSFF-VIT model of the present invention;

[0056] FIG5( b ) is a schematic diagram of the classification confusion matrix of the ResNet18 model of the present invention on the RSSCD47-GU dataset;

[0057] FIG5( c ) is a schematic diagram of the classification confusion matrix of the ResNet34 model of the present invention on the RSSCD47-GU dataset;

[0058] FIG5( d ) is a schematic diagram of the classification confusion matrix of the VIT model of the present invention on the RSSCD47-GU dataset;

[0059] FIG5( e ) is a schematic diagram of the classification confusion matrix of the MobileNetV2 model of the present invention on the RSSCD47-GU dataset;

[0060] FIG5( f ) is a schematic diagram of the classification confusion matrix of the Swin-Transformer model of the present invention on the RSSCD47-GU dataset;

[0061] FIG5( g ) is a schematic diagram of the classification confusion matrix of the ResNet50 model of the present invention on the RSSCD47-GU dataset;

[0062] FIG5( h ) is a schematic diagram of the classification confusion matrix of the Vgg16 model of the present invention on the RSSCD47-GU dataset;

[0063] FIG5( i ) is a schematic diagram of the classification confusion matrix of the GSFF-VIT model of the present invention on the RSSCD47-GU dataset;

[0064] Figure 6This is a visualization diagram showing the results of different models of the present invention for categories (0) to (4) in Table 2;

[0065] Figure 7 This is a visualization diagram showing the results of different models of the present invention for categories (5) to (9) in Table 2;

[0066] Figure 8 This is a visualization result display diagram of the effectiveness analysis of the GSFF-VIT model of the present invention. DETAILED DESCRIPTION

[0067] The present invention will be further described below in conjunction with the embodiments, but it does not constitute any limitation to the present invention. Any limited number of modifications made by anyone within the scope of the claims of the present invention are still within the scope of the claims of the present invention.

[0068] The present invention provides an automatic identification method of remote sensing images of solid waste landfills based on deep learning. The method mainly has six key core steps, and its process is as follows: Figure 1 As shown, the specific steps are as follows:

[0069] Step 1: Determine the scope of research.

[0070] This step aims to collect remote sensing images of landfills from different regions to expand the feature diversity of the data set and provide rich feature information for subsequent model learning. The data collection process follows a scientific and systematic approach. In terms of data distribution, in order to ensure the representativeness of the data and improve the generalization ability of the model, special consideration is given to the particularity, representativeness and data availability of different regions, and efforts are made to cover all regions of the country. In the selection process, provinces with large differences in topography, economic development level, and solid waste management practices are focused on, so that the data set contains landfills with different characteristics.

[0071] The detailed steps of data collection are shown in Figure 2(a). During the data collection process, the target site information was searched through various official public channels, news reports, Baidu Maps and Amaps using keywords such as "landfill", "disposal site" and "abandoned soil site" to ensure the authenticity and validity of the data. This information provides the research with key information such as the project name, location, scale and operation status of each target site, laying the foundation for subsequent site geographic coordinate query and image data collection. After the site information is basically determined, the longitude and latitude coordinates of each target site are queried through the Amap coordinate picker. At the same time, considering the actual situation that some site information is limited or invalid, so that accurate longitude and latitude coordinates cannot be obtained, the accurate longitude and latitude coordinate information of 70 solid waste landfills from 23 provinces in China was finally obtained.

[0072] Step 2: Collect remote sensing image data of solid waste landfill.

[0073] Based on these coordinate information, Google Earth (i.e., Google Maps) was used to sample each landfill. Figure 2(b) is a sample of a remote sensing image of a landfill, in which the main facilities of the landfill are marked. A solid waste landfill consists of several disposal units, structures, and sites, mainly including waste pretreatment facilities, waste landfill areas, garbage dams, leachate collection and treatment facilities, landfill gas drainage facilities, and auxiliary projects. The landfill area is the core area of ​​the landfill, usually presenting an irregular, high and low waste pile with a darker surface color. The top of the pile is covered with soil or synthetic film, and the surface of the covered area is flat without significant protrusions. The soil cover is mostly light brown or yellowish brown, while the synthetic film cover is dark black or dark gray. Garbage dams are generally set up outside the landfill area to stabilize the pile structure. The road system is clear, often circular or criss-crossed, which is convenient for large transport vehicles to transport waste or materials. Large landfills usually have leachate and gas treatment plants at the edges, and office, monitoring and management buildings in the entrance area. These buildings appear as small and regular geometric shapes in remote sensing images.

[0074] To make full use of the data, we conducted multiple time series sampling of the same site with significant changes over time as different samples. Finally, a total of 586 landfill remote sensing images were collected.

[0075] Step 3: Dataset construction.

[0076] This step combines the remote sensing image data collected in step 1 and integrates it into the RSD46-WHU dataset to complete the construction of the solid waste landfill category dataset.

[0077] In the present invention, the model learns data features from the image data by itself, and the quality of the data set directly affects the accuracy of subsequent model training and the final image recognition results, so the construction of the data set is a key task. The data set of the present invention has a wide coverage, covering landfill image data in many regions across the country. RSD46-WHU is a large-scale open data set, mainly used for remote sensing image scene classification. The data set is collected from multiple sources such as Google Earth and Tianditu. The ground resolution of most categories in the data set has reached a high precision of 0.5 meters, and other categories are about 2 meters. Each category contains 500 to 3000 images, totaling 117,000 images. The data set covers 46 different categories, providing rich data resources for remote sensing image analysis.

[0078] According to the construction mode of the RSD46-WHU dataset, the present invention screened 586 landfill images, constructed a new dataset Remote Sensing Scene Classification Dataset GuangzhouUniversity47 (RSSCD47-GU) in an 8:2 manner, and randomly allocated them into training and test sets: training set 467; test set 119. The dataset sample is as follows Figure 3 As shown, there are 47 categories in total. Figure 3In the figure, the serial numbers and their photos represent: (01) Airplane, (02) Airport, (03) Artificial dense forest land, (04) Artificial sparse forest land, (05) Bare land, (06) Basketball Court, (07) Blue Structured factory Building, (08) Building, (09) Construction Site, (10) Cross River bridge, (11) Crossroads, (12) Dense Tall building, (13) Dock, (14) Fish Pond, (15) Footbridge, (16) Graff, (17) Grassland, (18) Low Scattered building, (19) Regular Farmland, (20) Medium Density scattered building, (21) Sparse farmland, (22) Sparse farmland, (23) Sparse farmland, (24) Sparse farmland, (25) Sparse farmland, (26) Sparse farmland, (27) Sparse farmland, (28) Sparse farmland, (29) Sparse farmland, (30) Sparse farmland, (31) Sparse farmland, (32) Sparse farmland, (33) Sparse farmland, (34) Sparse farmland, (35) Sparse farmland, (36) Sparse farmland, (37) Sparse farmland, (38) Sparse farmland, (39) Sparse farmland, (40) Sparse farmland, (41) Sparse farmland, (42) Sparse farmland, (43) Sparse farmland, (44) Sparse farmland, (45) Sparse farmland, (46) Sparse farmland, (47) Sparse farmland, (48) Sparse farmland, (49) Sparse farmland, (50) Sparse farmland, (51) Sparse farmland, (52) Sparse farmland, (53) Sparse farmland, (54) Sparse farmland, (55) Sparse farmland, (56) Sparse farmland, (57) Sparse farmland Building, (21) Medium Density structured Building, (22) Natural Dense forest Land, (23) Natural Sparse forest Land, (24) 0il Tank, (25) 0verpass, (26) Parking Lot, (27) Plastic Greenhouse, (28) Playground, (29) Railway, (30) Red Structured Factory Building, (31) Refinery, (32) Regular Farmland, (33) Scattered Blue roof Factory building, (34) Scattered Red roof Factory building, (35) Sewage Plant type One, (36) Sewage Plant type Two, (37) Ship, (38) Solar Powerstation, (39) Sparse Residential area, (40) Square, (41) Steel Smelter, (42) Storage Land, (43) Tennis Court, (44) Thermal Power plant, (45) Vegetable Plot, (46) Waste Landfill, (47) Water.

[0079] Step 4: Training and development of deep learning models for identifying solid waste landfills.

[0080] This step aims to train and test based on classic deep learning models. The present invention uses 7 classic deep learning models and the GSFF-VIT model built based on VIT as the basis of the landfill automatic identification method, and applies them to the data set RSSCD47-GU for training and testing, thereby improving the recognition accuracy while realizing the automatic identification of solid waste landfills. Among them, the 7 classic deep learning models include two major categories, namely convolutional neural network and Transformer architecture models, specifically MobileNetV2, ResNet18, ResNet34, ResNet50, Vgg16, Swin-Transformer, and VIT.

[0081] The training and testing of these models can be performed on a platform with the following hardware configuration: Intel i7-12700H CPU, 32GB RAM and RTX 2080Ti GPU. All codes are programmed and executed in the PyTorch environment. During training, the batch size of the model is set to 16 during the freezing stage and 8 during the thawing stage. The initial learning rate is 1e-3 (when using the Adam optimizer). The model takes 256×256 size images as input and the total training epochs is 200. The Adam optimizer is used, combined with a learning rate scheduler. During training, cosine annealing is used as the learning rate decay strategy.

[0082] The MobileNetV2, ResNet18, Vgg16-bn, ResNet34, ResNet50, Swin-Transformer, VIT, and GSFF-VIT models were trained and tested on the RSSCD47-GU dataset, and finally the confusion matrix and confidence were obtained. The confusion matrix is ​​shown in Figure 5(b) to (i). In Figure 5(b) to (i), (01) to (47) correspond to the categories (01) to (47) in Figure 2. The recognition indicators of each model (including Top-1Accuracy, Top-5 Accuracy, Mean Recall, and Mean Precision) are calculated based on the results of the confusion matrix.

[0083] In the multi-category remote sensing image classification task, the output results of the model can be divided into four cases similar to the binary classification problem, namely, true positive (TP), true negative (TN), false positive (FP) and false negative (FN). For each category i, TP i Indicates that the model correctly predicts the sample whose true category is i as category i; TN i Indicates that the model correctly predicts samples whose true category is not i as non-i category; FP i Indicates that the model mistakenly predicts samples whose true category is not i as category i; FN i It means that the model mistakenly predicts samples whose true category is i as non-i category.

[0084] The comparison of model performance is carried out around several key indicators. The effects of different models on automatic identification of solid waste landfills are evaluated through quantitative analysis, mainly including the following indicators:

[0085] Top-1 Accuracy: It indicates the proportion of samples whose most likely category (i.e., the category with the highest probability) predicted by the model is consistent with the true category in the total samples. This indicator can well illustrate the accuracy of the model on the overall data set. The calculation formula is shown in Formula 1:

[0086]

[0087] Top-5 Accuracy: indicates the proportion of samples that contain the true category among the top five most likely categories predicted by the model to the total number of samples. This indicator works similarly to Top-1 Accuracy, with the only difference being that even if the highest probability in the prediction result is not the true category of the sample, as long as the true category of the sample is included in the top five most likely categories, the prediction result will be considered correct.

[0088] Mean Precision: Average precision. Precision indicates the ratio of the number of samples with true category i (TP) to the number of samples predicted to be category i among all samples predicted to be category i by the model. Mean Precision is the average of the Precision of all categories. The calculation formula is shown in Formula 2:

[0089]

[0090] Mean Recall: Average recall rate. Recall represents the proportion of samples with actual category i that can be correctly predicted as category i by the model. Mean Recall is the average of the recalls of all categories. The calculation formula is shown in Formula 3:

[0091]

[0092] In the present invention, the GSFF-VIT model design and construction method are as follows.

[0093] For the remote sensing scene classification task (RSSC), due to the variability of the scale of remote sensing targets, using VIT's patch embedding method to split remote sensing images may lose important information. Therefore, the present invention proposes to use the GSFF-VIT model to enhance VIT's fusion of the semantic feature information contained in the image patch and the spatial coordinates. The schematic diagram of the GSFF-VIT model is shown in the figure. Figure 4 As shown, the specific construction steps are as follows:

[0094] For an input image x, the shape is [B,C,H,W], where B is the batch size, C is the number of channels, and H and W are the height and width of the image, respectively.

[0095] Step 1: Patch Embedding.

[0096] The input image x is divided into P patches, and the size of each patch is Then it is mapped to the D-dimensional feature space. The specific formula is as follows: x′=PatchEmbed(x), where the size of x′ is [B, N, D], is the total number of patches and D is the feature dimension.

[0097] Step 2: Add Class Token.

[0098] Add a class token to the beginning of x′, the size of the class token is [1,1,D]. The specific formula is as follows:

[0099] x" = Concat(x',cls_token), where the size of x" is [B,N+1,D].

[0100] Step 3: Positional Embedding.

[0101] Add position embedding for each patch and class token in x″, as shown below:

[0102] cls_token_pe=pos_embed[:,0:1,:]; img_token_pe=pos_embed[:,1:,:].

[0103] Step 4: Interpolate Positional Embedding.

[0104] The position embedding of the image token in x″ is upsampled to match the new feature shape. This is shown in the following formula:

[0105] img_token_pe=Interpolate(img_token_pe, size=new_feature_shape, mode='bicubic').

[0106] Step 5: Concatenate Positional Embedding.

[0107] The position embedding of the sampled image token of x″ is concatenated with the category token of x″ to obtain a tensor. The specific formula is as follows:

[0108] pos_embed=Concat(cls_token_pe,img_token_pe,dim=1). Among them, pos_embed is a tensor that combines category (semantic) information and position information.

[0109] Step 6: Construct a bottle neck structure to fuse the information of the tensor. First, convert the tensor dimension: y = y.view(B,C,1). Then use two linear layers to achieve feature fusion and dimension expansion. Finally, convert the dimension back to the original dimension to achieve fusion.

[0110] Step 7: Residual learning.

[0111] The information fused tensor is added to the new block feature image x″, as shown in the following formula:

[0112] x″′=x″+pos_embed.

[0113] Step 8: Dropout.

[0114] The block feature image x″′ is subjected to a dropout operation, as shown in the following formula: x″″=Dropout(x″′, p=drop_rate).

[0115] Step 9: Transformer Blocks processing.

[0116] Process x″″ through L layers of Transformer Blocks. For a single Transformer block, it is as follows:

[0117] x i =Block(x i-1 ,num h eads,mlp_ratio,...), where x i is the output of the i-th layer Transformer block, x0=x″″.

[0118] Step 10: LayerNorm.

[0119] Perform LayerNorm normalization on the output of the last layer of Transformer Blocks. The specific formula is as follows:

[0120] x final =LayerNorm(x L ), where x final is the normalized feature.

[0121] Step 11: Extract Class Token.

[0122] Extract x final The features of the class token are used for classification. The specific formula is as follows: class =x final [:,0,:].

[0123] Finally, the GSFF-VIT model outputs the classtoken features processed and normalized by Transformer Blocks for subsequent classification tasks. Through the above 11 steps, the VIT model was improved and the GSFF-VIT model was constructed.

[0124] Step 5: Evaluation and comparison of model recognition performance.

[0125] This step aims to evaluate the accuracy of the model test results in the previous step to clarify the recognition ability of each model. Specifically, the recognition indicators are visualized as shown in Table 1.

[0126] Table 1 Results of performance evaluation indicators of each model

[0127] Top-1 Accuracy Top-5 Accuracy Mean Recall Mean Precision (b) ResNet1 8 84.38 98.62 84.41 83.99 (c) ResNet34 86.78 98.93 86.88 86.25 (d) VIT 87.39 99.18 86.9 87.57 (e) MobileNetV2 87.66 99.3 87.72 87.69 (f)Swin Transformer 88.53 99.34 88.28 88.28 (g) ResNet50 88.84 99.29 88.92 88.41 (h)Vgg16 89.36 99.38 89.29 89.05 (i) GSFF-VIT 93.51 99.78 93.53 93.52

[0128] ResNet18, ResNet34, ResNet50, MobileNetV2, Vgg16, Swin Transformer, VIT, and GSFF-VIT models were trained and tested on the RSSCD47-GU dataset, and their confusion matrices are shown in Figure 5. The recognition indicators of each model (including Top-1 Accuracy, Top-5 Accuracy, MeanRecall, and MeanPrecision) were calculated based on the confusion matrix results, and the results are shown in Table 1.

[0129] The analysis of the indicator results in Table 1 shows that the GSFF-VIT model has a significant advantage in image recognition tasks. Its Top-1Accuracy (highest accuracy) and Top-5Accuracy (top five accuracy) are 93.51% and 99.78% respectively, both of which are the highest values ​​among all models in the table, showing extremely high recognition accuracy and confidence. In addition, the average recall rate and average precision of GSFF-VIT also reached 93.53% and 93.52% respectively, which are also ahead of other models, indicating that it has excellent performance in the ability to identify positive samples and the accuracy of predicting positive samples. The comprehensive performance of these indicators shows that GSFF-VIT has strong competitiveness and application potential in the field of image recognition.

[0130] Step 6: Visual analysis of the performance of the solid waste landfill automatic identification model.

[0131] The first step is to visualize the confusing categories of each model.

[0132] Summarizing the confusion matrix results, we obtain the accuracy results shown in Table 2. As can be seen from Figure 5 and Table 2, there are 10 categories with an error rate exceeding 30%. This paper extracts the top 10 categories of targets with the highest errors for visual analysis to evaluate the performance of different models.

[0133] Table 2 Prediction error rate of each category

[0134]

[0135] It should be noted that the numbers before each category in Table 2 above are table numbers, which are not equivalent to Figure 3 The serial number corresponding to each category in .

[0136] Visual analysis results such as Figure 6 and Figure 7 For better analysis, we correspond the top ten categories with the highest errors (Medium density scattered building, Construction site, Artificial dense forestland, Building, Artificial sparse forest land, Thermal power plant, Sparseresidential area, Refinery, Bare land, Medium density structured building) to (0) to (9) in Table 2 above, and encode the model. Since the prediction results include prediction categories and confidence levels, this part of the information is summarized in Table 3. See Table 3 for details.

[0137] Table 3 Prediction results of each model for easily confused categories

[0138]

[0139] according to Figure 6 From the analysis, combined with Table 3, we can see that for the (0) Medium density scattered building category, except for GSFF-VIT, the classification results of each model for Medium density scattered building are wrong. The proposed GSFF-VIT not only captures the image information of the building part, but also achieves correct classification. Figure 6 (1)

[0140] In the Construction site category, all models are classified correctly and achieve high classification accuracy (arrows); Figure 6In the (2) Artificial dense forest land category, GSFF-VIT not only correctly classifies, but also has the highest confidence. Figure 6 In the (3) Building category, the distribution of key targets is captured, but classification errors are missing; Figure 6 In the (4) Artificial sparse forest land category, GSFF-VIT also performs superiorly.

[0141] according to Figure 7 From the analysis and combined with Table 3, we can see that for the (5) Thermal power plant category, all models classified correctly and the GSFF-VIT confidence is also one of the highest. For the (6) Sparse residential area category, the proposed GSFF-VIT performs best and has the highest confidence. For the (7) Refinery category, the performance of GSFF-VIT is only better than ResNet18. For the (8) and (9) categories, GSFF-VIT also has the highest classification performance for Bare land and Medium density structured building.

[0142] comprehensive Figure 6 , Figure 7 From Table 3, we can see that the proposed GSFF-VIT has the best overall performance and performs best in almost all high error categories.

[0143] Step 2: Visual analysis of GSFF-VIT on landfill categories.

[0144] The first step is to verify the effectiveness of the proposed model through all data sets. In order to further demonstrate the superiority of the proposed GSFF-VIT, this step organizes comparative ablation experiments to discuss the effectiveness of the proposed components. Since the classification object of the present invention is mainly Wastelandfill, the relevant data of the landfill are randomly selected, corresponding to (10) to (14) in Table 2, and VIT and GSFF-VIT are selected for classification and recognition, and the difference in the effects of the two models is discussed.

[0145] To further analyze the effectiveness of the proposed GSFF-VIT on landfill types, the present invention further demonstrates the performance of GSFF-VIT and VIT on different landfill images, as shown in the following figure. Figure 8 As shown, we summarize the corresponding classification accuracy in Table 4. From Table 4, we can see that GSFF-VIT is slightly better than the original VIT in terms of confidence. Therefore, the present invention uses visual analysis to explain the effectiveness of the proposed GSFF-VIT, as shown in Figure 8In column (10), VIT focuses on the surrounding farmland (arrow A), while GSFF-VIT focuses more on the landfill area in the middle and has a higher confidence level; Figure 8 In column (11), VIT also focuses on the surrounding wasteland, but GSFF-VIT captures the boundary information of key buildings and landfills (arrows BC); Figure 8 In column (12), both VIT and GSFF-VIT focus on the key area of ​​the landfill (EF area), but VIT allocates the attention weight to the land boundary on the other side of the river (arrow D); Figure 8 In column (13), VIT focuses on the surrounding farmland (arrow H), but GSFF-VIT focuses on the key target - the landfill area (arrow G); Figure 8 In column (14), VIT still focuses on the surrounding wasteland, while GSFF-VIT pays attention to the regional information of the landfill (arrow IJ).

[0146] Table 4 Classification results and confidence levels of landfill areas by VIT and GSFF-VIT

[0147]

[0148] Therefore, in the landfill identification task, compared with the traditional VIT model, the GSFF-VIT model performs better in the ability to focus on regions and facilities and the accuracy of feature capture, thereby improving the classification accuracy and practical application value.

[0149] The above is only a preferred embodiment of the present invention. It should be pointed out that for those skilled in the art, several modifications and improvements can be made without departing from the structure of the present invention, which will not affect the effect of the implementation of the present invention and the practicality of the patent.

Claims

1. A method for automatic identification of remote sensing images of solid waste landfills based on deep learning, characterized in that: The method comprises the following steps: Step 1: Collect remote sensing image data of solid waste landfill; Step 2: Integrate the solid waste landfill remote sensing image data into the dataset RSD46-WHU to construct a new dataset RSSCD47-GU; Step 3: construct a GSFF-VIT model improved based on the VIT model, train and test the data set RSSCD47-GU through the GSFF-VIT model and multiple deep learning models to obtain a confusion matrix, and calculate the recognition index of each model according to the confusion matrix; Step 4: Compare and analyze the models according to the confusion matrix and recognition index to determine the effectiveness of the GSFF-VIT model.

2. The method for automatic identification of remote sensing images of solid waste landfills based on deep learning according to claim 1 is characterized in that: The construction steps of the GSFF-VIT model are: Step 1: Take the image in the dataset RSSCD47-GU as the input image x, and divide the input image x into P small blocks, each of which has a size of Then it is mapped to the D-dimensional feature space to obtain the block feature image x′; where H and W are the height and width of the input image x respectively; Step 2: Add a category token to the beginning of the block feature image x′ to obtain a new block feature image x″; Step 3: Add position embedding for each small block and category token in the new block feature image x″; Step 4: Upsample the position embedding of the image token of the new block feature image x″ to match the new feature shape; Step 5: Concatenate the position embedding of the image token of the sampled new block feature image x″ with the category token of the new block feature image x″ to obtain a tensor; Step 6: Construct a bottle neck structure to fuse the information of the tensor; Step 7: Add the information-fused tensor to the new block feature image x″ to obtain the block feature image x″′; Step 8: Perform an exit operation on the block feature image x″′ to obtain an image feature x″″; Step 9: Process the image feature x″″ through L layers of Transformer blocks to obtain the output of L Transformer blocks; Step 10: Perform LayerNorm normalization on the output of the Transformer block of the last layer to obtain the normalized image feature x final ; Step 11: Extract the normalized image feature x final The features of the class token x class , according to the feature x class Classify the input image x.

3. The method for automatic identification of remote sensing images of solid waste landfills based on deep learning according to claim 2 is characterized in that: The information of the tensor is fused, specifically: The dimension of the tensor is transformed, and the features of the transformed tensor are fused and the dimension is amplified using two linear layers, and then the dimension is transformed back to the original dimension.

4. The method for automatic identification of remote sensing images of solid waste landfills based on deep learning according to claim 1 is characterized in that: The recognition indicators include Top-1 Accuracy, Top-5 Accuracy, Mean Precision and Mean Recall; Top-1 Accuracy indicates the proportion of samples whose predicted category with the highest probability is consistent with the true category to the total samples. Top-5 Accuracy indicates the proportion of samples in the top five most likely categories predicted by the model that contain the true category to the total number of samples; Mean Precision represents the average precision; Mean Recall represents the average recall rate.

5. The method for automatic identification of remote sensing images of solid waste landfills based on deep learning according to claim 4 is characterized in that: The Top-1 Accuracy is calculated by the following formula: Among them, TP i Indicates that the model correctly predicts the sample whose true category is i as category i; TN i Indicates that the model correctly predicts samples whose true category is not i as non-i category; FP i Indicates that the model mistakenly predicts samples whose true category is not i as category i; FN i It means that the model mistakenly predicts samples with true category i as non-i category; the category is the collection scene of remote sensing image data of solid waste landfill.

6. The method for automatic identification of remote sensing images of solid waste landfills based on deep learning according to claim 5 is characterized in that: The Mean Precision is calculated by the following formula: Here, k is the total number of categories.

7. The method for automatic identification of remote sensing images of solid waste landfills based on deep learning according to claim 6 is characterized in that: The Mean Recall is calculated by the following formula:

8. The method for automatic identification of remote sensing images of solid waste landfills based on deep learning according to claim 1, characterized in that: The multiple deep learning models include convolutional neural network models and Transformer architecture models, specifically: MobileNetV2, ResNet18, Vgg16, ResNet34, ResNet50, Swin Transformer, VIT.

9. The method for automatic identification of remote sensing images of solid waste landfills based on deep learning according to claim 1, characterized in that: Step four is as follows: visualize the recognition indicators of all models in graphs, analyze the visualized graphs to determine whether the GSFF-VIT model is superior to other models in processing remote sensing image data of solid waste landfills; analyze the confusion matrix, find the categories whose errors exceed the set threshold in each model as the target category, extract the confidence of the target category from the prediction results of each model for the data set RSSCD47-GU, and determine the effectiveness of the GSFF-VIT model through visual analysis of the confidence.

10. The method for automatic identification of remote sensing images of solid waste landfills based on deep learning according to claim 9, characterized in that: A model with the best confidence is selected from the convolutional neural network model as a comparison model, and multiple categories other than the target category are selected from all categories as comparison categories. The confidence of the comparison model and the GSFF-VIT model based on the comparison categories is visualized and analyzed to further determine the effectiveness of the GSFF-VIT model.

Citation Information

Patent Citations

  • Pedestrian small target detection method in video monitoring based on deep learning

    CN115240119A

  • Urban green land fine classification method and system based on GF-2 and open map data

    CN115984603A

  • Remote sensing image urban green land information extraction method and device and medium

    CN117611991A

  • Urban solid waste extraction method, system, equipment and medium

    CN117636044A

  • Prostate MRI image classification method based on deep learning image model

    CN117636076A

Cited By

  • Automatic segmentation method for functional area of solid waste landfill

    CN121788812A