Remote sensing mining area automatic labeling method based on deep learning and shallow feature statistics
Through methods based on deep learning and shallow feature statistics, the problems of inefficient and insufficient accuracy of traditional mining area identification methods are solved, efficient and accurate monitoring and management of mining areas are achieved, and sustainable utilization of resources is promoted.
Patent Information
- Application Number
- CN202510410080.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-01-22
- Filing Date
- 2025-04-02
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional mining area identification methods rely on manual annotation or simple machine learning algorithms, resulting in inefficient and insufficient accuracy when processing complex and large-scale data.
Automatic labeling method of remote sensing mining areas based on deep learning and shallow feature statistics is adopted. By obtaining remote sensing image sample sets, deep learning image segmentation model is trained, and combined with preset loss functions and evaluation indicators, automatic labeling of mining areas is achieved.
It has achieved efficient and accurate monitoring of mining areas, helping management departments to timely grasp mining areas dynamics, scientifically formulate resource development and environmental protection policies, optimize enterprise production management, improve resource utilization efficiency, and reduce environmental damage.
Smart Images

Figure CN120220152A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mining area identification, and particularly to a remote sensing mining area automatic annotation method based on deep learning and shallow feature statistics. Background Art
[0002] Remote sensing technology has important applications in geological exploration and resource management, and the automatic identification and annotation of mining areas is a key task of remote sensing technology. The accurate annotation of mining areas is of great significance for resource management and environmental protection. By accurately identifying and monitoring mining areas, it can help the government and enterprises better plan resource development and avoid resource waste and environmental damage caused by disorderly mining. At the same time, by real-time monitoring the changes in mining areas, illegal mining activities and potential environmental risks can be detected early, so as to take timely countermeasures. Traditional mining area identification methods rely on manual annotation or simple machine learning algorithms, and these methods often face problems of low efficiency and insufficient accuracy when dealing with complex and large-scale data. Summary of the Invention
[0003] Based on this, it is necessary to provide a remote sensing mining area automatic annotation method based on deep learning and shallow feature statistics for the above technical problems.
[0004] The remote sensing mining area automatic annotation method based on deep learning and shallow feature statistics includes the following steps:
[0005] Obtain a given set of remote sensing image samples; wherein, the set of remote sensing image samples includes: remote sensing sample images and reference mining area module segmentation images;
[0006] Taking each remote sensing sample image as the input of the network source and the reference mining area module segmentation image corresponding to the remote sensing sample image as the output, and combining a preset loss function and evaluation index, apply each remote sensing sample image to train a target deep learning image segmentation model to obtain a reference mining area module segmentation model;
[0007] Obtain a remote sensing image to be annotated, input the remote sensing image to be annotated into the reference mining area module segmentation model to obtain a mining area module segmentation image; wherein, the mining area module segmentation image includes a number of mining area patches;
[0008] According to the mining area module segmentation image, statistically analyze the object features of the mining area patches to obtain the corresponding land cover classes; wherein, the object features include: index features and geometric features; the land cover classes include: mineral classes and non-mineral classes;
[0009] Based on the mining area module segmentation image, retain the mining area patches with the land cover class of mineral classes to obtain the remote sensing mining area annotation result.
[0010] In one embodiment, before obtaining a given set of remote sensing image samples, it further includes:
[0011] Obtain remote sensing images;
[0012] Receive annotation information to obtain a reference mine area module segmentation image corresponding to the remote sensing image;
[0013] Cut the reference mine area module segmentation image according to a preset cutting step length to obtain a valid annotation image;
[0014] Perform screening processing on the valid annotation image and unify the image format to obtain a screened annotation image;
[0015] Perform random flipping, random cropping, and random color blurring on the screened annotation image to obtain a remote sensing sample image.
[0016] In one embodiment, cutting the reference mine area module segmentation image according to a preset cutting step length to obtain a valid annotation image includes:
[0017] Cut the reference mine area module segmentation image with a preset cutting step length to obtain a cut shp file;
[0018] Based on the cut shp file, perform coordinate conversion through the following formula to obtain a valid annotation image:
[0019]
[0020] where x is the abscissa of the valid annotation image, x geo represents the abscissa of the cut shp file, x′ represents the abscissa of the top - left vertex of the image, p represents the spatial resolution of the pixel, y is the ordinate of the valid annotation image, y′ is the ordinate of the top - left vertex of the image, y geo represents the ordinate of the cut shp file.
[0021] In one embodiment, the target deep learning image segmentation model includes:
[0022] A feature extraction backbone network, a neighbor feature aggregation module, a feature enhancement module, and a decoding module.
[0023] In one embodiment, obtaining a remote sensing image to be annotated and inputting the remote sensing image to be annotated into the reference mine area module segmentation model to obtain a mine area module segmentation image includes:
[0024] Obtain the remotely sensed image to be labeled, and input the remotely sensed image to be labeled into the benchmark mining area module segmentation model. The benchmark mining area module segmentation model receives the input remotely sensed image to be labeled, extracts preliminary image features through the feature extraction backbone network, and obtains features of five stages with different depths and different sizes;
[0025] Input the features into the neighbor feature aggregation module, and aggregate the features of adjacent stages through the neighbor feature aggregation module to obtain the first aggregated feature;
[0026] Input the first aggregated feature into the feature enhancement module. The feature enhancement module obtains the superimposed features of different receptive fields through the superimposition of dilated convolutions, and aggregates the superimposed features together through residual connections to obtain the second aggregated feature;
[0027] Input the second aggregated feature into the decoding module. The decoding module restores the second aggregated feature to the size of the input remotely sensed image to be labeled to obtain the mining area module segmentation image.
[0028] In one embodiment, according to the mining area module segmentation image, object features are statistically analyzed, and the obtained land cover classes include:
[0029] Obtain the mining area patches in the mining area module segmentation image, and calculate the index features and geometric features of each mining area patch; wherein, the index features include: normalized difference vegetation index, normalized difference water index, BG, BR, GR; the geometric features include: area and aspect ratio;
[0030] According to the index features and the geometric features, and according to the preset feature statistical results, obtain the land cover classes.
[0031] In one embodiment, obtain the mining area patches in the mining area module segmentation image, and calculate the index features and geometric features of each mining area patch; wherein, the index features include: normalized difference vegetation index, normalized difference water index, BG, BR, GR; the geometric features include: area and aspect ratio include:
[0032] Calculate the normalized difference vegetation index through the following formula:
[0033]
[0034] Wherein, NDVI represents the normalized difference vegetation index, NIR represents the mean value of the near-infrared band of the mining area patch, and R represents the mean value of the red band;
[0035] Calculate the normalized difference water index through the following formula:
[0036]
[0037] Among them, NDWI represents the Normalized Difference Water Index, G represents the average value of the green band, and NIR represents the average value of the near-infrared band of the mining area patch;
[0038] Calculate the index relationship between the blue band and the red band in the mining area patch according to the following formula:
[0039]
[0040] Among them, BR represents the index relationship between the blue band and the red band in the mining area patch, B represents the average value of the blue band, and R represents the average value of the red band;
[0041] Calculate the index relationship between the green band and the red band in the mining area patch according to the following formula:
[0042]
[0043] Among them, GR represents the index relationship between the green band and the red band in the mining area patch, G represents the average value of the green band, and R represents the average value of the red band;
[0044] Calculate the index relationship between the blue band and the green band in the mining area patch according to the following formula:
[0045]
[0046] Among them, BG represents the index relationship between the blue band and the green band in the mining area patch, B represents the average value of the blue band, and G represents the average value of the green band;
[0047] Calculate the area of the mining area patch through the following formula:
[0048]
[0049] Among them, A represents the area of the mining area patch, a i represents the actual area of the i-th pixel, and n represents the number of pixels contained in the mining area patch;
[0050] Calculate the aspect ratio of the mining area patch through the following formula:
[0051]
[0052] Among them, LW represents the aspect ratio of the mining area patch, l represents the length, w represents the width, A represents the area of the mining area patch, l0 represents the length of the circumscribed rectangle of the border, and w0 represents the width of the circumscribed rectangle of the border.
[0053] The remote sensing mining area automatic annotation system based on deep learning and shallow feature statistics is used to implement the remote sensing mining area automatic annotation method based on deep learning and shallow feature statistics as described above, including:
[0054] A device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the automatic annotation method for remote sensing mining areas based on deep learning and shallow feature statistics described in the above respective embodiments are implemented.
[0055] A storage medium stores a computer program which, when executed by a processor, implements the steps of the automatic annotation method for remote sensing mining areas based on deep learning and shallow feature statistics described in the above respective embodiments.
[0056] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: By developing an automatic annotation technology for remote sensing mining areas based on deep learning and shallow feature statistics, the present invention can achieve efficient and accurate monitoring of mining areas. It can not only help the management department timely master the dynamics of mining areas and scientifically formulate resource development and environmental protection policies, but also help enterprises optimize production management, improve resource utilization efficiency, and reduce environmental damage. In addition, real-time mining area monitoring technology can also promote the attention and participation of the public in mining area development and environmental protection, and jointly promote the sustainable utilization of resources. Accurate mining area annotation and monitoring can provide important data support for environmental impact assessment, help evaluate the short-term and long-term impacts of mining area development on the environment, and provide a scientific basis for the formulation and implementation of environmental protection measures. At the same time, real-time mining area monitoring technology can timely detect and warn of environmental risks, such as pollution of water bodies and soil around mining areas, prompting relevant departments to take timely countermeasures to reduce environmental risks.
[0057] The present invention can introduce deep learning technology, innovatively combine shallow feature statistical methods, improve the accuracy and efficiency of automatic annotation of remote sensing mining areas, promote the development of remote sensing technology, provide strong technical support for resource management and environmental protection, and promote the sustainable utilization of resources. It not only has important academic value, but also has significant social and economic significance. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a schematic flowchart of an automatic annotation method for remote sensing mining areas based on deep learning and shallow feature statistics in an embodiment;
[0059] Figure 2 It is a comparison diagram of the actual open-pit mining area range and the open-pit mining area range referred to in this project in an embodiment;
[0060] Figure 3 It is a schematic diagram of the sample flipping effect in an embodiment;
[0061] Figure 4 It is a schematic diagram of the random cropping effect in an embodiment;
[0062] Figure 5Schematic diagram of random color blur effect in an embodiment;
[0063] Figure 6 Schematic diagram of confusion matrix in an embodiment;
[0064] Figure 7 Schematic diagram of the structure of the benchmark mining area module segmentation model in an embodiment;
[0065] Figure 8 Schematic diagram of the structure of the neighbor feature aggregation module in an embodiment;
[0066] Figure 9 Schematic diagram of the structure of the feature enhancement module in an embodiment;
[0067] Figure 10 Schematic diagram of the structure of the decoding module in an embodiment;
[0068] Figure 11 Comparison diagram of the visualization results of a benchmark mining area module segmentation model in an embodiment;
[0069] Figure 12 Comparison diagram of the visualization results of another benchmark mining area module segmentation model in an embodiment;
[0070] Figure 13 Comparison diagram of the visualization results of the comparative experiment in an embodiment;
[0071] Figure 14 Schematic diagram of the segmentation visualization results of the benchmark mining area module segmentation model in an embodiment;
[0072] Figure 15 Schematic diagram of the visualization results of the benchmark mining area module segmentation model after segmentation combined with shallow feature screening in an embodiment;
[0073] Figure 16 Schematic diagram of the structure of the remote sensing mining area automatic annotation system based on deep learning and shallow feature statistics in an embodiment;
[0074] Figure 17 Schematic diagram of the internal structure of the device in an embodiment. Detailed implementation manners
[0075] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below through specific implementation manners in combination with the accompanying drawings.
[0076] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of this specification should have the ordinary meanings understood by those of ordinary skill in the art to which the present invention pertains. The "first", "second" and similar terms used in one or more embodiments of this specification do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0077] In one embodiment, as Figure 1 shown, a remote sensing mining area automatic annotation method based on deep learning and shallow feature statistics is provided, including the following steps:
[0078] Step S101, obtaining a given set of remote sensing image samples; wherein, the set of remote sensing image samples includes: remote sensing sample images and reference mining area module segmentation images.
[0079] Specifically, obtaining the constructed set of remote sensing image samples includes remote sensing sample images and corresponding reference mining area module segmentation images. For example, a set of multi-source image open-pit mine basic sample libraries is given, including 3979 pairs of samples. Each pair of samples includes a remote sensing sample image with a size of 256x256 and a binary mask image (reference mining area module segmentation image) of the mining area corresponding to the remote sensing sample image, providing a rich background data reference for the automatic recognition of mine remote sensing.
[0080] On this basis, before obtaining the given set of remote sensing image samples, it further includes:
[0081] Obtaining remote sensing images;
[0082] Receiving annotation information to obtain a reference mining area module segmentation image corresponding to the remote sensing image;
[0083] Cutting the reference mining area module segmentation image according to a preset cutting step length to obtain an effective annotation image;
[0084] Performing screening processing on the effective annotation image and unifying the image format to obtain a screened annotation image;
[0085] Performing random flipping, random cropping and random color blurring processing on the screened annotation image to obtain remote sensing sample images.
[0086] Specifically, since the coverage rate of open-pit mines in single-scene remote sensing images is low, that is, open-pit mines are sparse in remote sensing images, directly using the original images for training will result in a large machine load rate and a low positive sample ratio (the proportion of effective data is very low). Training based on such imbalanced sample ratios will affect the accuracy of feature extraction and class discrimination. Therefore, a series of data preprocessing such as annotating remote sensing images, selecting effective regions, cutting mining areas, and data augmentation is required to construct a remote sensing image sample set.
[0087] First, obtain remote sensing images. With the development of satellite software and hardware technologies, domestic satellite images have gradually replaced foreign satellite images and become the mainstream data source in the field of natural resources supervision. The remote sensing images used in this embodiment are provided by the domestic Gaofen-1, Gaofen-2, Gaofen-6, Gaofen-7, Ziyuan-1, Ziyuan-3, and CBERS-4 satellite images, and they are all result images after calibration, registration, and fusion. They include 4 bands: red, green, blue, and near-infrared, with resolutions of 2m, 0.8m, 2m, 0.7m, 2m, 2m respectively. The data depth is 8 bits for all, and the data format is TIFF. The data information is shown in Table 1.
[0088] Table 1 List of data source information
[0089]
[0090] In the experiment, remote sensing images successfully captured within the scope of Chongqing from 2021 to 2023 were selected as the data source. The most common types of open-pit mining areas in Chongqing, including limestone (such as limestone for facing, building stone limestone, cement limestone, etc.), sandstone (such as strip stone, building sandstone, sand for cement batching, etc.), shale (such as shale for bricks and tiles, shale for ceramsite, etc.), and other types (such as fluorite, calcite, etc.), were selected as the objects for mine identification.
[0091] It should be particularly noted that the open-pit mining areas or mining activity patches in actual work usually consist of a collection of ground features such as stope, industrial square, solid waste, mine buildings, and mine roads. Some mines only contain one or several of these ground features. Considering the influence of the complexity of ground feature combinations on the identification difficulty and the fact that the main goal of supervision work is to focus on information related to the "excavation activity area", the open-pit mining areas or mining activity patches in this embodiment only refer to the main carriers that can effectively reflect the mining situation - the stope, as well as the industrial square and solid waste connected to the stope, and do not include mine buildings or separate industrial squares, solid waste, etc. Figure 2 The object area range of the open-pit mining areas or mining activity patches in this project is shown.
[0092] Manual annotation of the mining area is carried out through ArcGIS software to generate specific annotation information, accept the annotation information, and obtain the segmented image of the benchmark mining area module corresponding to the remote sensing image. In one embodiment, the reference points of the open-pit mining areas issued by Chongqing are compared with the original remote sensing image for manual screening. The areas clearly identified as mines are interpreted as open-pit mines using ArcGIS software to generate vector annotations, classified according to the types of ore such as limestone, sandstone, and shale, and filled into the polygon attribute fields. The corner points of the vector polygons are all in CGCS2000 national geodetic coordinate information.
[0093] After that, considering the sparsity of the mining area polygons, the mining area is selected for image cutting. Taking the preset cutting step of 50 pixels as the step, 256×256 tif-format image data, annotation polygon shp files are cut and generated, and sliced images (tif) and class binary mask images (tif) are converted to obtain effective annotation images.
[0094] Then, considering that the sample size generated by cutting is too large and the sample similarity is too high, the effective annotation images of the slices are sampled and screened. Since the step size was set to 50 during the initial image cutting, for the mining areas with small areas, their proportion in the dataset can be enriched. However, for the mining areas with large areas, with a step size of 50, the change rate of their image features is not high, and increasing their proportion too much. Sampling and screening of data are carried out according to the mining area category number and the size of the mining area to achieve relative balance in the proportion of various types of mines.
[0095] Since the resolutions of GF-1, GF-6, ZY-1, ZY-3, and CBERS-4 remote sensing images are 2m, while the resolution of GF-2 image is 0.8m and the resolution of GF-7 image is 0.7m, the resolutions are not unified. To facilitate subsequent work such as shallow feature extraction, statistics, and deep learning experiments, it is necessary to unify the data formats of the images and binary mask images. Therefore, resampling of the GF-2 and GF-7 data is performed to make the pixel size uniformly 2m, and the screened annotation images are obtained.
[0096] Finally, due to the difficulty in obtaining remote sensing samples, in order to enrich the remote sensing image sample set, random flipping, random cropping, and random color blurring are performed on the screened annotation images.
[0097] 1) Random flipping
[0098] There are the following four cases through random horizontal and vertical flipping: no flipping, horizontal flipping, vertical flipping, and simultaneous horizontal and vertical flipping. That is, through the above operations, the samples in the dataset can be enriched four times. The specific flipping effects are as Figure 3 shown.
[0099] 2) Random cropping
[0100] Set a zoom range value [0.8 - 1.2], that is, the minimum zoom ratio is 0.8 (shrinking by 20%), the maximum zoom ratio is 1.2 (that is, enlarging by 20%), and 1 represents no zoom. Control the zoom ratio of the image through random numbers.
[0101] When the image is enlarged, the edges of the original image are sheared by the cropping box generated by random numbers. After shearing, in order to restore to the size of the original image, bilinear interpolation is used to restore the cropped image. When shrinking, the original image is shrunk proportionally. After shrinking, in order to match the input size, the edges of the shrunk image are padded to restore to the size of the original image. The specific cropping effect is as Figure 4 shown.
[0102] 3) Random color blur
[0103] The original image is blurred by using the method of Gaussian blur. A random floating-point number between 0 and 1 generated randomly is used as the radius of the Gaussian blur filter. The blur radius determines the intensity of the blur. The larger the value, the more obvious the blur effect. The specific color blur effect is as Figure 5 shown.
[0104] On this basis, the segmented image of the benchmark mining area module is cut according to a preset cutting step size to obtain an effective labeled image, including:
[0105] The segmented image of the benchmark mining area module is cut according to a preset cutting step size to obtain a cutting shp file;
[0106] Based on the cutting shp file, coordinate conversion is performed through the following formula to obtain an effective labeled image:
[0107]
[0108] where x is the abscissa of the effective labeled image, x geo represents the abscissa of the cutting shp file, x' represents the abscissa of the top left vertex of the image, p represents the spatial resolution of the pixel, y is the ordinate of the effective labeled image, y' is the ordinate of the top left vertex of the image, y geo represents the ordinate of the cutting shp file.
[0109] Specifically, since the coordinate system of the shp file of the mine sample is a spatial geographic coordinate system and the mask node coordinates in the generated binary image data are graphic coordinates, coordinate conversion needs to be performed through the formula:
[0110]
[0111] where x is the abscissa of the effective labeled image, xgeo represents the abscissa for cutting the shp file, x′ and y′ represent the coordinates of the upper left vertex of the image, which can be obtained from the xml file, p represents the spatial resolution of the pixel, y is the ordinate of the effective labeled image, y geo represents the ordinate for cutting the shp file.
[0112] Step S102: Using each remote sensing sample image as the network source input and the corresponding benchmark mining area module segmentation image of the remote sensing sample image as the output, combined with a preset loss function and evaluation metrics, apply each remote sensing sample image to train the target deep learning image segmentation model to obtain the benchmark mining area module segmentation model.
[0113] Specifically, the target deep learning image segmentation model OMSegNet is a remote sensing open-pit mining area semantic segmentation model based on CNN and Transformer. Train the target deep learning image segmentation model to obtain the benchmark mining area module segmentation model.
[0114] The model performance of the benchmark mining area module segmentation model is evaluated through evaluation metrics. Image semantic segmentation is to classify the image at the pixel level, and the evaluation of the recognition effect can be accurate to a single pixel. Therefore, the evaluation metrics selected for this model are the general evaluation criteria for most image segmentation algorithms, that is, taking each pixel point as the basic unit, calculating the accuracy (OP), precision, F1 score (F1), and recall rate (Recall) of the pixel point classification result.
[0115] As Figure 6 shown, each remote sensing image corresponds to a mask, and for different categories, it is a binary mask. If Positive represents that the predicted pixel point is a mine and Negative represents not a mine. Among them, TP represents the number of true positives, indicating the correct prediction of mine pixels. FP represents the number of false positives, indicating the wrong prediction of non-mine pixels. TN represents the number of true negatives, indicating the correct prediction of non-mine pixels. FN represents the number of false negatives, indicating the wrong prediction of mine pixels as non-mine pixels. Any pixel point must belong to one of the above four situations.
[0116] The accuracy rate represents the proportion of examples classified as mines that are actually mines, and is used to evaluate the precision of pixel points identified as mines in the generated results. The recall rate is a measure of coverage, measuring how many positive examples are classified as positive, and is used to evaluate how many foreground pixel points are found. Since the accuracy rate and the recall rate affect each other, ideally, both should be high, but generally, when the accuracy rate is high, the recall rate is low, and when the recall rate is low, the accuracy rate is high. To better compare the network performance, the comprehensive performance evaluation index F1 is introduced.
[0117] On this basis, the target deep learning image segmentation model includes:
[0118] A feature extraction backbone network, a neighbor feature aggregation module, a feature enhancement module, and a decoding module.
[0119] Specifically, the network structure of the target deep learning image segmentation model is as Figure 7 shown. It mainly includes four modules: a feature extraction backbone network, a neighbor feature aggregation module, a feature enhancement module, and a decoding module.
[0120] The feature extraction backbone network extracts preliminary image features and, through continuous deepening, obtains features at five stages with different depths and different sizes.
[0121] The role of the neighbor feature aggregation module is to fuse the preliminary features extracted from the backbone network. In one embodiment, in order to achieve a lightweight effect, the MobileNet is used for the feature extraction backbone network, and its number of parameters is much less than that of mainstream backbone networks such as ResNet and VGG. This module first unifies the sizes of three stages through pooling and bilinear interpolation methods, and then preliminarily fuses the features of the three stages through a concatenation operation. The first aggregated feature after fusion then obtains its long-range dependence through an attention mechanism to better aggregate the features of the three stages. The model structure is as Figure 8 shown.
[0122] The overall architecture of the feature enhancement module is as Figure 9 shown. This module mainly obtains features with different receptive fields through dilated convolution, and then fuses the features with different receptive fields together through residual connections to achieve the effect of feature enhancement.
[0123] In one embodiment, after receiving the first aggregated feature output by the neighbor feature enhancement module, the feature enhancement module first extracts features with a larger receptive field through a 3x3 dilated convolution with a dilation rate of 7, and after output, adds the features to the features after a 1x1 convolution operation on the original input features, then performs a 3x3 convolution with a dilation rate of 5, and then fuses with the features after a 1x1 convolution operation on the original input features. Repeat the above process to obtain the second aggregated feature. The dilation rate of each 3x3 dilated convolution is [7, 5, 3, 1].
[0124] After the first aggregation feature undergoes gradually deepening feature extraction operations, four second aggregation features of different scales are obtained. The shallower feature maps contain preliminary features such as lines, edges, and corners, while the deeper feature maps contain semantic features of the mining area. The decoding module fully fuses the second aggregation features of different scales. The decoding module starts from the deepest feature and continuously fuses with the shallower features upward, as shown in Figure 10. It receives the input of two second aggregation features at different levels. The deeper feature is first upsampled to the size of the shallower feature through bilinear interpolation and then preliminarily fused together through a channel-level concatenation operation. Then, a 1x1 convolution is used for channel information aggregation, and a 7x7 depthwise separable convolution is used for spatial-level feature fusion. The fused results are respectively obtained through max pooling and average pooling at the channel level and max pooling and average pooling at the spatial level to obtain a feature column rich in channel information and a feature map rich in spatial information. After performing fully connected operations and convolution operations on them respectively, and then cross-multiplying, two weight matrices are obtained. Finally, the two weight matrices are used to weight and sum the features of the two scales to obtain a feature map that fully fuses the feature information of the two different scales. Then, the fusion operation is performed from bottom to top in sequence to obtain the final segmentation binary map, that is, the mining area module segmentation image.
[0125] Step S103: Obtain the remotely sensed image to be labeled, and input the remotely sensed image to be labeled into the benchmark mining area module segmentation model to obtain a mining area module segmentation image; wherein, the mining area module segmentation image includes a plurality of mining area patches.
[0126] Specifically, input the remotely sensed image to be labeled into the benchmark mining area module segmentation model to obtain a mining area module segmentation image including a plurality of mining area patches.
[0127] On this basis, obtaining the remotely sensed image to be labeled and inputting the remotely sensed image to be labeled into the benchmark mining area module segmentation model to obtain a mining area module segmentation image includes:
[0128] Obtain the remotely sensed image to be labeled, and input the remotely sensed image to be labeled into the benchmark mining area module segmentation model. The benchmark mining area module segmentation model receives the input remotely sensed image to be labeled, extracts preliminary image features through a feature extraction backbone network, and obtains features of five stages with different depths and different sizes;
[0129] Input the features into the neighbor feature aggregation module, and aggregate the features of adjacent stages through the neighbor feature aggregation module to obtain a first aggregation feature;
[0130] Input the first aggregation feature into the feature enhancement module. The feature enhancement module obtains superimposed features with different receptive fields through the superimposition of dilated convolutions, and aggregates the superimposed features together through residual connections to obtain a second aggregation feature;
[0131] Input the second aggregated feature into the decoding module, and the decoding module restores the second aggregated feature to the size of the input remote sensing image to be labeled, obtaining the mining area module segmentation image.
[0132] Specifically, after the model receives an input remote sensing image to be labeled, it extracts preliminary image features through the feature extraction backbone network (MobileNet), and through continuous deepening, obtains features of five stages with different depths and different sizes. Then, the adjacent stage features are aggregated through the neighbor feature aggregation module to achieve the effect of preliminary feature enhancement, and the preliminarily enhanced features are input into the feature enhancement module. The feature enhancement module obtains features with different receptive fields through the stacking of dilated convolutions, and aggregates these features together through residual connections. Finally, through the decoding module, it is restored to the size of the input remote sensing image to be labeled, obtaining the final segmentation result and the mining area module segmentation image.
[0133] Step S104, statistically analyze the object features of the mining area patches according to the mining area module segmentation image to obtain the corresponding land cover classes; wherein, the object features include: index features and geometric features; the land cover classes include: mineral classes and non-mineral classes.
[0134] Specifically, taking the mining area patches as objects and the features of the mining area patches as the core, calculate the object features to identify the land cover classes to which the objects belong. The object features include: index features and geometric features; the land cover classes include: mineral classes and non-mineral classes. The index features mainly refer to the custom-related feature indices reflected by the land cover in the high-resolution image, which are used to describe the relevant land cover, and have the characteristics of strong pertinence and reflecting the specified land cover. For example, the normalized difference vegetation index (NDVI) that is sensitive to green land cover and is usually used to reflect the vegetation coverage, the normalized difference water index (NDWI) that is sensitive to cyan-blue land cover and is usually used to reflect the water coverage, the normalized house shape index (NHSI) that is sensitive to regular shapes and is usually used to reflect buildings, and so on.
[0135] On this basis, statistically analyze the object features according to the mining area module segmentation image, and the obtained land cover classes include:
[0136] Obtain the mining area patches in the mining area module segmentation image, and calculate the index features and geometric features of each mining area patch; wherein, the index features include: normalized difference vegetation index, normalized difference water index, BG, BR, GR; the geometric features include: area and aspect ratio;
[0137] According to the index features and the geometric features, obtain the corresponding land cover classes according to the preset feature statistical results.
[0138] Specifically, the index features of the imaging ground object targets are for a specific ground object, and describe the relevant features of the specific ground object through features such as spectrum and geometry, and are often used as the main features for extracting the corresponding ground objects. Commonly used index features include Normalized Difference Vegetation Index (NDVI), Normalized Difference Water Index (NDWI), Normalized House Shape Index (NHSI), etc. Since vegetation and water bodies are the most frequently occurring ground objects near mining areas, while houses appear less frequently, in this embodiment, the Normalized Difference Vegetation Index (NDVI), Normalized Difference Water Index (NDWI), BG, BR, and GR are selected as the index feature quantization indicators.
[0139] The geometric features reflect the geometric and morphological information of the ground object itself manifested on the image object. The image object is a set of adjacent pixels with similar gray levels, and the geometric shape features of the ground object can naturally be reflected on the two-dimensional image.
[0140] According to the shallow feature knowledge base of common open-pit mining activity patches, the 15 types of shallow feature values of spectral features, texture features, index features, and geometric features are statistically analyzed by ore type, and the screening ranges of 15 types of shallow features of 28 ore type patches are obtained, and the preset feature statistical results are obtained. The preset feature statistical results are the screening basis for the mining area patches. Among them, the screening ranges of the index features of all ore types are the mean ranges of NDVI, NDWI, BG, BR, and GR in the shallow feature screening table. The geometric feature screening ranges are uniformly set according to the area A≥300 and the aspect ratio 0.2≤LW≤5. The statistical analysis of the other 12 types of features is subdivided by ore type, and the details are shown in the attached table. Table 2 intercepts and shows the screening range settings of 12 types of shallow features of some ore types (limestone, sandstone).
[0141] Table 2 Partial Shallow Feature Screening Table for Limestone and Sandstone Patches
[0142]
[0143] On this basis, the mining area patches in the segmented image of the mining area module are obtained, and the index features and geometric features of each mining area patch are calculated; among them, the index features include: Normalized Difference Vegetation Index, Normalized Difference Water Index, BG, BR, GR; the geometric features include: area and aspect ratio include:
[0144] The Normalized Difference Vegetation Index is calculated by the following formula:
[0145]
[0146] where NDVI represents the Normalized Difference Vegetation Index, NIR represents the mean value of the near-infrared band of the mining area patch, and R represents the mean value of the red band;
[0147] Calculate the Normalized Difference Water Index (NDWI) using the following formula:
[0148]
[0149] Where NDWI represents the Normalized Difference Water Index, G represents the average value of the green band, and NIR represents the average value of the near-infrared band of the mining area patch;
[0150] Calculate the index relationship between the blue band and the red band in the mining area patch according to the following formula:
[0151]
[0152] Where BR represents the index relationship between the blue band and the red band in the mining area patch, B represents the average value of the blue band, and R represents the average value of the red band;
[0153] Calculate the index relationship between the green band and the red band in the mining area patch according to the following formula:
[0154]
[0155] Where GR represents the index relationship between the green band and the red band in the mining area patch, G represents the average value of the green band, and R represents the average value of the red band;
[0156] Calculate the index relationship between the blue band and the green band in the mining area patch according to the following formula:
[0157]
[0158] Where BG represents the index relationship between the blue band and the green band in the mining area patch, B represents the average value of the blue band, and G represents the average value of the green band;
[0159] Calculate the area of the mining area patch using the following formula:
[0160]
[0161] Where A represents the area of the mining area patch, a i represents the actual area of the i-th pixel, and n represents the number of pixels contained in the mining area patch;
[0162] Calculate the aspect ratio of the mining area patch using the following formula:
[0163]
[0164] Where LW represents the aspect ratio of the mining area patch, l represents the length, w represents the width, A represents the area of the mining area patch, l0 represents the length of the circumscribed rectangle of the border, and w0 represents the width of the circumscribed rectangle of the border.
[0165] Specifically, calculate the object characteristics using the following formula.
[0166] 1) Normalized Difference Vegetation Index (NDVI)
[0167] In high - resolution remote sensing images, it can usually be extracted through various vegetation indices composed of the red band and the near - infrared band. The commonly used one is the Normalized Difference Vegetation Index NDVI:
[0168]
[0169] NDVI represents the Normalized Difference Vegetation Index, NIR represents the mean value of the near - infrared band of the mining area map patch, and R represents the mean value of the red band.
[0170] 2) Normalized Difference Water Index
[0171] The Normalized Difference Water Index (NDWI) is used to extract water body information in remote sensing images by performing difference processing on the green band (G) and the near - infrared band (NIR) of remote sensing, and the effect is good. The specific calculation formula is as follows:
[0172]
[0173] NDWI represents the Normalized Difference Water Index, G represents the mean value of the green band, and NIR represents the mean value of the near - infrared band of the mining area map patch.
[0174] 3) BR Index
[0175] The BR Index calculates the relationship between the blue band (B) and the red band (R) in the mining area map patch. The specific calculation formula is as follows:
[0176]
[0177] BR represents the index relationship between the blue band and the red band in the mining area map patch, B represents the mean value of the blue band, and R represents the mean value of the red band.
[0178] 4) GR Index
[0179] The GR Index is similar to the BR Index and is used to calculate the index relationship between the green band (G) and the red band (R). The specific calculation formula is as follows:
[0180]
[0181] GR represents the index relationship between the green band and the red band in the mining area map patch, G represents the mean value of the green band, and R represents the mean value of the red band.
[0182] 5) BG Index
[0183] The BG Index is similar to the BR and GR Indices and is used to calculate the index relationship between the green band (G) and the blue band (B). The specific calculation formula is as follows:
[0184]
[0185] BG represents the exponential relationship between the blue band and the green band in the mining area patch. B represents the average value of the blue band, and G represents the average value of the green band.
[0186] 6) Area (Area, A)
[0187] The area A of the image object refers to the sum of the areas of the pixels that make up the image object. If the image has no geographic coordinates, the area of the object is the number of pixels that make up the image object. The specific calculation formula is as follows:
[0188]
[0189] A represents the area of the mining area patch, a i represents the true area of the i-th pixel, and n represents the number of pixels contained in the mining area patch.
[0190] 7) Aspect ratio (Length / Width, LW)
[0191] LW is the ratio of the length to the width of the image object, and is usually used to distinguish linear features. Usually, the aspect ratio of the image object can be approximated by the bounding box:
[0192]
[0193] LW represents the aspect ratio of the mining area patch, l represents the length, w represents the width, A represents the area of the mining area patch, and l0, w0 are the length and width of the circumscribed rectangle of the border box.
[0194] In step S105, based on the segmentation of the image by the mining area module, the mining area patches with the land cover type of minerals are retained to obtain the remote sensing mining area annotation result.
[0195] Specifically, the segmentation image of the mining area module is combined with the land cover type. Based on the segmentation image of the mining area module, the mining area patches with the land cover type of minerals are retained to obtain the remote sensing mining area annotation result.
[0196] To verify the effectiveness of the benchmark mining area module segmentation model, visual display and quantitative comparison of evaluation metrics were carried out. The images for visual display were all from the test machine, and the most comprehensive evaluation metric values were also from the mean of all remote sensing sample images in the test set.
[0197] Such as Figure 11 and Figure 12As shown. In the figure, the images in the first row are the remote sensing sample image data input into the benchmark mining area module segmentation model in the test set. The second row (P) represents the prediction results after being processed by the benchmark mining area module segmentation model, and the third row (GT) represents the true labels of the mining areas manually annotated in the dataset. It can be seen from the above images that the prediction results processed by the benchmark mining area module segmentation model are relatively close to the manually annotated labels, proving the effectiveness of the benchmark mining area module segmentation model.
[0198] To verify the effectiveness of the benchmark mining area module segmentation model, a quantitative evaluation was carried out by comparing the accuracies with current common deep learning image segmentation models such as UNet, SegNet, ST-Unet, and Easy-Net on the mine data test set. To reflect the overall recognition effect of various models on the mining area patches, the mining area patches are regarded as a whole here, without distinguishing the ore types, that is, only identifying and comparing according to "mining area patches" and "non-mining area patches". The calculation results of various model indicators are shown in Table 3.
[0199] Table 3 Comparison of overall recognition accuracies of several segmentation methods
[0200]
[0201] It can be seen from Table 3 that the method based on the shallow feature knowledge base has the worst indicators among all methods, with a large difference in almost all indicators compared with other deep learning methods. The differences in various indicators among these deep learning methods are relatively small. The method that only combines shallow feature statistics with traditional segmentation and classification has a much lower accuracy than deep learning methods. The OMSegNet method proposed in this application has significantly better performance in multiple indicators than other types of models, significantly better than the method that only uses shallow feature screening and models such as U-Net, SegNet, and DUSegNet, especially in its highest precision. It also has obvious advantages compared with other recent new semantic segmentation models. The comprehensive performance evaluation index F1 of the benchmark mining area module segmentation model reaches 0.9038, which is 3.26% higher than the F1 index (0.8712) of the second-place method, and far higher than other deep learning methods in the table.
[0202] Five remote sensing sample images were randomly selected from the test set for a comparative experiment visualization display to comprehensively evaluate the performance of various methods when processing different remote sensing sample images, such as Figure 13As shown in the figure. In the figure, TP, TN, FP, and FN are highlighted in white (the mining area is detected as the mining area), black (the non-mining area is detected as the non-mining area), dark gray (the non-mining area is detected as the mining area), and light gray (the mining area is detected as the non-mining area) respectively. The specific color representation can be set arbitrarily according to requirements. It can be clearly seen from the figure that the dark gray and light gray areas in the segmentation map of the benchmark mining area module segmentation model are less than those of other models. Especially when dealing with some mining areas with complex terrain and smaller mining areas, the benchmark mining area module segmentation model shows relatively excellent results. It can be seen from the overall visualization effect diagram that the performance result of the benchmark mining area module segmentation model in the remote sensing mining area segmentation task is more excellent than that of other models.
[0203] Due to the limitation of the GPU, the benchmark mining area module segmentation model can only process smaller image data. Therefore, when inputting a large remote sensing image, the large image will first be cut into small images (256x256) with the same size as the dataset images. This operation will cause some non-mining area targets with large color differences from the surrounding environment to be misdetected as mining areas. Such as Figure 14 shown, many bare lands in the upper right corner are misdetected as mining areas due to obvious color differences from the surrounding environment, as well as some roads below.
[0204] The above situation is significantly reduced through shallow feature screening, as Figure 15 shown. In Figure 15 , it can be clearly seen that many misdetected areas in the non-mining area are well screened out through exponential feature screening. In Figure 14 , some very small misdetection results in the upper left corner are also successfully excluded through area screening of geometric screening, and the misdetection results of the roads below are also successfully excluded through the aspect ratio screening condition of geometric screening. Shallow feature screening has a good correction effect on deep learning models, can greatly improve the segmentation effect of mining areas, and improve the segmentation accuracy of the entire mining area segmentation task.
[0205] To further improve the automatic classification accuracy of mine remote sensing images and deepen the practicality of intelligent means in mineral resource supervision work, this application combines the shallow features of remote sensing image mining activity patches, deep learning theory, and GIS technology to construct a refined recognition model for mining activity patches for multi-source remote sensing images. The advantages of the present invention are as follows:
[0206] 1) Constructed a set of basic sample libraries for open-pit mines of multi-source images, containing 3,979 pairs of samples. Each pair of samples contains a remote sensing sample image with a size of 256x256 and the binary mask map of the corresponding mining area in the image, providing rich background data references for the automatic recognition of mine remote sensing.
[0207] 2) Shallow quantization indicators for remote sensing images were established, and a knowledge base of shallow features of mines was constructed.
[0208] 3) A benchmark mining area module segmentation model that fuses shallow features and the OMSegNet deep network was proposed, which can effectively circle and classify mining area patches such as limestone, sandstone, and shale, effectively improving the efficiency of actual interpretation work.
[0209] In the model construction stage, a deep learning semantic segmentation model that fuses shallow features and the OMSegNet network was proposed and samples were input for training. After about 8 hours of training, the loss function stayed at about 0.15, and the training accuracy was about 98.34%. It can effectively circle mining area patches on the test data, and the extracted contour has a good fit with the actual mine boundary. After the initial extraction by the model, the mining patches are then carefully screened by humans. Compared with manual interpretation, it can greatly improve work efficiency while taking into account accuracy, and has higher practicality.
[0210] 4) The experimental results of mining area recognition show that the model of this project has high accuracy in multiple evaluation indicators, superior to traditional segmentation methods and other semantic segmentation models that do not fuse shallow features.
[0211] In the accuracy evaluation stage, an evaluation index system to describe the model effect was constructed using accuracy (AP), precision (P), recall (R), and F-measure (F). The classification and recognition effects of mining area patches were horizontally compared among the statistical method based on the feature knowledge base, U-Net, SegNet, DUSegNet, and the DUSegNet method that fuses shallow features. The advantages and disadvantages of various methods were qualitatively and quantitatively compared intuitively and by indicators. The experimental results show that the OMSegNet method that fuses shallow features proposed in this invention can effectively identify mining area patches, and the circled mining area edge contour is close to the manual interpretation effect, and its segmentation range has a good fit with the real boundary of the mining area. From the quantitative comparison of evaluation factors, this method has a relatively high overall recognition accuracy, superior to the traditional segmentation and classification methods based on shallow feature statistics, U-Net, SegNet, and the DUSegNet method that does not fuse shallow features. This model is applicable to the mining area patch recognition tasks of seven types of remote sensing image data, namely GF-1, GF-2, GF-6, GF-7, Ziyuan-1, Ziyuan-3, and CBERS-4, and has a better recognition effect on images with an original data resolution of 2m.
[0212] It should be noted that the method of the embodiment of the present invention can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiment of the present invention, and these multiple devices will interact with each other to complete the described method.
[0213] It should be noted that some embodiments of the present invention have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order from that in the above embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or consecutive order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0214] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present invention further provides a remote sensing mining area automatic annotation system based on deep learning and shallow feature statistics.
[0215] Reference Figure 16 , the remote sensing mining area automatic annotation system based on deep learning and shallow feature statistics includes:
[0216] A sample acquisition module 161, configured to acquire a given set of remote sensing image samples; wherein, the set of remote sensing image samples includes: remote sensing sample images and reference mining area module segmentation images;
[0217] A model training module 162, configured to use each remote sensing sample image as the input of the network source, and the reference mining area module segmentation image corresponding to the remote sensing sample image as the output, and in combination with a preset loss function and evaluation index, apply each remote sensing sample image to train a target deep learning image segmentation model to obtain a reference mining area module segmentation model;
[0218] A patch segmentation module 163, configured to acquire a remote sensing image to be annotated, input the remote sensing image to be annotated into the reference mining area module segmentation model, and obtain a mining area module segmentation image; wherein, the mining area module segmentation image includes a plurality of mining area patches;
[0219] A patch classification module 164, configured to statistically analyze the object features of the mining area patches according to the mining area module segmentation image to obtain the corresponding land cover classes; wherein, the object features include: index features and geometric features; the land cover classes include: mineral classes and non - mineral classes;
[0220] The mining area annotation module 165 is used to retain the mining area patches with the feature class of minerals based on the image segmented by the mining area module, so as to obtain the remote sensing mining area annotation result.
[0221] For the sake of convenience of description, when describing the above system, various modules are described separately according to their functions. Of course, when implementing the present invention, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0222] The system of the above embodiment is used to implement the corresponding remote sensing mining area automatic annotation method based on deep learning and shallow feature statistics in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0223] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the remote sensing mining area automatic annotation method based on deep learning and shallow feature statistics described in any of the above embodiments.
[0224] Figure 17 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0225] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0226] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0227] The input / output interface 1030 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure), or can be externally connected to the device to provide corresponding functions. Among them, the input devices can include keyboards, mice, touchscreens, microphones, various sensors, etc., and the output devices can include displays, speakers, vibrators, indicator lights, etc.
[0228] The communication interface 1040 is used to connect to the communication module (not shown in the figure) to achieve communication and interaction between this device and other devices. Among them, the communication module can achieve communication through wired means (such as USB, network cable, etc.), or can achieve communication through wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0229] The bus 1050 includes a passage for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0230] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0231] The electronic device of the above embodiment is used to implement the corresponding remote sensing mining area automatic annotation method based on deep learning and shallow feature statistics in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0232] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present invention also provides a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to make the computer execute the remote sensing mining area automatic annotation method based on deep learning and shallow feature statistics as described in any of the foregoing embodiments.
[0233] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0234] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the automatic annotation method for remote sensing mining areas based on deep learning and shallow feature statistics as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0235] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present invention (including the claims) is limited to these examples; under the concept of the present invention, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present invention as described above, which are not provided in detail for the sake of brevity.
[0236] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code that includes one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed. This should be understood by those skilled in the art to which the embodiments of the present invention belong.
[0237] In addition, for simplicity of explanation and discussion, and so as not to make the embodiments of the present invention difficult to understand, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the systems may be shown in block diagram form in order to avoid making the embodiments of the present invention difficult to understand, and this also takes into account the fact that details regarding the implementation of these block diagram systems are highly dependent on the platform on which the embodiments of the present invention are to be implemented (i.e., these details should be fully within the understanding of those of ordinary skill in the art). In cases where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present invention, it will be apparent to those of ordinary skill in the art that the embodiments of the present invention may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions should be regarded as illustrative rather than restrictive.
[0238] Although the present invention has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0239] Embodiments of the present invention are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the embodiments of the present invention shall be included within the protection scope of the present invention.
Claims
1. A remote sensing mining area automatic labeling method based on deep learning and shallow feature statistics, characterized in that: include: Acquire a given remote sensing image sample set; wherein the remote sensing image sample set includes: remote sensing sample images and benchmark mining area module segmentation images; Each remote sensing sample image is used as the network source input, and the benchmark mining area module segmentation image corresponding to the remote sensing sample image is used as the output. Combined with the preset loss function and evaluation index, each remote sensing sample image is applied to train the target deep learning image segmentation model to obtain the benchmark mining area module segmentation model; Acquire a remote sensing image to be annotated, input the remote sensing image to be annotated into the reference mining area module segmentation model, and obtain a mining area module segmentation image; wherein the mining area module segmentation image includes a plurality of mining area spots; According to the mining module segmentation image, the object features of the mining area patches are counted to obtain the corresponding ground feature categories; wherein the object features include index features and geometric features; the ground feature categories include mineral categories and non-mineral categories; Based on the segmented image of the mining area module, the mining area patches with mineral features are retained to obtain the remote sensing mining area annotation results.
2. The remote sensing mining area automatic labeling method based on deep learning and shallow feature statistics according to claim 1 is characterized in that: The step of obtaining a given remote sensing image sample set further comprises: Acquisition of remote sensing images; Accept the annotation information and obtain the benchmark mining area module segmentation image corresponding to the remote sensing image; Perform image cutting on the benchmark mining area module segmented image according to a preset cutting step length to obtain a valid annotated image; Screening the effective annotated images and unifying the image formats to obtain screened annotated images; The screened and annotated image is randomly flipped, randomly cropped, and randomly color blurred to obtain a remote sensing sample image.
3. The remote sensing mining area automatic labeling method based on deep learning and shallow feature statistics according to claim 2 is characterized in that: The step of performing image cutting on the benchmark mining area module segmented image according to a preset cutting step length to obtain a valid annotated image comprises: Performing image cutting on the benchmark mining area module segmentation image with a preset cutting step length to obtain a cutting shp file; Based on the cutting shp file, coordinate transformation is performed through the following formula to obtain a valid annotated image: Among them, x is the horizontal coordinate of the effective labeled image, x geo represents the horizontal coordinate of the cut shp file, x′ represents the horizontal coordinate of the vertex in the upper left corner of the image, p represents the spatial resolution of the pixel, y is the vertical coordinate of the effective labeled image, y′ represents the vertical coordinate of the vertex in the upper left corner of the image, y geo Indicates the vertical coordinate for cutting the shp file.
4. The remote sensing mining area automatic labeling method based on deep learning and shallow feature statistics according to claim 1 is characterized in that: The target deep learning image segmentation model includes: Feature extraction backbone network, neighbor feature aggregation module, feature enhancement module and decoding module.
5. The remote sensing mining area automatic labeling method based on deep learning and shallow feature statistics according to claim 4 is characterized in that: The step of acquiring a remote sensing image to be labeled and inputting the remote sensing image to be labeled into the reference mining area module segmentation model to obtain a mining area module segmentation image comprises: Acquire a remote sensing image to be labeled, input the remote sensing image to be labeled into the benchmark mining area module segmentation model, the benchmark mining area module segmentation model receives the input remote sensing image to be labeled, extracts preliminary image features through a feature extraction backbone network, and obtains features of five stages of different depths and sizes; Inputting the feature into a neighbor feature aggregation module, aggregating features of adjacent stages through the neighbor feature aggregation module to obtain a first aggregate feature; The first aggregated feature is input into a feature enhancement module, the feature enhancement module obtains superposition features of different receptive fields by superposition of dilated convolutions, and aggregates the superposition features together by residual connection to obtain a second aggregated feature; The second aggregated features are input into a decoding module, and the decoding module restores the second aggregated features to the size of the input remote sensing image to be annotated, so as to obtain a mining area module segmentation image.
6. The remote sensing mining area automatic labeling method based on deep learning and shallow feature statistics according to claim 1 is characterized in that: The object features are statistically analyzed based on the mining area module segmentation image to obtain the corresponding ground feature categories: Obtain the mining area spots in the mining area module segmented image, and calculate the index features and geometric features of each mining area spot; wherein the index features include: normalized vegetation index, normalized water index, BG, BR, GR; the geometric features include: area and aspect ratio; According to the index feature and the geometric feature, the category of the feature is obtained according to the preset feature statistics results.
7. The remote sensing mining area automatic labeling method based on deep learning and shallow feature statistics according to claim 6 is characterized in that: The mining area patches in the mining area module segmented image are obtained, and the index features and geometric features of each mining area patch are calculated; wherein the index features include: normalized vegetation index, normalized water index, BG, BR, GR; the geometric features include: area and aspect ratio include: The normalized difference vegetation index is calculated by the following formula: Among them, NDVI represents the normalized vegetation index, NIR represents the mean value of the near infrared band of the mining area, and R represents the mean value of the red band; The normalized water index is calculated by the following formula: Among them, NDWI represents the normalized water index, G represents the mean value of the green band, and NIR represents the mean value of the near infrared band of the mining area patch; The exponential relationship between the blue band and the red band in the mining area map is calculated according to the following formula: Among them, BR represents the exponential relationship between the blue band and the red band in the mining area, B represents the mean of the blue band, and R represents the mean of the red band; The exponential relationship between the green band and the red band in the mining area map is calculated according to the following formula: Among them, GR represents the exponential relationship between the green band and the red band in the mining area, G represents the mean of the green band, and R represents the mean of the red band; The exponential relationship between the blue band and the green band in the mining area map is calculated according to the following formula: Among them, BG represents the exponential relationship between the blue band and the green band in the mining area, B represents the mean of the blue band, and G represents the mean of the green band; The area of the mining area is calculated by the following formula: Among them, A represents the area of the mining area, a i represents the real area of the ith pixel, and n represents the number of pixels contained in the mining area patch; The aspect ratio of the mining area is calculated by the following formula: Among them, LW represents the aspect ratio of the mining area, l represents the length, w represents the width, A represents the area of the mining area, l0 represents the length of the circumscribed rectangle of the border, and w0 represents the width of the circumscribed rectangle of the border.
8. The remote sensing mining area automatic labeling system based on deep learning and shallow feature statistics is characterized by: The method for automatically labeling a remote sensing mining area based on deep learning and shallow feature statistics as claimed in any one of claims 1 to 7 comprises: A sample acquisition module is used to acquire a given remote sensing image sample set; wherein the remote sensing image sample set includes: remote sensing sample images and benchmark mining area module segmentation images; The model training module is used to use each remote sensing sample image as the network source input, and the benchmark mining area module segmentation image corresponding to the remote sensing sample image as the output. In combination with the preset loss function and evaluation index, each remote sensing sample image is applied to train the target deep learning image segmentation model to obtain the benchmark mining area module segmentation model; A spot segmentation module is used to obtain a remote sensing image to be labeled, input the remote sensing image to be labeled into the reference mining area module segmentation model, and obtain a mining area module segmentation image; wherein the mining area module segmentation image includes a plurality of mining area spots; The image patch classification module is used to count the object features of the mining area patches according to the mining area module segmentation image, and obtain the corresponding ground feature category; wherein the object features include: index features and geometric features; the ground feature category includes: mineral category and non-mineral category; The mining area labeling module is used to retain the mining area patches with mineral categories based on the mining area module segmentation image, and obtain the remote sensing mining area labeling results.
9. A device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Mining area land utilization identification method, device, equipment, medium and product
CN114549534A
Mining area image recognition method and device, server and storage medium
CN115953682A
Mining area image recognition model construction method and device
CN116052013A
Typical tailing pond remote sensing target identification method based on deep learning and random forest
CN116109935A
Cited By
Garlic identification system and method based on remote sensing deep learning
CN120563945A
Garlic identification system and method based on remote sensing deep learning
CN120563945B