Mining area typical ground feature recognition method based on improved YOLO11 algorithm
Through the improved YOLO11 algorithm, the problems of low recognition accuracy and serious missed detection in mining areas are solved, and high-precision, automated and intelligent land object recognition are achieved, suitable for complex backgrounds and multi-scale images.
Patent Information
- Application Number
- CN202510304396.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-24
AI Technical Summary
The existing mining area land object recognition methods have problems with low recognition accuracy and serious missed detection and missed detection when the background is complex, the image scale is diverse, the clustering of small targets and the differences between the targets and backgrounds.
The improved YOLO11 algorithm uses a multi-branch convolution module to replace the C3K2 residual block, a downsampled convolution with an Adown module, and a Detect detection head with an E_Detect structure to build a detection model that is more suitable for the mining area data set.
It realizes high-precision identification of mining areas, reduces false detection and missed detection, improves the automation, intelligence and scale capabilities of identification, and has a fast inference speed.
Smart Images

Figure CN120198644A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of mine ground object recognition methods, and relates to a method for recognizing typical mine ground objects based on an improved YOLO11 algorithm. Background Technique
[0002] As the main place for mineral resource development, it is of great significance to conduct a timely and comprehensive investigation of mines, carry out comprehensive mine rectification according to law, and strengthen mine ecological restoration to achieve sustainable development of mineral resources.
[0003] Traditional mine monitoring methods mainly summarize information through a combination of manual field surveys and statistics. Although they have extremely high accuracy, they consume a large amount of manpower, material resources and financial resources, and have a certain lag. In the early 20th century, with the continuous improvement of digital mine-related technologies, the means of ecological environment monitoring have become increasingly rich. Remote sensing technology has become an important means of mine monitoring due to its advantages such as intuitiveness, speed, efficiency and accuracy. However, it is still time-consuming and laborious to obtain information through manual visual interpretation of a large number of remote sensing images, and it cannot meet the need to quickly obtain a large amount of mine information.
[0004] In recent years, computer technology has developed vigorously. Deep learning has received more and more attention from scholars due to its advantages in multi-scale and multi-level feature extraction. The convolutional neural network it contains is widely used in various computer vision tasks because it is good at extracting local features in images. Currently, object detection algorithms based on deep learning can be divided into single-stage and two-stage. Among them, the single-stage object algorithm extracts feature information from the original image and then predicts the object category and anchor box, while the two-stage algorithm needs to first generate a series of potential object regions and then make predictions. The single-stage object algorithm is more suitable for real-time response application scenarios due to its faster processing speed and lower resource requirements.
[0005] At present, representative single-stage object detection algorithms include: SSD (Single Shot MultiBox Detector), YOLO (You Only Look Once) and other series of algorithms. Among them, the YOLO algorithm has laid a new foundation for the development of object detection algorithms. However, due to problems such as the difficulty in optimizing the non-maximum suppression (NMS) processing of the YOLO detector and its lack of robustness, the object detection accuracy is relatively low. As the first end-to-end algorithm based on the transformer, DRTR (Detection Transformer) successfully solves the problem of difficult NMS processing and improves the object recognition ability. Although the optimized RT-DETR (Real Time-Detection Transformer) has greatly improved the convergence and inference speed compared with the DRTR algorithm, it still cannot meet the real-time requirements. YOLO11 developed by Ultralytics has very excellent performance in terms of accuracy and speed, can capture complex details in images more accurately, and supports multiple tasks such as object detection, instance segmentation, and oriented object detection. However, the YOLO11 object detection algorithm has certain limitations when applied to the mining area detection task. As a single-stage object detection algorithm, YOLO11 converts the detection task into a regression problem, lacking pixel-level precise positioning, resulting in difficult accurate prediction when dealing with small objects or overlapping targets, and there are many false detections and missed detections in multi-scale and multi-direction detection tasks. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for identifying typical ground objects in mining areas based on an improved YOLO11 algorithm, which has the characteristics of high recognition accuracy and high reliability.
[0007] The technical solution adopted by the present invention is that the method for identifying typical ground objects in mining areas based on the improved YOLO11 algorithm is specifically implemented according to the following steps: Step 1: Collect mining area image data and process it to obtain an initial data set; Step 2: Label the initial data set obtained in Step 1 to construct a mining area target ground object data set; Step 3: Construct an improved YOLO11 algorithm, and use the data set obtained in Step 2 to train the algorithm to obtain a mining area target ground object detection model; Step 4: Use the model constructed in Step 3 to identify ground objects in the mining area.
[0008] The characteristics of the present invention also lie in: In Step 1, the image data uses Tianditu images with a resolution of 0.9m and Google images with a resolution of 1.8m.
[0009] In step 1, the image data includes complex samples in arid and semi-arid climate regions, single samples in arid and semi-arid climate regions, and single samples in humid and semi-humid climate regions.
[0010] The processing of the image data in step 1 includes cropping and expanding the original image.
[0011] The specific processing method of the image data in step 1 is as follows: First, use the sliding window method to cut the original image from the upper left to the lower right with 10% overlap. Secondly, expand the data set by rotating the image 90°, 180°, 270°, mirroring, and at the same time making corresponding changes to the coordinates of the corresponding label anchor boxes, as well as adjusting the brightness and darkness.
[0012] In step 2, use the SAM semi-automatic annotation software to annotate the initial data set obtained in step 1.
[0013] In step 3, the improved YOLO11 algorithm is obtained by improving the YOLO11 algorithm. The specific improvements include: replacing the C3K2 residual block in the baseline model with a multi-branch convolution module, replacing the downsampling convolution of YOLO11 with an Adown module, and replacing the Detect detection head with an E_Detect structure.
[0014] In step 3, randomly divide the data set obtained in step 2 into a training set, a test set, and a validation set according to the ratio of 7:2:1, and then use this data set to train the improved YOLO11 algorithm.
[0015] The beneficial effects of the present invention are: (1) The method of the present invention aims at the problems of complex mining area background, diverse image scales, small target aggregation, and low difference between the target and the background, and improves the YOLO11 algorithm. The improved YOLO11-DAE algorithm is more suitable for the mining area data set with diverse scales and complex backgrounds, realizing the automation, intelligence, and scale of the recognition of ground objects in the mining area; (2) The YOLO11-DAE algorithm in the method of the present invention has very excellent performance in terms of accuracy and measuring the inference speed of the model, captures complex details in the image more accurately, and the anchor boxes of this algorithm have a higher confidence with the target ground objects, are more fitting and have no redundant phenomenon. By comparing with the existing commonly used YOLOv5n, YOLOv8n, YOLOv10n, and YOLO11n algorithms, the YOLO11-DAE algorithm can detect the target ground objects missed by the other algorithms, improving the accuracy and reliability of the recognition of ground objects in the mining area. Description of the Drawings
[0016] Figure 1 is a flowchart of the method of the present invention; Figure 2 This is the structural diagram of the YOLO11-DAE model of the present invention; Figure 3 This is the schematic structural diagram of DBB in the YOLO11-DAE algorithm of the present invention; Figure 4 This is the schematic structural diagram of the ADown module in the YOLO11-DAE algorithm of the present invention; Figure 5 This is the schematic diagram of the grouped convolution operation divided into 3 groups in the method of the present invention; Figure 6 This is the comparison diagram of the recognition results of the mining area target ground objects obtained in the embodiment of the present invention. Detailed implementation manners
[0017] The present invention will be described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0018] The method for identifying typical ground objects in a mining area based on the improved YOLO11 algorithm of the present invention, as Figure 1 shown, is specifically implemented according to the following steps: Step 1, collect the image data of the mining area and process it.
[0019] Collect the image data of the mining area. The image data uses the Tianditu image with a resolution of 0.9m and the Google image with a resolution of 1.8m, and the image data includes complex background samples in arid and semi-arid climate regions, single-background samples in arid and semi-arid climate regions, and single-background samples in humid and semi-humid climate regions to enhance the generalization ability of the model.
[0020] After the image data is collected, due to the memory limitations of the computer and the graphics processing unit (GPU), it is necessary to crop the original image, select a fixed window size, and use the sliding window method to cut the original image from the upper left to the lower right with 10% overlap. In addition, in order to expand the available training data and improve the generalization ability of the network, the dataset is enhanced. The dataset is expanded by rotating the image 90°, 180°, 270°, mirroring and making corresponding changes to the coordinates of the corresponding label anchor boxes, and adjusting the brightness and darkness, to obtain the initial dataset.
[0021] Step 2, label the initial dataset obtained in Step 1 to construct the NOMTTD dataset.
[0022] Use the SAM (Segment Anything Model) semi-automatic annotation software to create a dataset of target ground objects in the mining area (National open-pit mine target terrain dataset, NOMTTD), including five types of targets: mining areas, waste dumps, mining area buildings, water bodies, and roads. The JSON files generated during the annotation process contain both the target category labels and bounding box information required for object detection (used to generate txt labels), and the pixel-level classification information required for semantic segmentation (used to generate png labels). SAM uses the generated yolo-format txt labels for data augmentation to construct the dataset of target ground objects in the mining area. In addition, the generated png labels can provide data support for the subsequent dynamic monitoring research of the mine scene.
[0023] Step 3: Based on the YOLO11-DAE algorithm, construct a detection model for target ground objects in the mining area on the NOMTTD dataset obtained in Step 2.
[0024] The method of the present invention improves the YOLO11 algorithm to obtain the YOLO11-DAE algorithm, and its framework is as Figure 2 shown. The specific improvement method is as follows: Replace the residual block of C3K2 in the baseline model with a multi-branch convolution module (Diverse Branch Block, DBB) to enhance the feature capture ability of the network and its adaptability to complex scenes; then replace the network downsampling convolution with an Adown module to improve the model's expression ability for multi-scale features and reduce the feature loss rate; finally, replace the original Detect detection head with an E_Detect structure, which helps to reduce the complexity and processing steps of the network, thereby achieving lightweight processing.
[0025] As Figure 3 shown, it is a schematic diagram of the structure of the DBB module. The convolutional network of the YOLO11 baseline model consists of a single branch and it is difficult to balance lightweight and high accuracy at the same time. The core idea of DBB is structure re-parameterization, which can decouple the convolutional neural network structure during model training and testing to improve the model performance without increasing the inference time. Therefore, the method of the present invention replaces the residual block in the C3k2 module with a DBB module to enhance the original backbone, so as to improve the network's feature capture ability for multi-scales and its adaptability to complex scenes.
[0026] The DBB module utilizes two important properties of convolution, namely homogeneity and additivity, to achieve a balance between lightweight and high accuracy. During the training phase, the original k×k convolution is enhanced by using different combinations of 1×1, 1×1−k×k, and 1×1−AVG, while during the inference phase, standard convolution is continued for calculation. Among them, homogeneity is to convolve a kernel with a scaling value, which is equivalent to convolving the original kernel and then scaling the resulting feature map; additivity is the sum of two convolutions, which is equivalent to convolving the input with the sum of their respective kernels. The specific formulas are as follows: (1) As Figure 4 shown, it is the structural schematic diagram of the ADown module. First, the feature map with C input channels is passed through a two-dimensional average pooling layer, and then the pooled feature map is split into two feature maps with C / 2 channels to achieve the model's perception ability for different features. Then, one of the feature maps is passed through a two-dimensional max pooling layer, and the two feature maps are respectively passed through a Conv module, and finally concatenated to output a feature map with the same number of channels as the original. This module combines average pooling and max pooling to obtain the smoothness and saliency features of the same module, and then retains more feature information through the Conv module. In addition, the concatenation of different features also enhances the module's representation ability for different features, which helps to improve the model's performance in different scenarios.
[0027] The structure of E_Detect is as Figure 1 shown. It reduces the computational complexity and the number of parameters by simplifying and optimizing the convolution operation to achieve the purpose of model lightweight. Its structure includes two layers of 3×3 grouped convolutions (GConv) and one layer of 1×1 convolution (Conv2d), which are used to extract features and generate bounding boxes and classification outputs. Among them, grouped convolution divides the input channels into multiple groups and assigns an independent set of convolution kernels to each group. Each group has its own convolution operation, which improves the computational efficiency through parallel computing. The operation process is as Figure 5 shown. Grouped convolution divides the C (number of channels), H (height), and W (width) of the input feature map into m groups in the channel dimension, expressed as , and then independent convolution operations are performed on each Xi.
[0028] Randomly divide the NOMTTD dataset obtained in step 2 according to the ratio of training set:test set:validation set of 7:2:1, and then use this dataset to train the YOLO11-DAE model to construct a mining area target ground object detection model.
[0029] Step 4: Use the mining area target ground object detection model constructed in step 3 to identify ground objects in the mining area.
[0030] Example 1: In this example, a dataset is constructed and selected using 0.9m Tianditu images and 1.8m Google images covering the Shendong Base, Xinjiang Base, Inner Mongolia East Base, Northern Shanxi Base, Northern Shaanxi Base, and Yunnan-Guizhou Base in China. The areas where the six major mining bases are located contain three different scenarios, namely complex samples in arid and semi-arid climate regions, single samples in arid and semi-arid climate regions, and single samples in humid and semi-humid climate regions. After annotation and processing, the NOMTTD dataset is formed.
[0031] This example runs on a computer with the operating system of windows10, the processor of Intel(R) Core(TM) i7-8700 CPU @3.20GHz 3.19GHz, the memory size of 32GB. The development tool is Pychram, the development language is Python3.8.18, the deep learning framework is PyTorch2.0.1, and the GPU acceleration library uses CUDA11.8. The training parameters are shown in Table 1 for details.
[0032] Table 1 Training Parameters
[0033] Select precision (Precision, P), recall (Recall, R), F1-Score (F1), Mean Average Precision (mAP), and the model operation parameter FPS (Frames Per Second) as evaluation indicators. The calculation formulas are as follows: (2) (3) (4) (5) (6) (7) In the formula: TP, FP, and FN are the number of positive samples predicted as positive samples, the number of negative samples predicted as positive samples, and the number of positive samples predicted as negative samples respectively; AP is the average precision, and the average of the average precisions of all classes is mAP@0.5. mAP@0.5 refers to the mAP when the IoU (Intersection over Union) threshold is 0.5; FPS is the number of frames of pictures that the algorithm can detect in 1s (including the pre-processing time Pre, the inference time Infer, and the post-processing time Post).
[0034] To verify the superiority and effectiveness of the improved YOLO11 algorithm (YOLO11-DAE algorithm) of the present invention, the YOLO11-DAE network model was compared with other mainstream YOLO series algorithms (YOLOv5n, YOLOv8n, YOLOv10n, YOLO11n) on the NOMTTD dataset, and the results are shown in Table 2.
[0035] Table 2 Comparison of ground object recognition results of different algorithms
[0036] As can be seen from Table 2, the YOLO11-DAE algorithm of the present invention has certain improvements in various evaluation indicators. Compared with the other four algorithms, the precision P of YOLO11-DAE has increased by 0.106, 0.06, 0.095, and 0.066 respectively; the recall rate R has increased by 0.15, 0.084, 0.108, and 0.081 respectively; the comprehensive evaluation index F1 and the mean average precision mAP@0.5 of YOLO11-DAE both reach above 0.9, which is about 0.07 higher than the second-highest model; the FPS is 528.1, which has decreased compared with the baseline model YOLO11n, but still remains at a good level and meets the requirements of real-time detection.
[0037] In summary, the YOLO11-DAE model of the present invention has a relatively fast inference speed, high inspection accuracy, and a decreased missed detection rate, showing obvious advantages compared with the other four algorithms.
[0038] Example 2: On the basis of Example 1, in order to more intuitively show the improvement effect, an image was randomly selected from the test sets of six major coal mine bases and input into five models for testing. The results are as Figure 6 shown. It can be seen that among the five models, YOLOv5n has the worst effect, with more missed detection phenomena, always having repeated predictions for a target, wasting computing resources, and the matching degree between the anchor box and the target ground object is poor, indicating that there is a certain ambiguity in the distinction between the target and the background of this model; YOLOv8n and YOLO11n perform relatively similarly, and the overall effect is better than YOLOv5n but lower than the YOLOv10n model; the YOLO11-DAE model has the best performance, and the confidence in judging the target ground object is significantly higher than that of the other four models, and the confidence is higher in the ground object prediction of the four major coal mining areas of Xinjiang, Shendong, northern Shanxi, and northern Shaanxi. The anchor box fits better with the target ground object and there is no redundancy. In addition, YOLO11-DAE also detected buildings and roads missed by the other four algorithms in the Inner Mongolia East and Yunnan-Guizhou bases, further verifying its superiority.
[0039] Example 3: On the basis of Example 1, the made NOMTTD dataset was input into the YOLO11-DAE model, and the classification results are shown in Table 3. It can be seen that the precision rates P of mine buildings, stope, waste dump, road, and water body are 0.88, 0.991, 0.978, 0.916, and 0.897 respectively. Among them, the precision performances of the stope, waste dump, and road are better, all higher than 0.9. The recall rate R of mine buildings is relatively low, being 0.795, and the rest are all higher than 0.85. Among them, the Rs of the waste dump and the mine stope are both higher than 0.94, and the missed detection rate is relatively low, and their mAP@0.5s are both higher than 0.98, indicating that the matching degree of the recognition effects of these two types of ground objects with the true labels manually marked is close. In terms of the four indicators, due to the relatively messy distribution of mine buildings, some buildings are too small and unclear in the pictures, resulting in relatively low precision and relatively more missed detection phenomena. From the overall results, the YOLO11-DAE algorithm of the present invention shows high detection precision and low missed detection rate for all five types of ground objects.
[0040] Table 3 Precision of Classification Results of YOLO11-DAE Algorithm
[0041] Example 4: On the basis of Example 1, ablation experiments were carried out.
[0042] In order to evaluate the effectiveness of the module, ablation experiments were carried out on the NOMTTD dataset to verify the impact of each improvement on the algorithm.
[0043] (1) Starting from the baseline model, first use the DBB module to change the residual blocks in C3K2; (2) Secondly, replace the downsampling parts in the backbone network and the feature pyramid of the model with the ADown module; (3) Thirdly, replace the detection head of the baseline model with the lightweight and efficient detection head E_Detect; (4) Finally, integrate the three modules to construct the YOLO11-DAE model of the present invention.
[0044] The results of the ablation experiments are shown in Table 4.
[0045] Table 4 Results of Ablation Experiments
[0046] As can be seen from Table 4, starting from the baseline model, first use the DBB module to change the residual blocks in C3K2. The P, R, F1, and mAP@0.5 of the model all increase, by approximately 0.03, and the FPS decreases slightly, by 1.4. The results show that relying solely on adding this one module, the improvement effect of the model is average. Secondly, replace the downsampling parts in the backbone network and feature pyramid of the model with the ADown module. The improvement in P, R, F1, and mAP@0.5 is obvious, increasing by approximately 0.05, but the FPS drops by 39.3, indicating that the inference speed of the model is affected to a certain extent. Thirdly, replace the detection head of the baseline model with the lightweight and efficient detection head E_Detect. The P, R, F1, and mAP@0.5 of the model remain almost unchanged, increasing by less than 0.01, but the FPS increases significantly, by 27.3, achieving the lightweight processing of the model. Finally, integrate the three modules to construct the YOLO11-DAE model of the present invention. The four evaluation indicators of P, R, F1, and mAP@0.5 are increased by 0.066, 0.081, 0.074, and 0.07 respectively, and the improvement effect is significant. The FPS drops by 16.4, but it increases by 24.7 compared with the model without adding the E_Detect detection head. In summary, although the FPS of the YOLO11-DAE model of the present invention drops to a certain extent, it still remains at a good level, and the other four indicators all increase significantly to varying degrees, indicating that each module added overall has a positive impact on the final result, and the model improvement is effective.
[0047] Example 5: The method for identifying typical ground objects in a mining area based on the improved YOLO11 algorithm in this embodiment is specifically implemented according to the following steps: Step 1, collect the image data of the mining area and process it to obtain the initial data set; Step 2, annotate the initial data set obtained in Step 1 to construct the data set of the target ground objects in the mining area; Step 3, construct the improved YOLO11 algorithm, and use the data set obtained in Step 2 to train this algorithm to obtain the detection model of the target ground objects in the mining area; Step 4, use the model constructed in Step 3 to identify the ground objects in the mining area.
[0048] Example 6: On the basis of Example 5, in Step 1, the image data uses the Tianditu image with a resolution of 0.9m and the Google image with a resolution of 1.8m, and this image data includes complex background samples in arid and semi-arid climate regions, single-background samples in arid and semi-arid climate regions, and single-background samples in humid and semi-humid climate regions.
Claims
1. A typical ground feature recognition method for mining areas based on the improved YOLO11 algorithm is characterized by: Follow the steps below to implement it: Step 1: Collect and process the mining area image data to obtain the initial data set; Step 2: annotate the initial data set obtained in step 1 to construct a target ground object data set in the mining area; Step 3: construct an improved YOLO11 algorithm, use the data set obtained in step 2 to train the algorithm, and obtain a mining area target object detection model; Step 4: Use the model constructed in step 3 to identify objects in the mining area.
2. The method for identifying typical objects in mining areas based on the improved YOLO11 algorithm according to claim 1 is characterized in that: In step 1, the image data uses 0.9m Tiandi map image and 1.8m Google image.
3. The method for identifying typical objects in mining areas based on the improved YOLO11 algorithm according to claim 1, characterized in that: In step 1, the image data includes complex background samples of arid and semi-arid climate zones, single background samples of arid and semi-arid climate zones, and single background samples of humid and semi-humid climate zones.
4. The method for identifying typical objects in mining areas based on the improved YOLO11 algorithm according to claim 1, characterized in that: The processing of the image data in step 1 includes cropping and expanding the original image.
5. The method for identifying typical objects in mining areas based on the improved YOLO11 algorithm according to claim 1 or 4, characterized in that: The specific method for processing the image data in step 1 is: First, the sliding window method is used to cut the original image from the upper left to the lower right with 10% overlap. Then, the dataset is expanded by rotating the image 90°, 180°, 270°, mirroring it, and making corresponding changes to the coordinates of the corresponding label anchor box, as well as adjusting the brightness.
6. The method for identifying typical objects in mining areas based on the improved YOLO11 algorithm according to claim 1, characterized in that: In step 2, the initial data set obtained in step 1 is annotated using the SAM semi-automatic annotation software.
7. The method for identifying typical landforms in mining areas based on the improved YOLO11 algorithm according to claim 1, characterized in that: In step 3, the improved YOLO11 algorithm is obtained by improving the YOLO11 algorithm. The specific improvements include: replacing the C3K2 residual block in the baseline model with a multi-branch convolution module, replacing the downsampling convolution of YOLO11 with an Adown module, and replacing the Detect detection head with an E_Detect structure.
8. The method for identifying typical landforms in mining areas based on the improved YOLO11 algorithm according to claim 1, characterized in that: In step 3, the data set obtained in step 2 is randomly divided into a training set: test set: validation set ratio of 7:2:1, and then the improved YOLO11 algorithm is trained using the data set.
Citation Information
Cited By
Training sample enhancement method and device for prospecting prediction of full convolutional neural network
CN121562695A