Semi-supervised tailing pond target detection method and device fused with self-supervised learning
Through the semi-supervised object detection method that integrates self-supervised learning, the problem of model confusion and data scarcity in tailings pond detection is solved, and efficient and accurate tailings pond detection is achieved.
Patent Information
- Application Number
- CN202510153878.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-30
AI Technical Summary
The existing tailings pond detection technology has problems such as model confusion, scarcity of data and high labeling costs, resulting in low detection efficiency and low accuracy.
The semi-supervised object detection method combined with self-supervised learning is adopted, and self-supervised learning and semi-supervised training is performed through high-resolution remote sensing images. The label-free data is used to improve model recognition efficiency, and the robustness of the model is improved through enhanced tasks.
It significantly reduces the use of labeled data, improves the accuracy and efficiency of tailings pond detection, and can efficiently identify tailings ponds in the context of large-scale remote sensing images, reducing the model's missed detection rate of large tailings ponds.
Smart Images

Figure CN120071184A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed by the present invention relates to remote sensing image information extraction. Specifically, it is a semi-supervised tailings pond object detection method and device that integrates self-supervised learning. Background Art
[0002] A tailings pond is a facility used to store tailings or other industrial waste generated during the processing of metal resources in the mining production process, usually located in flat areas or valley mouths. The waste in the tailings pond usually contains toxic, radioactive or other harmful substances, and most tailings ponds are distributed in mountainous areas with relatively weak supervision. According to the statistics of the National Mine Safety Supervision Bureau, as of 2024, the number of known tailings ponds in China is 4,919, and most of them exist in isolation. In recent years, tailings ponds have frequently triggered environmental safety accidents, causing serious casualties and irreversible environmental pollution. Therefore, improving the detection efficiency of tailings ponds is of great significance for environmental safety. Algorithms based on deep learning have made great breakthroughs in the field of tailings pond object detection. However, there are still the following problems in current tailings pond detection:
[0003] Except for the water body part, the tailings pond area usually presents as a light-colored or other monotonous-colored area, which is similar to the bare land, construction waste area or dry riverbed in the background, and is likely to cause model confusion. Some tailings ponds may be partially covered by vegetation, especially abandoned or long-unused tailings ponds, which will cause the model to be unable to fully capture the boundary of the tailings pond. The proportion of the part of a large tailings pond that is similar to the background is relatively large, resulting in difficulty for the detection model to identify the complete large tailings pond.
[0004] Currently, the data of tailings ponds is scarce, and the distribution of tailings ponds is relatively sparse. Manual annotation not only requires certain professional knowledge but also consumes a lot of time. The fully supervised method requires sufficient labeled data to optimize the model parameters. The lack of sufficient labeled data will cause the model to be unable to fully learn. The dependence of these methods on a large amount of labeled data limits the application of these methods in scenarios with high annotation costs or scarce data.
[0005] In view of the above problems, the present invention proposes a semi-supervised object detection framework that integrates self-supervised learning, and uses a large amount of unlabeled data to improve the model recognition efficiency. Summary of the Invention
[0006] The present invention aims to overcome the above-mentioned drawbacks of the prior art and provides a semi-supervised tailings pond object detection method and device that integrates self-supervised learning.
[0007] The first aspect of the present invention relates to a semi-supervised tailings pond object detection method that integrates self-supervised learning, specifically including the following steps:
[0008] Step 1: Select high-resolution remote sensing images in the area to be detected, and use the method of manual visual interpretation to mark no more than 50 tailing ponds, and the rest are used as unlabeled datasets.
[0009] Step 2: Use the unlabeled dataset in Step 1 as the input for the self-supervised learning stage. Each group of unlabeled data is respectively subjected to strong augmentation and weak augmentation of tailing pond data designed by the present invention. Here, different augmented views of the same picture are positive sample pairs with each other, and different samples are negative sample pairs with each other.
[0010] Step 3: Input the augmented data generated in Step 2 into the backbone network of the detector to extract global features, project and reduce the dimensions of the features, calculate the self-supervised loss, and train.
[0011] Step 4: Load the backbone network that has completed self-supervised training into the student model, and use a small amount of the marked data in Step 1 after mosaic augmentation to fine-tune the student model. At the same time, use the exponential moving average technique to update the parameters of the teacher network.
[0012] Step 5: Perform mosaic augmentation on the unlabeled dataset in Step 1, input it into the teacher network, predict the augmented data, calculate the pseudo-label score, and generate pseudo-labels.
[0013] Step 6: Based on the data after mosaic augmentation in Step 5, additionally add color jitter and noise block operations, input the secondarily augmented data into the student network, and calculate the semi-supervised loss between the model results and the corresponding pseudo-labels. The model stops training when it reaches the set number of iterations.
[0014] Further, the high-resolution remote sensing image in Step 1 is 5 meters or higher and has three bands of red, green, and blue.
[0015] Further, in Step 2, for the strong augmentation operation of tailing pond data, the image is first cropped to a size of 32×32. Then, it sequentially performs the following transformations: randomly flip horizontally with a probability of 0.55, adjust the brightness, contrast, and saturation of the image within the range of 0 to 0.45 with a probability of 0.85, and adjust the hue of the image within the range of 0 to 0.15 with a probability of 0.85. The weak augmentation operation includes random grayscale processing with a probability of 0.25 and random horizontal flipping with a probability of 0.55.
[0016] Further, the calculation formula of the self-supervised loss in Step 3 is the normalized temperature-scaled cross-entropy loss, aiming to pull positive sample pairs closer in space as much as possible while pulling negative sample pairs farther apart.
[0017] Further, in step 4, the parameters of the detector backbone network after self-supervised training are not loaded into the teacher network, but only into the student network. During the fine-tuning process, exponential moving average is used to initialize and update the parameters of the teacher network.
[0018] Further, the pseudo-label score in step 5 is composed of the objectness score and the classification score in the prediction result multiplied by a proportion. The pseudo-labels are selected based on the pseudo-label score. First, the candidate boxes are sorted according to the pseudo-label score, and the candidate box with the highest score is selected. Then, the intersection over union (IoU) between each box and it is calculated, and the top N optimal candidate boxes are selected as the pseudo-labels of the image.
[0019] Further, the secondary data augmentation operation in step 6 is to adjust the brightness, contrast, and saturation of the image within the range of 0 to 0.45 with a probability of 0.85, and add a 5×5 random color noise block with a probability of 0.75.
[0020] The second aspect of the present invention relates to a semi-supervised tailings pond object detection device integrating self-supervised learning, including a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the semi-supervised tailings pond object detection method integrating self-supervised learning of the present invention.
[0021] The third aspect of the present invention relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the semi-supervised tailings pond object detection method integrating self-supervised learning of the present invention.
[0022] The present invention has the following beneficial effects:
[0023] The present invention innovatively integrates self-supervised learning into the semi-supervised object detection framework. On the premise of ensuring high accuracy, it significantly reduces the use of labeled data. In the context of large-scale remote sensing images, only a small amount of labeled data is required to efficiently identify tailings ponds. In the self-supervised learning stage, by designing enhancement tasks to learn the consistency of features, the robustness of the model in dealing with data noise and distribution changes is effectively improved, and the omission rate of the model for large tailings ponds is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flowchart of the method of the present invention.
[0025] Figure 2 is an example diagram of the tailings pond label data of the present invention.
[0026] Figure 3 is a diagram of the self-supervised learning method of the present invention.
[0027] Figure 4 is a diagram of the semi-supervised training pseudo-label assignment method of the present invention.
[0028] Figure 5 (a) to Figure 5 (b) is the detection and prediction map of the tailings pond using the method of the embodiment of the present invention. Among them, Figure 5 (a) is the tailings pond label, Figure 5 (b) is the prediction result.
[0029] Figure 6 is the device diagram of the present invention. Detailed implementation manners
[0030] To further clarify the content of the present invention, the following combines Figure 1 the structural process and the specific embodiments of the present invention to further introduce the detailed description of the present invention. It should be understood that not all features of all actual implementation manners are the same as those of this embodiment, and the implementation details of the present invention may be changed according to specific conditions and objectives in actual engineering projects.
[0031] Embodiment 1
[0032] This embodiment provides a semi-supervised tailings pond object detection method integrating self-supervised learning:
[0033] Step 1: Use the tailings pond object detection dataset in Henan Province, China from 2016 to 2021 published on the China Science Data website, and additionally add remote sensing images in Sichuan Province from 2021 to 2023, totaling 1750 tailings pond data images and 2656 tailings pond object instances, with a resolution of 1 meter to 2.5 meters. Only 50 labeled data are used for training, 1600 of the remaining data are used as the unlabeled dataset, and 100 labeled data are used as the validation set. Typical tailings pond labels are as Figure 2 shown.
[0034] Step 2: Use the unlabeled dataset in Step 1 for self-supervised learning pre-training. The number of samples in each batch is 350.
[0035] Step 3: As Figure 3 shown, after data strong augmentation and data weak augmentation in Step 2, the number of samples finally input into the detector backbone network is 700. Different augmented views of each picture are positive samples with each other, and augmented views with other pictures are negative samples with each other. The detector backbone network extracts the global representations of these samples, reduces them from 1024 dimensions to 128 dimensions through the projection head, and then according to the formula:
[0036]
[0037] Calculate the self-supervised loss. Among them, N is the number of samples, 1 [k≠i]∈{0,1} is an indicator function, whose value is 1 if k = i, and τ is the temperature parameter. exp(similarity(z i ,z j )) / τ) is the similarity of the positive sample pair, and is the sum of the similarities between sample i and all other samples. Finally, taking the negative logarithm drives the similarity between positive sample pairs to increase, while further pulling the positive and negative sample pairs apart.
[0038] Step 4, load the backbone network that has completed self-supervised training into the backbone network of the student model. After performing mosaic data augmentation on the 50 labeled data in Step 1, input 16 pieces of data per batch into the student model for fine-tuning training, and train for 120 rounds while using the exponential moving average technique to update the teacher network parameters. The supervised loss formula in the fine-tuning stage of the present invention is defined as:
[0039]
[0040] CE represents the cross-entropy loss function, and CIoU is a variant of the IoU loss, which is used to further improve the accuracy of the model in the localization task. X and Y represent the output and the true label of the student model respectively, and cls, reg, and obj correspond to the classification score, regression score, and objectness score respectively.
[0041] Step 5, as Figure 5 (a)~ Figure 5 (b) shows, perform mosaic augmentation on the unlabeled dataset in Step 1, input it into the teacher network, make predictions on the augmented data, calculate the pseudo-label scores, and generate pseudo-labels. The specific screening process is as follows:
[0042] Let the input be the set of detection boxes and its corresponding score set S = {s 1 ,...,s N}. Given the suppression threshold N t and the maximum number of retained boxes The output is the screened detection box and its corresponding score S'. First, initialize the output set and the score set S' as {}. Then, sort the scores S of all detection boxes in descending order, and select the box b m , that is,
[0043]
[0044] Add the box b m to :
[0045]
[0046] Remove the bounding box b from the set of detection bounding boxes m and its corresponding score s m :
[0047]
[0048] For each remaining detection bounding box b i ∈ B, calculate its intersection over union (IoU) with the selected bounding box b m :
[0049]
[0050] If IoU(b m , b i ) ≥ N t , then remove this bounding box from and remove the corresponding score from S:
[0051]
[0052] This process is repeated continuously until is empty, or the number of selected bounding boxes reaches the maximum retention number Finally, return the filtered set of detection bounding boxes and their corresponding scores S'.
[0053] Step 6: Based on the data enhanced by mosaic in Step 5, additionally add color jittering and noise block operations, input the secondarily enhanced data into the student network, and calculate the semi-supervised loss between the model results and the corresponding pseudo-labels. Stop training when the semi-supervised training reaches 200 iterations, and save the model with the minimum loss.
[0054] To ensure the objectivity of the results, only unlabeled data is used for prediction. The model prediction results are as shown in Figure 5 (a)~ Figure 5 (b). Figure 5 (a) is the original tailings pond target box label of the input image of this implementation case. Figure 5 (b) is the semi-supervised object detection result of the corresponding implementation case.
[0055] The self-supervised learning technology and semi-supervised object detection technology of the present invention can significantly reduce the use of labeled data. In the context of large-scale remote sensing images, only a small amount of labeled data is required to efficiently identify tailings ponds. And in the self-supervised learning stage, by designing enhancement tasks to learn the consistency of features, the robustness of the model in dealing with data noise and distribution changes is effectively improved. It can quickly and effectively detect tailings ponds, which is of great significance for enhancing environmental safety.
[0056] Example 2
[0057] Reference Figure 6 , this embodiment relates to a semi-supervised tailings pond target detection device integrating self-supervised learning, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the semi-supervised tailings pond target detection method of integrating self-supervised learning in Embodiment 1.
[0058] Example 3
[0059] This embodiment relates to a computer-readable storage medium with a program stored thereon. When the program is executed by a processor, it implements the semi-supervised tailings pond target detection method of integrating self-supervised learning in Embodiment 1.
[0060] The above is only a description of the embodiments of the present invention. However, the protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art according to the concept of the present invention.
Claims
1. A semi-supervised tailings pond target detection method integrating self-supervised learning, characterized in that: The specific steps include: Step 1: select high-resolution remote sensing images in the area to be detected, and use manual visual interpretation methods to mark no more than 50 tailings ponds, and the rest are used as unlabeled data sets; Step 2: Use the unlabeled dataset in step 1 as input for the self-supervised learning phase; Each group of unlabeled data is subjected to strong and weak enhancement of the tailings pond data. Here, different enhanced views of the same image are considered as positive sample pairs, and different samples are considered as negative sample pairs. Step 3: Use the enhanced data generated in step 2 to input into the backbone network of the detector to extract global features, and then project the features to reduce the dimension and calculate the self-supervisory loss and train; Step 4: Load the self-supervised trained backbone network into the student model, fine-tune the student model using a small amount of labeled data after mosaic enhancement, and update the teacher network parameters using the exponential mean moving technique; Step 5: mosaic enhancement is performed on the unlabeled data set, which is input into the teacher network, and the enhanced data is predicted, the pseudo label score is calculated, and the pseudo label is generated; Step 6: Based on the mosaic-enhanced data, additional color jitter and noise block operations are added to perform secondary data enhancement; The secondary augmented data is input into the student network, and the model results and the corresponding pseudo labels are used to calculate the semi-supervised loss; The model training stops when the set number of iterations is reached.
2. The semi-supervised tailings pond target detection method integrating self-supervised learning according to claim 1, characterized in that: The high-resolution remote sensing image described in step 1 is higher than or equal to 5 meters and has three bands: red, green, and blue.
3. The semi-supervised tailings pond target detection method integrating self-supervised learning according to claim 1, characterized in that: The strong enhancement operation of the tailings pond data described in step 2 includes: first cropping the image to a size of 32×32; then, it performs the following transformations in sequence: randomly flipping horizontally with a probability of 0.55, adjusting the brightness, contrast and saturation of the image in the range of 0 to 0.45 with a probability of 0.85, and adjusting the hue of the image in the range of 0 to 0.15 with a probability of 0.85; the weak enhancement operation includes random grayscale processing with a probability of 0.25 and random horizontal flipping with a probability of 0.
55.
4. The semi-supervised tailings pond target detection method integrating self-supervised learning according to claim 1, characterized in that: The self-supervised loss calculation formula described in step 3 is the normalized temperature scale cross entropy loss, which brings positive sample pairs closer together in space while moving negative sample pairs further apart.
5. The semi-supervised tailings pond target detection method integrating self-supervised learning according to claim 1, characterized in that: In step 4, the parameters of the detector backbone network that completes the self-supervised training are not loaded into the teacher network, but only loaded into the student network. During the fine-tuning process, the exponential average move is used to complete the initialization and update of the teacher network parameters.
6. The semi-supervised tailings pond target detection method integrating self-supervised learning according to claim 1 is characterized in that: The pseudo-label score described in step 5 is composed of the object score and the classification score in the prediction result multiplied by their weights. The pseudo-label selection is based on the pseudo-label score. First, the candidate boxes with the highest scores are selected by sorting according to the pseudo-label scores. Then, the intersection and union ratio of each box with the pseudo-label scores is calculated, and the top N optimal candidate boxes are selected as the pseudo-labels of the image.
7. The semi-supervised tailings pond target detection method integrating self-supervised learning according to claim 1 is characterized in that: The secondary data enhancement operation described in step 6 is to adjust the brightness, contrast and saturation of the image in the range of 0 to 0.45 with a probability of 0.85, and add a 5×5 random color noise block with a probability of 0.
75.
8. A semi-supervised tailings pond target detection device integrating self-supervised learning, characterized in that: It comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the semi-supervised tailings dam target detection method integrating self-supervised learning as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the semi-supervised tailings pond target detection method integrating self-supervised learning as described in any one of claims 1-7 is implemented.