Intelligent extraction method for surface crack of coal mining subsidence area based on improved Transform model

By improving the Transformer model, combined with adaptive multi-scale patch mapping, dual attention fusion and breakpoint detection, the problem of insufficient noise interference and long-distance dependency identification in crack extraction in mining areas is solved, and high-precision and complete crack extraction is achieved, adapting to complex environments and conditions, and supporting mining area safety management.

CN120375231APending Publication Date: 2025-07-25LIAONING TECHNICAL UNIVERSITY
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510444082.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art has problems such as low efficiency, susceptible to noise interference, difficulty in identifying subtle or blocked cracks, and insufficient long-distance dependency recognition capabilities in the extraction of ground fractures in mining areas. Especially in the analysis of drone images, it is difficult to achieve high-precision and complete extraction.

Method used

Adaptive multi-scale patch mapping layer, dual attention fusion mechanism, dynamic sparse connection full connection layer and crack breakpoint detection module are adopted, combined with multi-scale pyramid decoding strategy and space-channel attention bottleneck mechanism, an improved Transformer model is built and improved, the convolution kernel size is dynamically adjusted, long-distance dependency modeling is strengthened, local features are accurately captured and incomplete cracks are repaired.

Benefits of technology

It realizes high-precision and complete crack extraction in complex mining areas, improves the generalization ability and computing efficiency of the model, adapts to different mining areas and shooting conditions, and provides reliable data support for mining area safety management and geological disaster prevention and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375231A_ABST
    Figure CN120375231A_ABST
Patent Text Reader

Abstract

The invention discloses a coal mining subsidence area surface crack intelligent extraction method based on an improved Transform model, and belongs to the technical field of remote sensing image processing. Firstly, an unmanned aerial vehicle carrying a high-resolution optical camera is used for collecting images, the image overlapping rate of 70%-80% is guaranteed, and a training data set is constructed through professional labeling, cutting screening and data enhancement. The encoder of the innovative model is very distinctive, and the adaptive multi-scale patch mapping layer can dynamically adjust the patch size according to the local complexity of the image and efficiently extract features; double-attention fusion is combined with optimization position coding, and long-distance dependency capture is enhanced; and the calculation amount and the overfitting are reduced by the dynamic sparse connection full-connection layer. Residual attention enhancement pyramid pooling and a space-channel attention bottleneck mechanism are adopted, key features are highlighted, and noise is suppressed; and a breakpoint detection and connection rule determination module is utilized to realize complete restoration of the ground fracture. After the data set is used for training a model, deployment is carried out, and through preprocessing, encoding and decoding and post-processing, ground fracture information can be accurately obtained, and a data foundation is built for mining area safety management and geological disaster prevention and control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and geological monitoring, and particularly relates to an intelligent extraction method for surface cracks in coal mining subsidence areas based on an improved Transformer model, which is particularly suitable for the scenario of analyzing ground fissures by using images collected by drones. Background Art

[0002] In mining activities in mining areas, the occurrence of ground fissures is extremely common. The causes of these ground fissures are complex and mainly stem from changes in formation stress and damage to rock structures caused by underground mining. They pose a serious threat to the infrastructure in mining areas, such as factories and transportation tracks, as well as the surrounding ecological environment, and may cause problems such as building collapses and surface soil and water loss. Therefore, accurately obtaining information such as the location, shape, and length of ground fissures is crucial for the safe operation of mining areas and disaster prevention.

[0003] Currently, there are significant defects in the methods for extracting ground fissures in mining areas: manual measurement is inefficient, and traditional image processing methods (such as threshold segmentation and edge detection) are affected by light and noise interference and are difficult to identify fine or occluded fissures; machine learning relies on artificial feature design and has weak generalization ability; although deep learning can extract features, it is difficult to capture long-distance dependencies, resulting in incomplete crack extraction and misjudgment points. Especially in images collected by drones, due to the influence of resolution, shooting angle, and terrain, existing methods are difficult to achieve high-precision and complete extraction.

[0004] In existing technologies, there are few studies on the complete extraction of cracks in mining areas. Scholars at home and abroad mainly use three types of methods to extract ground fissures in mining areas. The first type is traditional image processing methods, such as the fast search algorithm for extreme values of projection histograms based on wavelet transform by J. Zhang, etc., which is suitable for situations where the ratio of positive and negative samples varies greatly, but cannot cope with complex environments such as high noise, light, and shadows, and is not sensitive to small cracks; the second type is based on machine learning algorithms. Kaseko, etc. compared traditional classifiers such as Bayesian classifiers and K-nearest neighbor methods with feedforward neural network classifiers and two-stage piecewise linear neural network classifiers. The results show that the neural network classifier performs significantly better than traditional classifiers. The third type is based on deep learning methods. Liu Yuxiang, etc. constructed a hybrid model that integrates residual connections and max-pooling layers based on the residual learning mechanism of ResNet and the U-shaped encoder architecture of U-net to achieve feature learning from local to global, but there is a problem of feature information attenuation; W. Wang, etc. proposed an improved method that integrates an attention mechanism and a multi-scale input strategy to enhance the model's ability to capture global context information, but it also leads to an increase in computational complexity and the number of parameters, affecting the training and inference efficiency of the model.

[0005] In summary, previous studies have mainly focused on traditional image processing methods, linear models such as U-net and YOLO series based on deep learning, and machine learning algorithms. These methods have achieved certain results in the field of crack detection. However, traditional methods are vulnerable to noise interference in complex backgrounds, and existing deep learning models still have limitations in completely extracting cracks, especially fine cracks. Especially in a complex environment such as a mining area, where the crack background is complex and there is a lot of noise, existing methods are difficult to meet the requirements of high-precision and complete information extraction at the same time.

[0006] To solve this problem, this patent proposes a method for extracting mining area cracks that integrates dual attention and an adaptive patch mapping layer. This method uses an adaptive multi-scale patch mapping layer to dynamically adjust the size of the convolutional kernel to adapt to image regions of different complexities, thereby accurately capturing local features. At the same time, a dual attention fusion mechanism is introduced, combining spatial and semantic attention, which enhances the ability to model the global features and long-range dependencies of cracks. In addition, the risk of overfitting is reduced by a dynamically sparse connected fully connected layer, improving the generalization ability of the model. In the decoder, a residual attention enhanced pyramid pooling and a spatial-channel attention bottleneck mechanism are innovatively adopted to effectively restore crack details and suppress background noise. Finally, through a crack breakpoint and discontinuity point detection module, incomplete cracks are accurately repaired to achieve high-precision and complete crack extraction. Summary of the Invention

[0007] The purpose of this part is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this part, as well as in the abstract and title of the specification of this application, to avoid obscuring the purpose of this part, the abstract, and the title. However, such simplifications or omissions shall not be used to limit the scope of the present invention.

[0008] In view of the above problems existing in the prior art, the present invention is proposed.

[0009] Therefore, the present invention aims to provide an intelligent method for extracting surface cracks in coal mining subsidence areas based on an improved Transformer model. Through innovative model design and processing procedures, it overcomes the deficiencies of the prior art in terms of accuracy and integrity, realizes accurate and comprehensive extraction of ground cracks in mining areas, and provides reliable data support for mining area safety management and geological disaster prevention and control. Compared with the prior art, the core innovation point of the present invention is to propose a high-precision method for extracting ground cracks in mining areas that integrates adaptive multi-scale patch mapping, dual attention fusion mechanism, and crack breakpoint detection and repair. This method can effectively solve the problems of incomplete crack extraction, low accuracy, and insufficient ability to identify long-range dependencies in the complex mining area environment, and provides a new technical solution for mining area safety management and geological disaster prevention and control.

[0010] Preferably, a drone equipped with a high-resolution optical camera is used to fly at low altitude and at multiple angles along a predetermined route in the target mining area. The flight altitude is flexibly adjusted according to the terrain and the characteristics of ground fissures, and is maintained at 50-100 meters to ensure obtaining images with high clarity and multi-dimensional information. The overlap rate of adjacent images reaches 70%-80% to ensure the integrity of image stitching and subsequent analysis.

[0011] Furthermore, with the help of advanced image annotation software such as labelme, the collected images are finely annotated. For the fine branches and blurred areas of ground fissures, they are repeatedly confirmed by combining geological knowledge and image enhancement techniques to ensure that the outlines of ground fissures are accurately annotated. The foreground (ground fissure) is annotated as 1 and the background is annotated as 0. Using the OpenCV library of Python, according to the distribution of ground fissures in the image, the original image is cropped into sub-images of a fixed size, ensuring that each sub-image contains complete or partial ground fissure features, and the ground fissure is located near the center area of the image to reduce background interference.

[0012] Furthermore, according to criteria such as the proportion of ground fissures in the cropped images (the set proportion threshold is 10%-50%) and clarity (measured by indicators such as image entropy and gradient amplitude), high-quality ground fissure cropped images with obvious features are selected, and images with large noise and blurred ground fissure features are removed to improve the quality of the dataset. Using the TorchVision library, diverse data augmentation is implemented on the selected images. This includes random rotation (angle range [-30°, 30°]), flipping (horizontal and vertical flipping probabilities are each 0.5), scaling (scale range [0.8, 1.2]), and adjustment of brightness and contrast (adjustment factor range [0.7, 1.3]), expanding the original dataset by 5-10 times to enrich data diversity and improve the generalization ability of the model.

[0013] Furthermore, a unique adaptive multi-scale patch mapping layer is constructed under the PyTorch framework. By designing a convolutional kernel with dynamically adjustable parameters, the patch size is automatically adapted according to the complexity of the local area of the image (judged by calculating local entropy and variance). Calculate the local entropy H(X) and the pixel value variance σ 2 The formulas are as follows:

[0014]

[0015] Among them, P(x i) is the probability distribution of pixel gray values within the window, N is the number of window pixels, and μ is the window mean. In areas with rich texture and many details of ground fissures, small-sized patches of 3×3 pixels are used to finely extract features; for ordinary areas, a 5×5 convolutional kernel is used to balance accuracy and computational efficiency; in relatively smooth areas, the patch size is increased to 7×7 pixels to improve computational efficiency and comprehensively and efficiently capture local features of the image. During the forward propagation of this layer, the convolutional feature maps of 3×3, 5×5, and 7×7 are calculated in parallel, and the optimal features are selected and fused according to local complexity, and finally the optimal feature representation adapted to the characteristics of different regions is generated. This method takes into account both the retention of fissure details and the optimization of computational efficiency, effectively improving the integrity and accuracy of fissure extraction in mining areas and making it more robust in UAV image analysis.

[0016] Furthermore, a dual-attention fusion strategy is innovatively introduced in the Transformer encoding block. One attention branch focuses on the spatial position relationship of ground fissures, highlighting the ground fissure area, suppressing background interference, and enhancing the understanding of the long-distance continuity of fissures; the other branch focuses on semantic feature correlation rather than background noise, extracts channel-level features using global average pooling (GAP), and calculates attention weights. At the same time, the position encoding is optimized, and high-order sine and cosine functions are used to generate position encoding vectors. The formula is:

[0017]

[0018] (pos is the position, i is the dimension index, d model is the model dimension), which not only reflects the absolute position information but also can accurately capture relative position changes, strengthening the modeling ability for complex ground fissure trends and long-distance dependence relationships.

[0019] Furthermore, aiming at the problems of many parameters and easy overfitting in traditional fully connected layers, a dynamically sparse connected fully connected layer is designed. During training, the connection weights are dynamically adjusted according to the neuron activation frequency and the correlation with ground fissure features. For connections weakly associated with ground fissure features, the weights are reduced or the connections are disconnected to reduce the number of parameters and computational amount, and improve the model's learning ability and generalization for key features.

[0020] Furthermore, multiple above-mentioned Transformer encoding blocks are connected in series in sequence. Each encoding block contains an adaptive multi-scale patch mapping layer, a dual-attention fusion and position encoding optimization mechanism module, and a dynamically sparse connected fully connected layer. The output of the previous encoding block is used as the input of the next encoding block to gradually extract and fuse image features and build a powerful model encoder.

[0021] Furthermore, a multi-scale pyramid cascaded decoding strategy of convolutional neural network is adopted. The deconvolution operation is used to upsample the feature maps of different scales output by the model encoder, and the upsample kernel size and stride are reasonably set according to the scale of the feature maps. The upsampled feature maps are concatenated with the low-resolution feature maps of the corresponding levels of the encoder, and features are further extracted through a series of convolutional operations. With the help of skip connections, the detailed information in the encoder is retained, and the complete details of the ground fissures are gradually restored.

[0022] Furthermore, a residual attention enhanced pyramid pooling mechanism is innovatively proposed in the decoder. On the basis of traditional pyramid pooling, residual connections are added to prevent feature loss, and an attention mechanism is introduced to weight the features after pooling at different scales. The global information of the feature map is obtained through global average pooling and global max pooling, and the attention weights are generated through a fully connected layer and an activation function, and then multiplied by the pooled feature map to highlight the key features of the ground fissures at different scales, suppress background noise, and improve the decoding accuracy.

[0023] Furthermore, a spatial-channel attention bottleneck mechanism is constructed. Global average pooling and global max pooling are respectively performed on the input feature map in the spatial and channel dimensions to extract key information from different angles, and spatial and channel attention feature maps are obtained. After being respectively reduced and increased in dimension (the reduction dimension is set to 64) through a fully connected layer, and through non-linear transformation by ReLU and Sigmoid activation functions, spatial and channel attention weights are generated and multiplied by the original input feature map to highlight the key features of the ground fissures in the spatial and channel dimensions and suppress background interference.

[0024] Furthermore, a professional crack breakpoint and discontinuity detection module is added to the model decoder. Based on the Histogram of Oriented Gradients (HOG) algorithm, combined with the feature map extracted by the model, the directional gradient information of each position of the ground fissure is calculated to judge the directionality of the crack. By comparing the direction consistency and gradient amplitude change of adjacent regions, potential breakpoints and discontinuities are identified. For suspicious points, the local region curvature information and texture features are used for further confirmation to improve the detection accuracy.

[0025] Furthermore, according to the detected breakpoints and discontinuities, the crack connection rule determination module in the model formulates connection rules based on the overall trend of the ground fissure, the distance and angular relationship between adjacent crack segments. If the included angle between the directions of adjacent crack segments is within the threshold of [-30°, 30°], and the distance is less than the set maximum connection distance (determined according to the average width of the ground fissure and the image resolution), it is determined that the two segments should be connected to achieve precise repair and complete extraction of the incomplete crack.

[0026] Furthermore, the constructed ground fissure training dataset is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1. A training framework is built using PyTorch, the AdamW optimizer is adopted, the initial learning rate is set to 1e-4, and the cosine annealing learning rate adjustment strategy is used to dynamically adjust the learning rate. The loss function adopts a hybrid loss function combining cross-entropy loss and Dice loss to balance the class imbalance problem and improve the model's recognition ability for small targets (ground fissures). During the training process, the training set images are input into the model encoder to generate multi-scale feature maps, and then the probability feature maps are output through the model decoder, normalized by a preset activation function (such as Sigmoid) to obtain normalized feature maps, and binary feature maps are obtained through threshold division (the threshold is set to 0.5). Continuously optimize the model parameters to make the loss value and evaluation metrics (such as intersection over union, accuracy, recall) of the model on the validation set reach the optimal.

[0027] Furthermore, the trained ground fissure segmentation and recognition model is deployed to a high-performance server or an edge computing device. When new drone images of the mining area are input, the model automatically preprocesses the images (normalization, size adjustment, etc.), and then passes through the encoder and decoder in sequence to output the ground fissure information in the images, presented in the form of a binary image, where white pixels represent ground fissures and black pixels represent the background. Post-process the output results, such as morphological opening and closing operations, to remove noise points and connect small cracks, and finally obtain a complete and accurate ground fissure extraction result, providing reliable data for ground fissure monitoring and analysis in the mining area.

[0028] The present invention also provides a usage method of an intelligent extraction method for surface cracks in a coal mining subsidence area based on an improved Transformer model, which includes the following steps:

[0029] S1. Construct a ground fissure training dataset:

[0030] 1. Image acquisition: Use a drone equipped with a high-resolution optical camera to fly at low altitude from multiple angles, and flexibly adjust the height (50 - 100 meters) according to the terrain and ground fissure characteristics to ensure that the overlap rate of adjacent images reaches 70% - 80%.

[0031] 2. Construct label images: Organize geological personnel and image processing experts, and with the help of professional annotation software, combine geological knowledge and image enhancement technology to finely annotate the ground fissure contours, with the foreground being 1 and the background being 0.

[0032] 3. Crop images: Use the OpenCV library in Python to crop the original images into sub-images of a fixed size (such as 256×256 pixels), ensuring that the ground fissures are near the central area.

[0033] 4. Screen images: Select high-quality images based on the proportion of ground fissures (10%-50%) and clarity (indicators such as image entropy and gradient magnitude), and remove images with high noise and blurred features.

[0034] 5. Data augmentation: Use the TorchVision library to perform random rotation ([-30°, 30°]), flipping (horizontal and vertical probabilities are both 0.5), scaling ([0.8, 1.2]), and adjustment of brightness and contrast ([0.7, 1.3]) on the screened images to expand the dataset by 5-10 times.

[0035] S2. Build an innovative segmentation and recognition model:

[0036] 1. Generate the model encoder

[0037] (1) Adaptive multi-scale patch mapping layer: Under the PyTorch framework, dynamically adjust the convolution kernel size according to the local complexity of the image (entropy, variance) to adapt to the patch size and efficiently capture local features.

[0038] (2) Dual attention fusion and position encoding optimization mechanism: In the Transformer encoding block, adopt dual attention branches to respectively focus on spatial positions and semantic features, and at the same time optimize the position encoding to capture relative position changes.

[0039] (3) Dynamically sparse connected fully connected layer: During training, dynamically adjust the connection weights according to the neuron activation frequency and the correlation with ground fissure features to reduce the risk of overfitting.

[0040] (4) Serial connection of Transformer encoding blocks: Serialize multiple Transformer encoding blocks containing the above modules to gradually extract and fuse image features.

[0041] 2. Build the model decoder

[0042] (1) Multi-scale pyramid cascaded decoding strategy: Use deconvolution operations to upsample the feature maps output by the encoder, splice them with the corresponding low-resolution feature maps, and restore the ground fissure details through convolution operations and skip connections.

[0043] (2) Residual attention enhanced pyramid pooling: On the basis of traditional pyramid pooling, add residual connections and introduce an attention mechanism to highlight key features at different scales.

[0044] (3) Spatial-channel attention bottleneck mechanism: Perform global pooling on the input feature maps in the spatial and channel dimensions respectively to generate attention weights and highlight the key features of ground fissures.

[0045] (4) Crack breakpoint and discontinuity detection module: Based on the HOG algorithm, calculate the directional gradient by combining with the feature map, judge potential breakpoints and discontinuities by comparing adjacent regions, and further confirm using local curvature and texture.

[0046] (5) Crack connection rule determination module: According to the detected breakpoints, formulate connection rules based on the overall trend of the ground fissure, the distance and angle relationship between adjacent crack segments, and repair incomplete cracks.

[0047] S3. Model training and application:

[0048] 1. Model training: Divide the training set, validation set, and test set in a ratio of 8:1:1. Use the PyTorch framework, adopt the AdamW optimizer, and combine the cosine annealing learning rate adjustment strategy. Optimize the model with a mixed loss function of cross-entropy loss and Dice loss to make the model achieve the optimal performance on the validation set.

[0049] 2. Model application: Deploy the trained model to a high-performance server or edge computing device, preprocess the newly input drone images of the mining area, output the ground fissure information through the encoder and decoder in sequence, and then obtain the complete and accurate ground fissure extraction result through post-processing.

[0050] The beneficial effects of the present invention based on an intelligent extraction method for surface cracks in coal mining subsidence areas based on an improved Transformer model are as follows:

[0051] 1. High-precision feature extraction: The adaptive multi-scale patch mapping layer accurately captures the local features of the image. The dual-attention fusion and position encoding optimization mechanism deeply understands the global features and long-distance dependence relationship of the ground fissure. The dynamic sparse connection fully connected layer strengthens the learning of key features. The three work together to enable the model to extract the ground fissure features with high precision in a complex mining area environment, significantly improving the extraction accuracy.

[0052] 2. Complete crack restoration: The multi-scale pyramid cascaded decoding strategy, residual attention enhanced pyramid pooling, and spatial-channel attention bottleneck mechanism work together to effectively restore the details and edge information of the ground fissure. The crack breakpoint and discontinuity detection module and the crack connection rule determination module accurately detect and repair breakpoints, realizing the complete extraction of the ground fissure, reducing missed and false judgments, and providing comprehensive data for the analysis of ground fissures in the mining area.

[0053] 3. Strong generalization and adaptability: Rich data augmentation operations combined with an innovative model structure enable the model to learn a general ground fissure feature representation. The dynamic sparse connection fully connected layer reduces the risk of overfitting and improves the generalization ability. The model can adapt to drone images in different mining areas and different shooting conditions, and stably and accurately extract ground fissure information in new scenarios.

[0054] 4. Efficient Computation and Applications: The parallel computing feature of the Transformer architecture combined with the optimized design of the model significantly improves the efficiency of model training and inference. With the support of high-performance hardware, it can quickly process a large number of drone images, meet the real-time or near-real-time monitoring requirements of mining areas, and provide timely decision-making basis for mining area safety management. Brief Description of the Drawings

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0056] Figure 1 is the construction flow chart of the ground fissure training data set of the present invention;

[0057] Figure 2 is the SWOT structure schematic diagram of the hierarchical Transformer architecture network model in the present invention;

[0058] Figure 3 is the schematic diagram of the working principle of the module in the present invention;

[0059] Figure 4 is the overall architecture diagram of the present invention;

[0060] The realization of the object, functional features and advantages of the present invention will be further described with reference to the embodiments and the drawings. Detailed Embodiments

[0061] The following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the embodiments of the present invention. It should be understood that the specific embodiments described here are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0062] To make the above objects, features and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the embodiments of the specification.

[0063] Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0064] Second, the "one embodiment" or "embodiment" referred to herein means a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or selectively exclusive embodiments from other embodiments.

[0065] As a preferred solution of the present invention, the present invention provides a method for extracting mining area fractures by integrating dual attention and an adaptive patch mapping layer.

[0066] As a preferred solution of the present invention, for image acquisition of the present invention: The DJI Phantom 4 drone is selected, the camera resolution is set to 4000×3000 pixels, and the color mode is RGB optical image. According to the terrain and landform of the mining area and the prediction of possible ground fracture areas based on previous geological surveys, multiple "well"-shaped flight routes are planned to ensure full coverage of the target area. During the flight, the flight height is adjusted in real time between 50-100 meters according to the terrain undulation. Using the drone flight control software, the adjacent image overlap rate is set to 75%. The images are taken in the morning when the weather is clear and the light is uniform, and a total of 5000 high-quality drone images are obtained.

[0067] As a preferred solution of the present invention, for label image construction of the present invention: With the help of the labelme software, the collected images are labeled. For complex ground fracture areas, auxiliary means such as image magnification and enhancement are used, and combined with geological structure knowledge, the labeling results are determined after full discussion by the team. Finally, 1000 label images are accurately labeled, with the foreground (ground fracture) labeled as 1 and the background labeled as 0.

[0068] As a preferred solution of the present invention, for image cropping of the present invention: A special cropping script is written using the OpenCV library of Python. The original images and label images are uniformly cropped to a size of 256×256 pixels, ensuring that the central area of each cropped image contains ground fracture features. After the cropping operation, a total of 2000 cropped image pairs (ground fracture images and corresponding label images) are obtained.

[0069] As a preferred solution of the present invention, by calculating the proportion of ground fractures in the cropped images (using the label images to count the ratio of the number of ground fracture pixels to the total number of pixels) and the clarity (calculating the image entropy and the average gradient amplitude), the ground fracture proportion threshold is set to 18%-35%, and the clarity threshold is that the image entropy is greater than 3.2 and the average gradient amplitude is greater than 8. According to these criteria, 1500 high-quality cropped image pairs are selected, and low-quality images are removed, effectively improving the quality of the dataset.

[0070] As a preferred embodiment of the present invention, data augmentation of the present invention: Use the TorchVision library to perform data augmentation on the 1500 selected image pairs. The specific settings are that the random rotation angle range is [-25°, 25°], the horizontal and vertical flipping probabilities are both 0.5, the scaling ratio range is [0.85, 1.15], and the brightness and contrast adjustment factor range is [0.8, 1.2]. After data augmentation, the dataset is expanded to 4000 image pairs, greatly enriching the data diversity.

[0071] As a preferred embodiment of the present invention, the implementation of the adaptive multi-scale patch mapping layer of the present invention: Define the adaptive multi-scale patch mapping layer class in the PyTorch environment. By calculating the entropy and variance within the 5×5 neighborhood of the local region of the image, when the entropy is greater than 3.5 and the variance is greater than 25, the region is determined to be a complex region, and a 3×3 convolutional kernel is used for patch mapping; otherwise, a 7×7 convolutional kernel is used. Each patch is mapped to a vector with a dimension of 128 through a linear layer, thereby realizing the adaptive extraction of local features.

[0072] As a preferred embodiment of the present invention, the implementation of the dual attention fusion and position encoding optimization mechanism of the present invention: Construct a spatial attention branch and a semantic attention branch. The spatial attention branch obtains the spatial position attention weight by calculating the dot product of the query vector and the key vector; the semantic attention branch obtains the semantic attention weight by performing semantic analysis on the feature map. In terms of position encoding, the formula is:

[0073]

[0074] (where pos is the position, i is the dimension index, and d model is the model dimension) to generate a position encoding vector, add it to the input feature vector, and then input them into the two attention branches respectively. Finally, the output results of the two branches are added element-wise for fusion.

[0075] As a preferred embodiment of the present invention, the present invention constructs a dynamic sparse connection fully connected layer class. During the training process, every 100 batches are trained, and the activation frequency and the correlation score with the ground fissure feature of each neuron are calculated (by calculating the correlation between the neuron output and the ground fissure label). For neuron connections with an activation frequency lower than 0.12 and a correlation score lower than 0.35, set their weights to 0 to achieve dynamic sparsification of the connections, and re-evaluate and adjust the weights every 500 batches.

[0076] As a preferred solution of the present invention, the Transformer encoding blocks of the present invention are cascaded: five Transformer encoding blocks are constructed, and the adaptive multi-scale patch mapping layer, the dual attention fusion and position encoding optimization mechanism module, and the dynamic sparse connection fully connected layer are combined in each encoding block in sequence. These five encoding blocks are cascaded in sequence to construct the model encoder.

[0077] As a preferred solution of the present invention, the multi-scale pyramid cascaded decoding strategy of the present invention is implemented as follows: the feature map output by the model encoder is upsampled using a transposed convolution layer. For the feature map with the smallest scale, a 4×4 transposed convolution kernel with a stride of 2 is used for upsampling; for the feature map with a larger scale, an appropriate transposed convolution kernel and stride are selected according to the size of the feature map. The upsampled feature map is concatenated with the low-resolution feature map at the corresponding level of the encoder, and then three convolution operations are performed using a 3×3 convolution kernel. The features in the encoder are introduced into the decoder using skip connections, thereby restoring the detailed information of the ground fissure.

[0078] As a preferred solution of the present invention, the implementation of the residual attention enhanced pyramid pooling of the present invention is as follows: maximum pooling and average pooling layers are constructed, and the pooling kernel sizes are set to 2×2, 3×3, and 4×4 respectively. Pooling operations of different scales are performed on the input feature map, and the pooling results are added to the original input feature map through residual connections. Then, the attention mechanism is used to weight the features after pooling at different scales. The global information of the feature map is obtained through global average pooling and global maximum pooling, and the attention weights are generated through a fully connected layer (dimensionality reduction to 64 dimensions) and ReLU and Sigmoid activation functions, and then multiplied by the pooled feature map to highlight the key features.

[0079] As a preferred solution of the present invention, the implementation of the spatial-channel attention bottleneck mechanism of the present invention is as follows: global average pooling and global maximum pooling are performed on the input feature map in the spatial and channel dimensions respectively to obtain the spatial and channel attention feature maps. They are respectively reduced to 64 dimensions through a fully connected layer, and then upsampled to the dimension of the original feature map. The spatial and channel attention weights are generated through ReLU and Sigmoid activation functions, and multiplied by the original input feature map to highlight the key features of the ground fissure.

[0080] As a preferred solution of the present invention, the implementation of the crack breakpoint and discontinuity point detection module of the present invention is as follows: using the Histogram of Oriented Gradients (HOG) algorithm, the directional gradient information is calculated in units of 8×8 pixel blocks on the feature map extracted by the model. By comparing the directional consistency of adjacent blocks (the directional angle greater than 40° is regarded as inconsistent) and the change in gradient magnitude (the change exceeding 12 is regarded as abnormal), suspicious points are initially identified. For the suspicious points, the curvature information and texture features of the local area are further used for secondary confirmation to improve the accuracy of detection.

[0081] As a preferred embodiment of the present invention, the crack connection rule determination module of the present invention realizes: formulating connection rules based on the overall trend of the ground fissures, the distance and angular relationship between adjacent crack segments. If the included angle between the directions of adjacent crack segments is within the threshold range of [-30°, 30°], and the distance between them is less than the maximum connection distance determined according to the average width of the ground fissures in this mining area and the image resolution, it is determined that these two segments should be connected. In this way, the repair of incomplete cracks is realized, and complete ground fissure information is obtained.

[0082] As a preferred embodiment of the present invention, the model training of the present invention: The constructed dataset of 4000 image pairs is divided into a training set, a validation set and a test set according to the ratio of 8:1:1. Use PyTorch to build a training framework, select the AdamW optimizer, set the initial learning rate to 1e-4, and use the cosine annealing learning rate adjustment strategy to dynamically adjust the learning rate. The loss function adopts a hybrid loss function that combines cross-entropy loss and Dice loss to balance the class imbalance problem and improve the model's recognition ability for small targets (ground fissures). During the training process, the training set images are input into the model encoder to generate multi-scale feature maps, and then the probability feature maps are output through the model decoder, normalized by a preset Sigmoid activation function to obtain a normalized feature map, and a binary feature map is obtained through threshold division (the threshold is set to 0.5). Continuously optimize the model parameters. After multiple rounds of training, the loss value of the model on the validation set steadily decreases, and evaluation indicators such as intersection over union, accuracy, and recall reach a relatively high level.

[0083] As a preferred embodiment of the present invention, the model application of the present invention: Deploy the trained ground fissure segmentation and recognition model to a high-performance server equipped with an NVIDIA RTX 3090 GPU. When new UAV images of this mining area are input, the model first automatically performs preprocessing operations such as normalizing and resizing the images, and then passes through the encoder and decoder in sequence to output the ground fissure information in the images, presented in the form of a binary image, where white pixels represent ground fissures and black pixels represent the background. Perform post-processing operations such as morphological opening and closing on the output results to remove noise points and connect small cracks, and finally obtain a complete and accurate ground fissure extraction result. Verified by actual application, the model can accurately identify ground fissures of various forms in the mining area, providing reliable data support for the ground fissure monitoring and analysis in the mining area, and effectively assisting the safety production management and geological disaster prevention and control work in the mining area.

[0084] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An intelligent extraction method for surface cracks in coal mining subsidence areas based on an improved Transformer model, characterized in that, It includes the following steps: (1) Construct a ground fissure training dataset: Using a DJI Phantom 4 drone, fly at low altitude along a carefully planned "grid" flight path for shooting. The flight altitude is dynamically adjusted between 50 - 100 meters according to the complex terrain of the mining area to ensure that the overlap rate of adjacent images is stable at 75%. A total of 5000 high - definition optical remote sensing images with a resolution of 4000×3000 pixels are obtained; With the help of the labelme software, the collected images are finely annotated. A total of 1000 accurate label images are annotated, specifying that the ground fissure is the foreground (annotated as 1) and the background is 0; Using the OpenCV library in Python, the original images are uniformly cropped into sub - images of 256×256 pixels, ensuring that the ground fissure is near the center area of the image. After this operation, 2000 pairs of cropped images are obtained; According to indicators such as the proportion of the ground fissure in the cropped image (actually set to 18% - 35%) and clarity (image entropy greater than 3.2, average gradient magnitude greater than 8), 1500 pairs of high - quality cropped images are selected; Using the TorchVision library, data augmentation is performed on the selected images. The randomly rotated angle is set in the range of [-20°, 20°], the horizontal and vertical flip probabilities are both 0.5, the scaling ratio is controlled in the range of [0.88, 1.12], and the brightness and contrast adjustment factors are in the range of [0.85, 1.15], expanding the dataset to 4000 image pairs; (2) Construct an innovative segmentation and recognition model: 1) Generate the model encoder: Adaptive multi - scale patch mapping layer: An adaptive multi - scale patch mapping layer is constructed under the PyTorch framework. After a large number of experiments, for the images in this mining area, when the entropy in the 5×5 neighborhood of the local image area is greater than 3.5 and the variance is greater than 25, a 3×3 convolution kernel is used for patch mapping to finely extract features; In a relatively smooth area, a 7×7 convolution kernel is used to improve the calculation efficiency and achieve efficient capture of local features in different regions of the image; calculate the local entropy H(X) and the pixel value variance σ 2 The formulas are as follows: where P(x i ) is the probability distribution of the pixel gray values within the window, N is the number of window pixels, and μ is the window mean; Dual - attention fusion and position - encoding optimization mechanism: In the Transformer encoding block, an innovative dual - attention fusion strategy is adopted. The spatial attention branch obtains the spatial - position attention weight by calculating the dot product of the query vector and the key vector; the semantic attention branch obtains the semantic attention weight through semantic analysis of the feature map. At the same time, the position encoding is optimized, and a position - encoding vector based on high - order sine and cosine functions is introduced. Its calculation formula is: (pos is the position, i is the dimension index, and d model is the model dimension), effectively improving the modeling ability for complex ground fissure trends and long-distance dependency relationships; Dynamically sparse - connected fully - connected layer: Aiming at the problems of numerous parameters and easy over - fitting in the traditional fully - connected layer, a dynamically sparse - connected fully - connected layer is designed. When training the data of this mining area, every 100 batches, calculate the activation frequency of each neuron and the correlation score with the ground fissure feature. For neuron connections with an activation frequency lower than 0.12 and a correlation score lower than 0.35, set their weights to 0 to achieve dynamic sparsification of the connections, and re - evaluate and adjust the weights every 500 batches. While reducing the number of parameters and the amount of calculation, it enhances the model's learning ability for key features and improves the generalization of the model; Transformer Encoding Block Concatenation: Five Transformer encoding blocks each containing the above-mentioned adaptive multi-scale patch mapping layer, dual attention fusion and position encoding optimization mechanism module, and dynamic sparse connection fully connected layer are concatenated in sequence. The output of the previous encoding block serves as the input of the next encoding block to construct a powerful model encoder; 2) Construct the model decoder: Multi-scale Pyramid Cascade Decoding Strategy: Adopt the multi-scale pyramid cascade decoding strategy of the convolutional neural network. Use transposed convolution operations to upsample the feature maps of different scales output by the model encoder. For the smallest scale feature map, use a 4×4 transposed convolution kernel with a stride of 2 for upsampling; for larger scale feature maps, reasonably select the transposed convolution kernel and stride according to the feature map size. Concatenate the upsampled feature maps with the low-resolution feature maps of the corresponding levels in the encoder, and then perform 3 convolution operations using a 3×3 convolution kernel. Use skip connections to directly introduce the features in the encoder into the decoder to retain detailed information and gradually restore the complete details of the ground fissure; during the decoding process, perform skip connections between the upsampled feature maps and the low-resolution feature maps of the corresponding levels in the encoder to achieve cross-level information transfer and enhance the spatial detail expression ability. Subsequently, the concatenated feature maps perform convolution operations through a 3×3 convolution kernel to further extract key features, suppress noise, and enhance the model's attention to the ground fissure area. Residual Attention Enhanced Pyramid Pooling: Innovatively propose a residual attention enhanced pyramid pooling mechanism in the decoder. On the basis of traditional pyramid pooling, add residual connections to prevent feature loss, introduce an attention mechanism to weight the features pooled at different scales, obtain the global information of the feature map through global average pooling and global maximum pooling, generate attention weights through a fully connected layer (dimensionality reduction to 64 dimensions) and ReLU and Sigmoid activation functions, and multiply them with the pooled feature maps to highlight the key features of the ground fissure at different scales, suppress background noise, and improve the decoding accuracy; Spatial-Channel Attention Bottleneck Mechanism: Construct a spatial-channel attention bottleneck mechanism. Perform global average pooling and global maximum pooling operations on the input feature map in the spatial dimension and channel dimension respectively to obtain the spatial attention feature map and the channel attention feature map. Reduce the dimensions (dimensionality reduction to 64 dimensions) and then increase the dimensions of these two feature maps through a fully connected layer, and perform non-linear transformations through the ReLU activation function and the Sigmoid activation function to obtain the spatial attention weight and the channel attention weight. Finally, multiply the spatial attention weight and the channel attention weight with the original input feature map respectively to highlight the key features of the ground fissure in the spatial and channel dimensions and suppress the interference of background noise; Crack breakpoint and discontinuity detection module: A dedicated crack breakpoint and discontinuity detection module is added to the model decoder. Based on the Histogram of Oriented Gradients (HOG) algorithm, combined with the feature maps extracted by the model, the directional gradient information of ground fissures at different positions is calculated in units of 8×8 pixel blocks to determine the directionality of the fissures. By comparing the direction consistency of adjacent regions (a direction angle greater than 40° is considered inconsistent) and the change in gradient magnitude (a change exceeding 12 is considered abnormal), potential breakpoints and discontinuities are identified. For the detected suspicious points, the curvature information and texture features of the local region are further used for secondary confirmation to improve the detection accuracy; Crack connection rule determination module: According to the detected breakpoints and discontinuities, the crack connection rule determination module in the model formulates connection rules based on factors such as the overall trend of the ground fissure, the distance and angular relationship between adjacent crack segments. If the direction angle between adjacent crack segments is within the threshold range of [-25°, 25°], and the distance between them is less than the maximum connection distance (calculated to be 15 pixels) determined according to the average width of the ground fissures in this mining area (measured to be 0.2 meters) and the image resolution (4000×3000 pixels), then it is determined that these two crack segments should be connected, realizing the precise repair and complete extraction of incomplete cracks. (3) Model training and application: The constructed dataset of 4000 image pairs is divided into a training set, a validation set, and a test set in a ratio of 8:1:

1. A training framework is built using PyTorch, the AdamW optimizer is selected, the initial learning rate is set to 0.00012, and the cosine annealing learning rate adjustment strategy is used to dynamically adjust the learning rate. The loss function uses a hybrid loss function combining cross-entropy loss and Dice loss, with the cross-entropy loss weight set to 0.6 and the Dice loss weight set to 0.4; The trained model is deployed on a high-performance server equipped with an NVIDIA RTX 3090 GPU. The newly input UAV images of this mining area are preprocessed, and the ground fissure information is output through the encoder and decoder in sequence. After post-processing such as morphological opening and closing operations, a complete and accurate ground fissure extraction result is obtained.

2. The intelligent extraction method for surface cracks in coal mining subsidence areas based on the improved Transformer model according to claim 1, wherein, In the image acquisition step, the flight height of the UAV is adjusted in real-time between 50 - 100 meters to adapt to the terrain undulation.

3. The intelligent extraction method for surface cracks in coal mining subsidence areas based on the improved Transformer model according to claim 1, wherein In the image screening step, the threshold for the proportion of ground fissures is set to 18% - 35%, and the clarity threshold is that the image entropy is greater than 3.2 and the average gradient magnitude is greater than 8, which is applicable to image screening.

4. The intelligent extraction method for surface cracks in coal mining subsidence areas based on the improved Transformer model according to claim 1, characterized in that In the data augmentation step, the random rotation angle range is set to [-20°, 20°], the scaling ratio range is [0.88, 1.12], and the brightness and contrast adjustment factors range is [0.85, 1.15] for data expansion according to the actual situation.

5. The intelligent extraction method for surface cracks in coal mining subsidence areas based on the improved Transformer model according to claim 1, characterized in that, In the implementation process of the adaptive multi-scale patch mapping layer, for image data, when the entropy in the 5×5 neighborhood of the local region of the image is greater than 3.5 and the variance is greater than 25, a 3×3 convolution kernel is used for patch mapping, otherwise a 7×7 convolution kernel is used.

6. The intelligent extraction method for surface cracks in coal mining subsidence areas based on the improved Transformer model according to claim 1, wherein, During the implementation of the dynamic sparse connection fully connected layer, when training relevant data, the activation frequency and correlation score of neurons are calculated every 100 batches. The connection weights of neurons with an activation frequency lower than 0.12 and a correlation score lower than 0.35 are set to 0, and the weights are re-evaluated and adjusted every 500 batches.

7. The intelligent extraction method for surface cracks in coal mining subsidence areas based on the improved Transformer model according to claim 1, characterized in that During the implementation of the crack breakpoint and discontinuity detection module, for the characteristics of ground fissures, by comparing the direction consistency of adjacent blocks, a direction angle greater than 40° is considered inconsistent, and a gradient amplitude change exceeding 12 is considered abnormal.

Citation Information

Cited By

  • Intelligent analysis method and system for information extraction of intelligent electric meter

    CN120779323A

  • Range gating imaging noise elimination method for adaptive background suppression

    CN121008252A

  • Map generation method under support of customized large model

    CN121053250A

  • Intelligent detection method and system for differential settlement of super high-rise building foundation

    CN121071616A

  • Deep learning assisted acceleration fracturing construction parameter intelligent real-time optimization method

    CN121091692A