Method and device for identifying ground fissures in coal mining areas
By integrating adaptive feature fusion of visible light and thermal infrared dual-modal data with Transformer long-range modeling, a ground fissure identification model was constructed, which solved the identification problem in complex scenarios in coal mining areas and achieved high-precision and continuous ground fissure identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NORTHEASTERN UNIV CHINA
- Filing Date
- 2025-12-05
- Publication Date
- 2026-05-29
Smart Images

Figure CN121767867B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image segmentation technology, and in particular to a method and apparatus for identifying ground fissures in coal mining areas. Background Technology
[0002] Coal, as a major source of energy consumption and industrial raw materials, is regarded as the "ballast" and "stabilizer" of my country's energy security, and a key support for promoting economic development and ensuring energy security. However, large-scale underground coal seam mining damages the geological structure, causing surface movement and deformation, which in turn induces severe ground fissure disasters.
[0003] In related technologies, the identification of ground fissures induced by coal mining still relies on single-modal remote sensing imagery as input for deep learning models. Furthermore, long-distance spatial dependency modeling is insufficient, and performance in terms of continuity and hole repair is poor. Moreover, ground fissures in coal mining areas face unique scene constraints such as complex backgrounds, blurred edges, vegetation obstruction, and mulch film coverage, making it difficult to overcome scene adaptation limitations and achieve accurate and comprehensive identification of ground fissures. Summary of the Invention
[0004] In view of this, this application provides a method and device for identifying ground fissures in coal mining areas. By fusing visible light and thermal infrared dual-modal data and combining adaptive feature fusion and Transformer long-distance modeling mechanism, accurate identification of ground fissures is achieved in complex scenarios in coal mining areas.
[0005] According to one aspect of this application, a method for identifying ground fissures in coal mining areas is provided, comprising:
[0006] Acquire visible light and thermal infrared remote sensing images of the target coal mining area;
[0007] Based on the visible light remote sensing image and the thermal infrared remote sensing image, a ground fissure dataset is constructed;
[0008] A ground fissure identification model is constructed based on a neural network. The ground fissure identification model includes a dual-branch encoder, a feature extraction module, and a decoder. The dual-branch encoder includes a feature fusion module.
[0009] The ground fissure identification model is iteratively trained based on the ground fissure dataset.
[0010] The trained ground fissure identification model is used to identify ground fissures in the target coal mining area using both visible light remote sensing images and thermal infrared remote sensing images, resulting in a ground fissure identification map.
[0011] According to another aspect of this application, a coal mining area ground fissure identification device is provided, comprising:
[0012] The acquisition module is used to acquire visible light remote sensing images and thermal infrared remote sensing images of the target coal mining area;
[0013] The training module is used to construct a ground fissure dataset based on the visible light remote sensing image and the thermal infrared remote sensing image; and to construct a ground fissure recognition model based on a neural network, wherein the ground fissure recognition model includes a dual-branch encoder, a feature extraction module and a decoder, and the dual-branch encoder includes a feature fusion module; and to iteratively train the ground fissure recognition model according to the ground fissure dataset.
[0014] The identification module is used to identify ground fissures in the target coal mining area using the trained ground fissure identification model and the target visible light remote sensing image and the target thermal infrared remote sensing image, and to obtain the ground fissure identification result map.
[0015] According to another aspect of this application, a readable storage medium is provided having a program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method for identifying ground fissures in coal mining areas.
[0016] According to another aspect of this application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the steps of the above-described method for identifying ground fissures in coal mining areas.
[0017] By employing the aforementioned technical solutions, this application provides a method and apparatus for identifying ground fissures in coal mining areas. By designing a dual-branch encoder and an adaptive feature fusion mechanism in the ground fissure identification model, complementary information from visible light and thermal infrared dual-modal data is effectively integrated, overcoming the limitation of single visible light images in complex scenarios such as vegetation obstruction and plastic film coverage where it is difficult to extract hidden fissure information. Secondly, a feature extraction module centered on a Transformer is introduced into the ground fissure identification model, enhancing its ability to model long-distance continuous structures of ground fissures and solving the problems of fissure breakage and missed detection caused by the limited receptive field of traditional convolutional neural networks. Furthermore, the decoder in the ground fissure identification model adopts a multi-level structure that is inversely symmetrical to the encoder, and integrates shallow details and deep semantics through skip connections, effectively addressing the identification difficulties of long, thin ground fissure targets with extremely low pixel ratios and blurred edges, thus improving the accuracy of segmentation boundaries.
[0018] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0020] Figure 1 A flowchart illustrating the method for identifying ground fissures in coal mining areas provided in an embodiment of this application is shown.
[0021] Figure 2 A flowchart illustrating a method for identifying ground fissures in coal mining areas, according to another embodiment of this application, is shown.
[0022] Figure 3 A schematic diagram of the structure of a ground fissure identification model provided in another embodiment of this application is shown;
[0023] Figure 4 This paper shows a grayscale value profile difference diagram between a visible light remote sensing image and a thermal infrared remote sensing image in a crack region, according to another embodiment of this application.
[0024] Figure 5 The recognition result of the ground fissure recognition model provided in another embodiment of this application is shown;
[0025] Figure 6 A comparison diagram of confusion matrices provided in another embodiment of this application is shown;
[0026] Figure 7 This paper shows a comparison chart of ablation analysis provided in another embodiment of the present application;
[0027] Figure 8 A performance comparison chart of the model provided in another embodiment of this application is shown;
[0028] Figure 9 This illustrates another identification result of the ground fissure identification model provided in yet another embodiment of this application;
[0029] Figure 10 A structural block diagram of the coal mine area ground fissure identification device provided in an embodiment of this application is shown. Detailed Implementation
[0030] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.
[0031] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0032] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application’s specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “connected” to another element, it can be directly connected or connected to the other element, or there may be intermediate elements. Furthermore, “connected” or “connected” as used herein can include wireless connections or wireless bridging. The term “and / or” as used herein includes all or any unit and all combinations of one or more associated listed items.
[0033] Exemplary embodiments according to this application will now be described in more detail with reference to the accompanying drawings. However, these exemplary embodiments may be implemented in many different forms and should not be construed as being limited to the embodiments set forth herein. It should be understood that these embodiments are provided so that the disclosure of this application is thorough and complete, and that the concept of these exemplary embodiments is fully conveyed to those skilled in the art.
[0034] This application provides a method for identifying ground fissures in coal mining areas, such as... Figure 1 As shown, the method includes:
[0035] Step 101: Acquire visible light remote sensing images and thermal infrared remote sensing images of the target coal mining area.
[0036] Step 102: Construct a ground fissure dataset based on visible light remote sensing images and thermal infrared remote sensing images.
[0037] Step 103: Construct a ground fissure identification model based on a neural network.
[0038] The ground fissure identification model includes a dual-branch encoder, a feature extraction module, and a decoder. The dual-branch encoder includes a feature fusion module.
[0039] Step 104: Iteratively train the ground fissure identification model based on the ground fissure dataset.
[0040] Step 105: Using the trained ground fissure recognition model, ground fissures are identified in the target coal mining area using the visible light remote sensing image and the thermal infrared remote sensing image, resulting in a ground fissure recognition map.
[0041] The coal mine ground fissure identification method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the coal mine ground fissure identification method, but is not limited to the above forms.
[0042] This embodiment first acquires visible light and thermal infrared remote sensing images of the target coal mining area, and then constructs a sample group containing visible light images, thermal infrared images, and labeled images of the same geographical area, thus forming a ground fissure dataset. Subsequently, a ground fissure recognition model integrating a dual-branch encoder, a feature fusion module, a feature extraction module, and a decoder is constructed. In the ground fissure recognition model, the main branch and auxiliary branch of the dual-branch encoder extract multi-scale features from the visible light and thermal infrared images, respectively, and perform adaptive dynamic fusion through a feature fusion module connected to the bottleneck residual layer in the main branch. The fused depth features are enhanced with long-range spatial dependency modeling by a feature extraction module based on a Transformer architecture. Finally, the decoder outputs a binary probability map through stepwise upsampling and combining skip connection features from the dual-branch encoder, completing pixel-level recognition of ground fissures, comprehensively utilizing the high-resolution details of the visible light image and the sensitivity of the thermal infrared image to temperature anomalies.
[0043] Another embodiment of this application provides a method for identifying ground fissures in coal mining areas, such as... Figure 2 As shown, the method includes:
[0044] Step 201: Acquire visible light remote sensing images and thermal infrared remote sensing images of the target coal mining area; preprocess the visible light remote sensing images and thermal infrared remote sensing images respectively to obtain visible light orthophoto images and thermal infrared orthophoto images; annotate the visible light remote sensing images with ground fissures based on the thermal infrared orthophoto images to obtain labeled images; segment the visible light orthophoto images, thermal infrared orthophoto images and labeled images respectively to obtain visible light images, thermal infrared images and labeled images; construct sample groups based on the visible light images, thermal infrared images and labeled images corresponding to the same geographical area in the target coal mining area; construct a ground fissure dataset based on the sample groups.
[0045] It should be noted that visible light remote sensing imagery relies on the reflection of visible light by ground objects to capture the morphology and structure of ground fissures intuitively and clearly, and presents the texture details of ground fissures in greater detail. Thermal infrared remote sensing imagery, on the other hand, is based on the principle of thermal radiation from ground objects. Due to the significant difference in thermal radiation between ground fissures and surrounding ground objects in terms of their depth and morphological characteristics, ground fissures exhibit obvious temperature anomalies (troughs) in the fissure area.
[0046] In this step, a drone is used to acquire visible light and thermal infrared remote sensing images of the target coal mining area. Then, all visible light remote sensing images undergo aerial triangulation, orthorectification, and stitching preprocessing to generate a complete visible light orthorectified image covering the entire target coal mining area with accurate geographic information. Simultaneously, all thermal infrared remote sensing images undergo the same preprocessing to generate a complete thermal infrared orthorectified image covering the entire target coal mining area with accurate geographic information.
[0047] It is worth mentioning that, through aerial triangulation and orthorectification, geometric distortions caused by camera attitude, terrain undulations, etc., are eliminated, so that the visible light orthophoto and the thermal infrared orthophoto are unified into the same precise geodetic coordinate system. That is, for any geographical location in the target coal mining area, it has a unique corresponding geographical coordinate on both the visible light orthophoto and the thermal infrared orthophoto.
[0048] Because visible light orthophotos have high resolution, the shape, texture, edges, and other visual details of ground fissures are clearer, making it easier to accurately delineate their outlines. This step uses the visible light orthophoto as the primary image for annotation. During annotation, temperature information from the thermal infrared orthophoto is also referenced. For example, if a suspected ground fissure area on the visible light orthophoto also shows a clear linear temperature anomaly at the corresponding location on the thermal infrared orthophoto, then this is confirmed as a ground fissure and accurately delineated on the visible light orthophoto, forming a closed or polygonal target area. If almost no fissure traces are visible on a visible light orthophoto with dense vegetation cover, and a clear, low-temperature, dark line appears in a certain area of the thermal infrared orthophoto, and there are signs of fallen vegetation or abnormal vegetation arrangement at the corresponding location on the visible light orthophoto, then this is also confirmed as a ground fissure and accurately delineated on the visible light orthophoto, forming another closed or polygonal target area. Then, the target area on the visible light orthophoto is marked as "ground fissure", while other areas on the visible light orthophoto excluding the target area are marked as "non-ground fissure".
[0049] Here, a binary labeling method can be used, setting the pixel value of the target area marked as "ground fissure" to 1, and setting the pixel value of other areas marked as "non-ground fissure" to 0. The labeling process can be performed manually or using mature labeling software in related technologies; this embodiment does not impose specific limitations.
[0050] Therefore, the labeled visible light orthophoto is used as the tag image, distinguished from the visible light orthophoto and stored separately. This leverages the complementarity of multimodal information to generate more accurate and reliable tag data. It is understood that the tag image and the visible light orthophoto are exactly the same size and correspond one-to-one in pixels. Thus, three large-scale remote sensing images are obtained: the visible light orthophoto, the thermal infrared orthophoto, and the tag image.
[0051] Then, the visible light orthophoto image, thermal infrared orthophoto image, and tag image are segmented into visible light images, thermal infrared images, and tag images, respectively. Specifically, the large-scale visible light orthophoto image is segmented into visible light images according to a first size, and the large-scale tag image is similarly segmented into tag images of the first size. This ensures that each visible light image corresponds to a tag image of the same geographical area. Furthermore, since the spatial resolution of the infrared orthophoto image is lower than that of the visible light orthophoto image, this step segments the large-scale thermal infrared orthophoto image into thermal infrared images according to a second size smaller than the first size, ensuring that each visible light image also corresponds to a thermal infrared image of the same geographical area.
[0052] Here, the first and second dimensions can be specifically set according to the visible light resolution and the thermal infrared resolution. For example, if the ratio between the visible light resolution and the thermal infrared resolution is m:n, then the ratio of the first and second dimensions is n:m, where m is less than n. Specifically, the first dimension can be 224×224 pixels, and the second dimension can be 112×112 pixels.
[0053] Therefore, visible light images, label images, and thermal infrared images of the same geographical area are grouped into a sample group.
[0054] It should be noted that in real mining area images, ground fissures are extremely sparse. If all images are used directly to train the subsequent ground fissure recognition model, the model will be completely unable to identify the fissures.
[0055] This step randomly removes a first preset proportion of sample images that do not contain any ground fissure pixels to alleviate class imbalance and allow the ground fissure recognition model to effectively learn the features of ground fissures. Next, all remaining sample groups are randomly divided according to a second preset proportion to obtain training, validation, and test sets. Based on these training, validation, and test sets, a ground fissure dataset is obtained.
[0056] For example, the first preset ratio can be 80%, and the second preset ratio can be 8:1:1.
[0057] Step 202: Construct a ground fissure identification model based on a neural network.
[0058] In this embodiment, as Figure 3 As shown, the ground fissure recognition model includes a dual-branch encoder, a feature extraction module, and a decoder. The dual-branch encoder consists of a main branch and an auxiliary branch. The main branch extracts multi-scale features from high-resolution visible light images, while the auxiliary branch extracts multi-scale features from low-resolution thermal infrared images. The main branch reuses the weights of a pre-trained ResNet50 (Residual Network with 50 layers) on the ImageNet dataset to directly reuse the complete structure of the pre-trained ResNet50. The main branch includes a combined structure, a first spatial dimension alignment structure, and a residual layer structure. The auxiliary branch corresponds to the main branch, enabling rapid adaptation to the light and dark contrast features of fissures and soil in visible light images. The auxiliary branch includes a combined structure, a second spatial dimension alignment structure, and a residual layer structure.
[0059] The combined structure includes a 7×7 convolutional layer, a batch normalization layer, and a ReLU activation function (Linear Rectification function). The first spatial dimension alignment structure includes a 3×3 max-pooling downsampling layer (3×3 convolutional kernel, stride=2, padding=1), and the second spatial dimension alignment structure includes a 3×3 max-pooling downsampling layer and a bilinear interpolation upsampling layer. The residual layer structure includes four bottleneck residual layers connected sequentially: the first bottleneck residual layer, the second bottleneck residual layer, the third bottleneck residual layer, and the fourth bottleneck residual layer. The first bottleneck residual layer consists of 3 bottleneck residual blocks (the expansion of the bottleneck residual block is a multiple of the number of output channels to the number of input channels) = 4; the second bottleneck residual layer consists of 4 bottleneck residual blocks (the stride of the bottleneck residual block = 2, expansion = 4); the third bottleneck residual layer consists of 6 bottleneck residual blocks (the stride of the bottleneck residual block = 2, expansion = 4); and the fourth bottleneck residual layer consists of 3 bottleneck residual blocks (the stride of the bottleneck residual block = 2, expansion = 4).
[0060] Here, each bottleneck residual block in the residual layer structure includes a 1×1 convolution for dimensionality reduction, a 3×3 convolution for extracting ground fissure details, and a 1×1 convolution for dimensionality increase. This reduces computational cost and enables accurate extraction of ground fissure features in complex backgrounds, while avoiding noise interference from gravel, vegetation, etc.
[0061] Furthermore, a feature fusion module is connected after each bottleneck residual layer in the residual layer structure of the main branch. This module is used to input the main feature map and auxiliary feature map generated by the bottleneck residual layers of the same level in the residual layer structure of the main branch and the auxiliary branch into the feature fusion module connected to the bottleneck residual layer of the same level in the main branch. This allows the two feature maps to be dynamically fused to obtain the fused feature map generated by the feature fusion module connected to the bottleneck residual layer of the same level in the main branch.
[0062] It is worth mentioning that after the feature fusion module generates the fused feature map in each bottleneck residual layer connected in the main branch, the fused feature map is then input into the next level of bottleneck residual layer in the main branch to participate in the extraction of the main feature map of the bottleneck residual layer in the subsequent levels of the main branch. This allows the main branch to continue to extract higher-level semantic information based on the fused feature map until all levels of the residual layer structure in the main branch have been processed. This achieves the dynamic fusion of multi-scale visible light features and thermal infrared features, and solves the problem of the one-sidedness of single-branch features.
[0063] Furthermore, the feature extraction module is used as the bottleneck layer of the ground fissure recognition model, connected between the dual-branch encoder and decoder. It is used to enhance the deep detail information in the ground fissure image and model long-range dependencies. The feature extraction module includes multiple levels of Transformer layers. Each Transformer layer includes layer normalization, multi-head attention mechanism, and multilayer perceptron.
[0064] Furthermore, a skip connection is established between the decoder and the dual-branch encoder. The decoder, based on the classic UNet architecture, employs stepped upsampling and resampling feature integration to address key recognition challenges such as the elongated shape of ground fissures, low pixel ratio, fragile and easily broken edges, and complex background textures. The decoder can predict the probability that each pixel in the input visible light image belongs to the ground fissure category, ultimately generating a binary probability map.
[0065] Step 203: Input the visible light images and thermal infrared images from the sample group in the ground fissure dataset into the main branch and auxiliary branch of the dual-branch encoder, respectively, to obtain the fused feature map generated by the dual-branch encoder; input the target fused feature map from the fused feature map into the feature extraction module to obtain the target feature map generated by the feature extraction module; input the main feature map generated by the main branch and the target feature map into the decoder to obtain the binary probability map generated by the decoder, which includes the probability that the pixels in the visible light images and thermal infrared images belong to the ground fissure category as predicted by the decoder; determine the loss function of the ground fissure recognition model based on the binary probability map and the label images in the sample group; calculate the loss function and update the model parameters of the ground fissure recognition model until the loss function converges, and obtain the trained ground fissure recognition model.
[0066] In this step, the ground fissure recognition model is iteratively trained using the ground fissure dataset. The training set in the dataset is used to train the model, while the validation set is used for hyperparameter tuning during training and to prevent overfitting, ensuring the model has good generalization ability. The test set is used after the entire training process to simulate real-world application scenarios and calculate evaluation metrics such as recall and intersection-over-union ratio (IoU) to reliably verify the overall recognition performance of the ground fissure recognition model.
[0067] During training, the input to the ground fissure recognition model is any sample group from the ground fissure dataset. Visible light images from this sample group are input to the main branch of the dual-branch encoder in the model, while thermal infrared images are input to the auxiliary branch. This allows the dual-branch encoder to generate a fused feature map based on the main features extracted from the visible light images by the main branch and the auxiliary features extracted from the thermal infrared images by the auxiliary branch. Further, the feature extraction module, following the dual-branch encoder, processes the last fused feature map (i.e., the target fused feature map) generated by the dual-branch encoder to generate a target feature map. Thus, the decoder, based on the main feature map generated by the main branch of the dual-branch encoder and the target feature map, performs ground fissure recognition, generating a binary probability map. This binary probability map includes the probability predicted by the decoder that pixels in the visible light and thermal infrared images belong to the ground fissure category.
[0068] Based on the binary probability map predicted by the ground fissure recognition model and the real ground fissure results corresponding to the labeled images in the sample groups of the visible light image and the thermal infrared image, the loss function of the ground fissure recognition model is calculated, and the model parameters of the ground fissure recognition model are updated until the loss function converges, thus obtaining the trained ground fissure recognition model.
[0069] Furthermore, as a refinement and extension of the specific implementation of the above embodiments, in order to fully illustrate the specific implementation process of this embodiment, the visible light images and thermal infrared images within the sample groups of the ground fissure dataset are respectively input into the main branch and auxiliary branch of the dual-branch encoder to obtain the fused feature map generated by the dual-branch encoder. Specifically, this includes: inputting the visible light images and thermal infrared images into the combination structure in the main branch and auxiliary branch respectively to obtain the initial visible light feature map generated by the combination structure in the main branch and the initial thermal infrared feature map generated by the combination structure in the auxiliary branch; inputting the initial visible light feature map and the initial thermal infrared feature map into the first spatial dimension alignment structure and auxiliary branch respectively in the main branch. The second spatial dimension alignment structure in the branch is used to obtain the visible light feature map generated by the first spatial dimension alignment structure in the main branch, and the thermal infrared feature map with the same spatial scale as the visible light feature map generated by the second spatial dimension alignment structure in the auxiliary branch. The visible light feature map and the thermal infrared feature map are respectively input into the residual layer structure in the main branch and the auxiliary branch to obtain the main feature map generated by the bottleneck residual layer at different levels in the residual layer structure in the main branch and the auxiliary feature map generated by the bottleneck residual layer at different levels in the residual layer structure in the auxiliary branch. Based on the main feature map and the auxiliary feature map generated by the bottleneck residual layer at the same level in the residual layer structure in the main branch and the auxiliary branch, a fused feature map is generated.
[0070] In this step, during a single training iteration of the ground fissure recognition model, sample groups of size B from the ground fissure dataset can be processed simultaneously. Visible light images from the same sample group are input into the main branch of the dual-branch encoder in the ground fissure recognition model, while thermal infrared images from the same sample group are input into the auxiliary branch of the same dual-branch encoder. This allows the dual-branch encoder to process both visible light and thermal infrared images of the same geographical area simultaneously. Firstly, after the visible light image enters the main branch, it is first converted from the input three channels (RGB, Red, Green, Blue) to 64 channels through a 7×7 convolutional layer within the main branch's combined structure. Simultaneously, the visible light image size is downsampled from 224×224 to 112×112 to initially capture details such as the edges and textures of the ground fissures. Then, the visible light image, processed by the 7×7 convolutional layer in the main branch's combined structure, is sequentially processed by batch normalization layers within the main branch's combined structure for stabilization training and by ReLU activation functions to suppress background noise. This results in the combined structure in the main branch generating an initial visible light feature map of size [64, 112, 112]. Simultaneously, the auxiliary branch processes the input thermal infrared image. The 7×7 convolutional layer in the auxiliary branch's combined structure converts the input single-channel (grayscale) thermal infrared image into 64 channels and downsamples the thermal infrared image size from 112×112 to 56×56 to initially locate temperature anomaly regions. Then, the thermal infrared image, processed by the 7×7 convolutional layer in the auxiliary branch's combined structure, is sequentially processed by batch normalization layers within the auxiliary branch's combined structure for stabilization training and by ReLU activation functions to suppress background noise. This results in the combined structure in the auxiliary branch outputting an initial thermal infrared feature map of size [64, 56, 56]. By combining the structures, the dual-branch channel dimensions are unified, and preliminary features of visible light morphology and thermal infrared temperature are extracted respectively, laying the groundwork for subsequent spatial alignment and deep extraction.
[0071] Furthermore, on one hand, the initial visible light feature map enters the first spatial dimension alignment structure in the main branch. The first spatial dimension alignment structure performs a 3×3 max pooling downsampling on the initial visible light feature map of [64, 112, 112] to reduce computation and preserve key details, outputting a visible light feature map of size [64, 56, 56]. Here, the visible light feature map serves as the reference for spatial alignment. Simultaneously, on the other hand, the initial thermal infrared feature map also enters the second spatial dimension alignment structure in the auxiliary branch. The second spatial dimension alignment structure performs the same 3×3 max pooling downsampling on the initial thermal infrared feature map of [64, 56, 56], obtaining an intermediate feature map of [64, 28, 28]. Then, to match the intermediate feature map with the visible light feature map of the main branch, bilinear interpolation upsampling (with a magnification of 2) is used to enlarge the intermediate feature map to [64, 56, 56], obtaining the thermal infrared feature map. This ensures that the visible light feature map output by the first spatial dimension alignment structure in the main branch and the thermal infrared feature map output by the second spatial dimension alignment structure are completely identical in spatial dimensions (height and width) and number of channels, both being [64, 56, 56], thus achieving spatial dimension alignment.
[0072] Furthermore, the visible light feature map and the thermal infrared feature map are respectively fed into the residual layer structures in the main branch and the auxiliary branch for further feature extraction. Specifically, the four bottleneck residual layers in the residual layer structure of the main branch sequentially extract main feature maps at four scales, while the four bottleneck residual layers in the residual layer structure of the auxiliary branch sequentially extract auxiliary feature maps at four scales. Thus, a fused feature map is generated based on the main feature maps and auxiliary feature maps generated from the bottleneck residual layers of the same level in the residual layer structures of the main branch and the auxiliary branch.
[0073] Furthermore, as a refinement and extension of the specific implementation of the above embodiments, in order to fully illustrate the specific implementation process of this embodiment, a fused feature map is generated based on the main feature map and auxiliary feature map generated by the bottleneck residual layers of the same level within the residual layer structure of the main branch and auxiliary branch. Specifically, this includes: sorting the levels of the bottleneck residual layers to obtain a first level order, and taking the first level in the first level order as the first generation level; inputting the main feature map and auxiliary feature map generated by the bottleneck residual layers of the first generation level in the main branch and auxiliary branch into the feature fusion module connected to the bottleneck residual layers of the first generation level in the main branch, to obtain the fused feature map generated by the feature fusion module connected to the bottleneck residual layers of the first generation level in the main branch; if the first generation level is not in the first level order... For the last level, the fused feature map generated by the feature fusion module connected to the bottleneck residual layer of the first generation level in the main branch is input into the bottleneck residual layer of the next level of the first generation level in the main branch. The first generation level is updated according to the next level of the first generation level, which is the level located one level below the first generation level in the first level sequence. The main feature map and auxiliary feature map generated by the updated bottleneck residual layer of the first generation level in the main branch and auxiliary branch are input into the feature fusion module connected to the updated bottleneck residual layer of the first generation level in the main branch. This results in the fused feature map generated by the feature fusion module connected to the updated bottleneck residual layer of the first generation level in the main branch, until the updated first generation level is the last level in the first level sequence.
[0074] In this step, firstly, the bottleneck residual layers within the residual layer structure are sorted according to the neural network processing order to obtain the first-level order: layer 1, layer 2, layer 3, and layer 4. Then, the first layer in the first-level order is taken as the current first generation layer. When the residual layer structure in the main branch outputs the main feature map at its first bottleneck residual layer, the residual layer structure in the auxiliary branch simultaneously outputs the auxiliary feature map at its first bottleneck residual layer. These two feature maps are then input together into the feature fusion module connected to the first bottleneck residual layer in the main branch, so that the two feature maps are dynamically fused to obtain the fused feature map generated by the feature fusion module connected to the first bottleneck residual layer in the main branch.
[0075] Next, it is determined whether the current first generation level is the last level (i.e., the last level in the first level sequence). Since the current first generation level is the first layer, it is not the last level. Therefore, the fused feature map generated by the feature fusion module connected to the first bottleneck residual layer in the main branch is re-inputted into the next level within the residual layer structure in the main branch, namely the second bottleneck residual layer, to continue participating in the feature extraction of subsequent bottleneck residual layers in the main branch. Then, the second level is updated to the current first generation level. The main feature map output from the second bottleneck residual layer of the residual layer structure in the main branch and the auxiliary feature map output from the second bottleneck residual layer of the residual layer structure in the auxiliary branch are input into the feature fusion module connected to the second bottleneck residual layer in the main branch. This yields the fused feature map generated by the feature fusion module connected to the second bottleneck residual layer in the main branch. This process continues until the feature fusion module connected to the fourth bottleneck residual layer (i.e., the last level in the first level sequence) in the main branch generates the fused feature map.
[0076] For example, in the main branch and auxiliary branch, the main feature map and auxiliary feature map output by the bottleneck residual layer at the same level have the same scale, which are 112×112×256, 56×56×512, 28×28×1024, and 14×14×2048, respectively.
[0077] Furthermore, in this step, the feature fusion module updates the learnable parameters iteratively during training and calculates an adaptive ratio to dynamically weight and fuse the main feature map and the auxiliary feature map based on the learnable parameters. This solves the problem that fixed ratio fusion is difficult to adapt to the dynamic changes in the importance of main and auxiliary features under different scenarios.
[0078] For example, feature fusion is performed in the feature fusion module according to the following formula:
[0079] ,
[0080] ,
[0081] .
[0082] here, This step is to identify learnable parameters. An initial value, such as 0.75, is set, corresponding to the initial setting where visible light is the dominant feature and thermal infrared light is the auxiliary feature. In the first iteration of training the ground fissure identification model, a nonlinear mapping function is used... The initial values of the learnable parameters are compressed to the interval [0, 1] to obtain the main feature map of the input feature fusion module. weight Ensure weight It conforms to the mathematical definition of weights. And it is based on the main feature map of the input feature fusion module. weight Determine the auxiliary feature map of the input feature fusion module. The weight is Therefore, based on the weights of the main and auxiliary feature maps in the input feature fusion module, the two feature maps are weighted and fused within the feature fusion module to generate a fused feature map. Then, based on the loss function of the ground fissure identification model during the first iteration of training... Regarding learnable parameters gradient and the learning rate of the ground fissure identification model. For the current learnable parameters Update the parameters to obtain new learnable parameters. New learnable parameters The features are used for feature fusion in the next iteration of training of the ground fissure identification model, and so on.
[0083] in, The optimization of the ground fissure identification model can be set to control the speed at which the learnable parameters are updated each time. The loss function can be the dice loss function, but this embodiment does not impose any specific restrictions.
[0084] It should be noted that after the feature fusion module generates the fused feature map, this step will also use a combination structure to refine the fused feature map, remove redundant information in the fused feature map, enhance the discrimination ability of effective features, and update the fused feature map.
[0085] Furthermore, as a refinement and extension of the specific implementation of the above embodiments, in order to fully illustrate the specific implementation process of this embodiment, the target fused feature map in the fused feature map is input into the feature extraction module to obtain the target feature map generated by the feature extraction module. Specifically, this includes: globally flattening the target fused feature map into a one-dimensional feature sequence, where the target fused feature map is the fused feature map generated by the feature fusion module connected to the bottleneck residual layer of the last level in the first level order in the main branch, and the first level order is determined according to the level of the bottleneck residual layer; calculating the position encoding matrix of the one-dimensional feature sequence using a sine-cosine function; constructing a feature vector based on the one-dimensional feature sequence and the position encoding matrix; inputting the feature vector sequentially into the Transformer layer to obtain the enhanced feature tensor, where the Transformer layer includes layer normalization, multi-head attention mechanism, and multilayer perceptron; and reshaping the enhanced feature tensor to obtain the target feature map.
[0086] In this step, the fused feature map generated by the feature fusion module connected to the bottleneck residual layer in the main branch of the dual-branch encoder is used as the target fused feature map and input into the feature extraction module. Within the feature extraction module, the target fused feature map is first globally flattened from two-dimensional spatial features into a one-dimensional feature sequence. This one-dimensional feature sequence carries all the semantic information extracted from the fused visible light and thermal infrared images. Next, the position encoding matrix of the generated one-dimensional feature sequence is calculated using a sine-cosine function. This position encoding matrix systematically encodes the two-dimensional spatial coordinates (i.e., position encoding) corresponding to each position in the one-dimensional feature sequence. Then, the one-dimensional feature sequence is fused with the position encoding matrix by element-wise addition, i.e., the position encoding matrix is added to each corresponding position in the feature sequence, thus obtaining a complete feature tensor. This feature tensor serves as the direct input to the subsequent multi-head self-attention mechanism, enabling the Transformer layer to simultaneously consider the semantic association of features and their spatial layout in the image. This effectively establishes a long-range spatial dependency model for targets such as ground fissures, overcoming the inherent limitation of the limited receptive field in traditional convolutional neural networks.
[0087] Furthermore, the feature tensor is sequentially processed through multiple Transformer layers within the feature extraction module. In each Transformer layer, the feature tensor is first stabilized by layer normalization, then the relationships between all elements in the feature tensor are calculated using a multi-head attention mechanism to achieve long-distance spatial correlation modeling. Finally, a nonlinear feature transformation is performed using a multilayer perceptron. After this deep processing through multiple Transformer layers, an enhanced feature tensor is obtained. This enhanced feature tensor is then reshaped to obtain the target feature map, strengthening the ability to extract ground fissure features in complex backgrounds so that the subsequent decoder can process them smoothly. For example, eight Transformer layers can be set within the feature extraction module.
[0088] For example, the target fused feature map is globally flattened from two-dimensional spatial features into a one-dimensional feature sequence according to the following formula:
[0089] ,
[0090] ,
[0091] ,
[0092] in, It is the intermediate feature after flattening, which merges the "height H × width W" of the target fusion feature map into the sequence length; It is a one-dimensional feature sequence, with the dimension adjusted to the Transformer requirement of "[batch, sequence length, feature dimension]"; C is the number of feature channels; It is the length of the one-dimensional feature sequence.
[0093] The core parameters of the position encoding are determined according to the following formula:
[0094] ,
[0095] ,
[0096] in, This is the length of the positional encoding, which must be consistent with the length of the one-dimensional feature sequence. This is the dimension of the positional encoding, which must be consistent with the number of feature channels.
[0097] Based on the core parameters of the position encoding, the position code is generated according to the following formula:
[0098] ,
[0099] ,
[0100] ,
[0101] in, It is the dimension index of the position encoding, traversing each encoding dimension (corresponding to each dimension of the feature channel). It is the angular base, which varies with the dimension. Increase the exponential decay to control the sensitivity of different dimensions to changes in location; It is a rounding operation. It is a sequence index that iterates through each element of the one-dimensional feature sequence (corresponding to the spatial position of the pixel before flattening). It is an angle value, which will be used to index the sequence. With dimensional sensitivity Related. It is a fixed-position encoded tensor that carries pixel spatial location information (not updated during training).
[0102] The feature vector is generated according to the following formula:
[0103] ,
[0104] in, It is a complete feature tensor that contains bimodal feature information. and spatial location information .
[0105] For example, the scale of the target feature map is 14×14×4096.
[0106] Furthermore, as a refinement and extension of the specific implementation of the above embodiments, in order to fully illustrate the specific implementation process of this embodiment, the main feature map generated by the main branch and the target feature map are input into the decoder to obtain a binary probability map generated by the decoder. Specifically, this includes: sorting the levels of the decoding layer to obtain a second level order, and taking the first level in the second level order as the second generation level; inputting the target feature map and the main feature map generated by the bottleneck residual layer corresponding to the second generation level in the residual structure of the main branch into the decoding layer of the second generation level to obtain the decoding feature map generated by the decoding layer of the second generation level, wherein the levels of the decoding layer in the decoder correspond to the levels of the bottleneck residual layer in the main branch in reverse order; if the second generation level is not the last level in the second level order, the second generation level is... The decoded feature map generated by the decoding layer of the second generation level is input into the decoding layer of the next level of the second generation level in the decoder. The second generation level is updated according to the next level of the second generation level, which is the level located one position below the second generation level in the second generation level sequence. The master feature map generated by the bottleneck residual layer corresponding to the updated second generation level in the main branch is input into the decoding layer of the updated second generation level to obtain the decoded feature map generated by the decoding layer of the updated second generation level. This process continues until the updated second generation level is the last level in the second generation level sequence. The number of channels of the decoded feature map generated by the decoding layer of the last level in the second generation level sequence is compressed to a preset value to generate the original score map. A binary probability map is generated based on the original score map using a nonlinear mapping function.
[0107] In this step, the decoder corresponds to the residual layer structure in the dual-branch encoder, comprising four sequentially connected decoding layers: the first decoding layer, the second decoding layer, the third decoding layer, and the fourth decoding layer. It should be noted that the four decoding layers of the decoder correspond in reverse order to the four bottleneck residual layers within the residual layer structure of the main branch in the dual-branch encoder. That is, the first decoding layer corresponds to the fourth bottleneck residual layer in the residual layer structure of the main branch in the dual-branch encoder; the second decoding layer corresponds to the third bottleneck residual layer in the residual layer structure of the main branch in the dual-branch encoder; the third decoding layer corresponds to the second bottleneck residual layer in the residual layer structure of the main branch in the dual-branch encoder; and the fourth decoding layer corresponds to the first bottleneck residual layer in the residual layer structure of the main branch in the dual-branch encoder.
[0108] Specifically, firstly, the layers of the decoding layer within the decoder are ordered according to the processing order of the neural network, resulting in a second-level order: layer 1, layer 2, layer 3, and layer 4. Then, the first layer in this second-level order is used as the current second generation layer. After the feature extraction module outputs the target feature map, it enters the first decoding layer. The first decoding layer first upsamples the target feature map through a transposed convolutional layer to double its spatial size, resulting in an upsampled feature map generated by the first decoding layer. Simultaneously, the first decoding layer introduces the main feature map generated by the fourth bottleneck residual layer corresponding to the first decoding layer within the residual layer structure of the main branch of the dual-branch encoder via skip connections. Since this upsampled feature map and the main feature map have the same number of channels and the same spatial size, they are directly concatenated by channel dimension through resampling feature integration, resulting in a concatenated feature map generated by the first decoding layer. Then, the first decoding layer fuses the generated spliced feature map through a convolutional layer to obtain the decoded feature map generated by the first decoding layer, so that the decoded feature map has both semantic and detailed features.
[0109] Next, it is determined whether the current second generation layer is the last layer (i.e., the last layer in the second layer sequence). Since the current second generation layer is the first layer, it is not the last layer. Therefore, the decoded feature map generated by the first decoding layer is re-inputted into the next layer in the decoder, namely the second decoding layer, to continue participating in feature extraction in subsequent decoding layers. Then, the second layer is updated to the current second generation layer. The decoded feature map generated by the first decoding layer, along with the main feature map output from the third bottleneck residual layer corresponding to the second decoding layer within the residual layer structure of the main branch in the dual-branch encoder, are jointly input into the second decoding layer. The second decoding layer first upsamples the decoded feature map generated by the first decoding layer through a transposed convolutional layer to double the spatial size of the decoded feature map generated by the first decoding layer, resulting in the upsampled feature map generated by the second decoding layer. Simultaneously, the second decoding layer performs resampled feature integration, concatenating the upsampled feature map with the main feature map output by the bottleneck residual layer of the third layer corresponding to the second decoding layer within the residual layer structure of the main branch of the dual-branch encoder along the channel dimension, resulting in the concatenated feature map generated by the second decoding layer. Then, the second decoding layer fuses its generated concatenated feature map through a convolutional layer to obtain the decoded feature map generated by the second decoding layer. This process continues until the fourth decoding layer (the last layer in the second-level sequence) generates the decoded feature map. Here, the decoded feature map generated by the fourth decoding layer in the decoder is used as the target decoded feature map.
[0110] This embodiment first uses stepwise upsampling to progressively amplify the low-resolution deep semantic features (including category information) output by the dual-branch encoder (each step uses transposed convolution to increase resolution by a factor of 1). After each upsampling step, high-resolution shallow detail features (such as edges and textures) from the corresponding stage of the dual-branch encoder are brought in via multi-scale skip connections. Then, through resampling feature integration, the skipped shallow features (scaled if the size does not match) are concatenated / added with the upsampled deep features, and then fused using convolution to obtain features that retain both semantics and detail. This stepwise process of "upsampling → skip connections → integration" is repeated until the feature map is restored to the size of the input image. This approach avoids information loss caused by one-time amplification through stepwise upsampling, and through multi-scale skip connections and resampling integration, ensures that the decoder can always utilize full-scale information from shallow to high levels, ultimately outputting high-precision probability prediction results.
[0111] Finally, a 1×1 convolutional layer is used to compress the number of channels in the target decoded feature map to 1 (a preset value), so that each pixel in the target decoded feature map corresponds to a scalar score, resulting in the original score map. Then, a sigmoid activation function is used to map the score of each pixel in the original score map to the [0, 1] interval, obtaining the probability that each pixel in the visible light image and thermal infrared image predicted by the decoder belongs to the ground fissure category, ultimately generating a binary probability map. This binary probability map has the same size as the input visible light image and can be directly used to determine the location of the ground fissure. A position with a probability of 1 in the binary probability map indicates that the ground fissure identification model predicts that the geographical location within the corresponding target coal mining area is a ground fissure, while a position with a probability of 0 in the binary probability map indicates that the ground fissure identification model predicts that the geographical location within the corresponding target coal mining area is the background.
[0112] Step 204: Calculate the evaluation index of the trained ground fissure identification model.
[0113] It should be noted that, due to the extremely low pixel ratio of ground fissures, there is a serious class imbalance problem. Traditional evaluation metrics (such as accuracy) are easily affected by this problem and cannot objectively reflect the model's performance in identifying fissures. Therefore, this embodiment calculates multiple evaluation metrics based on the predicted ground fissure results (i.e., binary probability maps) of the ground fissure identification model and the actual ground fissure results corresponding to the labeled images in the ground fissure dataset.
[0114] For example, the evaluation indicators are expressed as follows:
[0115] ,
[0116] ,
[0117] ,
[0118] ,
[0119] ,
[0120] ,
[0121] Among them, TP (true positive) refers to the number of pixels that are actually ground fissures and are correctly predicted as ground fissures by the ground fissure identification model; TN (true negative) refers to the number of pixels that are actually not ground fissures and are correctly predicted as not ground fissures by the ground fissure identification model; FP (false positive) refers to the number of pixels that are actually not ground fissures but are incorrectly predicted as ground fissures by the ground fissure identification model; and FN (false negative) refers to the number of pixels that are actually ground fissures but are missed (not predicted as ground fissures) in the prediction by the ground fissure identification model. To improve recall, it can effectively avoid the problem of missing ground fissures in low-proportion scenarios and ensure a reasonable assessment of the ground fissure identification model's ability to capture hidden and small fissures. Intersection over Union (IUCN) is the crossover ratio used to quantify the degree of overlap between the predicted crack region and the actual crack region. The coefficient is used to measure the similarity between the prediction results of the ground fissure identification model and the actual results. The Kappa coefficient is used to eliminate random interference, quantify the non-accidental consistency of the ground fissure identification model's predictions, evaluate the true classification consistency of the ground fissure identification model, and comprehensively ensure the reliability of the evaluation results.
[0122] Step 205: Using the trained ground fissure recognition model, ground fissures are identified in the target coal mining area using the visible light remote sensing image and the thermal infrared remote sensing image, resulting in a ground fissure recognition map.
[0123] In this step, the trained ground fissure recognition model outputs a binary probability map of the same size as the input visible light remote sensing image and thermal infrared remote sensing image of the target coal mining area. The binary probability map includes the probability that each pixel in both images belongs to the ground fissure category. Positions with a probability of 1 in the binary probability map indicate that the model predicts the corresponding geographical location within the target coal mining area to be a ground fissure, while positions with a probability of 0 indicate that the model predicts the corresponding geographical location within the target coal mining area to be background. Thus, the binary probability map of the two images is used as the ground fissure recognition result image.
[0124] In yet another embodiment of this application, Figure 4This study demonstrates the differences in grayscale profiles between visible light remote sensing imagery and thermal infrared remote sensing imagery in the fracture region. Based on the characteristic differences caused by the imaging principles, combining the two images can provide multi-dimensional information support for the accurate identification of ground fissures in the Loess Plateau coal mining area, thereby effectively improving the accuracy of ground fissure identification.
[0125] This embodiment uses UAV photogrammetry to collect 3.02 GB of original visible light and thermal infrared remote sensing images of ground fissures in the Loess Plateau coal mining area. The images were acquired in summer, during the growing season of natural vegetation and crops. Some vegetation and farmland mulch film obscured the fissure area, and the complex background of vegetation and gravel caused the edges of the fissures to be blurred.
[0126] By creating the dataset, a ground fissure dataset for the Loess Plateau coal mining area was obtained, containing 1820 sample groups. Each sample includes one visible light image, one thermal infrared image, and one label for the same geographical area. After 200 rounds of training, the ground fissure recognition model in this embodiment can effectively identify fissures, and the overall accuracy is shown in Table 1.
[0127] Table 1
[0128]
[0129] This embodiment selects four representative images (a, b, c, and d) with different fissure widths and backgrounds. The recognition results of the fissure recognition model in this embodiment are as follows: Figure 5 As shown. Figure 6 The confusion matrix shows that the ground fissure identification model in this embodiment has an accuracy rate of over 90% on representative results, with the best performance reaching 99.2%.
[0130] Through ablation analysis, the ground fissure identification model in this embodiment shows significant performance improvements compared to single-modal methods, methods without adaptive scaling fusion, and methods without Transformer encoding modules. The results are as follows: Figure 7 As shown in Table 2. Specific accuracy indicators are shown in Table 2.
[0131] Table 2
[0132]
[0133] Through comparative analysis, the ground fissure identification model in this embodiment shows significant performance improvements compared to classic methods in semantic segmentation such as UNet and DeepLabV3+, as well as advanced methods such as TransUNet and SwinUNet. The results are as follows: Figure 8 As shown in Table 3. Specific accuracy indicators are shown in Table 3.
[0134] Table 3
[0135]
[0136] This embodiment selected typical scene images with complex backgrounds, blurred edges, vegetation obstruction, and mulch film coverage. The recognition results are as follows: Figure 9 As shown, the ground fissure recognition model in this embodiment has an accuracy rate of over 84% in typical scenarios, with the best performance reaching 99.4%.
[0137] In summary, the innovative method proposed in this embodiment effectively achieves accurate identification of ground fissures in the coal mining areas of the Loess Plateau. Its representative results and overall performance are superior to the current mainstream fissure identification methods.
[0138] It should be noted that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0139] Furthermore, such as Figure 10 As shown, as a specific implementation of the above-mentioned method for identifying ground fissures in coal mining areas, this application provides a ground fissure identification device 1000 for coal mining areas. The ground fissure identification device 1000 for coal mining areas includes: an acquisition module 1001, a training module 1002, and an identification module 1003.
[0140] Among them, the acquisition module 1001 is used to acquire visible light remote sensing images and thermal infrared remote sensing images of the target coal mining area.
[0141] Training module 1002 is used to construct a ground fissure dataset based on visible light remote sensing images and thermal infrared remote sensing images; and to construct a ground fissure recognition model based on a neural network, the ground fissure recognition model including a dual-branch encoder, a feature extraction module and a decoder, the dual-branch encoder including a feature fusion module; and to iteratively train the ground fissure recognition model based on the ground fissure dataset.
[0142] The identification module 1003 is used to identify ground fissures in the target coal mining area using the trained ground fissure identification model and the target visible light remote sensing image and the target thermal infrared remote sensing image, and obtain the ground fissure identification result map.
[0143] The acquisition module 1001 is specifically used to preprocess visible light remote sensing images and thermal infrared remote sensing images to obtain visible light orthophoto images and thermal infrared orthophoto images respectively; to annotate ground fissures on the visible light remote sensing images based on the thermal infrared orthophoto images to obtain label images; to segment the visible light orthophoto images, thermal infrared orthophoto images and label images respectively to obtain visible light images, thermal infrared images and label images; to construct sample groups based on the visible light images, thermal infrared images and label images corresponding to the same geographical area in the target coal mining area; and to construct a ground fissure dataset based on the sample groups.
[0144] The training module 1002 is specifically used to input visible light images and thermal infrared images from the sample groups in the ground fissure dataset into the main branch and auxiliary branch of the dual-branch encoder, respectively, to obtain the fused feature map generated by the dual-branch encoder; input the target fused feature map from the fused feature map into the feature extraction module to obtain the target feature map generated by the feature extraction module; input the main feature map generated by the main branch and the target feature map into the decoder to obtain the binary probability map generated by the decoder, which includes the probability that the pixels in the visible light images and thermal infrared images belong to the ground fissure category as predicted by the decoder; determine the loss function of the ground fissure recognition model based on the binary probability map and the label images in the sample groups; calculate the loss function and update the model parameters of the ground fissure recognition model until the loss function converges.
[0145] Training module 1002 is specifically used to input visible light images and thermal infrared images into the combination structures in the main branch and auxiliary branch, respectively, to obtain the initial visible light feature map generated by the combination structure in the main branch and the initial thermal infrared feature map generated by the combination structure in the auxiliary branch; input the initial visible light feature map and the initial thermal infrared feature map into the first spatial dimension alignment structure in the main branch and the second spatial dimension alignment structure in the auxiliary branch, respectively, to obtain the visible light feature map generated by the first spatial dimension alignment structure in the main branch and the thermal infrared feature map with the same spatial scale as the visible light feature map generated by the second spatial dimension alignment structure in the auxiliary branch; input the visible light feature map and the thermal infrared feature map into the residual layer structures in the main branch and the auxiliary branch, respectively, to obtain the main feature map generated by the bottleneck residual layers at different levels in the residual layer structure in the main branch and the auxiliary feature map generated by the bottleneck residual layers at different levels in the residual layer structure in the auxiliary branch; and generate a fused feature map based on the main feature map and the auxiliary feature map generated by the bottleneck residual layers at the same level in the residual layer structures in the main branch and the auxiliary branch.
[0146] The training module 1002 is specifically used to combine structures including 7×7 convolutional layers, batch normalization layers and linear rectified activation functions. The first spatial dimension aligned structure includes a 3×3 max pooling downsampling layer, and the second spatial dimension aligned structure includes a 3×3 max pooling downsampling layer and a bilinear interpolation upsampling layer.
[0147] Training module 1002 is specifically used to sort the levels of the bottleneck residual layer to obtain the first level order, and to take the first level in the first level order as the first generation level; the main feature map and auxiliary feature map generated by the bottleneck residual layer of the first generation level in the main branch and auxiliary branch are input into the feature fusion module connected to the bottleneck residual layer of the first generation level in the main branch to obtain the fused feature map generated by the feature fusion module connected to the bottleneck residual layer of the first generation level in the main branch; if the first generation level is not the last level in the first level order, the fused feature map generated by the feature fusion module connected to the bottleneck residual layer of the first generation level in the main branch is used. Input the bottleneck residual layer of the next level of the first generation level in the main branch, and update the first generation level according to the next level of the first generation level. The next level of the first generation level is the level located one level below the first generation level in the first level sequence. Input the main feature map and auxiliary feature map generated by the updated bottleneck residual layer of the first generation level in the main branch and auxiliary branch into the feature fusion module connected to the updated bottleneck residual layer of the first generation level in the main branch, and obtain the fused feature map generated by the feature fusion module connected to the updated bottleneck residual layer of the first generation level in the main branch, until the updated first generation level is the last level in the first level sequence.
[0148] Training module 1002 is specifically used to obtain the fused feature map according to the following formula:
[0149] ,
[0150] ,
[0151] ,
[0152] in, The learnable parameters for the current iteration of the ground fissure identification model are determined based on the initial values of the learnable parameters. It is a nonlinear mapping function; The main feature map is input to the feature fusion module during the current iteration training of the ground fissure recognition model. The weights; Auxiliary feature maps for the input feature fusion module; To fuse feature maps; The loss function for the current iteration training of the ground fissure identification model Regarding learnable parameters The gradient; The learning rate for the ground fissure identification model; These are the learnable parameters for the next iteration of the ground fissure identification model training.
[0153] Training module 1002 is specifically used to globally flatten the target fusion feature map into a one-dimensional feature sequence. The target fusion feature map is a fusion feature map generated by the feature fusion module, which connects the bottleneck residual layer at the end of the first-level sequence in the main branch. The first-level sequence is determined according to the level of the bottleneck residual layer. The position encoding matrix of the one-dimensional feature sequence is calculated using a sine-cosine function. Based on the one-dimensional feature sequence and the position encoding matrix, a feature vector is constructed. The feature vector is then sequentially input into a Transformer layer to obtain an enhanced feature tensor. The Transformer layer includes layer normalization, multi-head attention mechanism, and multilayer perceptron. The enhanced feature tensor is then reshaped to obtain the target feature map.
[0154] Training module 1002 is specifically used to sort the levels of the decoding layer to obtain the second level order, and to take the first level in the second level order as the second generation level; the target feature map and the main feature map generated by the bottleneck residual layer corresponding to the second generation level in the residual structure of the main branch are input into the decoding layer of the second generation level to obtain the decoding feature map generated by the decoding layer of the second generation level. The levels of the decoding layer in the decoder correspond to the levels of the bottleneck residual layer in the main branch in reverse order; if the second generation level is not the last level in the second level order, the decoding feature map generated by the decoding layer of the second generation level is input into the decoding layer of the next level of the second generation level in the decoder. The second generation level is updated according to the next level of the second generation level, which is the level that is one level below the second generation level in the second generation level sequence. The main feature map generated by the bottleneck residual layer corresponding to the updated second generation level in the main branch is input into the decoding layer of the updated second generation level to obtain the decoding feature map generated by the decoding layer of the updated second generation level, until the updated second generation level is the last level in the second generation level sequence. The number of channels of the decoding feature map generated by the decoding layer of the last level in the second generation level sequence is compressed to a preset value to generate the original score map. A binary probability map is generated based on the original score map using a nonlinear mapping function.
[0155] Specific limitations regarding the coal mine area ground fissure identification device can be found in the limitations of the coal mine area ground fissure identification method described above, and will not be repeated here. Each module in the aforementioned coal mine area ground fissure identification device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0156] Based on the above, Figures 1 to 2Accordingly, embodiments of this application also provide a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. Figures 1 to 2 The method for identifying ground fissures in coal mining areas is shown.
[0157] Based on this understanding, the technical solution of this application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, or portable hard drive), and includes several instructions to cause a computer device (such as a personal computer, server, or network device) to execute the methods described in the various implementation scenarios of this application.
[0158] Based on the above, Figures 1 to 2 The method shown, and Figure 10 To achieve the above objectives, the present application also provides a computer device, specifically a personal computer, server, network device, etc., as shown in the virtual device embodiment. This computer device includes a storage medium and a processor; the storage medium stores a computer program; the processor executes the computer program to achieve the above-described objectives. Figures 1 to 2 The method for identifying ground fissures in coal mining areas is shown.
[0159] Optionally, the computer device may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB ports, card reader ports, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Bluetooth interfaces, Wi-Fi interfaces), etc.
[0160] Those skilled in the art will understand that the computer device structure provided in this embodiment does not constitute a limitation on the computer device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0161] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages and stores the hardware and software resources of a computer device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software within the physical device.
[0162] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platform, or the embodiments of this application can be implemented by hardware.
[0163] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application. Those skilled in the art will understand that the modules in the apparatus of the embodiment can be distributed within the apparatus of the embodiment as described, or can be modified to be located in one or more apparatuses different from this embodiment. The modules of the above-described embodiment can be combined into one module, or further divided into multiple sub-modules.
[0164] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of any particular implementation scenario. The above disclosures are merely a few specific implementation scenarios of this application; however, this application is not limited thereto, and any variations conceived by those skilled in the art should fall within the protection scope of this application.
Claims
1. A method for identifying ground fissures in coal mining areas, characterized in that, The method includes: Acquire visible light and thermal infrared remote sensing images of the target coal mining area; Based on the visible light remote sensing image and the thermal infrared remote sensing image, a ground fissure dataset is constructed; A ground fissure identification model is constructed based on a neural network. The ground fissure identification model includes a dual-branch encoder, a feature extraction module, and a decoder. The dual-branch encoder includes a feature fusion module. The ground fissure identification model is iteratively trained based on the ground fissure dataset. The trained ground fissure identification model is used to identify ground fissures in the target coal mining area using both visible light remote sensing images and thermal infrared remote sensing images, resulting in a ground fissure identification map. Specifically, the iterative training of the ground fissure identification model based on the ground fissure dataset includes: The visible light image and thermal infrared image of the sample group in the ground fissure dataset are respectively input into the main branch and auxiliary branch of the dual-branch encoder to obtain the fused feature map generated by the dual-branch encoder. The target fused feature map in the fused feature map is input into the feature extraction module to obtain the target feature map generated by the feature extraction module; The main feature map generated by the main branch and the target feature map are input into the decoder to obtain a binary probability map generated by the decoder. The binary probability map includes the probability that the pixels in the visible light image and the thermal infrared image belong to the ground fissure category as predicted by the decoder. Based on the binary probability map and the label images within the sample group, the loss function of the ground fissure identification model is determined; Calculate the loss function and update the model parameters of the ground fissure identification model until the loss function converges; Accordingly, the main branch includes a residual layer structure, which includes multiple levels of bottleneck residual layers. Each bottleneck residual layer in the main branch is connected to a feature fusion module. The feature extraction module includes multiple Transformer layers. The step of inputting the target fused feature map from the fused feature map into the feature extraction module to obtain the target feature map generated by the feature extraction module specifically includes: The target fusion feature map is globally flattened into a one-dimensional feature sequence. The target fusion feature map is the fusion feature map generated by the feature fusion module that connects to the bottleneck residual layer at the end of the first level order in the main branch. The first level order is determined according to the level of the bottleneck residual layer. The position encoding matrix of the one-dimensional feature sequence is calculated using a sine-cosine function; Based on the one-dimensional feature sequence and the position encoding matrix, a feature vector is constructed; The feature vectors are sequentially input into the Transformer layer to obtain the enhanced feature tensor. The Transformer layer includes layer normalization, multi-head attention mechanism and multilayer perceptron. The enhanced feature tensor is reshaped to obtain the target feature map.
2. The method for identifying ground fissures in coal mining areas according to claim 1, characterized in that, The construction of the ground fissure dataset based on the visible light remote sensing image and the thermal infrared remote sensing image specifically includes: The visible light remote sensing image and the thermal infrared remote sensing image are preprocessed respectively to obtain a visible light orthophoto image and a thermal infrared orthophoto image. Based on the thermal infrared orthophoto image, the visible light remote sensing image is labeled with ground fissures to obtain a labeled image; The visible light orthophoto image, the thermal infrared orthophoto image, and the tag image are segmented to obtain a visible light image, a thermal infrared image, and a tag image; Based on the visible light images, thermal infrared images, and tag images corresponding to the same geographical area in the target coal mining area, a sample group is constructed. Based on the sample group, the ground fissure dataset is constructed.
3. The method for identifying ground fissures in coal mining areas according to claim 1, characterized in that, The main branch includes a combined structure, a first spatial dimension alignment structure, and a residual layer structure. The auxiliary branch includes a combined structure, a second spatial dimension alignment structure, and a residual layer structure. The residual structure includes multiple levels of bottleneck residual layers. The visible light images and thermal infrared images from the sample groups of the ground fissure dataset are respectively input into the main branch and auxiliary branch of the dual-branch encoder to obtain the fused feature map generated by the dual-branch encoder, specifically including: The visible light image and the thermal infrared image are respectively input into the combination structure in the main branch and the auxiliary branch to obtain the initial visible light feature map generated by the combination structure in the main branch and the initial thermal infrared feature map generated by the combination structure in the auxiliary branch. The initial visible light feature map and the initial thermal infrared feature map are respectively input into the first spatial dimension alignment structure in the main branch and the second spatial dimension alignment structure in the auxiliary branch to obtain the visible light feature map generated by the first spatial dimension alignment structure in the main branch, and the thermal infrared feature map with the same spatial scale as the visible light feature map generated by the second spatial dimension alignment structure in the auxiliary branch. The visible light feature map and the thermal infrared feature map are respectively input into the residual layer structure in the main branch and the auxiliary branch to obtain the main feature map generated by the bottleneck residual layer at different levels in the residual layer structure in the main branch and the auxiliary feature map generated by the bottleneck residual layer at different levels in the residual layer structure in the auxiliary branch. The fused feature map is generated based on the main feature map and the auxiliary feature map generated from the bottleneck residual layer at the same level in the residual layer structure of the main branch and the auxiliary branch.
4. The method for identifying ground fissures in coal mining areas according to claim 3, characterized in that, The combined structure includes a 7×7 convolutional layer, a batch normalization layer, and a linear rectified activation function. The first spatial dimension alignment structure includes a 3×3 max pooling downsampling layer, and the second spatial dimension alignment structure includes a 3×3 max pooling downsampling layer and a bilinear interpolation upsampling layer.
5. The method for identifying ground fissures in coal mining areas according to claim 3, characterized in that, The feature fusion module is connected after each bottleneck residual layer in the residual layer structure of the main branch. The generation of the fused feature map, based on the main feature map and auxiliary feature map generated from the bottleneck residual layers of the same level in the residual layer structures of the main branch and the auxiliary branch, specifically includes: The levels of the bottleneck residual layer are sorted to obtain a first level order, and the first level in the first level order is taken as the first generation level. The main feature map and the auxiliary feature map generated by the bottleneck residual layer of the first generation level in the main branch and the auxiliary branch are input into the feature fusion module connected to the bottleneck residual layer of the first generation level in the main branch to obtain the fused feature map generated by the feature fusion module connected to the bottleneck residual layer of the first generation level in the main branch. If the first generation level is not the last level in the first generation order, the fusion feature map generated by the feature fusion module connected to the bottleneck residual layer of the first generation level in the main branch is input into the bottleneck residual layer of the next level of the first generation level in the main branch, and the first generation level is updated according to the next level of the first generation level, where the next level of the first generation level is the level located one position below the first generation level in the first generation order. The main feature map and the auxiliary feature map generated by the updated bottleneck residual layer of the first generation level in the main branch and the auxiliary branch are input into the feature fusion module connected to the updated bottleneck residual layer of the first generation level in the main branch to obtain the fused feature map generated by the feature fusion module connected to the updated bottleneck residual layer of the first generation level in the main branch, until the updated first generation level is the last level in the first level sequence.
6. The method for identifying ground fissures in coal mining areas according to claim 5, characterized in that, In the feature fusion module, the fused feature map is obtained according to the following formula: , , , in, The learnable parameters for the current iteration training of the ground fissure identification model are determined based on the initial values of the learnable parameters. It is a nonlinear mapping function; The main feature map of the feature fusion module is input into the current iteration training of the ground fissure identification model. The weights; The auxiliary feature map is used as input to the feature fusion module; The fused feature map; The loss function for the current iteration training of the ground fissure identification model. Regarding the learnable parameters The gradient; The learning rate for the ground fissure identification model; These are the learnable parameters for the next iteration of the ground fissure identification model.
7. The method for identifying ground fissures in coal mining areas according to claim 1, characterized in that, The main branch includes a residual layer structure, which comprises multiple levels of bottleneck residual layers. Each bottleneck residual layer in the main branch is connected to a feature fusion module. The decoder includes multiple levels of decoding layers. The dual-branch encoder and the decoder are connected in a skip connection. The step of inputting the main feature map generated by the main branch and the target feature map into the decoder to obtain a binary probability map generated by the decoder specifically includes: The levels of the decoding layer are sorted to obtain a second level order, and the first level in the second level order is taken as the second generation level. The target feature map and the main feature map generated by the bottleneck residual layer corresponding to the second generation level in the residual structure of the main branch are input into the decoding layer of the second generation level to obtain the decoding feature map generated by the decoding layer of the second generation level. The levels of the decoding layer in the decoder correspond to the levels of the bottleneck residual layer in the main branch in reverse order. If the second generation level is not the last level in the second generation level sequence, the decoding feature map generated by the decoding layer of the second generation level is input into the decoding layer of the next level of the second generation level in the decoder, and the second generation level is updated according to the next level of the second generation level, which is the level located one position below the second generation level in the second generation level sequence. The main feature map generated by the bottleneck residual layer corresponding to the updated second generation level in the main branch is input into the decoding layer of the updated second generation level to obtain the decoding feature map generated by the decoding layer of the updated second generation level, until the updated second generation level is the last level in the second level sequence. The number of channels in the decoded feature map generated by the decoding layer of the last layer in the second level sequence is compressed to a preset value to generate the original score map; The binary probability map is generated based on the original fractional map using a nonlinear mapping function.
8. A device for identifying ground fissures in coal mining areas, characterized in that, The device includes: The acquisition module is used to acquire visible light remote sensing images and thermal infrared remote sensing images of the target coal mining area; The training module is used to construct a ground fissure dataset based on the visible light remote sensing image and the thermal infrared remote sensing image; and to construct a ground fissure recognition model based on a neural network, wherein the ground fissure recognition model includes a dual-branch encoder, a feature extraction module and a decoder, and the dual-branch encoder includes a feature fusion module; and to iteratively train the ground fissure recognition model according to the ground fissure dataset. The identification module is used to identify ground fissures in the target coal mining area using the trained ground fissure identification model and the target visible light remote sensing image and the target thermal infrared remote sensing image, and to obtain a ground fissure identification result map. Specifically, the training module is used for: The visible light image and thermal infrared image of the sample group in the ground fissure dataset are respectively input into the main branch and auxiliary branch of the dual-branch encoder to obtain the fused feature map generated by the dual-branch encoder. The target fused feature map in the fused feature map is input into the feature extraction module to obtain the target feature map generated by the feature extraction module; The main feature map generated by the main branch and the target feature map are input into the decoder to obtain a binary probability map generated by the decoder. The binary probability map includes the probability that the pixels in the visible light image and the thermal infrared image belong to the ground fissure category as predicted by the decoder. Based on the binary probability map and the label images within the sample group, the loss function of the ground fissure identification model is determined; Calculate the loss function and update the model parameters of the ground fissure identification model until the loss function converges; Accordingly, the main branch includes a residual layer structure, which includes multiple levels of bottleneck residual layers. Each bottleneck residual layer in the main branch is connected to a feature fusion module. The feature extraction module includes multiple Transformer layers. The training module is specifically used for: The target fusion feature map is globally flattened into a one-dimensional feature sequence. The target fusion feature map is the fusion feature map generated by the feature fusion module that connects to the bottleneck residual layer at the end of the first level order in the main branch. The first level order is determined according to the level of the bottleneck residual layer. The position encoding matrix of the one-dimensional feature sequence is calculated using a sine-cosine function; Based on the one-dimensional feature sequence and the position encoding matrix, a feature vector is constructed; The feature vectors are sequentially input into the Transformer layer to obtain the enhanced feature tensor. The Transformer layer includes layer normalization, multi-head attention mechanism and multilayer perceptron. The enhanced feature tensor is reshaped to obtain the target feature map.
Citation Information
Patent Citations
Mine surface damage crack extraction method and device, electronic equipment and medium
CN117788811A
Infrared and visible light image fusion method and device
CN117830118A