Intelligent interpretation method for coal mine power supply line geological disaster hidden danger based on multi-source satellite remote sensing image

By employing a cross-modal fusion deep learning method based on multi-source satellite remote sensing imagery, combined with high-resolution optical remote sensing and synthetic aperture radar data, the problems of low identification efficiency and insufficient accuracy in existing technologies have been solved. This has enabled full coverage and efficient, all-day interpretation of geological hazards along coal mine power supply lines, improving identification accuracy and inspection efficiency.

CN122135237APending Publication Date: 2026-06-02GUIZHOU COAL MINE DESIGN & RES INST +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUIZHOU COAL MINE DESIGN & RES INST
Filing Date
2026-05-08
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies for identifying geological hazards in coal mine power supply lines suffer from problems such as poor cross-domain model transfer adaptability, insufficient depth of heterogeneous data fusion, difficulty in balancing computing power and accuracy, poor local noise suppression, and mismatch in real-time performance. These issues result in low identification efficiency and insufficient accuracy, failing to meet the needs of routine monitoring.

Method used

A cross-modal fusion deep learning method based on multi-source satellite remote sensing imagery is adopted, which combines high-resolution optical remote sensing and synthetic aperture radar data. Through a dual-stream UNet architecture, a lightweight cross-modal Transformer, and a spatially perceptive SE denoising branch, pixel-level intelligent interpretation and risk classification of geological disaster hazards are achieved.

Benefits of technology

It achieves full coverage and efficient, 24/7 interpretation of geological hazards in coal mine power supply lines, improves inspection efficiency, reduces operation and maintenance costs, ensures the stability and timeliness of interpretation results, improves the accuracy of small-scale hazard identification, and solves many defects in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135237A_ABST
    Figure CN122135237A_ABST
Patent Text Reader

Abstract

This invention relates to the interdisciplinary field of remote sensing geological disaster identification, power facility safety operation and maintenance, and computer vision. It discloses an intelligent interpretation method for geological disaster hazards in coal mine power supply lines based on multi-source satellite remote sensing imagery, comprising: S1, acquiring and standardizing preprocessed optical remote sensing images and synthetic aperture radar (SAR) remote sensing images of the study area of ​​the coal mine power supply line to obtain optical feature sets and SAR feature sets; S2, constructing a multi-source geological disaster feature set specific to the coal mine power supply line based on the optical and SAR feature sets, and creating a multi-modal sample library; S3, constructing and training a cross-modal fusion deep learning model; S4, inputting the optical and SAR feature sets of the study area to be interpreted into the trained cross-modal fusion deep learning model, and outputting the interpretation results of geological disaster hazards within the line corridor. This method achieves pixel-level intelligent interpretation and risk classification of geological disaster hazards such as subsidence areas and landslides within the coal mine power supply line corridor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary fields of remote sensing geological disaster identification, power facility safety operation and maintenance, and computer vision, and particularly to an intelligent interpretation method for geological disaster hazards in coal mine power supply lines based on multi-source satellite remote sensing images. Background Technology

[0002] Coal mine power supply lines often traverse mining-affected areas and mountainous and hilly regions. Geological disasters such as mining subsidence and landslides can easily damage tower foundations and the line itself, posing a significant threat to coal mine safety. Existing multi-source remote sensing image fusion and interpretation technologies for identifying geological hazards in coal mine power supply lines suffer from problems such as poor cross-domain model transfer adaptability, insufficient depth of heterogeneous data fusion, difficulty in balancing computational power and accuracy, and poor local noise suppression. Furthermore, the inherent limitations of traditional inspection and single-sensor interpretation remain unresolved, as detailed below:

[0003] (1) The core limitations of traditional inspection and single remote sensing: manual on-site inspection is inefficient and covers dangerous areas. Single optical remote sensing cannot monitor the effects of light and fog around the clock. Single SAR remote sensing lacks detailed information about the ground surface, resulting in a high false alarm rate. Moreover, the existing model has not been optimized specifically for small-scale hidden dangers in coal mine power supply line corridors and high-risk areas around towers.

[0004] (2) Architectural contradictions in cross-domain model transfer: The existing method of rigidly transferring one-dimensional time series signal fusion model to two-dimensional remote sensing image segmentation task has a topological paradox between classification and segmentation - the global category probability (Logits vector) after dimensionality reduction is directly used for high-resolution pixel-level segmentation, and the hidden contours cannot be recovered after the spatial dimension information is lost; at the same time, the logic of "left and right limb asymmetry" is incorrectly applied to the fusion of optical and SAR homogeneous heterogeneous data, which violates the physical mechanism of geospatial observation.

[0005] (3) Imbalance between computing power and performance in cross-modal fusion: When the traditional Transformer cross-modal attention mechanism is directly applied to large-size remote sensing images, there is a problem of memory complexity explosion. The O(N2) complexity of the global attention matrix makes it impossible for existing hardware to support it. The receptive field of the simple CNN fusion architecture is limited and cannot capture the long-distance spatial correlation between optical texture and SAR deformation features.

[0006] (4) Insufficient depth and specificity of multi-source data fusion: Existing technologies are mostly shallow feature stitching of optical and SAR data, without designing modal-specific feature extraction paths for the characteristic laws of coal mine geological disasters. At the same time, there is a lack of comprehensive understanding of the global information of the image, resulting in insufficient exploration of the complementarity of heterogeneous features.

[0007] (5) Lack of spatial adaptability of noise suppression: Traditional noise reduction modules based on global average pooling (GAP) can only achieve channel-level global weight adjustment, and cannot accurately suppress local spatial noise such as cloud shadows and sensor noise, which affects the accuracy of hidden danger identification in complex backgrounds of mining areas.

[0008] (6) Performance indicators are out of touch with actual engineering: Existing related solutions contain false inference performance descriptions and fail to design a reasonable model architecture that takes into account the actual characteristics of large size and high dimension of remote sensing images, resulting in the technical solutions being unreproducible and unimplementable on existing general-purpose hardware.

[0009] (7) Mismatch between the real-time nature of SAR deformation data and interpretation requirements: Existing InSAR technologies (such as SBAS-InSAR) require offline processing of multiple historical data to generate deformation rate maps, which cannot be "instantly interpreted" like optical images. Furthermore, the fusion logic of offline and online features is not clearly defined, making it difficult to meet the operation and maintenance requirements of normalized and near real-time monitoring of hidden dangers in coal mine power supply lines.

[0010] Therefore, there is an urgent need for an intelligent interpretation method for geological disaster hazards that can deeply integrate multi-source remote sensing data, adapt to coal mine power supply line scenarios, and balance accuracy and efficiency, so as to meet the actual needs of safe operation and maintenance of coal mine power supply lines. Summary of the Invention

[0011] To address the shortcomings of the existing technologies, the core objective of this invention is to provide an implementable, highly adaptable, high-precision, and near-real-time intelligent interpretation method for geological hazard risks in coal mine power supply lines based on multi-source satellite remote sensing imagery. Specifically, it integrates high-resolution optical remote sensing and synthetic aperture radar (SAR) remote sensing data, and adopts a cross-modal fusion deep learning architecture adapted to two-dimensional remote sensing imagery to achieve pixel-level intelligent interpretation and risk classification of geological hazard risks such as subsidence areas and landslides within the corridors of coal mine power supply lines.

[0012] To achieve the above objectives, the following technical solution is adopted:

[0013] This invention provides an intelligent interpretation method for geological hazard risks in coal mine power supply lines based on multi-source satellite remote sensing imagery, comprising the following steps:

[0014] Step S1: Delineate the study area of ​​the coal mine power supply line, collect and standardize the optical remote sensing images and synthetic aperture radar remote sensing images of the study area to obtain optical feature sets and SAR feature sets.

[0015] Step S2: Based on the optical feature set and SAR feature set, construct a multi-source feature set of geological hazards specific to coal mine power supply lines, and create a multi-modal sample library adapted for pixel-level segmentation.

[0016] Step S3: Construct and train a cross-modal fusion deep learning model adapted to 2D remote sensing images; the cross-modal fusion deep learning model is an end-to-end architecture, used to receive the optical feature set and SAR feature set, and through the built-in encoder-decoder structure, perform multi-source heterogeneous feature extraction, deep fusion and noise suppression, and finally output the interpretation results of geological disaster hazards;

[0017] Step S4: Input the optical feature set and SAR feature set of the study area to be interpreted into the trained cross-modal fusion deep learning model, and output the interpretation results of the geological disaster hazards in the corridor, including pixel-level semantic segmentation map and risk classification map.

[0018] Furthermore, in step S1, delineating the study area includes: obtaining vector ledger data of power supply lines in the target coal mining area, setting up line corridor buffer zones based on voltage level differences, with voltage level and buffer zone width showing a step-like nonlinear positive correlation.

[0019] The standardized preprocessing of the synthetic aperture radar remote sensing imagery includes: employing a hybrid offline and online processing strategy, specifically including:

[0020] In the offline phase, time-series deformation rate maps and cumulative deformation maps are generated using long-term synthetic aperture radar remote sensing image data, serving as prior weight features characterizing long-term deformation trends.

[0021] During the online phase, for newly acquired synthetic aperture radar remote sensing images of a single scene, backscattering coefficient maps and polarization feature maps are quickly extracted as residual perturbation features characterizing recent sudden changes.

[0022] The prior weight features and residual perturbation features are spatiotemporally matched to form a dynamically updated SAR feature set.

[0023] Furthermore, in step S2, the construction of the geological hazard multi-source feature set specific to coal mine power supply lines is an optical-SAR complementary multi-dimensional feature set covering surface details, deformation features, and topographic attributes, including:

[0024] The optical feature subset includes multispectral reflectance, spectral index, texture features, and topographic features;

[0025] The SAR feature subset employs a combined strategy of static background and dynamic instantaneous data. It includes offline extracted temporal deformation rate and cumulative deformation as prior weight features characterizing long-term deformation trends, and online extracted backscattering coefficient and polarization decomposition components as residual perturbation features characterizing recent sudden changes.

[0026] Furthermore, in step S2, the creation of the multimodal sample library adapted for pixel-level segmentation adopts a combined annotation scheme:

[0027] Using typical geological hazard hazards in coal mining areas as positive samples and stable surfaces in the study area as negative samples, the outline boundary, type and scale of each hazard are marked by pixel-level polygon annotation, and the corresponding optical image blocks and SAR feature blocks are matched to form a one-to-one multimodal sample pair.

[0028] The sample data sources include historical geological disaster records of coal mines, on-site inspection records, results of manual visual interpretation, and mining subsidence monitoring data.

[0029] A combined annotation scheme of weakly supervised pre-annotation, semi-supervised optimization, and manual refinement was used to annotate the samples;

[0030] A two-dimensional remote sensing-specific data augmentation strategy was adopted to expand the sample library, and the sample library was divided into training set, validation set and test set according to a preset ratio.

[0031] Furthermore, in step S3, the cross-modal fusion deep learning model adopts an overall architecture of dual-stream input, dual-encoder fusion, and single-decoder output; the cross-modal fusion deep learning model specifically includes: a multi-scale feature extraction module, dual encoders, a lightweight cross-modal bidirectional block Transformer fusion module, a dynamic decision fusion module, a spatially aware SE denoising branch, and a segmentation and risk classification dual-task decoding module;

[0032] The cross-modal fusion deep learning model is based on a dual-stream UNet. It extracts multi-scale features from optical images and SAR features through a multi-scale feature extraction module, extracts modality-specific features and global fusion features through a dual encoder, achieves long-distance spatial correlation modeling of heterogeneous features through a lightweight cross-modal bidirectional block Transformer fusion module, achieves adaptive fusion of multi-scale features through a dynamic decision fusion module, suppresses local noise through a spatially aware SE denoising branch, and finally outputs pixel-level semantic segmentation maps and risk classification maps through a segmentation and risk classification dual-task decoding module.

[0033] Furthermore, the multi-scale feature extraction module includes: a convolution channel composed of convolution kernels of various sizes, used to extract features from optical input and SAR input respectively, and capture hidden danger features at different scales;

[0034] Each convolutional channel comprises a combination of dilated convolution, point convolution, activation function, and batch normalization. The output features of each convolutional channel are concatenated along the channel dimension to obtain optical multi-scale fusion features and SAR multi-scale fusion features.

[0035] Furthermore, the dual encoder includes a multimodal encoder and an auxiliary encoder;

[0036] The multimodal encoder has a dual-stream structure, corresponding to an optical branch and a SAR branch respectively. The optical branch and the SAR branch receive the multi-scale fusion features of the optical mode and the multi-scale fusion features of the SAR mode respectively, which are used to extract optical mode-specific features and SAR mode-specific features. A coordinate attention module is embedded at each level to enhance the feature response of small-scale hidden dangers and high-risk areas around the tower, and outputs optical features and SAR features at each level.

[0037] The auxiliary encoder uses DenseNet as its backbone network and takes as input a channel-level stitched image of optical and SAR features to extract and output global fusion features of optical and SAR features.

[0038] The optical features and SAR features output by the multimodal encoder at each level, and the global fusion features output by the auxiliary encoder at the corresponding level, are fused pixel by pixel to obtain multi-scale globally enhanced optical features and multi-scale globally enhanced SAR features.

[0039] Furthermore, the lightweight cross-modal bidirectional block Transformer fusion module is used to perform cross-modal interactive fusion of the multi-scale globally enhanced optical features and the multi-scale globally enhanced SAR features, specifically including:

[0040] The input multi-scale globally enhanced optical features and multi-scale globally enhanced SAR features are simultaneously divided into multiple non-overlapping feature blocks;

[0041] An attention mechanism for optically guided SAR is constructed, which uses the query of optical feature blocks to match the keys and values ​​of SAR feature blocks, and utilizes optical texture information to optimize the spatial contour of SAR features.

[0042] A SAR-guided optical attention mechanism is constructed, which uses the query of SAR feature blocks to match the keys and values ​​of optical feature blocks, and uses SAR deformation information to guide optical features to focus on the core area of ​​potential hazards.

[0043] The bidirectional guided feature blocks are fused to obtain cross-modal fused feature blocks;

[0044] The shallow, middle, and deep features output by the encoder are subjected to the above-mentioned cross-modal bidirectional block fusion, and the fused feature blocks are restored to complete feature maps to obtain multi-scale cross-modal deep fusion features.

[0045] The attention mechanism of optically guided SAR and the attention mechanism of SAR-guided optics are calculated only within the feature block at the corresponding spatial location and between cross-modal blocks.

[0046] Furthermore, the dynamic decision fusion module is used to adaptively weight and fuse the multi-scale cross-modal deep fusion features based on feature confidence, specifically including: calculating pixel-level feature confidence maps for optical and SAR deep fusion features at each scale; assigning adaptive fusion weights to the optical and SAR features at corresponding pixel locations based on the feature confidence maps; introducing learnable global weight parameters to globally fine-tune the pixel-level fused features to obtain dynamically fused features at each scale; and upsampling / downsampling the dynamically fused features at each scale to achieve multi-scale feature splicing and fusion, obtaining the final cross-modal deep fusion features.

[0047] Furthermore, the spatially aware SE denoising branch is used to suppress local spatial noise in the final cross-modal deep fusion features. Specifically, it includes: simultaneously capturing channel-level global feature dependencies and local spatial noise distribution features through hybrid pooling operations; adjusting the spatial distribution of attention weights through spatial convolution after channel attention calculation to specifically suppress feature weights in local noise regions; and using the denoising branch as a residual connection in the model, adding the output of the spatially aware SE denoising branch pixel-by-pixel with the final cross-modal deep fusion features through residual connections to achieve noise suppression while retaining effective features.

[0048] The segmentation and risk classification dual-task decoding module includes: a semantic segmentation head, used to output a pixel-level semantic segmentation map of geological disaster hazards, depicting the outline, range and type of hazards; and a risk classification head, used to extract the quantitative features of each hazard region based on the hazard mask output by the semantic segmentation map, and to classify the hazard region by risk through a multilayer perceptron, outputting a risk classification map with the same size as the input.

[0049] This invention addresses the two-dimensional characteristics of multi-source satellite remote sensing imagery, the processing characteristics of SAR data, and the need for identifying geological hazards along coal mine power supply lines. It constructs an implementable, highly adaptable, high-precision, and near-real-time cross-modal fusion deep learning model, resolving issues such as architectural contradictions, computational power consumption explosion, insufficient fusion depth, poor noise suppression, and real-time mismatch in existing technologies. This enables intelligent interpretation of geological hazard risks along coal mine power supply lines, offering the following significant advantages:

[0050] (1) By integrating high-resolution optical and SAR remote sensing data, this invention overcomes the limitations of single optical data by weather and lighting. Through the hybrid strategy of offline deformation background + online instantaneous features, it achieves full coverage and routine satellite survey of coal mine power supply line corridors. A single scene image can complete the interpretation of hidden dangers of hundreds of kilometers of lines. It breaks through the limitations of traditional inspection and single remote sensing, and achieves efficient interpretation of the whole area and all time. It solves the industry pain point that remote mining areas and dangerous mountainous areas cannot be inspected manually on a routine basis, greatly improving inspection efficiency and reducing operation and maintenance costs.

[0051] (2) This invention uses an offline background library as a priori weight and online features as residual perturbation as a fusion logic. This not only leverages the advantages of SBAS-InSAR high-precision deformation monitoring, but also captures recent sudden changes through online instantaneous features, clarifies the SAR feature fusion logic, solves the problem of balancing real-time performance and accuracy, and specifically addresses the industry pain point of mismatch between InSAR data and interpretation real-time performance, while ensuring the stability and timeliness of the interpretation results.

[0052] (3) The lightweight cross-modal bidirectional block Transformer (LS-BiCrossMT) proposed in this invention defines the bidirectional guidance logic of optical → SAR and SAR → optical through a clear Q / K / V cross-modal mapping mathematical formula. It is different from the isolated processing mode of single-modal block attention and reduces the computing power consumption of global attention to the hardware-bearable range. It realizes long-distance spatial correlation modeling of optical texture and SAR deformation features, fully explores the complementarity of heterogeneous data, and avoids being identified as a "pure algorithm combination".

[0053] (4) This invention improves the interpretation accuracy in complex backgrounds by using a dual encoder + spatial perception noise reduction design. The dual encoder architecture achieves the dual effect of modality-specific feature extraction and global information supplementation, thereby improving the model's ability to understand the complex background of the mining area. The improved spatial perception SE noise reduction branch achieves accurate suppression of local spatial noise such as cloud shadows and sensor noise through hybrid pooling and serial convolution, significantly reducing the false alarm rate of interpretation, and especially improving the recognition accuracy of small-scale hidden dangers (sinkholes, ground fissures).

[0054] (5) This invention designs a dynamic weighted multi-task loss function, which adaptively adjusts the weights of segmentation and classification tasks according to the training phase, dynamically weights multi-task learning, ensures the collaborative convergence of the two tasks, solves the convergence imbalance problem of tasks of different magnitudes, and ensures that the model has both high-precision hidden danger contour recognition and risk classification capabilities.

[0055] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of the present invention, nor is it intended to restrict the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0056] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the invention. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0057] Figure 1 This is a flowchart illustrating the intelligent interpretation method for geological hazard risks in coal mine power supply lines based on multi-source satellite remote sensing imagery, according to an embodiment of the present invention.

[0058] Figure 2 This is a schematic diagram of the overall architecture of the cross-modal fusion deep learning model according to an embodiment of the present invention;

[0059] Figure 3 This is the internal structure of the lightweight cross-modal bidirectional segmented Transformer fusion module in this embodiment of the invention;

[0060] Figure 4 This is a schematic diagram of the spatial perception SE noise reduction branch in an embodiment of the present invention. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0063] This invention presents an intelligent interpretation method for geological hazard hazards in coal mine power supply lines based on multi-source satellite remote sensing imagery. It employs a multi-scale feature extraction and dynamic decision fusion hierarchical architecture, combined with dual encoders and cross-modal Transformer technology, to deeply adapt and optimize the two-dimensional spatial characteristics of remote sensing imagery to the characteristic patterns of coal mine geological hazards. Specifically, it designs a dual-stream UNet infrastructure to adapt to pixel-level segmentation tasks, proposes a lightweight cross-modal bidirectional block Transformer to address computational power explosion and cross-modal collaboration issues, reconstructs a feature confidence-based dynamic decision fusion module to avoid topological contradictions, improves the spatial perception SE denoising branch to achieve precise local noise suppression, adopts an "offline + online" hybrid processing strategy, and clarifies the "prior weight + residual perturbation" fusion logic to solve the real-time problem of SAR data. Simultaneously, it constructs a dedicated sample library for coal mine power supply lines and a dynamically weighted multi-task loss function. Ultimately, it achieves all-weather, full-coverage, and high-precision intelligent interpretation of geological hazard hazards within coal mine power supply line corridors, providing practical technical support for the safe operation and maintenance of coal mine power supply lines.

[0064] Example 1:

[0065] Figure 1 This is a flowchart illustrating the intelligent interpretation method for geological hazard risks in coal mine power supply lines based on multi-source satellite remote sensing imagery, according to an embodiment of the present invention. Figure 1 As shown, the overall process of this method includes four main steps: delineation of the study area and standardization preprocessing of multi-source data, construction of a dedicated feature set and creation of a sample library, construction and training of a cross-modal fusion model adapted to remote sensing, and intelligent interpretation of hidden dangers and spatiotemporal evolution analysis. Specific details are as follows:

[0066] Step S1: Delineate the study area of ​​the coal mine power supply line, collect and standardize the optical remote sensing images and synthetic aperture radar (SAR) remote sensing images of the study area to obtain the optical feature set and SAR feature set.

[0067] Step S1 is used to delineate the study area for coal mine power supply lines and acquire and standardize preprocessing of multi-source satellite remote sensing data. Delineating the study area includes: acquiring vector ledger data of power supply lines in the target coal mine area; setting line corridor buffer zones based on voltage level differences; and establishing a step-like non-linear positive correlation between voltage level and buffer zone width. Specifically, it includes the following steps:

[0068] S11, Precise delineation of the research area

[0069] Obtain vector data of power supply lines in the target coal mining area (line path, tower locations, voltage level, etc.). Based on voltage level differences, set up line corridor buffer zones: a 500m buffer zone on one side for lines with voltage levels of 110kV and below, and a 1000m buffer zone on one side for lines with voltage levels of 220kV and above. Generate the boundary of the study area to achieve accurate coverage of the entire influence range of the lines while reducing the invalid interpretation area. The voltage level and buffer zone width have a step-like nonlinear positive correlation; the increase in buffer zone width increases proportionally with each increase in voltage level. Specific optimization examples are shown in Table 1.

[0070] Table 1

[0071]

[0072] Low-voltage lines have small loads and dense tower spacing, making them relatively less sensitive to surrounding disasters, and the increase in buffer zone width is relatively small. Medium and high voltage (110kV~220kV) lines are the main power supply lines for coal mines, with moderate tower spacing, requiring a significant expansion of the buffer zone to cover the impact range of mining. Ultra-high voltage (500kV and above) lines have large tower spacing and large foundation scale, and the increase in buffer zone tends to be stable, forming a step-like nonlinear positive correlation, which ensures the safe coverage of the line while avoiding an excessively large ineffective interpretation area.

[0073] S12, Multi-source satellite remote sensing data matching and acquisition

[0074] To ensure the geospatial consistency and temporal validity of the data, optical and SAR remote sensing data were collected simultaneously and spatiotemporally in the study area.

[0075] (1) High-resolution optical remote sensing data: panchromatic and multispectral satellite images such as Gaofen-2, Gaofen-7, and WorldView are selected, with spatial resolution better than 2m, covering the visible light to near-infrared band, and are used to extract detailed features such as surface texture, spectrum, and topography;

[0076] (2) SAR remote sensing data: Select C-band SAR data such as Gaofen-3 and Sentinel-1, including single-view complex (SLC) data and interferometric pair data, covering ascending / descending orbit imaging modes, and simultaneously acquire precise orbit data to extract features such as surface deformation, backscattering, and polarization.

[0077] S13. Multi-source data standardization preprocessing

[0078] Specialized preprocessing was performed on optical and SAR data respectively, and the spatial reference, resolution, and size were unified to provide standardized input for subsequent feature fusion. Specifically, this included:

[0079] S131. Optical data preprocessing: Radiometric calibration, atmospheric correction, orthorectification, panchromatic and multispectral image fusion, cloud shadow and noise suppression, image mosaicking and cropping are completed sequentially to obtain orthorectified optical images of the study area; Digital surface model (DSM) is extracted based on the panchromatic band of the optical image, and topographic factors such as slope, aspect, and elevation are calculated to generate a topographic factor map with the same spatial range and resolution as the optical image.

[0080] S132. SAR Data Preprocessing (Offline + Online Hybrid Strategy): A hybrid processing mode of "offline construction of a deformation background library + online updating of instantaneous features" is adopted, clearly defining the fusion logic of offline and online features. Specifically, this includes:

[0081] (1) Offline stage: Using long-term synthetic aperture radar remote sensing image data, a time-series deformation rate map and a cumulative deformation map are generated as prior weight features to characterize the long-term deformation trend. The specific process is as follows: Collect 15-20 or more long-term SAR data of the study area, and complete radiometric calibration, multi-view processing, Lee filtering denoising, orbit refinement, interferometric pair registration, and differential interferometric processing through SBAS-InSAR technology to generate a time-series deformation rate map and a cumulative deformation map of the study area. Construct a mining area deformation background feature library and update it regularly on a quarterly / semi-annual basis. This deformation background feature serves as the "prior weight" of the model, providing a long-term deformation trend reference for online interpretation and reducing the interference of instantaneous noise on the interpretation results.

[0082] (2) Online stage: For newly acquired synthetic aperture radar (SAR) remote sensing images of a single scene, the backscattering coefficient map and polarization feature map are quickly extracted as residual perturbation features that characterize recent sudden changes; the prior weight features and residual perturbation features are spatiotemporally matched to form a dynamically updated SAR feature set. The specific process is as follows: For newly acquired SAR remote sensing images of a single scene, radiometric calibration, geocoding, and polarization decomposition (Freeman-Durden three-component, Cloude-Pottier decomposition) are quickly completed, and instantaneous features such as backscattering coefficient map and polarization feature map are extracted. These instantaneous features are used as "residual perturbations" to capture the sudden changes and details of recent surface deformation. After spatiotemporal matching with the offline deformation background library, a dynamically updated SAR feature set is formed.

[0083] S133. Multi-source data unification: All preprocessed optical images, SAR dynamic feature sets, and topographic factor maps are uniformly registered to the CGCS2000 national geodetic coordinate system, resampled to 1m spatial resolution, and cropped into fixed-size image blocks (to adapt to model input), thus completing the spatial and dimensional matching of multi-source data.

[0084] Step S2: Based on the optical feature set and SAR feature set, construct a multi-source feature set of geological hazards specific to coal mine power supply lines, and create a multi-modal sample library adapted for pixel-level segmentation.

[0085] Step S2 is used to construct a multi-source feature set of geological hazards specific to coal mine power supply lines and create a sample library. Specifically, it includes the following steps:

[0086] S21. Construction of Multi-Dimensional Specific Feature Set for Coal Mine Geological Disasters

[0087] To investigate the occurrence and evolution patterns and remote sensing characteristics of mining-induced subsidence and landslides in coal mining areas, a multi-dimensional feature set complementary to optical and SAR technologies was constructed, covering surface details, deformation features, and topographic attributes, providing a basis for model feature extraction. The details are as follows:

[0088] (1) Optical feature subset: includes spectral indices such as multispectral band reflectance, normalized vegetation index (NDVI), normalized water index (NDWI), and normalized building index (NDBI), texture features such as contrast, entropy, and energy based on gray-level co-occurrence matrix, and topographic features such as slope, aspect, and elevation.

[0089] (2) SAR feature subset: includes offline extracted time-series deformation rate, cumulative deformation and other core deformation background features (prior weights) as prior weight features to characterize long-term deformation trends, and online extracted backscattering coefficient, polarization decomposition components and other instantaneous features (residual perturbations) as residual perturbation features to characterize recent sudden changes, forming a SAR feature combination of "static background + dynamic instantaneous".

[0090] S22, Creation of a Geological Disaster Sample Database for Coal Mine Power Supply Lines

[0091] To address the characteristics of geological hazards within coal mine power supply line corridors, a multimodal labeled sample library adapted to pixel-level segmentation was constructed to solve the problems of small sample size, class imbalance, and poor scene adaptability.

[0092] Furthermore, in step S2, a combined annotation scheme is used to create a multimodal sample library adapted for pixel-level segmentation: Typical geological hazard hazards in coal mining areas are used as positive samples, and stable surfaces within the study area are used as negative samples. Pixel-level polygon annotation is used to annotate the outline, type, and scale of each hazard, and the corresponding optical image blocks and SAR feature blocks are matched to form one-to-one multimodal sample pairs. Historical geological hazard ledgers of coal mines, on-site inspection records, manual visual interpretation results, and mining subsidence monitoring data are integrated as sample data sources. A combined annotation scheme of weakly supervised pre-annotation, semi-supervised optimization, and manual refinement is used to annotate the samples. A two-dimensional remote sensing-specific data augmentation strategy is used to expand the sample library, and the sample library is divided into training, validation, and test sets according to a preset ratio. Specifically, the following steps are included:

[0093] S221. Sample Types and Labeling Standards: Positive samples are typical geological hazards in coal mining areas (mining-induced subsidence areas: subsidence pits, ground fissures; landslides: traction landslides, push-moving landslides), and negative samples are stable surfaces in the study area (normal forest land, farmland, mountains without abnormal deformation, tower bodies, etc.). Pixel-level polygon labeling is used to label the outline boundary, type, and scale of each hazard, and at the same time, the corresponding optical image blocks and SAR feature blocks (including deformation background and instantaneous features) are matched to form a one-to-one multimodal sample pair.

[0094] S222. Sample data sources: integrating historical geological disaster records of coal mines, on-site inspection records, results of manual visual interpretation, and mining subsidence monitoring data to ensure the authenticity and representativeness of the samples;

[0095] S223. Sample Augmentation and Splitting: A two-dimensional remote sensing-specific data augmentation strategy is adopted, including geometric augmentation (random rotation, flipping, cropping, translation), radiometric augmentation (brightness / contrast fine-tuning, Gaussian noise perturbation), and feature augmentation (slight scaling of SAR deformation features, fine-tuning of optical spectral index) to expand the number of samples. The augmented sample library is randomly divided into training set, validation set, and test set in a 7:2:1 ratio, while retaining an independent test set across mining areas to verify the model's generalization ability.

[0096] Step S3: Construct and train a cross-modal fusion deep learning model adapted to 2D remote sensing images; the cross-modal fusion deep learning model is an end-to-end architecture, used to receive the optical feature set and SAR feature set, and through the built-in encoder-decoder structure, perform multi-source heterogeneous feature extraction, deep fusion and noise suppression, and finally output the interpretation results of geological disaster hazards;

[0097] Step S3 is used to build and train a cross-modal fusion deep learning model adapted to remote sensing images.

[0098] Step S3 is the core of this invention. It abandons the rigid copying of cross-domain models and uses a dual-stream UNet architecture (adapted to pixel-level segmentation of 2D remote sensing images). Through multi-scale feature extraction and dynamic decision fusion logic, combined with dual encoders and cross-modal Transformer technology, it performs deep adaptation and optimization for the fusion characteristics of optical and SAR homogeneous heterogeneous data. It designs an end-to-end model of multi-scale dual encoder - lightweight cross-modal bidirectional block Transformer - dynamic decision fusion. The model as a whole includes six core parts: multi-scale feature extraction module, dual encoder (multi-modal encoder + auxiliary encoder), lightweight cross-modal bidirectional block Transformer fusion module, dynamic decision fusion module adapted for segmentation, spatial perception SE noise reduction branch, and segmentation and risk classification dual-task decoding module. It realizes one-stop interpretation from multi-source feature extraction to hazard segmentation and risk classification. Moreover, all architecture designs are adapted to existing general-purpose hardware to ensure feasibility.

[0099] S31, Overall Model Architecture Design

[0100] The model adopts an overall architecture of dual-stream input, dual-encoder fusion, and single-decoder output. The input consists of preprocessed optical image blocks and SAR dynamic feature blocks (both 1m resolution, fixed size). The output is a semantic segmentation map of geological hazards and a hazard risk classification map with the same size as the input. Based on a dual-stream UNet, the model extracts specific features for optical and SAR modes in the encoder stage. An auxiliary encoder enhances the understanding of global image information. A lightweight cross-modal bidirectional block Transformer is used to model long-distance spatial associations of heterogeneous features. A reconstructed dynamic decision fusion module enables adaptive fusion of multi-scale features. Combined with a spatially aware SE denoising branch, local noise is suppressed. Finally, pixel-level segmentation and risk classification are completed in the decoder stage, preserving two-dimensional spatial dimensional information throughout the process and completely resolving the topological paradox of classification and segmentation. Figure 2 The diagram shown is a schematic representation of the overall architecture of the cross-modal fusion deep learning model according to an embodiment of the present invention.

[0101] Specifically, the cross-modal fusion deep learning model includes: a multi-scale feature extraction module, dual encoders, a lightweight cross-modal bidirectional block Transformer fusion module, a dynamic decision fusion module, a spatially aware SE denoising branch, and a segmentation and risk classification dual-task decoding module. Based on a dual-stream UNet, the model extracts multi-scale features from optical images and SAR features through the multi-scale feature extraction module, extracts modality-specific features and globally fused features through the dual encoder, achieves long-distance spatial correlation modeling of heterogeneous features through the lightweight cross-modal bidirectional block Transformer fusion module, achieves adaptive fusion of multi-scale features through the dynamic decision fusion module, suppresses local noise through the spatially aware SE denoising branch, and finally outputs pixel-level semantic segmentation maps and risk classification maps through the segmentation and risk classification dual-task decoding module. Details are as follows:

[0102] S32, Multi-scale Feature Extraction Module (Adapted for 2D Remote Sensing Imagery)

[0103] To address the characteristics of 2D remote sensing imagery and coal mine geological hazards, a modality-specific multi-scale convolutional feature extraction module (2D-MSCBlock) was designed. This module extracts features from both optical and SAR inputs, capturing hazard features at different scales (small-scale subsidence pits, large-scale landslides). Furthermore, the multi-scale feature extraction module includes convolutional channels composed of kernels of various sizes, used to extract features from both optical and SAR inputs, capturing hazard features at different scales. Each convolutional channel contains a combination of dilated convolution, point convolution, activation functions, and batch normalization. The output features of each convolutional channel are concatenated along the channel dimension to obtain the multi-scale fused features of the optical and SAR inputs. Details are as follows:

[0104] (1) The module consists of three different sizes of convolution channels: 3×3, 5×5, and 7×7. Small-sized convolution kernels extract detailed features of the ground surface (such as ground fissures and the edges of sinkholes), while large-sized convolution kernels extract macroscopic deformation features (such as the overall outline of a landslide).

[0105] (2) Each convolutional channel contains a combination of dilated convolution, point convolution, GELU activation, and batch normalization. Dilated convolution expands the receptive field without increasing the parameters, point convolution achieves channel dimension feature fusion and dimensionality reduction, and batch normalization suppresses overfitting and accelerates model training.

[0106] (3) The output features of the three convolutional channels are spliced ​​together by channel dimension to obtain the optical / SAR multi-scale fusion features, which are then input into the optical branch and SAR branch of the subsequent multimodal encoder.

[0107] S33, Dual Encoder Fusion Architecture (Adapted to Remote Sensing Imagery)

[0108] A dual encoder architecture consisting of a multimodal encoder and an auxiliary encoder is constructed to achieve the dual effects of modality-specific feature extraction and global information supplementation, thereby improving the model's ability to understand the complex background of the mining area.

[0109] Furthermore, the dual encoder includes a multimodal encoder and an auxiliary encoder. The multimodal encoder has a dual-stream structure, corresponding to the optical branch and the SAR branch respectively. The optical branch and the SAR branch receive multi-scale fusion features from the optical branch and the SAR branch respectively, used to extract optical mode-specific features and SAR mode-specific features. A coordinate attention module is embedded at each level to enhance the feature response of small-scale hazards and high-risk areas around towers, outputting optical and SAR features at each level. The auxiliary encoder uses DenseNet as its backbone network, and its input is a channel-level stitched image of optical and SAR features, used to extract and output global fusion features of optical and SAR features. The optical and SAR features output from the multimodal encoder at each level, and the corresponding global fusion features output from the auxiliary encoder, are fused pixel-by-pixel to obtain multi-scale globally enhanced optical features and multi-scale globally enhanced SAR features. Specifically:

[0110] S331, Multimodal Encoder: It has a dual-stream structure, corresponding to optical and SAR branches respectively. It uses an improved ResNet50 as the backbone network and embeds a coordinate attention (CA) module at each level. This module is designed to enhance the feature response of small-scale hidden dangers in coal mine power supply line corridors and high-risk areas around towers, thereby improving the recognition accuracy of small target hidden dangers. The optical branch focuses on extracting detailed features such as surface spectrum, texture, and terrain, while the SAR branch focuses on extracting deformation-related features such as surface deformation background, instantaneous scattering, and polarization. It outputs optical mode-specific features and SAR mode-specific features for each level.

[0111] S332, Auxiliary Encoder: With DenseNet as the backbone network, the input is a channel-level mosaic of optical and SAR features, used to extract global fusion features from multi-source data, supplementing the multimodal encoder's understanding of the global scene of the image (such as the overall terrain of the mining area, the direction of the route corridor, and the large-scale deformation trend), and outputting global fusion features.

[0112] S333, Encoder Feature Fusion: The optical / SAR features output by the multi-modal encoder at each level are fused pixel by pixel with the corresponding global features output by the auxiliary encoder to obtain multi-scale globally enhanced optical features and multi-scale globally enhanced SAR features, providing a foundation for subsequent cross-modal fusion.

[0113] S34, Lightweight Cross-Modal Bidirectional Block Transformer Fusion Module

[0114] To address the computational overhead of traditional Transformer global attention and the shortcomings of existing single-modal block attention (such as the Swin Transformer) in not considering cross-modal collaboration, a lightweight cross-modal bidirectional block Transformer (LS-BiCrossMT) is proposed. The core innovation lies in the "cross-modal bidirectional guided block attention mechanism" (defined by an explicit mathematical formula), rather than simple block processing, to achieve long-distance spatial correlation modeling of optical and SAR features, while ensuring the feasibility of the model on existing hardware.

[0115] Furthermore, the lightweight cross-modal bidirectional block-based Transformer fusion module is used to perform cross-modal interactive fusion of multi-scale globally enhanced optical features and multi-scale globally enhanced SAR features. Specifically, this includes: simultaneously dividing the input multi-scale globally enhanced optical features and multi-scale globally enhanced SAR features into multiple non-overlapping feature blocks; constructing an optically guided SAR attention mechanism, using queries from optical feature blocks to match the keys and values ​​of SAR feature blocks, and utilizing optical texture information to optimize the spatial contours of SAR features; constructing a SAR-guided optics attention mechanism, using queries from SAR feature blocks to match the keys and values ​​of optical feature blocks, and utilizing SAR deformation information to guide optical features to focus on the core area of ​​potential hazards; fusing the bidirectionally guided feature blocks to obtain cross-modal fused feature blocks; performing the above-mentioned cross-modal bidirectional block-based fusion on the shallow, middle, and deep features output by the encoder, and restoring the fused feature blocks to a complete feature map to obtain multi-scale cross-modal deep fused features; wherein, the optically guided SAR attention mechanism and the SAR-guided optics attention mechanism are only calculated within the feature blocks at corresponding spatial locations and between cross-modal blocks to reduce model computational consumption. Figure 3 The diagram shows the internal structure of the lightweight cross-modal bidirectional segmented Transformer fusion module according to an embodiment of the present invention. The construction of the lightweight cross-modal bidirectional segmented Transformer fusion module specifically includes the following steps:

[0116] S341, Blocking strategy: Divide the optical feature map output by the encoder into blocks. and Feature map The optical feature blocks are simultaneously divided into non-overlapping feature blocks of the same size (e.g., 16×16 pixels / block) to obtain a set of optical feature blocks. With SAR feature block set (M is the number of blocks); attention is computed only between and within cross-modal blocks at corresponding positions, reducing the complexity of the attention matrix from... Down to (S is the block size), which greatly reduces the consumption of video memory and computing power;

[0117] in, Optical feature map output by the encoder; SAR feature map output by the encoder; The data dimension representation of the feature map, where For real numbers, H, W, and C correspond to the height, width, and number of channels of the feature map, respectively; H: height of the optical feature map and SAR feature map; W: width of the optical feature map and SAR feature map; C: number of channels of the optical feature map and SAR feature map; 16×16 pixels: a preferred example of block size (can be adjusted within the range of 16×16~32×32 pixels); : Set of optical feature blocks, i is the feature block index, M is the number of a single modal feature block (i.e., the number of optical feature blocks and SAR feature blocks is M); : SAR feature block set, where the meanings of i and M are the same as those defined in the optical feature block set; The time complexity expression for a traditional global attention matrix. This represents the total number of pixels in the feature map. The time complexity expression for the block attention matrix in this invention, where S is the block size. S: Number of pixels in a single feature block; S: Size of the feature block (i.e., the pixel side length of a single feature block).

[0118] S342. Cross-modal bidirectional attention mechanism: Constructing bidirectional cross-modal attention interaction between "optics → SAR" and "SAR → optics", which differs from single-modal block attention. "All originate from the same mode", and their mathematical mapping relationship is as follows:

[0119] 1. Optical-guided SAR stage: querying optical feature blocks ( Keys that match SAR feature blocks ) and value ( By optimizing the spatial contour of SAR deformation features using boundary information from optical textures, the problem of blurred boundaries in SAR deformation maps can be solved; among them, : The query vector corresponding to the optical feature block; : The key vector corresponding to the SAR feature block; : The value vector corresponding to the SAR feature block;

[0120] (1) Optical query mapping: ,in For optical query projection matrix, For attention head dimension;

[0121] in, : The query vector of the i-th optical feature block; : The i-th optical feature block; : Query projection matrix of optical modes; The data dimension representation of the optical query projection matrix. The number of feature map channels. For attention head dimension (can be found) Adjustments within the scope);

[0122] (2) SAR key / value mapping: , ,in , These are the SAR key and value projection matrices, respectively.

[0123] in, : The key vector of the i-th SAR feature block; : The i-th SAR feature block; : The key projection matrix of the SAR mode; : The value vector of the i-th SAR feature block; : The projection matrix of SAR modes; The data dimension representation of the SAR key / value projection matrix, where C is the number of feature map channels. Value dimension (can be adjusted within the range of 64~128);

[0124] (3) Attention calculation: , thus obtaining the optically guided SAR feature block;

[0125] in, : The SAR feature block guided by the i-th optical feature block; : Activation function, used to normalize the attention weights so that the sum of the weights is 1; : Scaling factor, used to alleviate the problem of excessively large values ​​in the attention weight calculation process.

[0126] 2. SAR-guided optical stage: querying SAR feature blocks Key of matching optical feature blocks AND value By guiding optical features to focus on the core area of ​​potential hazards through SAR deformation anomaly regions, the interference of complex backgrounds on optical recognition is reduced; among them, : The query vector corresponding to the SAR feature block; : The key vector corresponding to the optical feature block; : The value vector corresponding to the optical feature block;

[0127] (1) SAR query mapping: ,in, Query the projection matrix for SAR;

[0128] in, : The query vector of the i-th SAR feature block; : Query projection matrix of SAR modes;

[0129] (2) Optical bond / value mapping: , ,in, , These are the optical bond and value projection matrices, respectively.

[0130] in, : The key vector of the i-th optical feature block; : Bond projection matrix of optical modes; : The value vector of the i-th optical feature block; : The projection matrix of the optical modes;

[0131] (3) Attention calculation: The optical feature blocks after SAR guidance are obtained;

[0132] in, : The optical feature block guided by the i-th SAR feature block.

[0133] 3. Cross-modal block fusion: Channel splicing and linear projection are performed on the guided feature blocks. ;

[0134] in, : No. A cross-modal fused feature block; : Channel stitching function, used to merge optically guided SAR feature blocks with optically guided SAR feature blocks by channel dimension; Cross-modal feature block fusion projection matrix Used to restore the original channel dimension. To fuse the data dimension representation of the projection matrix, The feature dimensions after concatenation ( (where C is the value dimension) and C is the number of channels in the original feature map.

[0135] S3423, Multi-scale cross-modal fusion: The shallow, middle and deep features output by the encoder are fused using the above lightweight cross-modal bidirectional block fusion method. The shallow features focus on the collaborative details of the edge of the hidden danger, the middle features focus on the texture / deformation correlation modeling of the hidden danger, and the deep features focus on the semantic category confirmation of the hidden danger. Finally, all the fused feature blocks are restored to the complete feature map to obtain multi-scale cross-modal deep fusion features, which fully explores the complementarity of optical and SAR features.

[0136] S35, Dynamic decision fusion module for adapting to heterogeneous data from the same source

[0137] To address the homogeneity and complementary features of optics and SAR, a dynamic decision fusion module based on feature confidence is designed to adaptively weight and fuse multi-scale optical and SAR features output by dual encoders, while preserving two-dimensional spatial dimensional information throughout the process to avoid topological contradictions.

[0138] Furthermore, the dynamic decision fusion module is used to adaptively weight and fuse the multi-scale cross-modal deep fusion features based on feature confidence. Specifically, it includes: calculating pixel-level feature confidence maps for the optical and SAR deep fusion features at each scale; assigning adaptive fusion weights to the optical and SAR features at corresponding pixel locations based on the feature confidence maps; introducing learnable global weight parameters to globally fine-tune the pixel-level fused features to obtain dynamically fused features at each scale; and upsampling / downsampling the dynamically fused features at each scale to achieve multi-scale feature splicing and fusion, resulting in the final cross-modal deep fusion features. Specifically, it includes the following steps:

[0139] S351. Feature confidence calculation: For optical fusion features and SAR fusion features at each scale, pixel-level feature confidence maps are calculated using convolutional layers and Sigmoid activation. The confidence value reflects the contribution of optical / SAR features to the identification of geological hazard risks at that pixel location (e.g., high confidence of optical features at the edge of a collapse pit, and high confidence of SAR deformation features inside a landslide body).

[0140] S352, Adaptive Weight Allocation: Based on the feature confidence map, pixel-level adaptive fusion weights are assigned to optical and SAR features. The higher the confidence, the greater the weight, thus achieving pixel-level dynamic feature fusion.

[0141] S353, Global Weight Fine-tuning: Two learnable global weight parameters are introduced to fine-tune the pixel-level fused features globally, adapting to the differences in feature contribution in different mining areas and different disaster types, and improving the model's generalization ability;

[0142] S354, Multi-scale feature fusion: Upsampling / downsampling is performed on the dynamically fused features at each scale to achieve multi-scale feature splicing and fusion, resulting in the final cross-modal deep fusion feature, which is then input into the subsequent decoding module.

[0143] S36, Spatial Awareness SE Noise Reduction Branch (Adapted to Local Noise Suppression)

[0144] To address the shortcomings of the traditional Squeeze-and-Excitation (SE) module in global noise reduction based on GAP, an improved spatial perception SE noise reduction branch is obtained. As a residual branch of the model, it can accurately suppress local spatial noise such as cloud shadows and sensor noise, without destroying effective hazard features.

[0145] Furthermore, the spatially aware SE denoising branch is used to suppress local spatial noise in the final cross-modal deep fusion features. Specifically, this includes: simultaneously capturing channel-level global feature dependencies and local spatial noise distribution features through hybrid pooling operations; adjusting the spatial distribution of attention weights through spatial convolution after channel attention calculation to specifically suppress feature weights in local noise regions; and adding the output of the spatially aware SE denoising branch pixel-by-pixel to the final cross-modal deep fusion features via residual connections, achieving noise suppression while preserving effective features. Figure 4 The diagram shown is a schematic representation of the spatially aware SE noise reduction branch according to an embodiment of the present invention. Details are as follows:

[0146] (1) Replace the global average pooling (GAP) of the traditional SE module with hybrid pooling (global average pooling + local max pooling), which preserves the global feature dependency at the channel level and captures the noise distribution features in the local space.

[0147] (2) After calculating the channel attention, add a 1×3+3×1 serial convolution. Adjust the spatial distribution of attention weights through two-dimensional spatial convolution. Reduce feature weights for local noise areas and increase feature weights for effective hidden danger areas to achieve spatial selective noise reduction.

[0148] (3) The noise reduction branch is connected between the encoder and decoder of the residual connection model. The noise reduction features are added pixel by pixel with the multi-scale cross-modal deep fusion features to achieve noise suppression while retaining effective features, thereby improving the interpretation accuracy in the complex background of the mining area.

[0149] S37, Dual-task decoding module for segmentation and risk grading

[0150] The UNet decoder architecture adopts progressive upsampling, which gradually fuses the shallow features output by the encoder during the upsampling process to restore the boundary details of the hidden dangers. The decoding end is set with dual task output heads to achieve the synchronous output of pixel-level semantic segmentation and regional risk classification of geological disaster hidden dangers, and retains the two-dimensional spatial dimension information throughout the process.

[0151] Furthermore, the segmentation and risk grading dual-task decoding module includes: a semantic segmentation head, used to output pixel-level semantic segmentation maps of geological hazard hazards, depicting the contours, extent, and type of hazards; and a risk grading head, used to extract quantitative features of each hazard region based on the hazard mask output from the semantic segmentation map, and to perform risk grading of the hazard regions through a multilayer perceptron, outputting a risk grading map with the same size as the input. Specifically, it includes:

[0152] S371, Semantic Segmentation Head: Composed of multiple 1×1 convolutions and activation functions, it achieves three-class pixel-level segmentation (collapsed area, landslide body, background), and outputs a semantic segmentation map with the same size as the input, accurately depicting the outline, range, and type of potential hazards;

[0153] S372, Risk Classification Head: Based on the hazard mask output by the segmentation head, extract the core quantitative features of each hazard area (deformation rate, minimum distance between the hazard and the line / tower, hazard area / volume), and realize the four-level risk classification (low, medium, high, and extremely high) of the hazard area through a multilayer perceptron (MLP), and output a risk classification map with the same size as the input, realizing the linkage between pixel-level segmentation and region-level classification;

[0154] S373, Feature Residual Connections: In the decoding process, cross-layer residual connections are introduced to solve the gradient vanishing problem in deep networks and ensure the effective transmission of hidden detail features.

[0155] S38. Model Training and Optimization

[0156] To address the specific characteristics of coal mine geological hazard segmentation and risk classification, a dynamically weighted multi-task joint loss function is designed. Appropriate training strategies and hyperparameters are employed to ensure model training stability and recognition accuracy. All training strategies are compatible with existing general-purpose GPU hardware (such as RTX 3090 / 4090, Tesla V100), specifically including:

[0157] S381. Dynamically weighted multi-task joint loss function: Total loss function This is a dynamic weighted sum of semantic segmentation loss, risk classification loss, and regularization term, addressing issues such as imbalanced samples, low accuracy in identifying high-risk hazards, and convergence imbalance in the dual-task approach. The formula is:

[0158] ;

[0159] in: It is a weighted combination of Dice loss and cross-entropy loss (to solve the class imbalance problem in segmentation tasks). Focal Loss (enhancing the accuracy of high-risk hazard classification); This is an L2 regularization term (to prevent the model from overfitting); , These are dynamic weighting coefficients that adaptively adjust with the number of training rounds t: in the early stages of training (t≤50 rounds). , The focus is on the convergence of the segmentation task, prioritizing the accuracy of hazard contour recognition; in the later stages of training (t>50 rounds). It decays linearly to 0.5. The linear improvement to 0.5 balances the convergence performance of segmentation and hierarchical tasks;

[0160] The regularization coefficient is fixed (set to 0.001) and determined through optimization using the validation set.

[0161] S382, Training Strategy and Hyperparameters: The AdamW optimizer is used, with an initial learning rate of 1e-4. A cosine annealing learning rate decay strategy and an early stopping mechanism are employed (training stops if the validation set accuracy does not improve for 10 consecutive epochs). The batch size is set to 8 / 16 based on the hardware memory, and the training epochs are 100. Gradient clipping (with a clipping threshold of 1.0) is used to address the gradient explosion problem.

[0162] S383. Model Validation and Optimization: The model performance is validated based on the test set and independent test sets across mining areas. Intersection over Union (IoU), mean Intersection over Union (mIoU), and accuracy (Acc) are used as evaluation indicators for segmentation tasks, while accuracy, precision, and recall are used as evaluation indicators for risk grading tasks. The model hyperparameters and architecture are fine-tuned based on the validation results, and finally an interpretation model that meets the requirements of engineering applications is obtained (see Example 2 for details).

[0163] Step S4: Input the optical feature set and SAR feature set of the study area to be interpreted into the trained cross-modal fusion deep learning model, and output the interpretation results of the geological disaster hazards in the corridor, including pixel-level semantic segmentation map and risk classification map.

[0164] Step S4 is used to realize intelligent interpretation and spatiotemporal evolution analysis of geological disaster hazards in coal mine power supply lines.

[0165] S41, Intelligent Interpretation of Routine Hidden Dangers

[0166] Standardized multi-source remote sensing data (optical images, SAR dynamic feature blocks) of the study area to be detected are input into the trained interpretation model. The model automatically completes multi-scale feature extraction, cross-modal bidirectional block fusion, local noise suppression, pixel-level segmentation and risk classification, and outputs core information such as the location, outline, type, scale and risk level of all geological disaster hazards in the corridor. It generates an intelligent interpretation list of hazards, which replaces traditional manual visual interpretation and greatly improves interpretation efficiency.

[0167] S42. Spatiotemporal Evolution Analysis of Hidden Dangers

[0168] Based on multi-temporal standardized multi-source remote sensing data, a temporal analysis was conducted on the interpreted hazard areas: by combining the offline SBAS-InSAR deformation background library with online SAR instantaneous features, the temporal deformation curves of the hazard areas were obtained. The changes in surface features (such as vegetation cover and terrain contours) of the hazard areas were extracted by combining multi-temporal optical images, and the expansion rate, deformation trend, and activity status (stable, slowly developing, and rapidly developing) of the hazards were analyzed, providing data support for early warning of geological disasters affecting coal mine power supply lines.

[0169] Example 2:

[0170] To more clearly illustrate the technical solution, inventiveness, and reproducibility of this invention, the following detailed explanation is provided using real data from three different types of coal mining areas, cross-mining area generalization experiments, and core module ablation experiments. All data are from real mining area remote sensing acquisition and model training experiments, and the experimental comparison object is clearly the existing mainstream technology (dual-stream CNN + shallow feature stitching model).

[0171] 1. Preparation of experimental data

[0172] 1.1 Selection of Experimental Mining Area and Data Parameters

[0173] Three typical coal mining areas were selected (covering different terrains, disaster types, and mining intensities), and the specific parameters are shown in Table 2.

[0174] Table 2

[0175]

[0176] The three mining areas represent three typical scenarios: mining in thick loose layers in plains, steep slopes in mountainous areas, and complex goaf areas in the Loess Plateau. The landforms, geological conditions, causes of disasters, mining methods, and route layouts are all fundamentally different. The model still maintains high accuracy in cross-mining area tests, which fully demonstrates that its generalization ability does not depend on the rules of a single scenario, but rather has the ability to interpret the remote sensing characteristics of different geological disasters universally.

[0177] Remote sensing data parameters:

[0178] (1) Optical remote sensing data: Gaofen-7 (spatial resolution 0.8m) and WorldView-3 (spatial resolution 0.3m), covering the visible light to near-infrared bands, 20 scenes were collected for each mining area (including different seasons and weather conditions, including 3-5 scenes of cloud shadow interference images).

[0179] (2) SAR data: Sentinel-1 C-band SAR data (ascending orbit + descending orbit, spatial resolution 10m), Gaofen-3 (spatial resolution 8m), 18 scenes of long time series data (time span 1 year) were collected in each mining area during the offline phase, and 1 scene was updated every quarter during the online phase;

[0180] (3) SBAS-InSAR processing parameters: time baseline ≤ 30 days (preferred value), spatial baseline ≤ 100m (preferred value), number of interferometric pairs ≥ 25 pairs, deformation rate inversion accuracy ± 2mm / year.

[0181] 1.2 Sample Library Construction and Annotation

[0182] 1.2.1 Sample collection range: Each mining area's line corridor buffer zone (a non-linear positive correlation corridor buffer zone is established according to the voltage level; for voltage levels of 110kV and below, a single-sided buffer zone of 500m is preferred, and for voltage levels of 220kV and above, a single-sided buffer zone of 1000m is preferred), cropped into image blocks of 512×512 pixels.

[0183] 1.2.2 Sample size: As shown in Table 3:

[0184] Table 3

[0185]

[0186] 1.2.3 Annotation Method: To address the engineering feasibility issue of large-scale pixel-level annotation of high-resolution remote sensing images, a combined approach of "weakly supervised pre-annotation + semi-supervised optimization + manual refinement" is adopted to reduce annotation costs and ensure the industrial-scale implementation of the sample library.

[0187] (1) Weak supervision pre-labeling: Using existing geological disaster ledgers, mining subsidence prediction maps, and historical inspection records, a rough bounding box (rather than pixel level) of the hidden danger area is generated through georegistration as the pre-labeling result;

[0188] (2) Semi-supervised optimization: Based on the pre-labeled bounding boxes, a "pseudo-label generation algorithm" (such as the Mean Teacher model) is used to automatically pre-fill the unlabeled areas at the pixel level to generate preliminary pixel-level labeling results. The pseudo-label accuracy is ≥85%.

[0189] (3) Manual refinement: Remote sensing geological engineers conduct sampling review and correction of the pseudo-label results, focusing on correcting errors in the determination of hazard boundaries and disaster types. The sampling ratio is 20% of the total sample, and the final labeling accuracy rate is ≥98.5%.

[0190] (4) Comparison of annotation efficiency: Traditional manual pixel-level annotation of a single 5000×5000 pixel image takes 8 to 10 hours, while this engineering solution only takes 2 to 3 hours, improving efficiency by 3 to 4 times and meeting the industrial needs of building a large-scale sample library.

[0191] 1.3 Training Environment and Comparison Scheme

[0192] (1) Hardware configuration: GPU (NVIDIA RTX 4090, 24GB video memory) CPU (Intel i9-13900K), 64GB RAM;

[0193] (2) Software environment: PyTorch 2.0, CUDA 11.8, Python 3.9;

[0194] (3) Comparison scheme (existing technology): The mainstream "dual-stream CNN + shallow feature concatenation" model is adopted. The encoder is ResNet50, the fusion method is channel-level concatenation, the decoder is UNet basic architecture, and the loss function is fixed-weight Dice + cross-entropy loss (without dynamic weighting).

[0195] 2. Model training parameters and core module ablation experiments

[0196] 2.1 Training Parameter Settings

[0197] (1) Hyperparameter settings: Batch size Training rounds Initial learning rate Gradient clipping threshold ;

[0198] (2) Dynamic weighting coefficient: Early training phase (to mid-training phase, optimal rounds) ), , The focus is on the convergence of the segmentation task; in the later stages of training (after the middle of training). linear decay to The linear improvement is increased to 0.5, balancing the tasks of segmentation and hierarchical classification;

[0199] (3) LS-BiCrossMT module parameters: optimal block size Pixels (available) Adjust within the pixel range), attention head count 8 (can be found) Adjustments within the scope) (can be found) Adjustments within the scope) (can be found) (Adjustments within the scope).

[0200] 2.2 Ablation experiment of LS-BiCrossMT module (to verify the technical effect of bidirectional guidance)

[0201] To verify the technical advantages of "bidirectional cross-modal guidance", three sets of ablation experiments were designed (based on the test set of mining area B), as shown in Table 4:

[0202] Table 4

[0203]

[0204] Compared to unidirectional guidance, bidirectional cross-modal guidance improved mIoU by 6.8 percentage points and landslide IoU by 9.2 percentage points. Compared to unguided block attention, mIoU improved by 10.1 percentage points and the recall rate for extremely high risks improved by 11.9 percentage points. This proves that "bidirectional guidance" is not a simple superposition, but rather solves the core problems of blurred SAR deformation boundaries and high optical background noise through the complementary interaction of optical and SAR features, producing unexpected technical effects.

[0205] 3. Experimental Results and Analysis

[0206] 3.1 Accuracy Indicators of Test Sets in the Same Mining Area

[0207] 3.1.1 Semantic segmentation task accuracy: As shown in Table 5:

[0208] Table 5

[0209]

[0210] 3.1.2 Accuracy of Risk Classification Task: As shown in Table 6:

[0211] Table 6

[0212]

[0213] 3.2 Cross-mining area generalization ability test (verification of non-overfitting)

[0214] The model's generalization ability was verified using a "single mining area training → cross-mining area testing" model, as shown in Table 7.

[0215] Table 7

[0216]

[0217] The mIoU in cross-mining area tests is ≥81.3%, and the classification accuracy is ≥86.5%, which is only 3-5 percentage points lower than that in tests within the same mining area. This is significantly better than existing cross-mining area tests (mIoU≤70%), proving that the model is not overfitted, has strong generalization ability, and can be adapted to different types of mining area scenarios.

[0218] 3.3 Visual Effects Comparison Description

[0219] (1) Interpretation of landslide body B in mining area: Existing technologies only use CNN shallow fusion, which cannot capture the long-distance texture and deformation correlation of the landslide edge. The interpreted contour is jagged and deviates from the actual surface crack direction by ≥15°. This invention uses the optical → SAR guidance of the LS-BiCrossMT module to optimize the SAR deformation boundary by utilizing the crack texture features of the optical image. The interpreted contour is smooth and continuous, and the deviation from the actual crack direction is ≤3°, accurately matching the actual range of the landslide body.

[0220] (2) Interpretation of the C composite subsidence area in the mining area: Existing technologies fail to distinguish between offline and online SAR features, resulting in a false negative rate of ≥30% for small subsidence pits (diameter ≤5m) that have recently occurred. This invention reduces the false negative rate to below 5% by fusing "offline deformation background (prior) + online instantaneous features (residual)" and can accurately distinguish the boundary between multi-layer mining subsidence and surface subsidence.

[0221] (3) Interpretation of cloud shadow interference scene: In the image of mining area A containing cloud shadow, the existing technology does not use spatial perception noise reduction, and the false alarm rate in the cloud shadow area is ≥25%; the present invention reduces the false alarm rate in the cloud shadow area to below 8% through the spatial perception SE noise reduction branch of "hybrid pooling + serial convolution", while retaining the effective features of the sinkhole under the cloud.

[0222] It should also be noted that, in the embodiments of this application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0223] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined in the embodiments of this application may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown in this application, but is to be accorded the widest scope consistent with the principles and novel features disclosed in the embodiments of this application.

Claims

1. A method for intelligent interpretation of geological hazard risks in coal mine power supply lines based on multi-source satellite remote sensing imagery, characterized in that, Includes the following steps: Step S1: Delineate the study area of ​​the coal mine power supply line, collect and standardize the optical remote sensing images and synthetic aperture radar remote sensing images of the study area to obtain optical feature sets and SAR feature sets. Step S2: Based on the optical feature set and SAR feature set, construct a multi-source feature set of geological hazards specific to coal mine power supply lines, and create a multi-modal sample library adapted for pixel-level segmentation. Step S3: Construct and train a cross-modal fusion deep learning model adapted to 2D remote sensing images; The cross-modal fusion deep learning model is an end-to-end architecture used to receive the optical feature set and SAR feature set, and through the built-in encoder-decoder structure, to perform multi-source heterogeneous feature extraction, deep fusion and noise suppression, and finally output the interpretation results of geological disaster hazards. Step S4: Input the optical feature set and SAR feature set of the study area to be interpreted into the trained cross-modal fusion deep learning model, and output the interpretation results of the geological disaster hazards in the corridor, including pixel-level semantic segmentation map and risk classification map.

2. The intelligent interpretation method according to claim 1, characterized in that, In step S1, delineating the study area includes: obtaining vector ledger data of power supply lines in the target coal mining area, setting up line corridor buffer zones based on voltage level differences, with voltage level and buffer zone width showing a step-like nonlinear positive correlation; The standardized preprocessing of the synthetic aperture radar remote sensing imagery includes: employing a hybrid offline and online processing strategy, specifically including: In the offline phase, time-series deformation rate maps and cumulative deformation maps are generated using long-term synthetic aperture radar remote sensing image data, serving as prior weight features characterizing long-term deformation trends. During the online phase, for newly acquired synthetic aperture radar remote sensing images of a single scene, backscattering coefficient maps and polarization feature maps are quickly extracted as residual perturbation features characterizing recent sudden changes. The prior weight features and residual perturbation features are spatiotemporally matched to form a dynamically updated SAR feature set.

3. The intelligent interpretation method according to claim 1, characterized in that, In step S2, the construction of a multi-source geological hazard feature set specific to coal mine power supply lines is a complementary multi-dimensional optical-SAR feature set covering surface details, deformation features, and topographic attributes, including: The optical feature subset includes multispectral reflectance, spectral index, texture features, and topographic features; The SAR feature subset employs a combined strategy of static background and dynamic instantaneous data. It includes offline extracted temporal deformation rate and cumulative deformation as prior weight features characterizing long-term deformation trends, and online extracted backscattering coefficient and polarization decomposition components as residual perturbation features characterizing recent sudden changes.

4. The intelligent interpretation method according to claim 3, characterized in that, In step S2, the creation of the multimodal sample library adapted for pixel-level segmentation adopts a combined annotation scheme: Using typical geological hazard hazards in coal mining areas as positive samples and stable surfaces in the study area as negative samples, the outline boundary, type and scale of each hazard are marked by pixel-level polygon annotation, and the corresponding optical image blocks and SAR feature blocks are matched to form a one-to-one multimodal sample pair. The sample data sources include historical geological disaster records of coal mines, on-site inspection records, results of manual visual interpretation, and mining subsidence monitoring data. A combined annotation scheme of weakly supervised pre-annotation, semi-supervised optimization, and manual refinement was used to annotate the samples; A two-dimensional remote sensing-specific data augmentation strategy was adopted to expand the sample library, and the sample library was divided into training set, validation set and test set according to a preset ratio.

5. The intelligent interpretation method according to claim 1, characterized in that, In step S3, the cross-modal fusion deep learning model adopts an overall architecture of dual-stream input, dual-encoder fusion, and single-decoder output; The cross-modal fusion deep learning model specifically includes: a multi-scale feature extraction module, dual encoders, a lightweight cross-modal bidirectional block Transformer fusion module, a dynamic decision fusion module, a spatially aware SE denoising branch, and a segmentation and risk classification dual-task decoding module. The cross-modal fusion deep learning model is based on a dual-stream UNet. It extracts multi-scale features from optical images and SAR features through a multi-scale feature extraction module, extracts modality-specific features and global fusion features through a dual encoder, achieves long-distance spatial correlation modeling of heterogeneous features through a lightweight cross-modal bidirectional block Transformer fusion module, achieves adaptive fusion of multi-scale features through a dynamic decision fusion module, suppresses local noise through a spatially aware SE denoising branch, and finally outputs pixel-level semantic segmentation maps and risk classification maps through a segmentation and risk classification dual-task decoding module.

6. The intelligent interpretation method according to claim 5, characterized in that, The multi-scale feature extraction module includes: a convolution channel composed of convolutional kernels of various sizes, used to extract features from optical input and SAR input respectively, and capture hidden danger features at different scales; Each convolutional channel comprises a combination of dilated convolution, point convolution, activation function, and batch normalization. The output features of each convolutional channel are concatenated along the channel dimension to obtain optical multi-scale fusion features and SAR multi-scale fusion features.

7. The intelligent interpretation method according to claim 6, characterized in that, The dual encoders include: a multimodal encoder and an auxiliary encoder; The multimodal encoder has a dual-stream structure, corresponding to an optical branch and a SAR branch respectively. The optical branch and the SAR branch receive the multi-scale fusion features of the optical mode and the multi-scale fusion features of the SAR mode respectively, which are used to extract optical mode-specific features and SAR mode-specific features. A coordinate attention module is embedded at each level to enhance the feature response of small-scale hidden dangers and high-risk areas around the tower, and outputs optical features and SAR features at each level. The auxiliary encoder uses DenseNet as its backbone network and takes as input a channel-level stitched image of optical and SAR features to extract and output global fusion features of optical and SAR features. The optical features and SAR features output by the multimodal encoder at each level, and the global fusion features output by the auxiliary encoder at the corresponding level, are fused pixel by pixel to obtain multi-scale globally enhanced optical features and multi-scale globally enhanced SAR features.

8. The intelligent interpretation method according to claim 7, characterized in that, The lightweight cross-modal bidirectional block Transformer fusion module is used to perform cross-modal interactive fusion of the multi-scale globally enhanced optical features and the multi-scale globally enhanced SAR features, specifically including: The input multi-scale globally enhanced optical features and multi-scale globally enhanced SAR features are simultaneously divided into multiple non-overlapping feature blocks; An attention mechanism for optically guided SAR is constructed, which uses the query of optical feature blocks to match the keys and values ​​of SAR feature blocks, and utilizes optical texture information to optimize the spatial contour of SAR features. A SAR-guided optical attention mechanism is constructed, which uses the query of SAR feature blocks to match the keys and values ​​of optical feature blocks, and uses SAR deformation information to guide optical features to focus on the core area of ​​potential hazards. The bidirectional guided feature blocks are fused to obtain cross-modal fused feature blocks; The shallow, middle, and deep features output by the encoder are subjected to the above-mentioned cross-modal bidirectional block fusion, and the fused feature blocks are restored to complete feature maps to obtain multi-scale cross-modal deep fusion features. The attention mechanism of optically guided SAR and the attention mechanism of SAR-guided optics are calculated only within the feature block at the corresponding spatial location and between cross-modal blocks.

9. The intelligent interpretation method according to claim 8, characterized in that, The dynamic decision fusion module is used to adaptively weight and fuse the multi-scale cross-modal deep fusion features based on feature confidence. Specifically, it includes: calculating pixel-level feature confidence maps for optical and SAR deep fusion features at each scale; assigning adaptive fusion weights to the optical and SAR features at corresponding pixel locations based on the feature confidence maps; introducing learnable global weight parameters to globally fine-tune the pixel-level fused features to obtain dynamically fused features at each scale; and upsampling / downsampling the dynamically fused features at each scale to achieve multi-scale feature splicing and fusion, resulting in the final cross-modal deep fusion features.

10. The intelligent interpretation method according to claim 9, characterized in that, The spatially aware SE noise reduction branch is used to suppress local spatial noise in the final cross-modal deep fusion features. Specifically, it includes: capturing both channel-level global feature dependencies and local spatial noise distribution features through hybrid pooling operations; adjusting the spatial distribution of attention weights through spatial convolution after channel attention calculation to specifically suppress feature weights in local noise regions; and adding the output of the spatially aware SE noise reduction branch to the final cross-modal deep fusion features pixel by pixel through residual connections to achieve noise suppression while preserving effective features. The segmentation and risk classification dual-task decoding module includes: a semantic segmentation head, used to output a pixel-level semantic segmentation map of geological disaster hazards, depicting the outline, range and type of hazards; and a risk classification head, used to extract the quantitative features of each hazard region based on the hazard mask output by the semantic segmentation map, and to classify the hazard region by risk through a multilayer perceptron, outputting a risk classification map with the same size as the input.