Multi-dimensional feature inversion and risk diagnosis method and system for coal-related industrial area

By constructing a hierarchical feature sample set and cascaded visual inversion, and combining white-box decision-making and multi-path risk quantification, the accuracy and automation issues of remote sensing monitoring methods in coal-related industrial areas were solved, achieving high-precision intelligent identification, documentation, and risk diagnosis.

CN121661491APending Publication Date: 2026-03-13CHINA UNIV OF MINING & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing remote sensing monitoring methods suffer from low classification accuracy, insufficient automation, opaque decision-making processes, and a lack of systematic sample set construction in coal-related industrial areas, making it difficult to achieve high-precision, automated intelligent identification and documentation.

Method used

A hierarchical feature sample set is constructed, and cascaded visual inversion and white-box decision-making are adopted. Combined with a multi-scheme fusion decision-making mechanism, a multi-dimensional feature data cube is constructed through multi-source geographic environment data, and a dual-path risk quantification model is used for risk diagnosis.

Benefits of technology

It has achieved high-precision, automated, and interpretable intelligent identification and documentation of coal-related industrial zones, enhanced the credibility of AI models in key decision-making areas, and provided structured risk assessment data input.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661491A_ABST
    Figure CN121661491A_ABST
Patent Text Reader

Abstract

The invention discloses a coal-related industrial area multi-dimensional feature inversion and risk diagnosis method and system. The method comprises the following steps: firstly, constructing a hierarchical multi-dimensional feature sample set serving macroscopic-microscopic-stress source three-level linkage; then, a cascade vision inversion strategy from coarse to fine is adopted; four types of coal-related sites including an open-air coal site, an underground coal site, a coal power site and a coal chemical field are judged according to the extracted ground feature combination and a multi-scheme fusion mechanism through a white box logical reasoning engine, and end-to-end black box recognition is decomposed into transparent traceable decisions; constructing a coal-related site multi-dimensional feature data cube containing physical attributes and environmental backgrounds; and finally, implementing dual-path risk quantification and prototype diagnosis based on the data cube, and decoupling a collaborative and event risk driving mechanism. According to the method, the problems of opaque decision-making process and single risk representation in the prior art are solved, and high-precision identification, total-factor filing and deep risk portraying of the coal-related industrial area are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing and artificial intelligence technology, specifically relating to a method and system for multi-dimensional feature inversion and risk diagnosis of coal-related industrial areas. Background Technology

[0002] Accurate and dynamic monitoring of coal-related industrial areas has become an urgent need. Existing remote sensing monitoring methods have the following significant drawbacks: ① Limitations of traditional remote sensing classification methods: Classification methods based on pixel spectral features suffer from low accuracy when dealing with phenomena such as "different spectra for the same object, and similar spectra for different objects" within coal-related industrial areas (e.g., cement buildings and ash / slag fields have similar spectra), making it difficult to meet application requirements. ② Shortcomings of object-oriented methods (OBIA): Although OBIA methods incorporate information such as shape and texture, they heavily rely on manually designed complex segmentation and classification rules, which are not only time-consuming and labor-intensive, but also have sensitive model parameters and poor portability between images from different regions and time periods, failing to meet the needs of large-scale, long-term automated monitoring. ③ The "black box" problem of end-to-end deep learning models: Existing deep learning models mostly adopt an "end-to-end" model, directly mapping from images to classification results. While this may achieve a certain level of accuracy, its decision-making process is opaque and cannot explain "why a certain site is identified as a specific type." ④ Lack of systematic sample set construction: Existing research sample sets are often aimed at a single target or macro category, lacking a sample system that can support the transition from macro-level coarse screening to micro-level fine identification, structuring, and hierarchy. This results in a single model training objective, making it difficult to achieve a deep understanding of complex industrial scenarios. Summary of the Invention

[0003] The purpose of this invention is to provide a method and system for multi-dimensional feature inversion and risk diagnosis of coal-related industrial zones, which can achieve high-precision, automated, and fully interpretable intelligent identification and filing of various coal-related industrial zones.

[0004] To achieve the above objectives, this invention provides a method for multi-dimensional feature inversion and risk diagnosis of coal-related industrial zones, comprising the following steps: S1: Hierarchical Sample Construction: Construct a three-level hierarchical feature sample set of "macro positioning - micro components - stress source" to serve multi-stage AI identification, and clarify the key diagnostic ground feature maps and risk entity characteristics of various coal-related industrial areas; S2: Cascaded Visual Inversion: Based on the hierarchical feature sample set of S1, a coarse-to-fine cascaded strategy is adopted. First, the spatial distribution of suspected targets is inverted using medium-resolution remote sensing images and a candidate list is generated. Then, high-resolution remote sensing images are used to perform instance segmentation and feature extraction on key diagnostic features inside the candidate targets. S3: White Box Decision and Reconstruction: The key diagnostic feature combination information extracted in S2 is input into the white box logic reasoning engine based on expert knowledge. The final site type is determined through a multi-scheme fusion decision-making mechanism, and multi-source geographic environmental data are coupled to construct a multi-dimensional feature data cube of coal-related sites. S4: Risk Diagnosis Profile: Based on the S3 coal-related site multi-dimensional feature data cube, a dual-path risk quantification model is used to decouple the site's collaborative risk and event-related risk, identify global risk prototypes and calculate the prototype driving force index PDI, and achieve in-depth diagnosis of risk patterns.

[0005] As a further aspect of the present invention: the hierarchical feature sample set in S1 includes: S1.1 Macro-scale positioning sample subset: bounding boxes of suspected open-pit coal mining sites and suspected complex industrial sites based on medium spatial resolution image annotations; S1.2, Microscale Diagnostic Feature Sample Subset: Based on high spatial resolution image annotations, key diagnostic features that can clearly indicate the functional type of the site are systematically divided into four feature classes: ① Open-pit coal mine feature class OCMS: including mining faces / mine edges, large spoil heaps; ② Underground coal mine feature class UCMS: including coal yard transfer conveyor belts, coal preparation plant roof features, railway coal loading platforms; ③ Coal-fired power plant feature class CPS: including chimneys, cooling towers, switchyard areas; ④ Coal chemical plant feature class CCS: including distillation / reaction tower groups, large spherical or horizontal storage tanks, enclosed conveyor belts, and coal bunkers; S1.3, Sample subset of stress source exposure indicators: six types of physical entities with potential environmental risks based on high spatial resolution image annotations.

[0006] As a further aspect of the present invention: the cascaded visual inversion in S2 includes: 1) Automated initial screening: The first target detection model is trained using a macro-scale localized sample subset to generate a list of candidate targets on medium spatial resolution images; 2) Feature Interpretation: The second target detection model is trained using a subset of diagnostic ground feature samples at the microscale. The model performs internal scanning on high spatial resolution images of candidate targets, detects and outputs key diagnostic ground feature instances and their confidence scores.

[0007] As a further aspect of the present invention: the multi-scheme fusion decision-making mechanism in S3 includes: 1) Parallel execution decision schemes: including the maximum confidence priority scheme, the cumulative confidence maximization scheme, the weighted confidence scheme, the hybrid constraint scheme, and the optimal feature threshold scheme; 2) Priority Arbitration: Based on the preset category identifiability priority P=["OCMS", "CPS", "CCS", "UCMS"], the output results of the above multiple schemes are voted on by majority vote to determine the final site type.

[0008] As a further aspect of the present invention: the multi-dimensional feature data cube of the coal-related site in S3 includes: 1) Site basic information layer: includes the identified site type, vector boundary, and six types of original stress source exposure indicators identified by the model trained by applying a subset of stress source exposure indicator samples; 2) Site environment attribute layer: Integrates multi-source time-series remote sensing data, and generates a 13-dimensional environmental feature vector after time series feature extraction of mean, trend, volatility, peak value, principal component analysis dimensionality reduction and robust standardization.

[0009] As a further aspect of the present invention: the dual-path risk quantification model in S4 includes: 1) Collaborative risk path: Geographically weighted principal component analysis is used to generate local weights for each site, and the collaborative risk components are calculated by the approximation of ideal solution ranking method to characterize the risk combination of spatial heterogeneity; 2) Event-based risk path: The global weight is generated by the entropy weight method, and the event-based risk component is calculated by combining the multi-criteria compromise solution ranking method to characterize the sudden risk driven by a few extreme values.

[0010] As a further aspect of the present invention: the prototype driving force index PDI calculation step in S4 includes: 1) Construct a four-dimensional feature space consisting of collaborative risk components, event-based risk components, and their sub-indicators, and define the cluster center of the lowest risk prototype as the collaborative pole and the cluster center of the highest risk prototype as the event pole. 2) Calculate the Prototype Driving Force Index (PDI) for any site: The closer the PDI value is to 1, the more the risk pattern is biased towards event-driven.

[0011] To achieve the above-mentioned objectives, this invention also provides a multi-dimensional feature inversion and risk diagnosis system for coal-related industrial zones, used to implement the aforementioned multi-dimensional feature inversion and risk diagnosis method for coal-related industrial zones, comprising: Sample and Data Management Module: Used to store and manage hierarchical feature sample sets and multi-source remote sensing data; Cascaded Inversion Module: Equipped with first and second target detection models, it performs cascaded operations from macroscopic positioning to microscopic feature interpretation, and outputs candidate target and internal ground feature information; Logical reasoning and cube generation module: Built-in white-box logical reasoning engine and data fusion algorithm, used to perform multi-scheme fusion decision to determine site type and generate multi-dimensional feature data cubes of coal-related sites; Risk Calculation and Diagnosis Module: Integrates the GWPCA-TOPSIS algorithm and the EWM-VIKOR algorithm to output dual-path risk components and perform risk prototype identification and PDI calculation.

[0012] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention constructs a hierarchical sample dataset that links macroscopic, microscopic, and stress source levels, and combines it with a multi-stage AI identification process that progresses from coarse to fine, feature interpretation, and logical reasoning. This enables high-precision, automated, and fully interpretable intelligent identification and documentation of various coal-related industrial zones.

[0013] This invention effectively balances computational efficiency and recognition accuracy through a multi-stage strategy of "macroscopic initial screening + microscopic precise judgment." By employing a white-box recognition paradigm of "feature interpretation + rule reasoning," it decomposes the end-to-end black-box problem into two transparent steps: identifying visible components and making logical judgments based on component combinations. This ensures that each judgment result has clear, traceable, and expert-compliant decision-making basis, greatly enhancing the credibility of the AI ​​model in critical decision-making areas. The final output, a multi-dimensional feature data cube of coalfield sites, is provided in standard GIS format. It not only includes the site type and boundaries but also integrates internal risk source information and external environmental background features, providing structured and information-rich direct data input for downstream risk assessment, environmental supervision, dynamic monitoring, and policy formulation. Attached Figure Description

[0014] Figure 1 This is a flowchart illustrating the multi-dimensional feature inversion and risk diagnosis method for coal-related industrial zones according to the present invention.

[0015] Figure 2 This is a schematic diagram of the hierarchical feature sample set system of the present invention.

[0016] Figure 3 This is a visual schematic diagram of the refined identification logic process based on rule-based reasoning in this invention.

[0017] Figure 4 This is a schematic diagram showing the distribution of the six Global Risk Assimilation (GRAs) identified in this embodiment of the invention in a two-dimensional risk space.

[0018] Figure 5 This is a schematic diagram of the driving force pattern distribution of each Global Risk As (GRAs) in an embodiment of the present invention. Detailed Implementation

[0019] The invention will now be further described with reference to the accompanying drawings.

[0020] like Figure 1As shown, a method for multi-dimensional feature inversion and risk diagnosis of coal-related industrial zones includes the following steps: S1: Hierarchical Sample Construction: Construct a three-level hierarchical feature sample set of "macro positioning - micro components - stress source" to serve multi-stage AI identification, and clarify the key diagnostic ground feature maps and risk entity characteristics of various coal-related industrial areas.

[0021] Furthermore, such as Figure 2 As shown, the hierarchical feature sample set includes: S1.1, Macro-scale localization sample subset: bounding boxes of suspected open-pit coal mining sites and suspected mixed industrial sites based on medium spatial resolution image annotations; specifically including: ①Data basis: Based on medium spatial resolution remote sensing images, Sentinel-2 Level-2A (surface reflectance) multispectral images were preferred, and median composite technology was used to generate cloud-free images of the study area during the growing season (June-September) of 2024.

[0022] ② Labeling objects: Perform coarse classification and bounding box labeling for suspected targets, including at least two categories: Suspected open-pit coal mine (S-OCMS): Visual characteristics include areas with obvious mining patterns, large areas of exposed surface, dark gray tones, or stepped landforms.

[0023] Suspected Complex Industrial Site (SCIS): Visually characterized as a complex area containing a dense cluster of industrial buildings, large structures (such as chimneys and cooling towers), material yards, railways, and other ancillary facilities.

[0024] ③ Data format: The labeled data is organized in the COCO (Common Objects in Context) format, which includes image information, category labels and bounding box coordinates, and is used to train a preliminary screening model for a large range of targets.

[0025] S1.2, Microscale Diagnostic Feature Sample Subset: Key diagnostic features based on high spatial resolution image annotations that clearly indicate the site's functional type. These key diagnostic features are systematically divided into four feature classes: ① Open-pit coal mine feature class (OCMS): including mining faces / mine edges and large spoil heaps; ② Underground coal mine feature class (UCMS): including coal yard conveyor belts, coal preparation plant roof features, and railway coal loading platforms; ③ Coal-fired power plant feature class (CPS): including chimneys, cooling towers, and switchyard areas; ④ Coal chemical plant feature class (CCS): including distillation / reaction tower groups, large spherical or horizontal storage tanks, enclosed conveyor belts, and coal bunkers; specifically including: ①Data basis: Based on sub-meter high spatial resolution remote sensing imagery, with Jilin-1 satellite imagery being the preferred choice.

[0026] ② Labeling objects: Refined instance segmentation or bounding box labeling is performed on 11 types of key diagnostic features that can clearly indicate the site function type. The features are systematically divided into four groups according to their site type: open-pit coal mine features, underground coal mine surface facility features, coal-fired power plant features, and coal chemical site features. Detailed descriptions are shown in Table 1.

[0027] Table 1. Detailed Description of Key Diagnostic Features at Detailed Scale

[0028] ③ Data format: Also follows the COCO format and is used to train the core model for feature interpretation.

[0029] S1.3, Sample subset of stress source exposure indicators: Six categories of physical entities with potential environmental risks, based on high spatial resolution image annotations. Specifically, these include: ①Data basis: Same as S1.2, based on sub-meter high spatial resolution remote sensing imagery.

[0030] ② Labeling Objects: Six categories of physical entities with potential environmental risks are labeled, including: mining face / mine edge (R1), large / steep spoil heap (R2), large coal pile (R3), ash dump (R4), industrial wastewater / sedimentation pond (R5), and adjacent bare / disturbed land (R6). Detailed feature descriptions are shown in Table 2.

[0031] Table 2. Detailed Description of Risk Scale Identification Features

[0032] ③ Technical role: This sample subset is used to train the risk factor identification model. Its labeling results (present / absent) directly serve as a raw physical attribute of the site, forming the basic information layer of the subsequent data cube.

[0033] S2: Cascaded Visual Inversion: Based on the hierarchical feature sample set of S1, a coarse-to-fine cascaded strategy is adopted. First, the spatial distribution of suspected targets is inverted using medium-resolution remote sensing images to generate a candidate list. Then, high-resolution remote sensing images are used to perform instance segmentation and feature extraction on key diagnostic features inside the candidate targets.

[0034] Furthermore, cascaded visual inversion includes: S2.1 Automated Initial Screening: The first target detection model is trained using a macro-scale localization sample subset to generate a candidate target list on medium spatial resolution images; specifically including: ① Model Training: Using the macro-scale localization sample subset described in S1.1, a deep learning model for object detection is trained. The Mask R-CNN architecture is preferred, with ResNet-101 combined with a Feature Pyramid Network (FPN) as its backbone network. During training, a stochastic gradient descent (SGD) optimizer is used, a dynamic learning rate strategy is configured, and data augmentation techniques such as random flipping, rotation, and color dithering are applied. Its multi-task loss function L is defined as: ;in, For classifying losses, For bounding box regression loss, This represents the loss due to mask segmentation.

[0035] ② Model Application: The trained model is applied to medium-resolution remote sensing images (such as Sentinel-2 annual composite images) covering the entire study area. Through sliding window inference, non-maximum suppression (NMS), and confidence threshold filtering, all suspected targets (S-OCMS and SCIS) are automatically identified and located. The IoU threshold for non-maximum suppression is set to 0.5, and the confidence threshold is set to 0.3, generating a list of candidate targets containing geographic coordinates and preliminary classification.

[0036] S2.2 Feature Interpretation: A second target detection model is trained using a subset of microscale diagnostic feature samples. This model performs internal scanning of high spatial resolution images of candidate targets, detects, and outputs key diagnostic feature instances and their confidence scores. Specifically, this includes: ① Data preparation: Based on the candidate target list and its geographical range generated in S2.1, acquire high-resolution remote sensing images of the corresponding areas as needed.

[0037] ② Model Training: Using the microscale diagnostic feature sample subset described in S1.2, a deep learning model specifically designed to identify 11 key diagnostic features was trained, preferably using the Mask R-CNN architecture. To enable the model to learn more refined features, each larger site image was further sliced ​​into 512×512 pixel sub-images for training and detection.

[0038] ③ Feature detection: Apply the model to the high-resolution image of each candidate target, perform an internal scan, detect and output all key diagnostic ground feature instances contained therein and their confidence scores.

[0039] S3: White Box Decision and Reconstruction: The key diagnostic feature combination information extracted in S2 is input into the white box logic reasoning engine based on expert knowledge. The final site type is determined through a multi-scheme fusion decision mechanism, and multi-source geographic environmental data are coupled to construct a multi-dimensional feature data cube of coal-related sites.

[0040] Furthermore, S3.1, the multi-solution fusion decision-making mechanism includes: 1) Parallel execution decision schemes: including: ① Maximum Confidence Priority Scheme: Based on a preset category priority, the site is identified as the first category to meet its maximum feature confidence threshold. Its decision function is: ;in, This is the identification result of site i. The maximum confidence level for the feature corresponding to class C. This is the confidence threshold for this category.

[0041] ② Cumulative Confidence Maximization: This method selects the class with the highest cumulative confidence among those classes that meet their respective cumulative confidence thresholds. Its decision function is: ;in, Let be the sum of the confidence scores of all detection instances j in class C. This is the cumulative confidence threshold.

[0042] ③ Weighted Confidence Scheme: This scheme integrates feature significance and spatial distribution breadth through linear weighting to calculate the category with the highest overall score. ;in, and These are weighting coefficients. It represents the coverage of Class C features across all sub-images.

[0043] ④ Hybrid Constraints: These schemes require each class to simultaneously satisfy both maximum confidence and cumulative confidence constraints, and select the class with the highest overall score. ;in, It is the average confidence level. and These are weighting coefficients; Optimal Feature Threshold: This method makes judgments at the finest sub-feature level, determining whether a site contains any sub-feature that exceeds its optimal threshold based on category priority. ;in, It is the set of sub-features corresponding to class C. It is a certain sub-feature. It is the optimal threshold for this sub-feature.

[0044] 2) Priority Arbitration: Based on the preset category identifiability priority P=["OCMS", "CPS", "CCS", "UCMS"], the output results of the above multiple schemes are subject to majority voting. This determines the final site type.

[0045] like Figure 3 As shown, by establishing a logical decision-making system based on expert knowledge, the engine receives detected key diagnostic feature combinations and performs inference based on a preset, configurable rule base. The rule base contains a series of conditional statements, such as: "IF (F-CPS-2 AND F-CPS-1 detected in the site) THEN (site type = CPS)" and "IF (F-OCMS-2 AND F-OCMS-1 detected in the site) THEN (site type = OCMS)". The engine also integrates a multi-scheme fusion decision-making mechanism, applying multiple classification schemes with different design principles in parallel, determining the final category through majority voting, and handling fuzzy cases by setting a classification priority (P = ["OCMS", "CPS", "CCS", "UCMS"]). The engine uses a multi-scheme fusion decision-making mechanism to determine the site type.

[0046] Furthermore, S3.2, the multi-dimensional feature data cube for coal-related sites, creates a comprehensive, multi-dimensional dataset indexed by its unique ID (sample_id) for each site successfully identified in S3.1, including: 1) Site Basic Information Layer: This layer includes the identified site type, vector boundaries, and six types of original stress source exposure indicators identified by a model trained using a subset of stress source exposure indicator samples; specifically, it includes: ① Physical attributes: including the site type identified by S3.1 (such as OCMS, UCMS, etc.) and the vector boundary (Polygon geometric information) generated by the model.

[0047] ② Original stress source: Includes the presence or absence (binary label) of the original 6 types of SSE indicators (R1-R6) of the site. This information comes from the results of the identification model trained on the sample subset of stress source exposure (SSE) indicators described in S1.3 of the site. The presence or absence information of the 6 types of indicators is integrated into a comprehensive stress source exposure count indicator SSE_Count, which has a value range of 0 to 6.

[0048] 2) Site Environmental Attribute Layer: Integrating multi-source time-series remote sensing data, a 13-dimensional environmental feature vector is generated after time-series feature extraction (mean, trend, volatility, peak values), principal component analysis for dimensionality reduction, and robust standardization. Specifically, it includes: ① Stressor Environmental Characterization (SEC) data: derived from time series data from remote sensing satellites such as Sentinel-5P (NO2, SO2), MODIS (AOD, LST, fire point), and Sentinel-2 (NDVI, BSI, NDTI).

[0049] ② Socio-Ecological Sensitivity (SES) data: derived from land use products (CLCD), population data (GHSL), nighttime light (VIIRS), protected areas (WDPA), and local river networks, etc.

[0050] ③ Feature Engineering and Processing Flow: A rigorous processing flow is implemented for the integrated data, including: 1) Time series feature extraction: Aggregate SEC time series data (such as AOD, NDVI, etc.) into four static features: mean, trend (using Theil-Sen slope estimation), volatility (using standard deviation), and peak (using maximum value).

[0051] 2) Multicollinearity Treatment: For the four highly correlated features within each topic, principal component analysis (PCA) was performed separately for each of the seven topic groups (AOD, BSI, LST, NDTI, NDVI, NO2, SO2). The score of the first principal component (PC1) that best explains the original variance was extracted and retained as the comprehensive dynamic index of that topic, named SEC_. <theme>_Dynamic.

[0052] 3) Distribution optimization: For features with an absolute skewness exceeding 2.0, the Winsorization method is used to process extreme values. Then, the Yeo-Johnson transform is applied to all numerical features to normalize their distribution, making it closer to a Gaussian distribution.

[0053] 4) Robust Standardization: To eliminate dimensional differences between different features, a robust standardization method based on the median and interquartile range (IQR) is used for numerical features. The formula is as follows: ;in, These are the standardized eigenvalues; It is the original value of a certain site in terms of specific characteristics; It is the median of this feature across all samples; It is the interquartile range of this feature (i.e., the difference between the upper quartile and the lower quartile).

[0054] At the same time, a special proportional mapping is applied to the SSE_Count comprehensive count metric: This ensures that its value range falls within the interval [0, 1].

[0055] Final output: Generates a 13-dimensional, highly processed, collinear, approximately normally distributed, and scaled final feature vector for each site.

[0056] S4: Risk Diagnosis Profile: Based on the S3 coal-related site multi-dimensional feature data cube, a dual-path risk quantification model is used to decouple the site's collaborative risk and event-related risk, identify global risk prototypes and calculate the prototype driving force index PDI, and achieve in-depth diagnosis of risk patterns.

[0057] Furthermore, S4.1, the dual-path risk quantification model, in order to decouple and independently quantify risks driven by different mechanisms, this invention adopts a dual-path parallel modeling strategy, representing the risk of each site as a two-dimensional vector, specifically including: 1) Collaborative risk path: Geographically weighted principal component analysis is used to generate local weights for each site, and the collaborative risk components are calculated by the approximation of ideal solution ranking method to characterize the risk combination of spatial heterogeneity; Specifically, to capture the synergistic patterns of spatially heterogeneous risk factor combinations, a combined approach of Geographically Weighted Principal Component Analysis (GWPCA) and the Top-Ordered Solution Approximation Method (TOPSIS) is adopted. First, the GWPCA model is applied to generate a unique, localized set of index weights for each site, which are obtained by normalizing the loading vector of the local first principal component. ;in, It is a venue First principal component pair index The load.

[0058] Subsequently, the comprehensive stress index (CSS) and comprehensive sensitivity index (CSN) for each site are calculated using this local weight. Finally, the relative proximity is calculated using the TOPSIS method. This value is defined as the synergistic risk component, with a range of [0, 1].

[0059] 2) Event-driven risk path: The entropy weight method is used to generate global weights, and the event-driven risk components are calculated by combining the multi-criteria compromise solution ranking method to characterize sudden risks driven by a few extreme values; Specifically, to identify event-driven risks driven by a few indicators with extreme values ​​or high variability, a combined approach of Entropy Weight Method (EWM) and Multi-Criterion Compromise Ranking Method (VIKOR) is adopted. First, EWM is applied to calculate a set of global weights applicable to all sites based on the information entropy of each indicator. ; ;in, As an indicator Information entropy.

[0060] Then, the CSS and CSN are calculated using this global weight. Finally, the VIKOR method is applied to calculate the compromise ranking value. After normalization, the event-based risk component is obtained, and its value range is also [0, 1].

[0061] S4.2: Driving mechanism diagnosis based on Prototype Driving Force Index (PDI): ① Global Risk Prototype (GRAs) Identification: For the feature space composed of CSS and CSN obtained by event-based risk quantification in S4.1, the K-Means clustering algorithm is executed, and the optimal number of clusters K is determined by combining the elbow rule and silhouette coefficient analysis. The most typical K risk patterns at the macro scale are identified and defined as global risk prototypes.

[0062] ② Local Risk Microtypes (LRMs) Identification: For the feature space composed of CSS and CSN obtained by synergistic risk quantification in S4.1, the K-Means clustering algorithm is executed to determine the optimal number of clusters K', and K' types of local risk microtypes reflecting the fine structural differences between sites are identified.

[0063] S4.3: Risk Driver Diagnosis and Heterogeneity Measurement: To diagnose whether the risk of each site is more "synergistic" or "event-driven", a Prototype Driver Index (PDI) is constructed.

[0064] ① Constructing a fusion four-dimensional space: Constructing a fusion feature space consisting of four Z-score standardized dimensions: CSS and CSN of collaborative paths and CSS and CSN of event-driven paths.

[0065] ② PDI Index Calculation: In this four-dimensional space, the collaborative poles (cluster centers of the prototypes with the lowest risk) are defined in a data-driven manner. And the event extremes (cluster centers of the prototypes with the highest risk). The formula for calculating the PDI index for any site is as follows: ;in, and The venues are respectively coordinates The Euclidean distance to the cooperating pole and the event pole. The range of PDI is [0, 1], and the closer the value is to 1, the more the risk pattern is biased towards event-driven.

[0066] ③ Internal heterogeneity measurement: The Herfindahl-Hirschman Index (HHI) is introduced to quantify the complexity of the LRM within each GRA. For the i-th GRA, the HHI index is calculated as follows: ;in, It is the first The first GRA within The proportion of LRMs. The smaller the HHI value, the more diverse the internal composition.

[0067] A multi-dimensional feature inversion and risk diagnosis system for coal-related industrial zones, used to implement the aforementioned multi-dimensional feature inversion and risk diagnosis method for coal-related industrial zones, including: Sample and Data Management Module: Used to store and manage hierarchical feature sample sets and multi-source remote sensing data; Cascaded Inversion Module: Equipped with first and second target detection models, it performs cascaded operations from macroscopic positioning to microscopic feature interpretation, and outputs candidate target and internal ground feature information; Logical reasoning and cube generation module: Built-in white-box logical reasoning engine and data fusion algorithm, used to perform multi-scheme fusion decision to determine site type and generate multi-dimensional feature data cubes of coal-related sites; Risk Calculation and Diagnosis Module: Integrates the GWPCA-TOPSIS algorithm and the EWM-VIKOR algorithm to output dual-path risk components and perform risk prototype identification and PDI calculation.

[0068] The specific implementation method is as follows: This embodiment is being tested in the Yellow River Basin, which has 13 large coal-fired power bases.

[0069] I. Data Preparation and Sample Set Construction (corresponding to S1) First, Sentinel-2 multispectral images covering the study area were systematically acquired using the Google Earth Engine (GEE) platform, and 10-meter resolution RGB true-color mosaic maps for the 2024 growing season (June-September) were generated as the basis for macro-scale analysis. Simultaneously, Jilin-1 sub-meter resolution images covering the candidate areas were acquired as needed, serving as core data for micro-scale analysis.

[0070] Based on the above data, a hierarchical sample dataset was constructed: 1. Macro-scale location sample set: On Sentinel-2 imagery, bounding boxes of 1053 suspected open-pit coal mines (S-OCMS) and suspected complex industrial sites (SCIS) were marked.

[0071] 2. Microscale diagnostic feature sample set: 4,879 instances were finely annotated on the Jilin-1 image, covering 11 key diagnostic features, such as mining faces, cooling towers, and large storage tanks.

[0072] 3. Stress Source Exposure (SSE) Index Sample Set: Also on the Jilin-1 image, 2020 instances were marked, corresponding to 6 types of potential environmental risk sources, such as spoil heaps, ash dumps, and industrial wastewater ponds.

[0073] All sample sets are divided into training set, validation set and test set in a ratio of 70% / 15% / 15%.

[0074] II. Cascaded Visual Inversion (corresponding to S2) 1. Automated initial screening: A Mask R-CNN model (ResNet-101+FPN backbone) was trained using a macro-scale localization sample set. The trained model was applied to Sentinel-2 imagery covering the entire Yellow River basin. Through sliding window inference, Non-Maximum Score (NMS) (IoU threshold 0.5), and confidence filtering (threshold 0.3), a total of 237 S-OCMS and 1783 SCIS candidate targets were detected. On an independent test set, the model achieved an overall mean average precision (mAP) of 0.823, with an AP of 0.824 for S-OCMS and 0.822 for SCIS, effectively achieving preliminary screening across a wide range.

[0075] 2. Feature Interpretation: Using a microscale diagnostic ground feature sample set, a Mask R-CNN model was retrained specifically to identify 11 key ground features. This model was applied to high-resolution imagery of all candidate targets, detecting and outputting all ground feature instances contained within them. On independent test sets, the model's recognition performance varied across different ground features. For uniquely shaped and large ground features (such as the coal preparation plant roof feature F-UCMS-2 and large spoil heap F-OCMS-2), the F1 scores reached 0.843 and 0.829, respectively. However, for small, complex, and sparsely sampled ground features (such as the distillation / reaction tower group F-CCS-1), it failed to learn effectively (F1=0), and therefore this feature was excluded from subsequent rules.

[0076] III. Decision Making and Cube Construction (corresponding to S3) Based on the results of feature interpretation, a rule-based reasoning engine integrating five different decision-making schemes (maximum confidence priority, cumulative confidence maximization, etc.) and setting a class priority order (OCMS>CPS>CCS>UCMS) is used to perform final type identification of candidate targets. Through a majority voting mechanism, an overall classification accuracy of 0.8512 is achieved on an independent validation set. For the OCMS and UCMS categories, due to their highly unique diagnostic features, the optimal F1 scores reach 0.9739 and 0.8696, respectively. For CPS and CCS, the optimal F1 scores also reach 0.6122 and 0.6667, demonstrating that the method has effective discriminative ability for all categories.

[0077] For each of the 1927 successfully identified sites, a multidimensional feature data cube was constructed. ① Site Basic Information Layer: This layer includes the identified site type, vector boundaries, and the presence or absence of six original SSE indicators obtained from a model trained using the SSE sample set. ② Site Environmental Attribute Layer: This layer integrates multi-source data from Sentinel series, MODIS, CLCD, GHSL, VIIRS, etc., and constructs a three-dimensional data system of SSE (Site State), SEC (Environmental Constraints), and SES (Socio-Ecological Sensitivity) based on the "stress-state-exposure" logic. Through a series of processing steps, including time series feature extraction, principal component analysis (PCA) dimensionality reduction, Yeo-Johnson transform, and robust normalization, a final feature vector consisting of multiple highly processed features was generated for each site, laying the foundation for subsequent risk analysis.

[0078] IV. Risk Diagnosis Profile (corresponding to S4) This embodiment utilizes the constructed data cube to systematically identify and interpret risk patterns in 1927 coal-related sites in the Yellow River Basin. ① Risk Quantification and Prototype Identification: Through a dual-path risk quantification method, the risk of each site is represented as a two-dimensional vector (cooperative risk component, event-based risk component). Based on this, multi-level clustering successfully interprets the risk patterns of sites such as... Figure 4 The diagram shows six Global Risk Assemblages (GRAs) and fifteen Local Risk Microtypes (LRMs). ② Results Analysis and Validation: Global Risk Assemblage Profile: The analysis revealed significant differences among the six archetypes in risk severity, driving patterns, and internal composition. For example, GRA-I was named "Event-Driven - Extremely High Risk," with a mean PDI of 0.635, significantly higher than the equilibrium point of 0.5, indicating that its risk is mainly driven by sudden, event-driven factors. GRA-III, on the other hand, was named "Low-Stress - Moderate Risk," with a mean PDI of 0.342, indicating that its risk is mainly caused by the long-term synergistic effect of multiple factors. These findings confirm the complexity of risk etiology.

[0079] Verification of the risk duality hypothesis: Through cross-analysis, it was found that the macro-prototype (GRA) exhibits strong selectivity in its internal micro-types (LRMs). For example, the internal composition of GRA-II (sensitive-driven - high-risk) is highly concentrated, with over 60% of member sites exhibiting two specific LRMs; while the internal composition of GRA-I (event-driven - extremely high-risk) is highly diverse, with an HHI index of 0.154 (6.5 effective types), indicating that event-driven risk can act on a variety of different local collaborative contexts. This systematically verifies the "dualistic" characteristic of risk, which simultaneously possesses macro-patternization and micro-diversity.

[0080] Driving force pattern diagnosis: Distribution analysis of the PDI index reveals that risk severity, driving force pattern, and internal diversity are three independent dimensions that can be decoupled. For example... Figure 5 As shown, GRA-I and GRA-II, both of which are high-risk, have drastically different PDI distributions; while GRA-VI, which has the lowest risk, has the largest standard deviation of its internal PDI distribution (0.108), indicating that the driving force patterns are the most diverse under low-risk conditions.

[0081] This embodiment demonstrates that the method can not only identify coal-related industrial areas with high accuracy, but also use the final generated data cube to achieve a deep interpretation of the internal structure and driving mode of regional environmental risks, providing a solid decision-making basis for scientific supervision and precise governance, and realizing a cognitive paradigm leap from "identification and filing" to "diagnosis and profiling".< / theme>

Claims

1. A method for multi-dimensional feature inversion and risk diagnosis of coal-related industrial zones, characterized in that, Includes the following steps: S1: Hierarchical Sample Construction: Construct a three-level hierarchical feature sample set of "macro positioning - micro components - stress source" to serve multi-stage AI identification, and clarify the key diagnostic ground feature maps and risk entity characteristics of various coal-related industrial areas; S2: Cascaded Visual Inversion: Based on the hierarchical feature sample set of S1, a coarse-to-fine cascaded strategy is adopted. First, the spatial distribution of suspected targets is inverted using medium-resolution remote sensing images and a candidate list is generated. Then, high-resolution remote sensing images are used to perform instance segmentation and feature extraction on key diagnostic features inside the candidate targets. S3: White Box Decision and Reconstruction: The key diagnostic feature combination information extracted in S2 is input into the white box logic reasoning engine based on expert knowledge. The final site type is determined through a multi-scheme fusion decision-making mechanism, and multi-source geographic environmental data are coupled to construct a multi-dimensional feature data cube of coal-related sites. S4: Risk Diagnosis Profile: Based on the S3 coal-related site multi-dimensional feature data cube, a dual-path risk quantification model is used to decouple the site's collaborative risk and event-related risk, identify global risk prototypes and calculate the prototype driving force index PDI, and achieve in-depth diagnosis of risk patterns.

2. The method for multi-dimensional feature inversion and risk diagnosis of coal-related industrial zones according to claim 1, characterized in that, The hierarchical feature sample set in S1 includes: S1.1 Macro-scale positioning sample subset: bounding boxes of suspected open-pit coal mining sites and suspected complex industrial sites based on medium spatial resolution image annotations; S1.2, Microscale Diagnostic Feature Sample Subset: Based on high spatial resolution image annotations, key diagnostic features that can clearly indicate the functional type of the site are systematically divided into four feature classes: ① Open-pit coal mine feature class OCMS: including mining faces / mine edges, large spoil heaps; ② Underground coal mine feature class UCMS: including coal yard transfer conveyor belts, coal preparation plant roof features, railway coal loading platforms; ③ Coal-fired power plant feature class CPS: including chimneys, cooling towers, switchyard areas; ④ Coal chemical plant feature class CCS: including distillation / reaction tower groups, large spherical or horizontal storage tanks, enclosed conveyor belts, and coal bunkers; S1.3, Sample subset of stress source exposure indicators: six types of physical entities with potential environmental risks based on high spatial resolution image annotations.

3. The method for multi-dimensional feature inversion and risk diagnosis of coal-related industrial zones according to claim 2, characterized in that, The cascaded visual inversion in S2 includes: 1) Automated initial screening: The first target detection model is trained using a macro-scale localized sample subset to generate a list of candidate targets on medium spatial resolution images; 2) Feature Interpretation: The second target detection model is trained using a subset of diagnostic ground feature samples at the microscale. The model performs internal scanning on high spatial resolution images of candidate targets, detects and outputs key diagnostic ground feature instances and their confidence scores.

4. The method for multi-dimensional feature inversion and risk diagnosis of coal-related industrial zones according to claim 2, characterized in that, The multi-option fusion decision-making mechanism in S3 includes: 1) Parallel execution decision schemes: including the maximum confidence priority scheme, the cumulative confidence maximization scheme, the weighted confidence scheme, the hybrid constraint scheme, and the optimal feature threshold scheme; 2) Priority Arbitration: Based on the preset category identifiability priority P=["OCMS", "CPS", "CCS", "UCMS"], the output results of the above multiple schemes are voted on by majority vote to determine the final site type.

5. The method for multi-dimensional feature inversion and risk diagnosis of coal-related industrial zones according to claim 2, characterized in that, The multidimensional feature data cube of coal-related sites in S3 includes: 1) Site basic information layer: includes the identified site type, vector boundary, and six types of original stress source exposure indicators identified by the model trained by applying a subset of stress source exposure indicator samples; 2) Site environment attribute layer: Integrates multi-source time-series remote sensing data, and generates a 13-dimensional environmental feature vector after time series feature extraction of mean, trend, volatility, peak value, principal component analysis dimensionality reduction and robust standardization.

6. The method for multi-dimensional feature inversion and risk diagnosis of coal-related industrial zones according to claim 1, characterized in that, The dual-path risk quantification model in S4 includes: 1) Collaborative risk path: Geographically weighted principal component analysis is used to generate local weights for each site, and the collaborative risk components are calculated by the approximation of ideal solution ranking method to characterize the risk combination of spatial heterogeneity; 2) Event-based risk path: The global weight is generated by the entropy weight method, and the event-based risk component is calculated by combining the multi-criteria compromise solution ranking method to characterize the sudden risk driven by a few extreme values.

7. A method for multi-dimensional feature inversion and risk diagnosis of coal-related industrial zones according to claim 1 or 6, characterized in that, The steps for calculating the Prototype Driving Force Index (PDI) in S4 include: 1) Construct a four-dimensional feature space consisting of collaborative risk components, event-based risk components, and their sub-indicators, and define the cluster center of the lowest risk prototype as the collaborative pole and the cluster center of the highest risk prototype as the event pole. 2) Calculate the Prototype Driving Force Index (PDI) for any site: The closer the PDI value is to 1, the more the risk pattern is biased towards event-driven.

8. A multi-dimensional feature inversion and risk diagnosis system for coal-related industrial zones, characterized in that, The method for implementing the multi-dimensional feature inversion and risk diagnosis of coal-related industrial zones as described in any one of claims 1-7 includes: Sample and Data Management Module: Used to store and manage hierarchical feature sample sets and multi-source remote sensing data; Cascaded Inversion Module: Equipped with first and second target detection models, it performs cascaded operations from macroscopic positioning to microscopic feature interpretation, and outputs candidate target and internal ground feature information; Logical reasoning and cube generation module: Built-in white-box logical reasoning engine and data fusion algorithm, used to perform multi-scheme fusion decision to determine site type and generate multi-dimensional feature data cubes of coal-related sites; Risk Calculation and Diagnosis Module: Integrates the GWPCA-TOPSIS algorithm and the EWM-VIKOR algorithm to output dual-path risk components and perform risk prototype identification and PDI calculation.