An industrial carbon emission scenario simulation system and method based on multi-source data fusion
By integrating multi-source data and intelligent algorithms, the STIRPAT model was optimized. Combined with urban development factors and random forest algorithms, the problem of industrial carbon emission accounting errors and the disconnect between policy formulation was solved, achieving high-precision carbon emission measurement and risk warning, and supporting scientific low-carbon governance decisions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN UNIV OF TECH
- Filing Date
- 2026-03-05
- Publication Date
- 2026-06-26
AI Technical Summary
In urban carbon emission accounting, existing technologies suffer from large errors in industrial carbon emission accounting and a lack of alignment between data prediction and policy guidance, making it difficult to achieve high-precision city-level carbon emission measurement and effective policy support.
An industrial carbon emission scenario simulation system based on multi-source data fusion is constructed. By integrating high-resolution remote sensing, geographic information and deep learning, the STIRPAT model is optimized, and urban development factors and random forest algorithms are combined to achieve accurate calculation of carbon emissions at multiple scales. A multi-modal data fusion framework is also constructed to improve the accuracy of land use identification.
It has achieved high-precision measurement of carbon emissions at multiple scales at the provincial, municipal, and county levels, and provided a dynamic threshold monitoring and risk early warning mechanism to support scientific decision-making for differentiated low-carbon governance and policy formulation.
Smart Images

Figure CN122287302A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental protection, specifically to an industrial carbon emission scenario simulation system and method based on multi-source data fusion. Background Technology
[0002] As research on urban carbon emission accounting continues to develop, numerous studies are exploring different methodologies: at the theoretical modeling level, scholars have constructed multidimensional analytical frameworks including the STIRPAT model and the LMDI factor decomposition method; at the data application level, nighttime light remote sensing data and energy statistics methods have become important supports. However, focusing on the field of municipal carbon emission accounting, significant research bottlenecks still exist: 1. Significant errors exist in industrial carbon emission accounting. In the industrial sector, the accuracy of correlated data is limited when calculating city-level carbon emissions, making it difficult to adapt to diverse industrial scenarios. When applying macro-level accounting methods to the micro-scale, even though some scholars have attempted to extrapolate using provincial data downscaling, the lack of a systematic spatial allocation model stems from neglecting localized characteristics such as urban socio-economic conditions and industrial structure, resulting in a persistent dilemma of "clear provincial totals but ambiguous city-level structures." Simultaneously, the adaptability of measurement techniques in carbon emission accounting also faces challenges. Existing applications of light data are limited to construction land identification and linear regression modeling, and the spatiotemporal resolution potential of high-resolution remote sensing imagery has not been fully explored.
[0003] 2. Inadequate Connection Between Data Forecasting and Policy Guidance. Currently, there is a disconnect between carbon emission data forecasting and policy formulation: on the one hand, data forecasting models, lacking policy variables, struggle to reflect carbon emission trends after policy adjustments; on the other hand, policy formulation lacks accurate data support, leading to implementation deviations in industrial carbon emission control policies and hindering the efficient construction of a coordinated advancement mechanism. Some studies model solely based on historical trends in economic growth and energy consumption, failing to consider policy factors such as carbon taxes and carbon emission trading. This lack of consideration provides ineffective references for policy formulation. When addressing complex scenarios such as industrial restructuring and energy policy adjustments, forecast accuracy drops significantly, and the absence of dynamic threshold monitoring and risk warning mechanisms leaves potential over-emission risks in a state of "unknowability and difficulty in intervention" for extended periods. Summary of the Invention
[0004] The purpose of this invention is to provide both an industrial carbon emission scenario simulation system and a method based on multi-source data fusion. Both systems and methods aim to improve the accuracy of industrial carbon verification, emphasizing the crucial role of accurate industrial land identification in carbon verification. Based on this, a collaborative system of multi-source heterogeneous data fusion and intelligent algorithms is constructed. In carbon emission calculation, high-resolution remote sensing and geographic information are integrated, the STIRPAT model is optimized using deep learning, and a spatiotemporal adaptive mechanism is employed to achieve accurate calculation of carbon emissions at multiple scales (province, city, county). In the industrial land identification stage, high-resolution remote sensing image interpretation combined with deep learning object detection technology and semantic segmentation algorithms effectively reduces the missed detection rate of small industrial land. Simultaneously, a multimodal fusion framework integrating nighttime light data, POI information, and land use data is constructed, and further improved in accuracy through feature cross-validation and information entropy weighted fusion.
[0005] To achieve this objective, the present invention provides an industrial carbon emission scenario simulation system based on multi-source data fusion, comprising: The industrial carbon emission inversion module is used to construct urban development factors through different dimensions of comprehensive nighttime light data of the city, add a quadratic term of GDP to the existing STIRPAT model, and introduce urban development factors into the STIRPAT model to obtain the ISTIRPAT carbon emission inversion model. The city-level industrial carbon emissions are calculated through the ISTIRPAT carbon emission inversion model. The industrial land information extraction module is used to identify and calculate the density of industrial land based on a preset multi-source fusion dataset using the random forest algorithm, thereby obtaining industrial land density distribution data. Based on the industrial land density distribution data and the city-level industrial carbon emissions, the city-level industrial carbon emissions are decomposed into spatial grid units through spatial proxy and gridded allocation mechanisms to obtain gridded industrial carbon emission distribution data. The scenario simulation and early warning model is used to simulate and solve the early warning probability model based on historical and predicted municipal industrial carbon emission data, gridded industrial carbon emission distribution data, and preset growth rates of carbon emission driving factors. Multiple different carbon emission development paths are preset through scenario analysis. Based on multiple different carbon emission development paths, an early warning probability model is established. The early warning probability model is solved by Monte Carlo simulation to obtain the predicted values of municipal industrial carbon emissions under multiple different carbon emission development paths and the probability that the predicted values of municipal industrial carbon emissions exceed the preset peak values.
[0006] The beneficial effects of this invention are: To address the limitations of traditional carbon emission accounting methods in terms of spatiotemporal dimensions, this invention constructs a three-dimensional accounting system based on spatiotemporal geographic big data technology, featuring a "two-dimensional spatiotemporal structure and four-level deconstruction." Through the fusion of multi-source heterogeneous data (including 500m resolution nighttime light data, Sentinel-2 remote sensing imagery, and POI locations), a four-level spatial analysis framework of "national-provincial-municipal-500m grid" is built. The invention innovatively introduces the City Development Factor (CDE), combined with an improved STIRPAT downscaling model and a random forest algorithm, to achieve high-precision spatial inversion of provincial carbon emission data to the municipal and grid levels. This model, through a nonlinear spatial allocation mechanism, solves the problem of cross-scale data discontinuity, reveals spatiotemporal differentiation characteristics, and provides a scientific basis for differentiated regional governance.
[0007] To address the dynamic disconnect between policy formulation and carbon emission evolution, this invention develops a three-in-one intelligent decision-making system integrating policy, simulation, and optimization. In scenario analysis, by integrating ISTIRPAT driving factors with multi-source spatiotemporal emission reduction policies and data, a multi-scenario linkage analysis module is developed, pre-setting three frameworks: baseline scenario, economy-oriented scenario, and green transition scenario. In simulation optimization, this invention uses simulation and iterative feedback mechanisms to interactively adjust carbon emission simulation results with policy parameters, achieving dynamic simulation of policy intervention. Coupled with carbon emission thresholds and industry-related factors, Monte Carlo simulation generates a probability distribution of abnormal carbon emission risks under multiple production scenarios, providing falsifiable decision support for differentiated low-carbon governance paths. Attached Figure Description
[0008] Figure 1 This is a framework diagram for provincial and municipal industrial downscaling estimation in this invention; Figure 2 This is a system functional diagram of the present invention; Figure 3 This is the carbon emission reduction decision simulation interface of the present invention; Figure 4 This is a schematic diagram of the structure of the present invention. Detailed Implementation
[0009] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to represent selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0010] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments: Example 1 like Figure 4 As shown, an industrial carbon emission scenario simulation system based on multi-source data fusion includes: The industrial carbon emission inversion module is used to construct urban development factors through different dimensions of comprehensive nighttime light data of the city, add a quadratic term of GDP to the existing STIRPAT model, and introduce urban development factors into the STIRPAT model to obtain the ISTIRPAT carbon emission inversion model. The city-level industrial carbon emissions are calculated through the ISTIRPAT carbon emission inversion model. The industrial land information extraction module is used to identify and calculate the density of industrial land based on a preset multi-source fusion dataset using the random forest algorithm, thereby obtaining industrial land density distribution data. Based on the industrial land density distribution data and the city-level industrial carbon emissions, the city-level industrial carbon emissions are decomposed into spatial grid units through spatial proxy and gridded allocation mechanisms to obtain gridded industrial carbon emission distribution data. The scenario simulation and early warning model is used to simulate and solve the early warning probability model based on historical and predicted municipal industrial carbon emission data, gridded industrial carbon emission distribution data, and preset growth rates of carbon emission driving factors. Multiple different carbon emission development paths are preset through scenario analysis. Based on multiple different carbon emission development paths, an early warning probability model is established. The early warning probability model is solved by Monte Carlo simulation to obtain the predicted values of municipal industrial carbon emissions under multiple different carbon emission development paths and the probability that the predicted values of municipal industrial carbon emissions exceed the preset peak values.
[0011] In some embodiments, the preset multi-source fusion dataset includes, but is not limited to, the following data types: The carbon emission factor database is 2024 industry enterprise emission factor CO2 factor data, sourced from the National Greenhouse Gas Emission Factor Database, obtained via web download; GOSAT data (i.e., satellite-observed CO2 flux data) comes from the GOSAT official website, obtained via web download; Boundary data (including but not limited to industrial carbon emissions, population, total GDP, and the proportion of secondary industry) includes .shp files for the People's Republic of China (nationwide), Hubei Province, and counties within Hubei Province, sourced from the Center for Resources and Environment Science and Technology, Chinese Academy of Sciences; the 2008-2020 Hubei Provincial Statistical Yearbook is provided by the Hubei Provincial Bureau of Statistics, obtained via web download; POI data (used for marking in electronic maps or geographic information systems). Point data (POI data specifically refers to industrial enterprise points of interest data) representing entities with specific functions or names are in the form of POI.xlsx files, obtained from Amap via web crawling; CEADs (China Carbon Emission Database) lists include multiple related lists from the CEADs China Carbon Accounting Database, downloaded from a webpage; nighttime light data is 2000-2023.zip (original TIF format), from the National Earth System Science Data Center, downloaded from a webpage; the original data of hubeiclass and wuhanclass (from satellite remote sensing images of Hubei Province and Wuhan City) are in TIF format, existing as Hubei_result.zip and Wuhan_result.zip respectively. The former was classified and exported online using code written on the Google Earth Engine platform, while the latter was processed by ArcGIS software; Hubei_day data (a dataset generated by daily-scale dynamic monitoring and short-term forecasting in Hubei Province) is a Hubei.csv file from the CarbonMonitor platform, downloaded from a webpage.
[0012] Table 1 Data Description Table Based on geographic spatiotemporal big data (i.e., the multi-source fusion dataset), this invention employs ISTIRPAT and panel regression models, as well as random forest algorithms, to achieve high-resolution extraction of industrial land use and calculation of industrial carbon emissions in terms of high-precision accounting. In terms of intelligent control, scenario analysis and Monte Carlo simulation are used to classify early warning levels.
[0013] In some embodiments, the present invention is based on GOSAT fossil fuel combustion CO2 flux data (L4A), using provincial boundary data of China to extract fossil fuel combustion CO2 flux of each province (excluding Tibet Autonomous Region, Hong Kong Special Administrative Region, Macao Special Administrative Region and Taiwan Province) from 2010 to 2020, and using energy consumption data in the statistical yearbooks of each province and IPCC carbon emission factor library data, to calculate the industrial carbon emissions of each province using the carbon emission factor method, and to construct a panel data econometric model using the carbon flux of each province and the industrial carbon emissions.
[0014] like Figure 1 As shown, the prediction framework for obtaining municipal carbon emissions from provincial carbon emissions based on the ISTIRPAT model is as follows: When constructing the city development factor CDE, the regional light intensity value DNQ, regional average brightness MDN, and the total area of regional nighttime light grids SA, obtained through data preprocessing of nighttime light images, are used to construct the city development factor CDE. At the provincial carbon emission level, the provincial carbon emission inventory, provincial GDP, provincial total construction output value CI, total population of each province People, provincial city development factor CDE, and provincial secondary industry output value SeGDP are input into the ISTIRPAT model for construction, and the regression coefficients of the ISTIRPAT model are obtained through panel regression. In the process of obtaining municipal carbon emissions, the total population of each city, municipal GDP, municipal total construction output value CI, municipal secondary industry output value SeGDP, and municipal city development factor CDE are combined with the aforementioned ISTIRPAT model regression coefficients to finally calculate the municipal carbon emissions.
[0015] The STIRPAT basic model only uses statistical data to calculate carbon emissions, but the actual statistical data has errors in the statistical process. Therefore, there is a large gap between the actual carbon emissions calculated only by provincial statistical data and the city-level carbon emissions
[29] . Given that the accuracy of nighttime light remote sensing data is high at both the provincial and city levels, and that urbanization is closely related to urban energy consumption and its carbon emissions, the process of urban development is the main driving factor for changes in energy carbon emissions. Based on the extended STIRPAT model, this invention proposes a new downscaling method, which introduces the urban development factor (CDE). The urban development factor constructed from nighttime light remote sensing data is extended to the STIRPAT model as a dimensionless variable CDE, resulting in an improved STIRPAT (ISTIRPAT) carbon emission inversion model. Then, the model parameters are trained using provincial data to invert city-level carbon emissions.
[0016] Finally, taking Hubei Province as an example, the industrial carbon emission data of each city in the province were estimated by substituting the carbon emission data of each city into the corresponding panel data econometric model of Hubei Province through the construction of the ISTIRPAT carbon emission inversion model.
[0017] Some technical solutions utilize different dimensions of comprehensive urban nighttime illumination data to construct urban development factors, including: The different dimensions of comprehensive urban nighttime illumination data include average brightness, total brightness, and luminous intensity per unit area. The method for constructing the urban development factor CDE using these different dimensions of comprehensive urban nighttime illumination data is as follows: ; Wherein, MDN is the average brightness of the target city's comprehensive nighttime illumination data in the satellite remote sensing image, SUML is the total brightness of the target city's comprehensive nighttime illumination data in the satellite remote sensing image, and DNQ is the luminous intensity per unit area of the target city's comprehensive nighttime illumination data in the satellite remote sensing image. The first regression coefficient, The second regression coefficient, The third regression coefficient, The total number of pixels in the satellite remote sensing image. The first in satellite remote sensing image Each pixel area The first in satellite remote sensing image The light brightness value of each pixel; MDN is to achieve the same results in urban and rural areas. The brightness is averaged under different values to correct for carbon emission estimation biases caused by differences in development levels. ; SUML is the sum of the light intensity values of all pixels in a satellite remote sensing image, directly reflecting the overall development level and total energy consumption of a city or region. ; DNQ further eliminates the impact of urban-rural area differences on carbon emissions by standardizing the luminous intensity per unit area using the ratio of total brightness to total area based on comprehensive nighttime illumination data of the target city. ; SA represents the sum of the nighttime light pixel areas of the target city.
[0018] By integrating three dimensions of nighttime light data—average brightness, total brightness, and luminous intensity per unit area—and introducing adjustable regression coefficients, the constructed urban development factor simultaneously reflects the average intensity, overall scale, and spatial agglomeration of the target city's development. This overcomes the limitations of a single light indicator and improves the accuracy and stability of quantitative representation of complex characteristics such as urban economic development, energy consumption, and spatial structure, enabling a more comprehensive and accurate representation of the overall level of urban development.
[0019] Some technical solutions involve adding a quadratic term of GDP to the existing STIRPAT model and incorporating urban development factors into the STIRPAT model to obtain the ISTIRPAT carbon emission inversion model. As one of the environmental impact factor models, IPAT (a foundational conceptual model for analyzing the impact of human activities on environmental stress) and its extended versions have been widely used to analyze the impact of socioeconomic factors on pollutant emissions. However, this model requires a unified dimension of environmental impact factors and cannot accurately reflect the environmental impact when these factors change. Dietz improved the IPAT model by proposing a stochastic regression model called STIRPAT.
[0020] The existing STIRPAT model is as follows: ; in, Environmental conditions are driving factors. For population driving factors, As a driving force for economic development As a driving force for technological innovation, These are the estimated coefficients for population drivers. These are the estimated coefficients for the driving factors of economic development. The estimated coefficients of the driving factors of technological innovation. Here are the STIRPAT model parameters, e is the error term; I, P, A, and T represent environmental conditions, population, economic development, and technological innovation, respectively. In some embodiments of this invention, environmental conditions... Represents carbon emissions (Tan), and its driving factors Represents the population (People) Represents Gross Domestic Product (GDP) This represents the added value of the secondary sector (SeGDP). Carbon emissions primarily involve energy combustion in the cement industry, with energy mainly related to the use of coal, diesel, and other combustion materials. The use of these fuels is closely linked to the production of urban heavy industry, light industry, and manufacturing. Data collection for the cement industry at the city level is relatively difficult. However, the construction industry constitutes a major component of cement production and consumption, with most cement consumption occurring at construction sites. Therefore, the gross output value of the secondary sector (SeGDP) and the gross output value of the construction industry (JianZhu) are included in the model.
[0021] The Kuznets curve, proposed by Nobel laureate economist Simon Smith Kuznets in the 1950s, is used to analyze the relationship between income level and economic growth. Research shows that with economic growth, social income inequality first increases and then decreases, exhibiting an inverted U-shaped curve relationship. Considering the environmental Kuznets curve, this paper extends the existing STIRPAT model by adding a quadratic term to GDP. Finally, the urban development factor constructed from nighttime light remote sensing data is used as a dimensionless variable (CDE) as a downscaling factor to extend the STIRPAT model, resulting in an improved ISTIRPAT carbon emission inversion model.
[0022] By adding a quadratic term of GDP to the existing STIRPAT model and incorporating urban development factors into the STIRPAT model, the resulting ISTIRPAT carbon emission inversion model is as follows: ; Wherein, CDE represents urban development factors, Tan represents provincial industrial carbon emissions, People represents population, GDP represents regional gross domestic product, SeGDP represents the added value of the secondary industry, JianZhu represents the total output value of the construction industry, e represents the random error term, d represents the population elasticity coefficient, f represents the linear elasticity coefficient of GDP, g represents the quadratic coefficient of GDP, h represents the elasticity coefficient of the added value of the secondary industry, and j represents the elasticity coefficient of the total output value of the construction industry.
[0023] Adding a quadratic term to GDP to the existing STIRPAT model can capture the nonlinear relationship between carbon emissions and economic development, especially the inverted U-shaped trend described by the Environmental Kuznets Curve (EKC), thus more accurately reflecting the marginal effect of economic growth on carbon emissions. At the same time, introducing the Urban Development Factor (CDE) as a direct driving factor enables the model to more specifically quantify the impact of construction, energy, and structural transformation in the urbanization process on carbon emissions, improve the model's fitting accuracy for provincial carbon emission inversion, and provide a more reliable parameter basis for downscaling estimation at the city level.
[0024] Some technical solutions use the ISTIRPAT carbon emission inversion model to calculate municipal-level industrial carbon emissions, including: Estimating provincial industrial carbon emissions using the carbon emission factor method: ; The provincial industrial carbon emissions estimated using the carbon emission factor method were then calibrated using a panel regression model. The panel regression model is as follows: ; in, The provincial industrial carbon emissions are estimated using the carbon emission factor method. For the industry China Energy Consumption volume The total number of industries, The total energy consumption in an industry; This is the standard coal conversion factor; The carbon emission factor is 12, where 12 represents carbon. The relative atomic mass is 44, and the relative molecular mass is 44, which is the relative molecular mass of carbon dioxide (CO2). The provincial industrial carbon emissions were estimated using the carbon emission factor method. For satellite observation of CO2 flux data, This is a provincial coefficient. This is the error term. Used to convert carbon emissions into carbon dioxide emissions.
[0025] The provincial industrial carbon emissions calculated using the carbon emission factor method are substituted into the ISTIRPAT carbon emission inversion model to obtain the random error term, population elasticity coefficient, linear elasticity coefficient of GDP, quadratic elasticity coefficient of GDP, elasticity coefficient of secondary industry added value, and elasticity coefficient of construction industry total output value of the ISTIRPAT carbon emission inversion model. Based on the random error term, population elasticity coefficient, linear elasticity coefficient of GDP, quadratic elasticity coefficient of GDP, elasticity coefficient of secondary industry added value, and elasticity coefficient of construction industry total output value of the ISTIRPAT carbon emission inversion model, and the urban development factor constructed based on comprehensive nighttime light data of the city, the city-level industrial carbon emissions are calculated by inversion using pre-set population, regional GDP, secondary industry added value, and construction industry total output value. The city-level industrial carbon emissions calculated by inversion using the pre-set population, regional GDP, secondary industry added value, and construction industry total output value are derived from the historical statistical data of the target city for carbon emission prediction.
[0026] A quantitative relationship between provincial industrial carbon emissions and satellite-observed CO2 flux data was established using a panel regression model. This relationship was then used as a bridge to accurately downscale provincial industrial carbon emissions retrieved from the STIRPAT model to the municipal level. This improved the accuracy, spatial continuity, and scientific rigor of municipal-level industrial carbon emission estimates, and solved the accounting challenges at the municipal level caused by the scarcity of direct observation data and inconsistent statistical standards.
[0027] In some embodiments, it is necessary to verify the accuracy of the ISTIRPAT carbon emission retrieval model. Taking Hubei Province as an example, the city-level carbon emission estimates are compared with the "2010 Carbon Dioxide Emission Inventory of 182 Chinese Cities" published by CEADS, and the coefficient of determination of the ISTIRPAT carbon emission retrieval model is calculated. The formula is as follows: in, The total number of samples used for verification or testing. For the first The actual observed values of each sample The model is for the first The predicted value for each sample.
[0028] In some embodiments, this invention uses the Google Earth Engine (GEE) cloud platform as a carrier, calling functions such as cloud removal and cropping to process Sentinel-2 imagery, while simultaneously importing VIIRS nighttime light imagery data, adding its nighttime light features as new bands to form a multi-feature, multi-band combination, thereby improving the accuracy of industrial land extraction. A random forest classifier method is employed, using multiple decision trees to vote or average the results for prediction. POI data is used as an aid, visually interpreting the classification results and comprehensively validating the POIs to achieve intelligent and batch extraction of industrial land across the country.
[0029] The extracted industrial land raster data focuses on Hubei Province. The total industrial carbon emissions were spatially decomposed into a grid, and industrial land density was selected as a spatial proxy. ArcGIS Pro software was used for reclassification and raster aggregation, converting the industrial land distribution raster data into industrial land density raster data.
[0030] By using an industrial land density grid, industrial carbon emissions in various cities of Hubei Province are allocated proportionally, thus decomposing industrial carbon emissions to each pixel unit. This method can reflect the spatial distribution of industrial carbon emissions and has a higher spatial resolution than GOSAT's CO2 flux data, enabling it to more accurately reflect the spatial distribution of industrial carbon emissions.
[0031] In terms of result verification, the gridded industrial carbon emission results of this work are compared with the MIXv2 inventory to verify the reliability of the gridded industrial carbon emission data.
[0032] In some technical solutions, a method for identifying and calculating the density distribution of industrial land using a random forest algorithm based on a pre-set multi-source fusion dataset to obtain industrial land density distribution data includes: the pre-set multi-source fusion dataset includes satellite remote sensing image data of the target city; all original pixels with artificial labels in the satellite remote sensing images are used as training sample sets to train the random forest algorithm; each original pixel sample contains the image features of the original pixel and the corresponding artificially labeled classification label; during training of the random forest algorithm, based on the Bagging strategy, a single decision tree in the random forest algorithm randomly selects multiple samples with replacement from the training sample set as a training subset; the decision tree randomly selects multiple image features from all image features of the training subset as an image feature subset; based on the image feature subset and the corresponding classification label, the split point that minimizes the Gini impurity is selected to perform a binary partition on all samples in the training subset; all samples in the training subset are divided into two child nodes; the training process is repeated for the two child nodes until a pre-set termination condition is met to obtain a trained single decision tree; the Gini impurity calculation formula is as follows: ; in, The Gini impurity of the current split node. This refers to the node in the decision tree that is currently to be evaluated or split. For nodes Category tags in The sample proportion, where C is the total number of categories for the classification label; Random Forest is a supervised classification algorithm based on decision tree ensemble. It improves classification accuracy by constructing multiple differentiated decision trees and aggregating the results. Its core advantage lies in combining the Bagging ensemble strategy and random feature selection, effectively reducing the risk of overfitting, and making it suitable for high-dimensional remote sensing data classification.
[0033] The decision tree performs binary splits on samples based on features (such as night light intensity and NDBI), selects the split point that minimizes Gini impurity, and repeats the split until the preset termination condition is met (such as node purity ≥ 95% or reaching the maximum depth).
[0034] Because single decision trees are prone to overfitting, sensitive to noise, and have weak generalization ability, this invention introduces the random forest algorithm to address the limitations of single trees. The algorithm implements a backbone Bagging strategy, which involves sampling with replacement from the sample set to generate B subsets, with each decision tree trained independently. Classification results are aggregated using majority voting. The trained single decision trees are then used to identify industrial land use in the image to be predicted. Finally, the classification results from all trained decision trees are aggregated using majority voting to obtain the industrial land identification data, as shown in the following formula: ; in, This represents the final predicted category of the random forest algorithm for the current training sample set. For the total number of decision trees, For the first The predicted results for each tree, For indicator functions, For a maximum value operator, Candidate category labels; Density calculations were performed based on industrial land identification data to obtain industrial land density distribution data: The industrial land identification data contains multiple raw pixels, each with multiple corresponding classification labels. The extracted industrial land data is known to be raster data with a raw pixel spatial resolution of 10m x 10m. This land classification data is reclassified using ArcGIS software, and each raw pixel is assigned a value: ; The spatial resolution of the original pixels of the reclassified industrial land identification data is aggregated into pixels with a set spatial resolution to obtain a set pixel. The spatial resolution of the set pixel is divided by the spatial resolution of the original pixels to obtain the total number of original pixels contained in the set pixel. The ratio of the number of original pixels with a value of 1 to the total number of original pixels is used as the industrial land density distribution data in the set pixel.
[0035] In some embodiments, the set spatial resolution pixel is a 500m x 500m spatial resolution pixel. Original pixels with a spatial resolution of 10m x 10m are aggregated into a set 500m x 500m spatial resolution pixel. In this case, in the industrial land distribution raster data, the pixel value of each pixel is the number of 10m x 10m industrial land grids contained in the 500m x 500m grid. Through scale transformation, the refined industrial land identification results (10m) are integrated into a statistical unit (500m) suitable for macro-regional analysis, so that the value of each low-resolution pixel directly represents the spatial proportion or total area of industrial land within it. This significantly simplifies the data volume and improves computational efficiency while retaining key spatial distribution information.
[0036] By integrating the classification results of multiple decision trees using the random forest algorithm, the accuracy and robustness of industrial land identification are effectively improved, and the risk of overfitting is reduced. The Bagging strategy and random feature selection enhance the model's generalization ability to high-dimensional remote sensing image features, ensuring the stability of the identification results. Finally, through spatial aggregation and density calculation, continuous and high-resolution industrial land density distribution data are generated, providing a highly reliable data foundation for urban industrial activity monitoring, accurate carbon emission inversion, and spatial planning.
[0037] In some technical solutions, based on industrial land density distribution data and municipal industrial carbon emissions, a method is used to decompose municipal industrial carbon emissions into spatial grid units through spatial proxy and gridded allocation mechanisms to obtain gridded industrial carbon emission distribution data. This method includes: dividing the target city into regions using set pixels; using the ratio of the industrial land density distribution data of each set pixel to the sum of the industrial land density distribution data of all set pixels as the allocation weight of the corresponding set pixel; and multiplying the municipal industrial carbon emissions of the target city by the allocation weight of the corresponding set pixel to obtain the carbon emissions of the corresponding set pixel.
[0038] By effectively decomposing the macro-level total industrial carbon emissions at the municipal level into a defined grid scale and using industrial land density as a spatial proxy variable, the distribution of carbon emissions is closely correlated with the geographical location of actual industrial activities, thereby generating high spatial resolution gridded industrial carbon emission distribution data. This transformation of carbon emissions from administrative statistical units to continuous spatial units significantly improves the accuracy and rationality of the spatial representation of carbon emissions.
[0039] Scenario analysis is a crucial decision-making tool for addressing future uncertainties. Its core lies in identifying key driving factors and constructing a logically consistent hypothetical framework to systematically extrapolate various possible future scenarios. This method breaks away from the reliance of traditional forecasting models on single trends, emphasizing the interaction of economic, technological, policy, and social factors within complex systems. Taking the six-step method proposed by the Stanford Research Institute as an example, scenario analysis forms a closed-loop analysis process from defining the research scope and identifying key variables to scenario narrative and strategy verification. This ensures scientific rigor while providing visual support for cross-disciplinary collaborative decision-making, and has now become a standard method for long-term strategic research such as urban planning and climate governance.
[0040] Applying scenario analysis to carbon emission prediction has advantages in three key aspects: First, it comprehensively integrates multiple driving factors. Based on the STIRPAT model, it can simultaneously quantify the nonlinear impact of factors such as population expansion, per capita GDP growth, and changes in carbon intensity on carbon emissions, and clearly reveal the mechanism by which technological innovation plays a role in offsetting the carbon emission pressure brought about by economic growth.
[0041] Second, it provides precise support for dynamic policy simulation. By flexibly adjusting key parameters such as GDP growth rate and secondary industry GDP growth rate, scenario analysis can accurately determine the threshold for technology promotion and identify the optimal window for policy intervention, thus contributing to the scientific and timely nature of policy formulation.
[0042] Third, it significantly enhances risk early warning capabilities. Scenario analysis can help policymakers anticipate potential crises and design flexible response plans based on different scenarios. This mapping relationship of "multi-scenario and multi-strategy" provides solid scientific support for coordinating economic growth and low-carbon goals, becoming a key scientific anchor for balancing the two.
[0043] This invention uses scenario analysis as its foundational framework, integrating improved nighttime light data with the STIRPAT model to construct a dynamic carbon emission prediction system. By establishing a quantitative model to analyze the impact of population size, economic development, and technological level on carbon emissions, three typical development paths are defined: a baseline scenario continuing existing trends, a high-speed scenario simulating rapid economic growth, and a green scenario focusing on the context of low-carbon transformation. By dynamically adjusting the parameters of each driving factor, carbon emissions under different scenario paths are simulated, thus providing data-driven decision-making support for local governments to formulate differentiated carbon control policies.
[0044] Some technical solutions, based on historical and predicted municipal industrial carbon emission data, gridded industrial carbon emission distribution data, and preset growth rates of carbon emission driving factors, use scenario analysis to pre-determine multiple different carbon emission development paths; and establish early warning probability models based on these multiple different carbon emission development paths, including: Under the pre-defined carbon emission development path, each carbon emission driving factor variable annual growth rate It is not fixed, but a variable is set. It follows a normal distribution with a given expected growth rate as the mean and a given variance. ,in, For the first Annual growth rate of each carbon emission driver variable For the first The mean of the annual growth rates of the individual carbon emission driving factors. For the first The standard deviation of the annual growth rate of each carbon emission driver variable. It follows a normal distribution. The representation follows a normal distribution, after... After the new year, the carbon emission driving factors variables are expressed as follows: ; in, For the first One carbon emission driving factor variable, For the first The year's first One carbon emission driving factor variable, For the first Baseline year observations of the carbon emission driving factor variables, For the first The carbon emission driving factor variables are in Annual random growth rate This represents the total brightness of the lights at night. Light intensity per unit area The total output value of the construction industry; The carbon emission driving factor variables are decomposed into a sum of constant and random terms after taking the logarithm: ;remember , ; Substituting the carbon emission driving factor variables into the carbon emission inversion model, we obtain: ; in, For municipal-level industrial carbon emissions, For the natural constant An exponential function with base 0. For the first Regression coefficients of individual carbon emission driving factor variables, For the first The regression coefficient of annual GDP value For the first Annual GDP This is the scaling factor; Carbon emission inversion model by substituting the log-linearized carbon emission driving factor variables: ; in, For the certainty of GDP, The stochastic component of GDP Target probability: Assume the government sets a threshold for peak carbon emissions in a given year. Based on this threshold, the government wants to know whether carbon emissions will exceed it under current policies. Target probability prediction is used to derive the warning level for each policy scenario, helping the government formulate appropriate strategies. The basic idea is to calculate the probability of peak carbon emissions exceeding the threshold under each policy scenario, with different probabilities resulting in different warning levels.
[0045] ; in, For the city-level industrial carbon emissions in the nth year in the future Exceeding the preset carbon emission safety threshold The probability, The preset carbon emission safety threshold, This represents the deterministic portion of all carbon emission driver variables.
[0046] By setting the annual growth rate of carbon emission drivers as a normally distributed random variable and taking its logarithmic decomposition into deterministic and random components, the model can systematically quantify the uncertainty of future carbon emissions. Combining scenario analysis to simulate multiple development paths, and constructing an early warning probability model based on a log-linearized carbon emission inversion model, the model ultimately calculates the probability of carbon emissions exceeding the safety threshold, enabling dynamic and probabilistic assessment of municipal industrial carbon emission risks. (Due to the quadratic term...) The existence of these factors makes directly calculating the target probability difficult. This invention uses Monte Carlo simulation to solve for this probability. The carbon emission drivers include, but are not limited to, industrial carbon emission thresholds (million tons), population growth rate, GDP growth rate, secondary industry GDP growth rate, and construction industry value-added growth rate.
[0047] In some technical solutions, the method of simulating and solving the early warning probability model using the Monte Carlo method to obtain the predicted values of municipal industrial carbon emissions under multiple different carbon emission development paths and the probability that the predicted values of municipal industrial carbon emissions exceed a preset peak includes: ; in, The total number of simulations for Monte Carlo is calculated. In the first In the Monte Carlo simulation, the predicted future carbon emissions are calculated by the early warning probability model. As an indicator function, it indicates that for the th A Monte Carlo simulation is valid if and only if the carbon emission results of that simulation are... Greater than the threshold When the time is right, the value of this function is 1; otherwise, its value is 0, based on the city-level industrial carbon emissions in the nth year. Exceeding the preset carbon emission safety threshold probability Carbon emission control measures will be implemented for target urban areas corresponding to pixels whose carbon emissions exceed the threshold.
[0048] By transforming deterministic carbon emission forecasts into probabilistic risk assessments, and through large-scale simulations of multiple possible development paths, the objective probability of future carbon emissions exceeding safe thresholds is quantified. This achieves a decision-making upgrade from point estimation to risk distribution, enabling control measures to shift from passive, reactive emission reduction management to proactive, preventative risk control, significantly improving the scientific rigor, foresight, and resource allocation efficiency of urban carbon emission early warning and control systems.
[0049] The Monte Carlo method, also known as the random sampling method or statistical experiment method, is a numerical calculation method based on probability and statistics theory. Its main idea is to solve problems by utilizing the statistical laws of random variables through a large number of random experiments. In this work, for each policy option, 10,000 carbon emission predictions were simulated based on the above model (with relevant random variables simulated by computer). The number of times carbon emissions exceeded a given threshold by 2030 was counted, and the probability was estimated using frequency.
[0050] Example 2 An industrial carbon emission scenario simulation method based on multi-source data fusion according to the system includes: Urban development factors are constructed by using different dimensions of comprehensive nighttime light data. A quadratic term of GDP is added to the existing STIRPAT model. The urban development factors are then introduced into the STIRPAT model to obtain the ISTIRPAT carbon emission inversion model. The city-level industrial carbon emissions are calculated using the ISTIRPAT carbon emission inversion model. Based on a pre-set multi-source fusion dataset, the random forest algorithm is used to identify and calculate the density of industrial land to obtain industrial land density distribution data. Based on the industrial land density distribution data and the city-level industrial carbon emissions, the city-level industrial carbon emissions are decomposed into spatial grid units through spatial proxy and gridded allocation mechanisms to obtain gridded industrial carbon emission distribution data. Based on historical and predicted municipal industrial carbon emission data, gridded industrial carbon emission distribution data, and preset growth rates of carbon emission driving factors, multiple different carbon emission development paths are preset using scenario analysis. Based on these multiple carbon emission development paths, an early warning probability model is established. The model is then solved using the Monte Carlo method to obtain the predicted municipal industrial carbon emission values under multiple different carbon emission development paths and the probability that the predicted municipal industrial carbon emission values will exceed the preset peak values.
[0051] In some embodiments, the present invention is implemented in the form of a web page system. Figure 2 The system functional architecture diagram of this invention is as follows: The architecture of this invention consists of two main parts: the "calculation module" and the "regulation module". The calculation module is responsible for generating a national / provincial industrial carbon emission map and performing intelligent daily-scale industrial carbon emission prediction. Its function extends to displaying carbon emission trends and carbon emission maps, and further refines the spatial distribution of carbon emissions into planar spatial distribution, 3D distribution display, and spatial analysis. The regulation module is responsible for carbon emission reduction decision simulation. Through self-defined scenario simulation and setting thresholds, it predicts the peak carbon emissions for future specified years to achieve intelligent detection of the probability of exceeding standards.
[0052] The web-based system of this invention ("Carbon Detection Number" one-stop service platform) mainly includes a calculation module and a control module, such as... Figure 2 As shown: This project utilizes information visualization plugins such as Echarts and MapBox, and selects different ways to display data.
[0053] The “Carbon Detection” one-stop service platform’s interface uses green elements that conform to carbon emission reduction, and mainly includes the system name, system logo, and function buttons to guide users.
[0054] The Policy and Information module showcases the latest relevant policies and environmental quality documents, including scrollable national and local policies on industrial carbon emissions, political news, environmental news, local news updates, and video news. Users can browse relevant information and select functional modules according to their needs on the main interface.
[0055] The main interface of the "Carbon Detection" one-stop service platform The main interface is divided into two main functional areas: carbon accounting and carbon regulation, as follows: The carbon accounting system mainly includes four sub-functions: "National Industrial Carbon Emission Map", "Provincial Industrial Carbon Inventory Map", "Intelligent Prediction of Industrial Carbon Emissions", and "Carbon Emission Spatial Analysis".
[0058] The carbon emission reduction decision simulation module displays changes in carbon emissions under three development scenarios: baseline, green, and rapid development, by adjusting the growth rate indicators of factors influencing industrial carbon emissions. Users can interactively change the growth rate of each driving factor to display carbon emission trends under customized scenarios. The module also allows users to view the industrial carbon emission values at which emissions peak under a given scenario. Users can set a desired carbon emission threshold for a future year when emissions peak. The system calculates the probability that carbon emissions will exceed the predetermined threshold in a given future year under the current growth rate and provides an early warning level, thus providing a reference for the intensity and direction of carbon emission reduction policies (e.g., ...). Figure 3 ).
[0059] Figure 3The diagram shows an interactive system interface for carbon emission reduction decision simulation: After selecting the target region (Hubei Province), users can select different modules through the function menu on the left. In the scenario parameter setting area on the right side of the target region (Hubei Province), users can input or adjust key driving factors, including preset values for industrial carbon emission threshold (million tons), population growth rate, GDP growth rate, secondary industry GDP growth rate, and construction industry added value growth rate. After clicking the calculation button, the system will perform Monte Carlo simulation based on the built-in early warning probability model (such as the ISTIRPAT model) and dynamically present carbon emission prediction trend curves under multiple different development paths in the central line graph. At the same time, the probability of carbon emissions exceeding the preset safety threshold is directly output in the results module at the bottom right in the form of visual charts and percentages, thereby realizing integrated analysis from parameter setting, scenario simulation, trend visualization to risk quantification.
[0065] Warning level classification Based on relevant information, the warning levels are classified as follows: Table 2 Warning Level Table This invention focuses on improving the accuracy of industrial carbon verification, emphasizing the crucial role of accurate industrial land identification in carbon verification. Based on this, it constructs a collaborative system of multi-source heterogeneous data fusion and intelligent algorithms. In carbon emission accounting, it integrates high-resolution remote sensing, geographic information, and enterprise production data, optimizes the ISTIRPAT model using deep learning, and utilizes a spatiotemporal adaptive mechanism to achieve accurate calculation of carbon emissions at multiple scales (province, city, county). In the industrial land identification stage, it employs high-resolution remote sensing image interpretation combined with deep learning object detection technology and semantic segmentation algorithms to effectively reduce the missed detection rate of small industrial facilities. Simultaneously, it constructs a multimodal fusion framework integrating nighttime light data, POI information, and land use data, further improving the accuracy of industrial land identification through feature cross-validation and information entropy weighted fusion.
[0066] This invention innovatively constructs a two-way dynamic interaction mechanism between forecasting and policy. By deeply embedding a policy simulation module into the core model, it tracks the policy implementation effects in real time, synchronously importing feedback data into the model to trigger intelligent comparison between predicted and actual carbon emission values. Once a significant deviation is detected, the system immediately activates an early warning function, visually presenting the warning prompts and interactively optimizing carbon emission thresholds and characteristic growth rates, forming a closed-loop iteration of "prediction-decision-feedback-optimization." This continuously evolving operational mode effectively improves the matching degree between forecast accuracy and policy effectiveness, significantly enhances the scientific nature of policy formulation and implementation efficiency, and provides strong support for the efficient advancement of carbon emission management.
[0067] Example 3 The present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the method described in Embodiment 2.
[0068] This invention can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented in whole or in part as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0069] It will be readily understood by those skilled in the art that the above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, combinations, substitutions, improvements, etc., made under the spirit and principles of the present invention are included within the protection scope of the present invention.
[0070] The contents not described in detail in this specification are existing technologies known to those skilled in the art.
Claims
1. An industrial carbon emission scenario simulation system based on multi-source data fusion, characterized in that, It includes: The industrial carbon emission inversion module is used to construct urban development factors through different dimensions of comprehensive nighttime light data of the city, add a quadratic term of GDP to the existing STIRPAT model, and introduce urban development factors into the STIRPAT model to obtain the ISTIRPAT carbon emission inversion model. The city-level industrial carbon emissions are calculated through the ISTIRPAT carbon emission inversion model. The industrial land information extraction module is used to identify and calculate the density of industrial land based on a preset multi-source fusion dataset using the random forest algorithm, thereby obtaining industrial land density distribution data. Based on the industrial land density distribution data and the city-level industrial carbon emissions, the city-level industrial carbon emissions are decomposed into spatial grid units through spatial proxy and gridded allocation mechanisms to obtain gridded industrial carbon emission distribution data. The scenario simulation and early warning model is used to simulate and solve the early warning probability model based on historical and predicted municipal industrial carbon emission data, gridded industrial carbon emission distribution data, and preset growth rates of carbon emission driving factors. Multiple different carbon emission development paths are preset through scenario analysis. Based on multiple different carbon emission development paths, an early warning probability model is established. The early warning probability model is solved by Monte Carlo simulation to obtain the predicted values of municipal industrial carbon emissions under multiple different carbon emission development paths and the probability that the predicted values of municipal industrial carbon emissions exceed the preset peak values.
2. The industrial carbon emission scenario simulation system based on multi-source data fusion according to claim 1, characterized in that: Methods for constructing urban development factors based on different dimensions of comprehensive urban nighttime illumination data include: The different dimensions of comprehensive urban nighttime illumination data include average brightness, total brightness, and luminous intensity per unit area. The method for constructing the urban development factor CDE using these different dimensions of comprehensive urban nighttime illumination data is as follows: ; Wherein, MDN represents the average brightness of the target city's comprehensive nighttime illumination data in the satellite remote sensing image, SUML represents the total brightness of the target city's comprehensive nighttime illumination data in the satellite remote sensing image, and DNQ represents the luminous intensity per unit area of the target city's comprehensive nighttime illumination data in the satellite remote sensing image. The first regression coefficient, The second regression coefficient, The third regression coefficient, The total number of pixels in the satellite remote sensing image. The first in satellite remote sensing image Each pixel area The first in satellite remote sensing image The light brightness value of each pixel; MDN is to achieve the same results in urban and rural areas. Average the brightness at the given value: ; SUML is the sum of the light intensity values of all pixels in a satellite remote sensing image: ; DNQ standardizes the luminous intensity per unit area by using the ratio of total luminance to total area from comprehensive nighttime illumination data of a target city. ; SA represents the sum of the nighttime light pixel areas of the target city.
3. The industrial carbon emission scenario simulation system based on multi-source data fusion according to claim 2, characterized in that: Methods for obtaining the ISTIRPAT carbon emission inversion model by adding a quadratic term of GDP to the existing STIRPAT model and incorporating urban development factors include: The existing STIRPAT model is as follows: ; in, Environmental conditions are driving factors 、 Population driving factors 、 As a driving force for economic development 、 As a driving force for technological innovation, Estimated coefficients of population drivers , Estimated coefficients of economic development drivers , The estimated coefficients of the driving factors of technological innovation. For STIRPAT model parameters, e This is the error term; By adding a quadratic term of GDP to the existing STIRPAT model and incorporating urban development factors into the STIRPAT model, the resulting ISTIRPAT carbon emission inversion model is as follows: ; Among them, CDE represents urban development factors. Tan For provincial industrial carbon emissions, People For population, GDP For regional GDP, SeGDP For the added value of the secondary industry, JianZhu The total output value of the construction industry e denoted as the random error term, d is the population elasticity coefficient, f is the linear elasticity coefficient of GDP, g is the quadratic coefficient of GDP, h is the elasticity coefficient of the added value of the secondary industry, and j is the elasticity coefficient of the total output value of the construction industry.
4. The industrial carbon emission scenario simulation system based on multi-source data fusion according to claim 3, characterized in that: Methods for calculating municipal-level industrial carbon emissions using the ISTIRPAT carbon emission inversion model include: Estimating provincial industrial carbon emissions using the carbon emission factor method: ; The provincial industrial carbon emissions estimated using the carbon emission factor method were then calibrated using a panel regression model. The panel regression model is as follows: ; in, The provincial industrial carbon emissions are estimated using the carbon emission factor method. For the industry China Energy Consumption volume The total number of industries, The total energy consumption of an industry; This is the standard coal conversion factor; The carbon emission factor is 12, where 12 represents carbon. The relative atomic mass is 44, and the relative molecular mass is 44, which is the relative molecular mass of carbon dioxide (CO2). The provincial industrial carbon emissions were estimated using the carbon emission factor method. For satellite observation of CO2 flux data, This is a provincial coefficient. This is the error term; Substituting the provincial industrial carbon emissions calculated using the carbon emission factor method into the ISTIRPAT carbon emission inversion model yields the random error term, population elasticity coefficient, linear elasticity coefficient of GDP, quadratic elasticity coefficient of GDP, elasticity coefficient of secondary industry added value, and elasticity coefficient of total construction output value of the ISTIRPAT carbon emission inversion model. Based on the random error term, population elasticity coefficient, linear elasticity coefficient of GDP, quadratic elasticity coefficient of GDP, elasticity coefficient of secondary industry added value, and elasticity coefficient of total construction output value of the ISTIRPAT carbon emission inversion model, and the urban development factor constructed based on comprehensive nighttime light data of the city, the city-level industrial carbon emissions are calculated by inversion based on the pre-set population, regional GDP, secondary industry added value, and total construction output value.
5. The industrial carbon emission scenario simulation system based on multi-source data fusion according to claim 1, characterized in that: The method for identifying and calculating the density distribution of industrial land using a random forest algorithm based on a pre-defined multi-source fusion dataset includes: The pre-defined multi-source fusion dataset includes satellite remote sensing image data of the target city. All original pixels with manual labels in the satellite remote sensing images are used as training samples to train the random forest algorithm. Each original pixel sample contains the image features of that original pixel and a corresponding manually labeled classification label. During training, based on the Bagging strategy, a single decision tree in the random forest algorithm randomly selects multiple samples with replacement from the training sample set as a training subset. This decision tree then randomly selects multiple image features from all image features in the training subset as image feature subsets. Based on the image feature subsets and their corresponding classification labels, the split point that minimizes Gini impurity is selected to perform a binary partitioning of all samples in the training subset. All samples in the training subset are divided into two child nodes. The training process is repeated for the two child nodes until a pre-defined termination condition is met, resulting in a trained single decision tree. The Gini impurity calculation formula is as follows: ; in, The Gini impurity of the current split node. This refers to the node in the decision tree that is currently to be evaluated or split. For nodes Category tags in The sample proportion, where C is the total number of categories for the classification label; The trained single decision trees are used to identify industrial land in the image to be predicted. The classification results of all trained decision trees are aggregated through majority voting to obtain industrial land identification data, as shown in the following formula: ; in, This represents the final predicted category of the random forest algorithm for the current training sample set. For the total number of decision trees, For the first The predicted results for each tree, For indicator functions, For a maximum value operator, Candidate category labels; Density calculations were performed based on industrial land identification data to obtain industrial land density distribution data: The industrial land identification data contains multiple original pixels, each with multiple corresponding classification labels. The industrial land identification data is reclassified by assigning a value to each original pixel: ; The spatial resolution of the original pixels of the reclassified industrial land identification data is aggregated into pixels with a set spatial resolution to obtain a set pixel. The spatial resolution of the set pixel is divided by the spatial resolution of the original pixels to obtain the total number of original pixels contained in the set pixel. The ratio of the number of original pixels with a value of 1 to the total number of original pixels is used as the industrial land density distribution data in the set pixel.
6. The industrial carbon emission scenario simulation system based on multi-source data fusion according to claim 5, characterized in that: Based on industrial land density distribution data and municipal industrial carbon emissions, a method is used to decompose municipal industrial carbon emissions into spatial grid units through spatial proxy and gridded allocation mechanisms to obtain gridded industrial carbon emission distribution data. The method includes: dividing the target city into regions using set pixels; using the ratio of the industrial land density distribution data of each set pixel to the sum of the industrial land density distribution data of all set pixels as the allocation weight of the corresponding set pixel; and multiplying the municipal industrial carbon emissions of the target city by the allocation weight of the corresponding set pixel to obtain the carbon emissions of the corresponding set pixel.
7. An industrial carbon emission scenario simulation system based on multi-source data fusion according to claim 4 or 6, characterized in that: Based on historical and projected municipal industrial carbon emission data, gridded industrial carbon emission distribution data, and the preset growth rate of carbon emission driving factors, multiple different carbon emission development paths are preset through scenario analysis. Methods for establishing early warning probability models based on multiple different carbon emission development paths include: Under the pre-defined carbon emission development path, each carbon emission driving factor variable annual growth rate It is not fixed, but a variable is set. It follows a normal distribution with a given expected growth rate as the mean and a given variance. ,in, For the first Annual growth rate of each carbon emission driver variable For the first The mean of the annual growth rates of the individual carbon emission driving factors. For the first The standard deviation of the annual growth rate of each carbon emission driver variable. It follows a normal distribution. The representation follows a normal distribution, after... After the new year, the carbon emission driving factors variables are expressed as follows: ; in, For the first One carbon emission driving factor variable, For the first The year's first One carbon emission driving factor variable, For the first Baseline year observations of the carbon emission driving factor variables, For the first The carbon emission driving factor variables are in Annual random growth rate This represents the total brightness of the lights at night. Light intensity per unit area The total output value of the construction industry; The carbon emission driving factor variables are decomposed into a sum of constant and random terms after taking the logarithm: ;remember , ; Substituting the carbon emission driving factor variables into the carbon emission inversion model, we obtain: ; in, For municipal-level industrial carbon emissions, For the natural constant An exponential function with base 0. For the first Regression coefficients of individual carbon emission driving factor variables, For the first The regression coefficient of annual GDP value For the first Annual GDP This is the scaling factor; Carbon emission inversion model by substituting the log-linearized carbon emission driving factor variables: ; in, For the certainty of GDP, The stochastic component of GDP Target probability ; in, For the city-level industrial carbon emissions in the nth year in the future Exceeding the preset carbon emission safety threshold The probability, The preset carbon emission safety threshold, This represents the deterministic portion of all carbon emission driver variables.
8. The industrial carbon emission scenario simulation system based on multi-source data fusion according to claim 7, characterized in that: The method of solving the early warning probability model using the Monte Carlo method to obtain the predicted values of municipal industrial carbon emissions under multiple different carbon emission development paths and the probability that the predicted values of municipal industrial carbon emissions will exceed a preset peak includes: ; in, The total number of simulations for Monte Carlo. In the first In the Monte Carlo simulation, the predicted future carbon emissions are calculated by the early warning probability model. As an indicator function, it indicates that for the th A Monte Carlo simulation is valid if and only if the carbon emission results of that simulation are... Greater than the threshold When the time is right, the value of this function is 1; otherwise, its value is 0, based on the city-level industrial carbon emissions in the nth year. Exceeding the preset carbon emission safety threshold probability Carbon emission control measures will be implemented for target urban areas corresponding to pixels whose carbon emissions exceed the threshold.
9. A method for simulating industrial carbon emission scenarios based on multi-source data fusion according to the system of claim 1, characterized in that, include: Urban development factors are constructed by using different dimensions of comprehensive nighttime light data. A quadratic term of GDP is added to the existing STIRPAT model. The urban development factors are then introduced into the STIRPAT model to obtain the ISTIRPAT carbon emission inversion model. The city-level industrial carbon emissions are calculated using the ISTIRPAT carbon emission inversion model. Based on a pre-set multi-source fusion dataset, the random forest algorithm is used to identify and calculate the density of industrial land to obtain industrial land density distribution data. Based on the industrial land density distribution data and the city-level industrial carbon emissions, the city-level industrial carbon emissions are decomposed into spatial grid units through spatial proxy and gridded allocation mechanisms to obtain gridded industrial carbon emission distribution data. Based on historical and predicted municipal industrial carbon emission data, gridded industrial carbon emission distribution data, and preset growth rates of carbon emission driving factors, multiple different carbon emission development paths are preset using scenario analysis. Based on these multiple carbon emission development paths, an early warning probability model is established. The model is then solved using the Monte Carlo method to obtain the predicted municipal industrial carbon emission values under multiple different carbon emission development paths and the probability that the predicted municipal industrial carbon emission values will exceed the preset peak values.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method of claim 9.