Carbon emission fusion model modeling method and system based on multi-dimensional geographic data

By constructing a carbon emission fusion model based on multi-dimensional geographic data, combining SAR, GWR and random forest models, the shortcomings of the existing models in spatial distribution accuracy and heterogeneity analysis are solved, and high-precision carbon emission prediction and support for low-carbon policy formulation are achieved.

CN119988820APending Publication Date: 2025-05-13GUIZHOU POWER GRID CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510007357.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing carbon emission models have shortcomings in spatial distribution accuracy and heterogeneity analysis, and it is difficult to support the formulation of low-carbon policies at the urban or regional level.

Method used

By combining carbon emission data with multidimensional geographical feature data, a spatial autoregression model (SAR), geo-weighted regression model (GWR) and random forest model are used to construct a refined carbon emission distribution model to conduct a fusion analysis of multidimensional geographical factors and carbon emissions.

Benefits of technology

It significantly improves the accuracy and interpretability of carbon emission forecasts, can more truly reflect the characteristics of carbon emissions in various regions, and provides scientific data support and analysis tools for low-carbon management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988820A_ABST
    Figure CN119988820A_ABST
Patent Text Reader

Abstract

The invention discloses a carbon emission fusion model modeling method and system based on multi-dimensional geographic data, and relates to the technical field of carbon emission analysis and management. Comprising the following steps: collecting multi-dimensional geographic feature data and carbon emission data, and preprocessing the data; constructing a multi-dimensional geographic data carbon emission correlation model; training and calibrating the model, and evaluating the performance of the model; fusion analysis of multi-dimensional geographic factors and carbon emission is carried out, and decision support is provided for regional low-carbon management. According to the method, the carbon emission data and the multi-dimensional geographic feature data are collected and preprocessed, so that standardization and space matching of the data are realized; by constructing a multi-model combination of a spatial autoregression model and a geographically weighted regression model, the spatial dependency relationship between regions and the regional heterogeneity of geographic features can be captured at the same time; and carrying out quantitative analysis on the importance of the multi-dimensional geographic features through a random forest model, and identifying key factors influencing carbon emission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of carbon emission analysis and management, and in particular to a carbon emission fusion modeling method and system based on multi-dimensional geographic data. Background Art

[0004] In terms of geographic data modeling methods, methods such as spatial regression models and geographically weighted regression models (GWR) have been widely used in spatial data analysis. However, since spatial regression models have certain advantages in capturing spatial correlations between regions, while geographically weighted regression models are more prominent in revealing regional heterogeneity, it is necessary to use these two types of models in combination to improve the adaptability of carbon emission models to spatial dependence and regional differences. In addition, the relative importance of geographical characteristics to carbon emissions has not been fully analyzed in existing models, resulting in insufficient targeting of low-carbon policies. In view of the above status quo, how to construct a refined carbon emission distribution model based on multidimensional geographic data and then optimize low-carbon policies has become an urgent problem to be solved. Summary of the invention

[0005] In view of the problems existing in the prior art, the present invention is proposed.

[0006] Therefore, the problem to be solved by the present invention is how to establish a refined carbon emission distribution model by combining carbon emission data with multi-dimensional geographic feature data (such as population density, land use, transportation network, etc.), aiming to support the formulation of low-carbon policies at the city or regional level and provide scientific data support and analysis tools for low-carbon management.

[0007] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0008] In a first aspect, an embodiment of the present invention provides a carbon emission fusion modeling method based on multidimensional geographic data, which includes collecting multidimensional geographic feature data and carbon emission data, and preprocessing the data;

[0009] Construct a carbon emission correlation model based on multi-dimensional geographic data;

[0010] Train and calibrate the model and evaluate model performance;

[0011] Conduct an integrated analysis of multi-dimensional geographical factors and carbon emissions to provide decision-making support for regional low-carbon management.

[0012] As a preferred solution of the carbon emission fusion modeling method based on multidimensional geographic data of the present invention, the steps of collecting multidimensional geographic feature data and carbon emission data and preprocessing the data include:

[0013] Carbon emission data collection: Collect time series data on carbon emissions in the region, including emission amounts, emission sources and spatial distribution information;

[0014] Multi-dimensional geographic feature data collection: including geographic distribution data of population density, land use types including the proportion and distribution data of industrial areas, commercial areas, residential areas, and green areas, and transportation network data including road distribution, major traffic flows, and commuting patterns;

[0015] Data preprocessing: Use normalization method to standardize multidimensional data;

[0016] Spatial data matching: Align carbon emission data and geographic data at the same spatial resolution using a geographic coordinate system.

[0017] As a preferred solution of the carbon emission fusion modeling method based on multidimensional geographic data of the present invention, wherein: constructing a multidimensional geographic data carbon emission association model includes:

[0018] The spatial autoregression model (SAR) is used to construct the spatial dependence model between adjacent regions. The formula is:

[0019]

[0020] Among them, E(x,y) represents the carbon emissions of region (x,y), ρ is the spatial autoregression coefficient, and w i is the weight of the neighboring region, X(x,y) is the geographic feature matrix of region (x,y), β is the regression coefficient, and ∈ is the error term;

[0021] The geographically weighted regression model (GWR) was used to construct the regional heterogeneity model, and the formula is:

[0022] E(x,y)=α(x,y)+β(x,y)P(x,y)+γ(x,y)L(x,y)+δ(x,y)T(x,y)+∈

[0023] Among them, P(x,y) is the population density, L(x,y) is the land use type, T(x,y) is the transportation network information, α(x,y), β(x,y), γ(x,y), δ(x,y) are the regression coefficients that vary with spatial position, and ∈ is the error term.

[0024] As a preferred solution of the carbon emission fusion modeling method based on multi-dimensional geographic data of the present invention, the steps of training and calibrating the model include:

[0025] Dataset division: Divide the training set and test set by region to ensure that the model has generalization ability in the regional dimension;

[0026] SAR model training: set the spatial weight matrix and calculate the autoregressive coefficient and regression coefficient by the least squares method;

[0027] GWR model training: assign weights to each region based on geographical location and obtain location-dependent regression coefficients through optimization algorithms;

[0028] Random forest model training: Evaluating the importance coefficient of each geographic feature.

[0029] As a preferred solution of the carbon emission fusion modeling method based on multidimensional geographic data of the present invention, the step of evaluating the model performance includes:

[0030] The importance of geographical features is evaluated using the random forest model, and the calculation formula is:

[0031]

[0032] in, Indicates the removal of feature X i The mean square error increment after , N is the number of decision trees;

[0033] The mean square error (MSE) is calculated to evaluate the deviation between the model prediction value and the true value. The MSE formula is:

[0034]

[0035] Among them, E i is the true value, is the predicted value, n is the number of samples;

[0036] The K-fold cross validation method was used to measure the performance of the model.

[0037] As a preferred solution of the carbon emission fusion modeling method based on multi-dimensional geographic data of the present invention, the step of providing decision support for regional low-carbon management includes:

[0038] Carbon emission spatial distribution simulation: Generate spatial distribution map of carbon emissions through SAR and GWR models to identify high emission areas;

[0039] Multidimensional geographic feature importance analysis: Based on the importance coefficient output by the random forest model, the contribution of each geographic feature to carbon emissions is analyzed;

[0040] Provide low-carbon policy recommendations: Develop differentiated low-carbon policies based on the analysis results, such as taking vehicle flow control measures in densely populated areas or increasing green space in industrial areas.

[0041] As a preferred solution of the carbon emission fusion modeling method based on multi-dimensional geographic data of the present invention, it also includes a model optimization step:

[0042] Optimize the regression parameters of SAR and GWR models based on the gradient descent method;

[0043] Genetic algorithms are used to optimize the tree structure configuration of the random forest model to improve the accuracy of identifying the importance of geographical features;

[0044] The particle swarm algorithm is used to accelerate the training of the model on large-scale data sets to ensure the training efficiency and computational stability of the model.

[0045] In a second aspect, an embodiment of the present invention provides a carbon emission fusion modeling system based on multi-dimensional geographic data, which includes a data processing module, a model building module, a training and evaluation module, and an analysis and decision-making module;

[0046] The data processing module is used to collect carbon emission data and multi-dimensional geographic feature data, and perform pre-processing and spatial data matching;

[0047] The model building module is used to establish spatial autoregression models, geographically weighted regression models and random forest models to achieve spatial distribution prediction of carbon emissions;

[0048] The training and evaluation module is used to train, calibrate and evaluate the performance of the model;

[0049] The analysis and decision-making module is used to simulate the spatial distribution of carbon emissions, analyze the importance of multi-dimensional geographical features, and provide regional low-carbon policy recommendations.

[0050] In a third aspect, an embodiment of the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program instructions are executed by the processor, the steps of the carbon emission fusion modeling method based on multidimensional geographic data as in the first aspect of the present invention are implemented.

[0051] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program instructions are executed by a processor, the steps of the carbon emission fusion modeling method based on multidimensional geographic data as in the first aspect of the present invention are implemented.

[0052] The beneficial effects of the present invention are as follows: the carbon emission fusion model modeling method based on multidimensional geographic data provided by the present invention realizes data standardization and spatial matching by collecting carbon emission data and multidimensional geographic feature data and performing preprocessing, ensures the accuracy and consistency of model input data, provides a reliable data basis for subsequent analysis, and improves the refinement of carbon emission analysis; by constructing a multi-model combination of a spatial autoregression model (SAR) and a geographically weighted regression model (GWR), it can simultaneously capture the spatial dependence between regions and the regional heterogeneity of geographical features, significantly improve the accuracy of carbon emission prediction, and enable the model to more truly reflect the characteristics of carbon emissions in each region; through the random forest model, the importance of multidimensional geographical features is quantitatively analyzed, and the key factors affecting carbon emissions are identified, which provides a scientific basis for formulating differentiated low-carbon policies and improves the pertinence and effectiveness of low-carbon management decisions; by establishing a carbon emission spatial distribution model and conducting a multidimensional geographic feature importance analysis, refined management of carbon emissions is achieved, high-emission areas can be identified and targeted low-carbon policy recommendations can be put forward, the effectiveness of regional low-carbon management is improved, and it is helpful to achieve carbon emission reduction goals. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0054] Figure 1 A flowchart of the modeling method for the carbon emission fusion model based on multi-dimensional geographic data;

[0055] Figure 2 The collection and preprocessing flow chart of the carbon emission fusion modeling method based on multi-dimensional geographic data;

[0056] Figure 3 A flowchart for constructing a carbon emission correlation model based on a carbon emission fusion model modeling method based on multi-dimensional geographic data;

[0057] Figure 4 A schematic diagram of the importance analysis of geographical factors using the random forest model of the carbon emission fusion model based on multidimensional geographic data;

[0058] Figure 5 This is a flow chart of carbon emission spatial distribution simulation and importance analysis based on the carbon emission fusion model modeling method of multidimensional geographic data. DETAILED DESCRIPTION

[0059] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0060] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0061] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0062] Example 1

[0063] Reference Figure 1 to Figure 5 , which is the first embodiment of the present invention, and provides a carbon emission fusion modeling method based on multi-dimensional geographic data, including:

[0064] S1, collect multi-dimensional geographic feature data and carbon emission data, and pre-process the data;

[0065] In the embodiment of the present application, the step of collecting multi-dimensional geographic feature data and carbon emission data and pre-processing the data includes:

[0066] Carbon emission data collection: Collect time series data on carbon emissions in the region, including emission amounts, emission sources and spatial distribution information;

[0067] Multi-dimensional geographic feature data collection: including geographic distribution data of population density, land use types including the proportion and distribution data of industrial areas, commercial areas, residential areas, and green areas, and transportation network data including road distribution, major traffic flows, and commuting patterns;

[0068] Data preprocessing: Use normalization method to standardize multidimensional data;

[0069] Spatial data matching: Align carbon emission data and geographic data at the same spatial resolution using a geographic coordinate system.

[0070] It should be noted that carbon emission data collection: In the present invention, carbon emission data is one of the basic data. The collection content includes carbon emissions in each region, sources of carbon emissions (such as industrial emissions, transportation emissions, etc.), time series changes in carbon emissions and spatial distribution. Carbon emission data can be obtained through monitoring sites, remote sensing images and carbon emission statistics from local governments, and integrated according to regional divisions;

[0071] Population density data collection: In the present invention, population density data is an important component of multidimensional geographic feature data. The data source of population density can be statistical data released by urban planning departments, satellite remote sensing images, and urban big data systems. The collection content includes the geographical distribution and temporal changes of the population in the region, especially in high-density urban areas, which require intensive sampling;

[0072] Land use type data collection: Land use data is an important parameter for building a carbon emission model, which includes the functional division of each area, such as industrial area, commercial area, residential area, green space, transportation land, etc. Land use type data can be collected through remote sensing data and GIS system, and aligned with the administrative area division layer for subsequent data matching;

[0073] Traffic network data collection: Traffic network data plays an important role in the analysis of urban carbon emissions. Traffic network data includes regional road distribution, main traffic flow, commuting mode, traffic density, etc. The data source can be road network data from the transportation planning department, traffic flow detection equipment or satellite traffic images.

[0074] It should also be noted that data standardization: In the embodiment of the present invention, the collected multidimensional data is first standardized or normalized to ensure that the dimensions of different feature data are consistent. For example, carbon emission data is expressed as emissions per square kilometer, population density data is expressed as the number of people per square kilometer, and road traffic volume is expressed as the number of vehicles per unit time. After standardization, the data has the same dimension, which is convenient for input into the model for analysis;

[0075] Spatial data matching: The spatial data matching process spatially aligns carbon emission data and geographic feature data according to geographic coordinates. The specific steps include using the GIS system to map carbon emission data and geographic feature data to the same geographic coordinate system to ensure the consistency of the data in the spatial dimension. For example, each carbon emission data point is matched with the corresponding land use type, population density, transportation network and other information to form a geographic feature matrix for each region, laying the foundation for subsequent model construction;

[0076] After data collection and preprocessing steps, all data were integrated into a unified regional feature matrix with the same spatial resolution to ensure the coordination of carbon emission data and geographic feature data.

[0077] S2, constructing a carbon emission correlation model with multi-dimensional geographic data;

[0078] It should be noted that after the data preprocessing is completed, the carbon emission correlation model is constructed. The construction of the correlation model aims to link the carbon emission data with the multi-dimensional geographic feature data so as to analyze the impact of different geographic features on carbon emissions through the model.

[0079] Furthermore, the present invention uses a spatial autoregression model (SAR) to capture the spatial dependence characteristics of carbon emission data. The SAR model reveals the interdependence of carbon emissions between different regions by defining the weight relationship between neighboring regions. The model expression is as follows:

[0080]

[0081] Among them, E(x,y) represents the carbon emissions of region (x,y), ρ is the spatial autoregression coefficient, and w i is the weight of the neighboring region, X(x,y) is the geographic feature matrix of region (x,y), β is the regression coefficient, and ∈ is the error term;

[0082] It should be noted that the modeling steps of the SAR model are:

[0083] 1. Determine the spatial weight matrix: Define the spatial weight matrix between regions through the proximity relationship, for example, define the weights between regions through the k-neighbor method. Regions that are close are given higher weights, and regions that are far away are given lower weights.

[0084] 2. Model parameter estimation: The spatial autoregression coefficients and regression coefficients are estimated using the least squares method to minimize the model residual sum of squares (RSS).

[0085] 3. Model calibration and optimization: Iteratively calibrate the SAR model to ensure that the model’s performance in terms of the spatial dependence of carbon emissions is consistent with the actual regional distribution characteristics.

[0086] It should also be noted that the present invention further introduces a geographically weighted regression model (GWR) to model the regional heterogeneity of geographical factors on carbon emissions. The GWR model allows different regression coefficients at different spatial locations, so it can more accurately simulate the spatial variability of geographical factors. The expression of the GWR model is:

[0087] E(x,y)=α(x,y)+β(x,y)P(x,y)+γ(x,y)L(x,y)+δ(x,y)T(x,y)+∈

[0088] Among them, P(x,y) is the population density, L(x,y) is the land use type, T(x,y) is the transportation network information, α(x,y), β(x,y), γ(x,y), δ(x,y) are the regression coefficients that vary with spatial position, and ∈ is the error term.

[0089] The modeling steps of the GWR model are:

[0090] 1. Location weight assignment: Assign location weights to each region so that the GWR model can capture the spatial differences in regional characteristics. Usually, a Gaussian weight function is used to assign location weights, with larger weights given to regions that are close and smaller weights given to regions that are far away.

[0091] 2. Regression coefficient estimation: For each region, local regression is used to estimate its regression coefficients α(x,y), β(x,y), γ(x,y), and δ(x,y) to capture the differences in the impact of different geographical characteristics on carbon emissions.

[0092] 3. Model fitting and optimization: Optimize the GWR model to ensure that the model's performance in capturing regional heterogeneity is consistent with the actual situation. The accuracy of the model can be improved by increasing the number of regional sampling points.

[0093] S3, train and calibrate the model and evaluate the model performance;

[0094] In the embodiment of the present application, the steps of training and calibrating the model include:

[0095] Dataset division: Divide the training set and test set by region to ensure that the model has generalization ability in the regional dimension;

[0096] SAR model training: set the spatial weight matrix and calculate the autoregressive coefficient and regression coefficient by the least squares method;

[0097] GWR model training: assign weights to each region based on geographical location and obtain location-dependent regression coefficients through optimization algorithms;

[0098] Random forest model training: Evaluating the importance coefficient of each geographic feature.

[0099] It should be noted that in order to evaluate the relative importance of different geographical features to carbon emissions, the present invention introduces a random forest model. The random forest model is a method based on decision tree integration, which quantifies the contribution of each geographical factor to carbon emissions by training multiple decision trees and calculating the importance coefficient of each feature. The formula of the model is:

[0100]

[0101] in, Indicates the removal of feature Xi The mean square error increment after , N is the number of decision trees;

[0102] Modeling steps of random forest model:

[0103] 1. Feature selection: Select the main variables in carbon emission data and geographic feature data as the input feature set of the model, including population density, land use type, transportation network, etc.

[0104] 2. Training model: Use the random forest algorithm to train multiple decision trees and record the importance score of each feature for each tree during the training process.

[0105] 3. Calculate feature importance: By comparing the contribution of each feature to the model prediction, quantify the relative importance of different geographical features to carbon emissions, and provide data support for subsequent carbon emission management.

[0106] After completing the construction of the carbon emission correlation model, the embodiment of the present invention optimizes the model parameters through training and calibration steps, so that it can more accurately predict the spatial distribution characteristics of carbon emissions and provide scientific support for the contribution analysis of geographical features.

[0107] It should also be noted that this embodiment first divides the processed data to ensure the generalization ability of the model in the regional dimension and maintain data consistency:

[0108] 1. Dataset division: The integrated multidimensional geographic data and carbon emission data are divided into training sets and test sets according to regions. The training set is used for model parameter fitting and optimization, while the test set is used to evaluate the prediction effect of the model to verify the generalization ability of the model. When dividing the data set, ensure that each data set covers different regions to ensure representativeness.

[0109] 2. Standardization: Standardize all input features (such as population density, land use type, transportation network, etc.) and unify the dimensions to avoid the influence of feature data magnitude differences on the model training process. Standardization usually uses zero mean normalization or interval scaling to unify the data to the same scale, so that the model can maintain consistent sensitivity on different features.

[0110] The spatial autoregressive model SAR expresses the interaction between regions through the spatial weight matrix. Therefore, during the model training process, it is necessary to set the spatial weight matrix and optimize the autoregressive coefficient. The training process of the SAR model is as follows:

[0111] 1. Spatial weight matrix setting: According to the geographical proximity of the regions, the spatial weight matrix of the SAR model is set. The weights of neighboring regions are usually defined using the k-nearest neighbor method or the distance decay function. Neighboring regions are given higher weights, while distant regions are given lower or zero weights to capture the spatial dependencies between regions.

[0112] 2. Calculation of autoregressive coefficients: After setting the spatial weight matrix, the least squares method is used to calculate the autoregressive coefficient ρ and regression coefficient β of the model to minimize the fitting error of the model. During the training process, the objective function is to minimize the residual sum of squares RSS, and the formula is as follows:

[0113]

[0114] Among them, E(x i ,y i ) represents the area (x i ,y i ) is the carbon emission, ρ is the autoregressive coefficient, w j is the weight of the neighboring area, X(x i ,y i ) is the regional feature matrix, and β is the regression coefficient.

[0115] 3. Model calibration and optimization: The autoregressive coefficient and regression coefficient of the model are adjusted through multiple iterations to ensure that the SAR model can accurately reflect the spatial dependence of carbon emissions among regions. At the same time, the model parameters can be further optimized through cross-validation methods to ensure the adaptability of the model under different regional conditions.

[0116] It should be noted that the Geographically Weighted Regression (GWR) model allows the regression coefficients in a region to vary with geographical location, which is suitable for capturing the spatial heterogeneity of geographical factors. The training steps of the GWR model are as follows:

[0117] 1. Location weight setting: Set location weights for each region based on geographical distance so that the model can reflect the different impacts of different geographical features on carbon emissions in the region. The weight calculation usually uses the Gaussian weight function, with closer regions being given higher weights and farther regions being given lower weights. The formula for the Gaussian weight function is:

[0118]

[0119] Among them, W ij is the weight between region i and region j, d ij is the geographical distance between the two regions, and b is the bandwidth parameter.

[0120] 2. Regression coefficient estimation: The regression coefficient of the GWR model is optimized according to location dependence. The regression coefficient in the region is fitted using the least squares method or the weighted least squares method (WLS) to optimize the prediction effect of the model at each geographical location. The regression equation of the model is:

[0121] E(x,y)=α(x,y)+β(x,y)P(x,y)+γ(x,y)L(x,y)+δ(x,y)T(x,y)+∈

[0122] Among them, α(x,y), β(x,y), γ(x,y), and δ(x,y) are regression coefficients that vary with location and represent the spatial influence of different geographical features.

[0123] 3. Model calibration: The GWR model is calibrated multiple times to ensure that the model can reflect the spatial heterogeneity of carbon emissions. The bandwidth parameters are adjusted to optimize the regression coefficients of each region to enhance the model's adaptability to the spatial heterogeneity of geographical factors.

[0124] It should be noted that the random forest model integrates multiple decision trees to evaluate the relative importance of different geographical features to carbon emissions, thereby providing a reference for carbon emission policies. The training steps are as follows:

[0125] 1. Multiple decision tree ensemble training: Use training data to build multiple decision trees. Each tree uses the bootstrap method to randomly select samples and perform feature selection on each sample. By integrating multiple decision trees, the robustness of the model is improved.

[0126] 2. Importance coefficient calculation: By removing each geographic feature, calculate the change in mean square error (MSE). The formula for the feature importance coefficient is:

[0127]

[0128] in, Indicates the removal of feature X i The mean square error increment after the calculation, N is the number of decision trees, V i is the importance coefficient of feature i;

[0129] 3. Feature screening and importance ranking: Calculate the importance coefficient of each feature and sort them to screen out the geographical features that have the greatest impact on carbon emissions, providing a scientific basis for subsequent carbon emission management and policy formulation.

[0130] S4, conduct integrated analysis of multi-dimensional geographical factors and carbon emissions to provide decision support for regional low-carbon management;

[0131] In order to ensure the reliability and generalization ability of the model, the present invention adopts cross-validation method and mean square error (MSE) calculation to verify the model and perform error analysis.

[0132] The present invention uses K-Fold Cross Validation to verify the performance of the model. The specific method is:

[0133] 1. Dataset division: Divide the dataset into K subsets. In each training process, select K-1 subsets as training sets and the remaining 1 subset as validation set. Repeat K times to ensure that each subset is used as a validation set once.

[0134] 2. Calculate the average error: Get an error value for each validation, and finally take the average of all validation errors as the overall error of the model to measure the stability of the model on different data sets.

[0135] The present invention uses mean square error (MSE) to evaluate the deviation between the model prediction value and the actual carbon emission value. The MSE formula is:

[0136]

[0137] Among them, E i is the true value, is the predicted value, n is the number of samples, and the smaller the MSE, the higher the accuracy of the model prediction.

[0138] It should be noted that after completing the model training and calibration, the present invention further explores the correlation characteristics between carbon emissions and multidimensional geographical factors through a fusion analysis step and proposes an optimization strategy.

[0139] Based on the SAR and GWR models, the carbon emission distribution characteristics of different regions are simulated to provide a geographical basis for the formulation of low-carbon policies. The steps of spatial distribution simulation are as follows:

[0140] Spatial dependency simulation of SAR model: The SAR model captures the spatial dependency of neighboring regions through a spatial weight matrix to reveal the carbon emission characteristics of different geographical locations.

[0141] Regional heterogeneity simulation of GWR model: The regional heterogeneity of carbon emissions is captured by the GWR model, so that the model can generate a spatial distribution map of carbon emissions, identify high-emission areas, and provide a reference for regional differentiated management.

[0142] It should also be noted that the contribution of each geographical feature to carbon emissions is analyzed based on the importance coefficients generated by the random forest model, and targeted policy recommendations are put forward. The specific steps include:

[0143] 1. Importance coefficient analysis: The importance coefficient of each geographical feature (such as population density, land use type, and transportation network) is obtained through the random forest model to identify the factors that have the greatest impact on carbon emissions.

[0144] 2. Policy recommendation formulation: Propose differentiated low-carbon policy recommendations based on the analysis results, such as controlling vehicle flow in high-density population areas, increasing green space in industrial areas, and optimizing regional low-carbon management.

[0145] 3. Identification of high-emission areas: Through spatial distribution simulation, areas with high carbon emissions are located, and carbon reduction measures are deployed in these areas first to ensure the accuracy and effectiveness of low-carbon policies.

[0146] Furthermore, this embodiment also provides a carbon emission fusion modeling system based on multi-dimensional geographic data, including a data processing module, a model building module, a training and evaluation module, and an analysis and decision-making module;

[0147] The data processing module is used to collect carbon emission data and multi-dimensional geographic feature data, and perform pre-processing and spatial data matching;

[0148] The model building module is used to establish a spatial autoregression model, a geographically weighted regression model and a random forest model to achieve spatial distribution prediction of carbon emissions;

[0149] The training and evaluation module is used to train, calibrate and evaluate the performance of the model;

[0150] The analysis and decision-making module is used to simulate the spatial distribution of carbon emissions, analyze the importance of multi-dimensional geographical features, and provide regional low-carbon policy recommendations.

[0151] This embodiment also provides a computer device, which is suitable for the case of a carbon emission fusion model modeling method based on multidimensional geographic data, including a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions to implement the carbon emission fusion model modeling method based on multidimensional geographic data as proposed in the above embodiment.

[0152] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0153] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the method for modeling a carbon emission fusion model based on multi-dimensional geographic data as proposed in the above embodiment is implemented.

[0154] In summary, the present invention provides a modeling method for a carbon emission fusion model based on multidimensional geographic data. By combining carbon emission data with multidimensional geographic feature data such as population density, land use type, and transportation network, the spatial distribution simulation of carbon emissions and the importance analysis of geographic features are realized using the spatial autoregression model (SAR), geographically weighted regression model (GWR), and random forest model. This method can refine the spatial distribution characteristics of carbon emissions, effectively reveal the impact of different geographical factors on carbon emissions, and provide data support for regional differentiated low-carbon management.

[0155] This paper solves the shortcomings of existing carbon emission models in spatial distribution accuracy and heterogeneity analysis through multi-model combination analysis, significantly improves the accuracy and interpretability of carbon emission prediction, and has the following advantages and application prospects:

[0156] 1. Improve model accuracy: By capturing regional spatial dependence and heterogeneity characteristics through SAR and GWR models, the carbon emission model can more realistically reflect the spatial impact of geographical characteristics and improve model accuracy and applicability.

[0157] 2. Support refined policy making: This invention uses the random forest model to evaluate the contribution of various geographical factors to carbon emissions, which can provide a scientific basis for low-carbon policy making, support precise intervention and management in high-emission areas, and help regions achieve low-carbon goals.

[0158] 3. Adapt to multiple scenarios: The present invention is applicable to various scenarios such as cities, industrial areas, and traffic-intensive areas. It can conduct refined analysis and management of carbon emissions in different regions and has broad application potential.

[0159] Example 2

[0160] This is the second embodiment of the present invention. This embodiment provides a carbon emission fusion modeling method based on multi-dimensional geographic data. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0161] This embodiment selects 11 prefecture-level cities in a province as research objects, and collects carbon emission data and multi-dimensional geographic feature data from 2019 to 2023. First, the carbon emission data of each prefecture-level city is collected through the monitoring network of the Provincial Environmental Protection Bureau, including emission source data in the fields of industry, transportation, and construction. At the same time, population density distribution data is obtained from the Statistics Bureau and the Urban Planning Department, land use type data is obtained using remote sensing images and GIS databases, and transportation network data is obtained from the transportation management department. All data are spatially matched and aligned using a 1km×1km grid.

[0162] The collected data were standardized, and the min-max normalization method was used to unify the data of each dimension to the interval [0,1]. Then the SAR and GWR models were constructed. The SAR model used the inverse distance weighted method to construct the spatial weight matrix, and the GWR model used the Gaussian kernel function to calculate the geographic weight. In the model training stage, the data was divided into training set and test set in a ratio of 7:3. The data from 2019 to 2022 were used to train the model, and the data from 2023 were used for testing and verification.

[0163] The random forest model sets the number of decision trees to 100, and uses the grid search method to optimize the model parameters. When integrating the three models, different weights are assigned to each model based on its performance on the validation set. The model validation uses 5-fold cross validation, and the model performance is evaluated by calculating the MSE. Finally, the trained model is applied to the 2023 data forecast and compared with the traditional single model method.

[0164]

[0165] By comparing the experimental data, it can be clearly seen that the method of the present invention has significant advantages over the traditional single model method. In terms of prediction error, the average prediction error of the traditional method is 8.46%, while the average prediction error of the multi-model fusion method of the present invention is only 3.29%, an increase of about 61.1%. Among them, City I achieved the best prediction effect, with a prediction error of only 2.89%, thanks to the spatial autoregressive model of the present invention effectively capturing the mutual influence between regions.

[0166] In terms of computational efficiency, thanks to the particle swarm optimization strategy adopted by the present invention, the average computational time is controlled between 11 and 13 minutes, meeting the actual application requirements. In terms of spatial accuracy, the present invention achieves fine-grained predictions of 1.0-1.2 kilometers. This high-precision spatial analysis capability provides strong support for the formulation of accurate regional low-carbon policies.

[0167] It is particularly noteworthy that the present invention has outstanding performance in terms of accuracy of policy recommendations, reaching an average of 92.94%. This shows that the multi-dimensional geographic feature importance analysis of the present invention can effectively identify key factors affecting carbon emissions. For example, in City B, which has the highest population density, the present invention accurately identified traffic congestion as the main contributing factor to carbon emissions, and the traffic flow control policy proposed based on this achieved significant results, with a policy recommendation accuracy of 94.1%.

[0168] The feature importance analysis of the random forest model found that there are obvious differences in the main factors affecting carbon emissions in different cities. Cities with a high proportion of industrial areas (such as City G) are mainly affected by industrial layout, while cities with dense commercial areas (such as City I) are mainly affected by population density and transportation network. This differentiated analysis result provides a scientific basis for local governments to formulate targeted low-carbon policies. In addition, the GWR model of the present invention successfully captures the impact of seasonal changes on carbon emissions, which is often overlooked in traditional methods.

[0169] In summary, the present invention demonstrates obvious advantages in prediction accuracy, spatial precision, and policy guidance through multi-model fusion and multi-dimensional geographic data analysis, providing strong technical support for the refined management of regional carbon emissions.

[0170] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A carbon emission fusion modeling method based on multidimensional geographic data, characterized by: include, Collect multi-dimensional geographic feature data and carbon emission data, and pre-process the data; Construct a carbon emission correlation model based on multi-dimensional geographic data; Train and calibrate the model and evaluate model performance; Conduct an integrated analysis of multi-dimensional geographical factors and carbon emissions to provide decision-making support for regional low-carbon management.

2. The carbon emission fusion modeling method based on multidimensional geographic data according to claim 1 is characterized in that: The step of collecting multi-dimensional geographic feature data and carbon emission data and pre-processing the data includes: Carbon emission data collection: Collect time series data on carbon emissions in the region, including emission amounts, emission sources and spatial distribution information; Multi-dimensional geographic feature data collection: including geographic distribution data of population density, land use types including the proportion and distribution data of industrial areas, commercial areas, residential areas, and green areas, and transportation network data including road distribution, major traffic flows, and commuting patterns; Data preprocessing: Normalization method is used to standardize multidimensional data; Spatial data matching: Align carbon emission data and geographic data at the same spatial resolution using a geographic coordinate system.

3. The carbon emission fusion modeling method based on multidimensional geographic data according to claim 2 is characterized in that: The construction of a multidimensional geographic data carbon emission association model includes: The spatial autoregression model (SAR) is used to construct the spatial dependence model between adjacent regions. The formula is: Among them, E(x,y) represents the carbon emissions of region (x,y), ρ is the spatial autoregression coefficient, and w i is the weight of the neighboring region, X(x,y) is the geographic feature matrix of region (x,y), β is the regression coefficient, and ∈ is the error term; The geographically weighted regression model (GWR) was used to construct the regional heterogeneity model, and the formula is: E(x,y)=α(x,y)+β(x,y)P(x,y)+γ(x,y)L(x,y)+δ(x,y)T(x,y)+∈ Among them, P(x,y) is the population density, L(x,y) is the land use type, T(x,y) is the transportation network information, α(x,y), β(x,y), γ(x,y), δ(x,y) are the regression coefficients that vary with spatial position, and ∈ is the error term.

4. The carbon emission fusion modeling method based on multidimensional geographic data according to claim 3 is characterized in that: The steps of training and calibrating the model include: Dataset division: Divide the training set and test set by region to ensure that the model has generalization ability in the regional dimension; SAR model training: set the spatial weight matrix and calculate the autoregressive coefficient and regression coefficient by the least squares method; GWR model training: assign weights to each region based on geographical location and obtain location-dependent regression coefficients through optimization algorithms; Random forest model training: Evaluating the importance coefficient of each geographic feature.

5. The carbon emission fusion modeling method based on multidimensional geographic data according to claim 4 is characterized in that: The steps of evaluating model performance include: The importance of geographical features is evaluated using the random forest model, and the calculation formula is: in, Indicates the removal of feature X i The mean square error increment after , N is the number of decision trees; The mean square error (MSE) is calculated to evaluate the deviation between the model prediction value and the true value. The MSE formula is: Among them, E i is the true value, is the predicted value, n is the number of samples; The K-fold cross validation method was used to measure the performance of the model.

6. The carbon emission fusion modeling method based on multidimensional geographic data according to claim 5, characterized in that: The steps of providing decision support for regional low-carbon management include: Carbon emission spatial distribution simulation: Generate spatial distribution map of carbon emissions through SAR and GWR models to identify high emission areas; Multidimensional geographic feature importance analysis: Based on the importance coefficient output by the random forest model, the contribution of each geographic feature to carbon emissions is analyzed; Provide low-carbon policy recommendations: Develop differentiated low-carbon policies based on the analysis results, such as taking vehicle flow control measures in densely populated areas or increasing green space in industrial areas.

7. The carbon emission fusion modeling method based on multidimensional geographic data according to claim 6 is characterized in that: It also includes a model optimization step: Optimize the regression parameters of SAR and GWR models based on the gradient descent method; Genetic algorithms are used to optimize the tree structure configuration of the random forest model to improve the accuracy of identifying the importance of geographical features; The particle swarm algorithm is used to accelerate the training of the model on large-scale data sets to ensure the training efficiency and computational stability of the model.

8. A carbon emission fusion modeling system based on multidimensional geographic data, based on the carbon emission fusion modeling method based on multidimensional geographic data according to any one of claims 1 to 7, characterized in that: It also includes a data processing module, a model building module, a training evaluation module, and an analysis and decision-making module; The data processing module is used to collect carbon emission data and multi-dimensional geographic feature data, and perform pre-processing and spatial data matching; The model building module is used to establish a spatial autoregression model, a geographically weighted regression model and a random forest model to achieve spatial distribution prediction of carbon emissions; The training and evaluation module is used to train, calibrate and evaluate the performance of the model; The analysis and decision-making module is used to simulate the spatial distribution of carbon emissions, analyze the importance of multi-dimensional geographical features, and provide regional low-carbon policy recommendations.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the carbon emission fusion modeling method based on multidimensional geographic data according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the carbon emission fusion modeling method based on multidimensional geographic data according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Enterprise nitrogen and phosphorus pollution discharge flux calculation method based on multi-source heterogeneous data fusion

    CN120373675A

  • Carbon emission tracking and source analysis method and device combined with geographic information system

    CN120387596A

  • Sensory evaluation method for taste and texture of fresh chili

    CN120892857A