A method for assessing carbon neutrality in eco-parks
By combining time series analysis, geographic information system, and multiple regression analysis with a random forest model, the problems of data quality and spatiotemporal distribution in the carbon asset assessment of eco-parks were solved, enabling dynamic assessment and prediction of carbon assets and improving the carbon neutrality level.
Patent Information
- Application Number
- CN202510371130.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-03-27
AI Technical Summary
In the assessment of carbon assets in eco-parks, there are many types of carbon assets, uneven spatial and temporal distribution, and inconsistent data quality. Existing technologies are insufficient to reasonably quantify and weigh various types of carbon assets, and data updates are not timely, resulting in a lack of real-time and traceability in the assessment results.
By employing time series analysis, geographic information system technology, multiple regression analysis, and random forest models, combined with basic carbon asset data, spatial distribution models of carbon storage and carbon flux are constructed, data cleaning and prediction are performed, and dynamic assessment results are generated.
It enables a comprehensive assessment and prediction of carbon assets in the eco-park, improves the level of carbon neutrality, and provides a scientific basis for carbon asset management and decision-making.
Smart Images

Figure CN120258612B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information technology, and in particular to a method for assessing carbon neutrality in eco-parks. Background Technology
[0002] Constructing a carbon asset assessment model for eco-parks presents numerous technical challenges. First, the carbon assets within these parks are diverse, encompassing natural ecosystems such as forests, grasslands, and wetlands, as well as human activities like energy consumption and industrial production. The carbon sequestration capacity and emission intensity vary significantly among different types of carbon assets. A reasonable quantification and balance of these various carbon assets within the model requires a comprehensive consideration of their dynamic characteristics and interaction mechanisms. Second, the spatiotemporal distribution of carbon assets is uneven; the patterns of carbon storage and flux changes differ across regions and periods. The model needs to characterize this spatiotemporal heterogeneity and reveal its underlying driving mechanisms. Third, changes in carbon assets are influenced by both natural and human factors. The model needs to identify and differentiate the direction and magnitude of the effects of different factors and quantify their impact on carbon assets. Finally, the quality of carbon asset monitoring data is inconsistent, with issues such as inconsistent time frequencies, missing data, and outliers. The model needs to be able to filter and repair low-quality data and appropriately fill data gaps to ensure the accuracy and continuity of the model's input data. At the same time, the model also needs to address issues such as untimely data updates and poor connectivity between different data sources to ensure the real-time nature and traceability of the evaluation results. Summary of the Invention
[0003] This invention provides a method for assessing carbon neutrality in eco-parks, mainly comprising:
[0004] Acquire basic data on various carbon assets within the eco-park, including the carbon sink capacity of the natural ecosystem and the carbon emission intensity of human activities;
[0005] Based on carbon asset data, time series analysis is used to model the changing trends of carbon sink capacity and carbon emission intensity, and obtain a trend model.
[0006] Using geographic information system technology and combining trend models, the spatial distribution of carbon storage and carbon flux is modeled to obtain a spatial distribution model;
[0007] Key variables were extracted from the spatial distribution model, and multiple regression analysis was used to quantify the effects of different factors on carbon assets.
[0008] Acquire carbon asset monitoring data, assess data quality, and correct outliers using data cleaning algorithms.
[0009] Based on the revised carbon asset data, a random forest model is used to predict carbon sink capacity and carbon emission intensity, generating dynamic assessment results.
[0010] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0011] This invention discloses a method for assessing carbon neutrality in eco-parks. The method first acquires basic data such as the carbon sink capacity of the natural ecosystem and the intensity of carbon emissions from human activities within the park, and then establishes a trend model through time series analysis. Combining geographic information system (GIS) technology, a spatial distribution model of carbon storage and carbon flux is constructed, and multiple regression analysis is used to quantify influencing factors. This invention also performs quality assessment and cleaning of the monitoring data, employs a random forest model to predict carbon sink capacity and carbon emission intensity, and generates dynamic assessment results. This method integrates spatiotemporal analysis, data mining, and machine learning techniques, enabling a comprehensive assessment and prediction of carbon assets in eco-parks, providing a scientific basis for carbon asset management and decision-making, and contributing to improving the carbon neutrality level of eco-parks. Attached Figure Description
[0012] Figure 1 This is a flowchart of a carbon neutrality assessment method for an eco-park according to the present invention.
[0013] Figure 2 This is a schematic diagram of a carbon neutrality assessment method for an eco-park according to the present invention. Detailed Implementation
[0014] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0015] like Figure 1-2 This embodiment of a carbon neutrality assessment method for an eco-park may specifically include:
[0016] S101. Obtain basic data on various carbon assets within the eco-park, including data on the carbon sink capacity of natural ecosystems and the intensity of carbon emissions from human activities.
[0017] This study acquires data on the types and areas of natural ecosystems within the ecological park. An object-oriented classification method is used to interpret remote sensing images of the ecological park, yielding spatial distribution vector data for forest, grassland, and wetland ecosystems. The interpretation results are then validated for accuracy to determine if the data quality meets the requirements for subsequent analysis. If so, for the forest ecosystem, an allometric growth model is used to calculate biomass based on input tree species diameter at breast height (DBH) and tree height data. For the grassland and wetland ecosystems, a regression model is established based on the statistical relationship between the normalized difference in vegetation index (NDE) and biomass to estimate biomass. Based on the biomass estimation results and a pre-set carbon content conversion coefficient, the carbon storage and annual carbon sink of various natural ecosystems are calculated. Data on the types and scale of human activities related to energy consumption, transportation, and waste disposal within the ecological park are also acquired. An inventory of human-induced carbon emissions is established through field surveys and data collection. The emission factor method is used to calculate carbon emissions from various human activities based on energy consumption and pre-set emission factors. Finally, the carbon sink data of natural ecosystems and the carbon emission data of human activities are integrated to analyze the carbon balance within the ecological park, yielding the park's carbon balance analysis results.
[0018] Specifically, object-oriented classification is an advanced remote sensing image interpretation technique that considers not only the spectral information of individual pixels but also the shape, texture, and spatial relationships of objects in the image, thus more accurately identifying land cover types. For example, when identifying forests, traditional pixel-based classification methods may misclassify shrubs with similar spectral characteristics as forests, while object-oriented classification methods can distinguish forest patches from shrubs based on their typically larger area and regular shape. In remote sensing image interpretation of ecological parks, the image can be segmented into multiple homogeneous objects, and then the spectral, shape, and texture features of each object can be extracted. Classifiers such as decision trees and support vector machines can be used to classify these objects, thereby obtaining the spatial distribution of different ecosystems such as forests, grasslands, and wetlands. Allometric growth models describe the proportional relationships between the growth of different parts of a tree and are often used to estimate forest biomass. For example, the allometric growth model for a certain type of tree is B = a * D^b * H^c, where B represents biomass, D represents diameter at breast height (DBH), H represents tree height, and a, b, and c are model parameters. Through field surveys, the average diameter at breast height (DBH) of this tree species was measured to be 20 cm, and the average tree height was 15 m. Combined with existing model parameters, the average biomass per tree of this species can be calculated. Assuming that the average biomass per tree of a dominant tree species in the park is calculated to be 0.5 tons using an allometric growth model, and there are 10,000 trees of this species, then the total biomass of this species is 5,000 tons. The Normalized Difference Vegetation Index (NDVI) is an important indicator reflecting vegetation growth and has a significant correlation with biomass. For grassland and wetland ecosystems, a regression model between the two can be established using field-measured biomass data and corresponding remote sensing image NDVI values. For example, through field surveys, biomass was measured in multiple grassland quadrats at different locations within the park, and the corresponding NDVI values for these quadrats were obtained. Using this data, a linear regression model between NDVI and biomass can be established, such as B = m * NDVI + n, where B represents biomass, NDVI is the Normalized Difference Vegetation Index, and m and n are model parameters. Assuming that the average biomass of grasslands in the park is calculated to be 3 tons per hectare and the average biomass of wetlands is 4 tons per hectare, the carbon conversion factor (CCF) refers to the mass of carbon contained in a unit of dry biomass. The CCF varies among different ecosystems. For example, the CCF of forest ecosystems is typically around 0.5, while the CCFs of grassland and wetland ecosystems are slightly lower. By combining the biomass estimation results with the CCF, the carbon storage of various ecosystems can be calculated. For example, if the forest biomass in the park is 5000 tons and the CCF is 0.5, then the carbon storage of the forest ecosystem is 2500 tons. Annual carbon sink refers to the amount of carbon dioxide absorbed by an ecosystem each year, which can be estimated using the annual biomass growth and the CCF.Establishing a carbon emission inventory for human activities requires a detailed investigation of the types and scale of various human activities within the park. Through on-site investigations and data collection, detailed data on these activities can be obtained. For example, the annual electricity consumption of the office building is 100,000 kWh, the annual mileage of employee commutes and official vehicles is 50,000 kilometers, and the annual generation of domestic waste is 100 tons. The emission factor method is a commonly used method for calculating carbon emissions from human activities. Its basic principle is to calculate carbon emissions based on the consumption volume of various activities and the corresponding emission factors. For example, the emission factor for electricity consumption is 0.6 kg CO2 equivalent per kWh, the emission factor for gasoline is 2.3 kg CO2 equivalent per liter, and the emission factor for landfill is 0.5 tons CO2 equivalent per ton. Based on the consumption volume and corresponding emission factors of various activities within the park, the carbon emissions of various human activities can be calculated. For example, if the annual electricity consumption of the office building in the park is 100,000 kWh and the emission factor for electricity consumption is 0.6 kg CO2 equivalent per kWh, then the annual carbon emissions of the office building are 60 tons CO2 equivalent. By integrating carbon sink data from natural ecosystems with carbon emission data from human activities, the carbon balance within an ecological park can be analyzed. For example, if the total carbon sink of the park's forest, grassland, and wetland ecosystems is 100 tons of carbon per year, while human activities emit 80 tons of CO2 equivalent per year, then the park's net carbon sink is 20 tons of carbon per year. This indicates that the park as a whole acts as a carbon sink and plays a positive role in mitigating climate change. If human activities emit more carbon than the carbon sink of the natural ecosystems, then the park acts as a carbon source, requiring measures to reduce emissions or increase carbon sinks. Carbon balance analysis can provide data support for ecological park management, enabling the development of scientifically sound carbon reduction and carbon sink enhancement strategies, and promoting the green and low-carbon development of the park.
[0019] S102. Based on the basic data of carbon assets, time series analysis is used to model the carbon sink capacity and the variability trend of carbon emissions, and obtain the trend model.
[0020] Specifically, based on fundamental carbon asset data, the Pandas library in Python was used for data cleaning, deduplication, and missing value imputation preprocessing to ensure data quality and consistency. The preprocessed carbon asset data was then segmented along the time dimension, and key indicators such as carbon sequestration capacity and carbon emission intensity were extracted to construct a time-series dataset. The ARIMA model from Python's statsmodels library was used to model and analyze the constructed time-series dataset. The model hyperparameters were optimized using grid search and cross-validation to obtain the final trend prediction model. This trend prediction model was then used to predict future trends in carbon sequestration capacity and carbon emission intensity, yielding predictions for a specific future timeframe. The Matplotlib and Seaborn libraries in Python were used to visualize the prediction results, presenting future changes in carbon sequestration capacity and carbon emission intensity in the form of line charts and area plots, providing decision support for carbon asset management.
[0021] Specifically, preprocessing of carbon asset baseline data is a crucial step in ensuring the accuracy of subsequent analysis. The Pandas library can efficiently handle large-scale data, such as cleaning years of carbon emission data from a company within an eco-park. Data quality can be improved by removing duplicate records and filling in missing values (e.g., using averages or interpolation). Splitting the data along a time dimension allows for better capture of trends in carbon sink capacity and carbon emission intensity. Constructing time-series datasets is fundamental for predictive analysis. For example, carbon sink data from the eco-park over the past three years can be extracted, including annual afforestation area and tree growth, forming a continuous time series. Such datasets reflect the changing patterns of carbon sink capacity over time, providing a basis for subsequent modeling. ARIMA models are suitable for handling data with clear trends and seasonality, such as monthly carbon emissions from factories. Optimizing model parameters through grid search and cross-validation can improve predictive accuracy. Visualizing the prediction results is essential for decision support. Matplotlib line charts can visually display future trends in carbon emission intensity, while Seaborn area plots better represent the cumulative effect of carbon sink capacity. For example, a forecast curve of the carbon sequestration capacity of the eco-park over the next three years can be plotted, and the expected effects of different afforestation policies can be overlaid to help decision-makers choose the optimal solution.
[0022] S103. Using geographic information system technology and combining it with a trend model, model the spatial distribution of carbon storage and carbon flux to obtain a spatial distribution model.
[0023] Geographic information data of the study area, including land use type and vegetation cover, was acquired to construct a geographic information database. Historical monitoring data of carbon storage and carbon flux in the study area were also collected. The data was preprocessed, and outliers were identified and removed using methods such as box plots. Missing values were appropriately filled in. Carbon flux refers to the amount of carbon exchanged between an ecosystem and the atmosphere per unit time, including carbon absorption (positive flux) and carbon emissions (negative flux). The geographic coordinates of carbon storage and carbon flux monitoring points were spatially correlated with the geographic information data. ArcGIS's spatial connectivity tools were used to assign the carbon storage and carbon flux values of each monitoring point to the corresponding geographic location. ArcGIS's trend analysis tools were used to analyze the changing trends of carbon storage and carbon flux over time. Based on the trend analysis results, an appropriate time series model, such as the ARIMA model, was selected to establish a trend model for carbon storage and carbon flux. Using the established trend model, future changes in carbon storage and carbon flux were predicted. Using ArcGIS spatial analysis tools, the predicted changes in carbon storage and carbon flux were spatially matched with geographic information data to obtain predicted spatial distribution data of carbon storage and carbon flux within the study area. Kriging interpolation was employed using the Geostatistical Wizard module, selecting the spherical variogram model. Interpolation parameters were determined based on cross-validation results (minimizing root mean square error) to interpolate the predicted spatial distribution data of carbon storage and carbon flux, resulting in a continuous spatial distribution layer. The interpolation accuracy under different parameter combinations was evaluated, and the optimal parameters were selected for final interpolation. Based on the spatial distribution layer generated by interpolation, ArcGIS spatial statistics tools were used to calculate the global Moran's I index of carbon storage and carbon flux, assessing their spatial autocorrelation. Using ArcGIS's GeostatisticalAnalyst tool, empirical variograms of carbon storage and carbon flux were fitted to analyze their spatial heterogeneity characteristics. Combining spatial distribution characteristics and trend models, a spatial distribution model of carbon storage and carbon flux was constructed. A random forest algorithm was used to train the spatial distribution model, and grid search was employed to optimize the hyperparameters of the random forest, including the number of decision trees and the maximum number of features. Cross-validation was used to evaluate the predictive performance of the spatial distribution model, and the optimal model parameters were selected. The optimized spatial distribution prediction model was then used to predict carbon storage and carbon flux in the study area. Different future scenarios, such as land-use change and climate change, were set to simulate changes in carbon storage and carbon flux under these scenarios. The carbon balance under different scenarios was evaluated to provide a quantitative basis for formulating carbon neutrality strategies.
[0024] Specifically, the construction of a geographic information database is fundamental to carbon storage and carbon flux analysis. Taking a forest ecosystem as an example, data including land use type, vegetation cover, and topography can be collected. By combining remote sensing image interpretation with field surveys, detailed land use classification maps of the study area can be obtained, such as coniferous forests, broad-leaved forests, and shrublands. Vegetation cover can be calculated using the Normalized Difference Vegetation Index (NDVI). This data provides spatial background information for subsequent analysis. Preprocessing historical monitoring data of carbon storage and carbon flux is crucial to ensuring the reliability of the analysis results. Taking three years of monitoring data from an ecological park as an example, box plots are first used to identify outliers. Suppose that in a certain summer, extreme weather events caused abnormally high carbon flux values; these data points would appear as outliers in the box plot and need to be removed or corrected. For missing values, appropriate imputation methods can be selected based on the data characteristics. For example, for carbon flux data with obvious seasonality, the historical average value of the same season can be used for imputation. Spatial correlation analysis can combine carbon storage and carbon flux data with geographic location information. For example, carbon storage data from forest plots can be correlated with information such as land use type and altitude. This step lays the foundation for subsequent spatial analysis, enabling the exploration of the relationship between carbon storage and environmental factors. Time series analysis helps reveal long-term trends in carbon storage and carbon flux. Taking a forest ecosystem as an example, the ARIMA model can capture the seasonal fluctuations and long-term growth trends of carbon storage. The model may show that carbon storage increases year by year with the age of the forest, but the rate of increase gradually slows down. This trend analysis provides a basis for predicting future carbon sink potential. Spatial interpolation is a key technique for extending discrete monitoring point data to the entire study area. Taking Kriging interpolation as an example, the carbon storage distribution of the entire forest area can be estimated based on the carbon storage data of known monitoring points. By comparing the interpolation results of different variogram models (such as the spherical model and the exponential model), the optimal model is selected. The interpolation results may show that carbon storage is higher in areas with higher altitudes and gentler slopes, and lower in peripheral areas with more human disturbance. Spatial autocorrelation analysis can reveal the spatial distribution patterns of carbon storage and carbon flux. The global Moran's I index can quantify the spatial clustering of the entire study area. For example, a significantly positive Moran's I index for carbon storage indicates that areas with high carbon storage tend to cluster together, which may be related to specific terrain or vegetation types. The construction of spatial distribution models comprehensively considers geographical factors and temporal trends. Taking the random forest algorithm as an example, land use type, vegetation cover, and altitude can be used as predictor variables, with carbon storage as the target variable. Optimizing parameters through grid search, such as setting the number of decision trees to 500 and the maximum number of features to the square root of the total number of variables, can improve the model's predictive accuracy. Scenario analysis provides a scientific basis for formulating carbon neutrality strategies. For example, it can simulate the changes in carbon storage in the study area over the next three years under different afforestation policies.The results may show that under an active afforestation scenario, regional carbon storage could increase by 30%, while under a maintain-status scenario, it could increase by only 10%. These quantitative results provide important reference for policymakers to weigh the effects of different policies.
[0025] S104. Extract key variables from the spatial distribution model and use multiple regression analysis to quantify the effects of different factors on carbon assets.
[0026] Based on the spatial distribution model, multi-dimensional data related to carbon assets, including environmental, geographical, and climatic data, were acquired to construct a dataset of carbon asset influencing factors. Exploratory analysis was conducted on the dataset to understand its distribution characteristics, missing values, and outliers, followed by necessary data cleaning and preprocessing. Missing values were handled using methods such as mean imputation or KNN imputation, and outliers were identified and addressed using box plots or Z-scores, while the data was standardized. Pearson correlation coefficient analysis was used to analyze the correlation between each influencing factor and carbon asset quantity. Based on the absolute value and significance level of the correlation coefficient, key variables significantly correlated with carbon asset quantity were selected. For the selected key variables, multicollinearity was checked using the variance inflation factor (VIF). When the VIF value was greater than 10, principal component analysis was used to extract a comprehensive variable. A multiple linear regression model was constructed, with carbon asset quantity as the dependent variable and the key variables as independent variables, establishing a regression equation. The least squares method was used to estimate the parameters of the regression model, obtaining the regression model coefficients, and a t-test was used to test the significance and determine the degree of influence of each key variable. Diagnostic analysis was performed on the regression model to examine its goodness of fit, residual distribution, and heteroscedasticity, thus assessing its effectiveness and reliability. Simultaneously, 10-fold cross-validation was employed, and the predictive power and generalization performance of the model were evaluated using the root mean square error (RMSE) and coefficient of determination (R²). Based on the regression analysis, the impact of spatial autocorrelation was considered, and Moran's I was used to test the spatial autocorrelation of the residuals. A significantly greater than 0 Moran's I value indicates positive spatial autocorrelation, and a spatial lag model was chosen; a significantly less than 0 Moran's I value indicates negative spatial autocorrelation, and a spatial error model was chosen. The estimated values of the regression model coefficients were adjusted based on the test results. Using the adjusted regression model coefficients as weights, combined with actual data from each region, a weighted summation was performed to obtain the quantitative assessment results of carbon assets in each region. A spatial distribution map of carbon assets was then generated using ArcGIS to visually display the spatial distribution characteristics of carbon assets.
[0027] Specifically, constructing a dataset of carbon asset influencing factors is a crucial step in assessing carbon assets. Taking a forest ecosystem as an example, multi-dimensional data can be collected, including vegetation cover, average annual temperature, annual precipitation, soil type, and topographic slope. During the data exploration phase, histograms may reveal that average annual temperature follows a normal distribution, while vegetation cover exhibits a right-skewed distribution. For missing precipitation data, the average values from nearby meteorological stations can be used to fill in the gaps. When using box plots to identify outliers, abnormally high carbon storage at certain sampling points may be discovered. After data preprocessing, correlation analysis is performed. Assuming the Pearson correlation coefficient shows a correlation coefficient of 0.85 between vegetation cover and carbon asset quantity, -0.62 for average annual temperature, and 0.58 for annual precipitation, all significant at the 0.05 significance level. This indicates that vegetation cover has a strong positive impact on carbon asset quantity, while rising temperatures may lead to a decrease in carbon asset quantity. In the multicollinearity check, if the VIF values of both soil organic matter content and soil type exceed 10, principal component analysis can be used to merge these two variables into a single "soil fertility index." When constructing the multiple linear regression model, the assumed equation is: Carbon asset quantity = 2.5 * Vegetation cover - 1.8 * Average annual temperature + 1.2 * Annual precipitation + 0.8 * Soil fertility index + 0.5 * Topographic slope + Constant term. The t-test shows that, except for topographic slope, the coefficients of all other variables are significant at the 0.01 level. This implies that topographic slope has a relatively small impact on carbon asset quantity in this ecosystem. Model diagnostics show that the residuals are normally distributed, but the QQ plot shows slight deviations at high and low values, suggesting that the model may have some bias in its predictions under extreme conditions. The 10-fold cross-validation results show that the model's average RMSE is 0.35 and R² is 0.78, indicating that the model has good predictive ability and generalization performance. Considering the spatial characteristics of carbon assets, a spatial autocorrelation test is performed. Assuming Moran's I value is 0.32 and the p-value is less than 0.01, it indicates a significant positive spatial autocorrelation. This suggests that high-carbon asset areas tend to cluster, possibly related to specific geographical features. Based on this, a spatial lag model is selected for correction, yielding correction coefficients that consider spatial effects. Finally, using the corrected model coefficients and actual data from each region, a quantitative assessment of carbon assets is calculated. The spatial distribution map of carbon assets generated by ArcGIS may show that high-carbon asset areas are mainly concentrated in areas with high forest vegetation cover, while carbon assets are relatively low in areas with concentrated grassland. This visualization not only intuitively demonstrates the spatial distribution characteristics of carbon assets but also provides important evidence for formulating regional carbon neutrality strategies.
[0028] S105. Obtain carbon asset monitoring data, assess the data quality, and correct any outliers using data cleaning algorithms.
[0029] Carbon asset monitoring data is acquired through a data interface and stored in a MySQL database. This carbon asset monitoring data includes, but is not limited to: natural carbon sink data: vegetation cover (based on Landsat-8 remote sensing imagery, 30-meter resolution, quarterly updates), soil organic carbon content (laboratory measurements, annual sampling); and anthropogenic emission data: electricity consumption within the park (real-time monitoring by smart meters), and transportation fuel consumption (GPS mileage statistics, monthly summaries). For the stored carbon asset monitoring data, Python's Pandas library is used for data preprocessing, including converting string data to numeric data and filling missing values with the mean to ensure data format consistency and integrity. According to preset data quality assessment rules, Python's NumPy library is used to assess the quality of the preprocessed carbon asset monitoring data. The 3σ principle is used to calculate the mean and standard deviation of the data, determining whether each data point is within the range of the mean ± 3 times the standard deviation; data points outside this range are considered outliers. For outliers outside the range, different processing methods are selected based on the degree of anomaly. Extreme outliers that significantly deviate from the normal range are directly deleted. Minor outliers are replaced with the median to reduce their impact on the overall data distribution. Simultaneously with outlier processing, Z-score normalization is used to calculate the Z-score for each data point (the difference between the data point and the mean divided by the standard deviation). Data points with Z-scores exceeding a preset threshold (e.g., ±3) are marked as outliers. For these outliers, the median is used to replace them, ensuring data continuity and integrity. The processed carbon asset monitoring data is compared with the original data, and the mean, median, and standard deviation are calculated using Python's SciPy library. By comparing the data distribution before and after processing, the effectiveness of the outlier handling and normalization methods is evaluated. Based on the evaluation results, the threshold parameter for Z-score normalization is adjusted to optimize the accuracy and recall of outlier detection and processing, improving the robustness and adaptability of the data cleaning algorithm. The optimized data cleaning algorithm is then applied again to the carbon asset monitoring data to obtain high-quality cleaned data.
[0030] Specifically, the raw data obtained through the data interface may contain various types, such as carbon dioxide concentration, vegetation cover, and soil carbon content. After this data is stored in the MySQL database, it needs to undergo systematic preprocessing. For example, the string "45.2%" is converted to the floating-point number 0.452 to ensure data type consistency. For missing data, such as data gaps at certain monitoring stations due to equipment failure, the mean of historical data for that station can be used to fill the gaps, ensuring data continuity. Data quality assessment is a crucial step in ensuring the reliability of the analysis results. When using the 3σ principle for outlier detection, assuming the mean carbon dioxide concentration at a certain monitoring point is 400 ppm and the standard deviation is 20 ppm, data outside the range of 340 ppm to 460 ppm will be considered outliers. For data with significant deviations, such as a sudden appearance of 900 ppm, which may be caused by equipment failure or human interference, it should be deleted directly. For data with slight deviations, such as 430 ppm, the median can be used to replace it to reduce the impact on the overall distribution. The Z-score normalization method can further refine the judgment of outliers. For example, after standardizing the data from all monitoring points, a data point with a Z-score of 2.8, although within the 3σ range, may still be considered a potential outlier. Adjusting the Z-score threshold can balance the stringency of outlier detection with the amount of data retained. This method is particularly suitable for situations where there are significant differences in data ranges between different monitoring points, such as comparing the carbon absorption capacity of forests and grasslands. Evaluating the effectiveness of data processing is crucial for optimizing the algorithm. Assuming the original data has a mean of 410 ppm and a standard deviation of 25 ppm for carbon dioxide concentration, after processing, the mean becomes 405 ppm, and the standard deviation decreases to 20 ppm. This change indicates that the impact of outliers has been effectively reduced, and the data distribution is more concentrated. However, if the processed mean deviates significantly from the original data, such as decreasing to 380 ppm, it may indicate overprocessing, requiring readjustment of the algorithm parameters. The optimized data cleaning algorithm should be adaptable to different types of carbon asset monitoring data. For example, for vegetation cover data with significant seasonal variations, the algorithm can adjust the outlier judgment criteria according to different seasons. During the growing season, larger numerical fluctuations are allowed; while during the dormant period, a stricter standard is used. This dynamic adjustment enhances the algorithm's adaptability, ensuring that the cleaned data retains accurate information about environmental changes while removing unreasonable outliers. Through this series of data processing and optimization steps, the resulting high-quality data provides a reliable foundation for carbon asset assessment. This cleaned data not only more accurately reflects actual carbon emissions and absorption but also provides strong support for developing carbon neutrality strategies. For example, analyzing the processed data can identify areas with the highest carbon absorption efficiency, providing a scientific basis for enhancing the carbon sequestration capacity of eco-parks.At the same time, this data can also be used to build predictive models to help policymakers prepare for potential peak carbon emissions and develop more targeted emission reduction measures.
[0031] S106. Based on the revised carbon asset data, a random forest model is used to predict carbon sink capacity and carbon emission intensity, generating dynamic assessment results.
[0032] The original dataset was obtained from carbon asset-related data. Missing values were imputed using the mean method, and outliers were identified and removed using box plots. The processed data was then min-max standardized to obtain a cleaned carbon asset dataset. Features related to carbon sink capacity and carbon emission intensity were extracted from the cleaned dataset, and a subset of key features with significant impact on prediction results was selected using the information gain ratio method. A random forest regression model was constructed using this subset of key features as input, with 100 trees and a minimum sample size of 5 per node. The model's hyperparameters were optimized using grid search. The model was trained and evaluated using 5-fold cross-validation, and the mean absolute error and mean squared error were recorded as model performance metrics. The trained model was applied to new carbon asset data to predict carbon sink capacity and carbon emission intensity. The predicted results were compared with actual values, and the R² value was calculated to evaluate the model's prediction accuracy. Based on the prediction results and considering the dynamic changes in carbon assets, a weighted average method was used to calculate a comprehensive score for carbon assets and classify them into levels. Matplotlib is used to plot the changes in carbon sink capacity and carbon emission intensity, generating a carbon asset assessment report, and displaying the overall score and grade distribution using a heatmap.
[0033] Specifically, firstly, raw datasets are obtained from multiple sources, such as satellite remote sensing data and ground monitoring data. For missing values, the mean imputation method is used. For example, if carbon sequestration data for a monitoring point in June 2023 is missing, the average value of the same period data from the past five years can be used to impute it. For outliers, box plots are used to identify and remove them. For instance, if abnormally high values far exceeding the normal range appear in the carbon emission intensity data of an eco-park, this may be due to equipment malfunction and should be removed. Data standardization is a crucial step in ensuring the comparability of indicators with different dimensions. The min-max standardization method is used to map each indicator value to the 0-1 range. For example, carbon sequestration capacity is converted from the original 0-500 tons / hectare / year to a standardized value of 0-1. This process contributes to the stability and accuracy of subsequent model training. Feature selection is key to improving model efficiency. The information gain ratio method is used to screen key features, such as vegetation cover and soil organic matter content, which have a significant impact on carbon sequestration capacity. This not only reduces model complexity but also improves prediction accuracy. The construction of a random forest regression model is central to predicting carbon asset changes. A set of 100 decision trees is used, with a minimum sample size of 5 per node. Hyperparameters such as maximum tree depth and number of features are optimized using grid search. Five-fold cross-validation is used to evaluate model performance, recording the mean absolute error and mean squared error. For example, when predicting the carbon sink capacity of a region, the model has a mean absolute error of 0.5 tons / hectare / year and a mean squared error of 0.3, indicating good predictive ability. When applied to new data, the model can predict future carbon sink capacity and carbon emission intensity. For instance, it predicts the carbon emission intensity of an eco-park in 2025 to be 4.2 tons of CO2 equivalent, with an R-squared value of 0.95 compared to the actual value of 4.3 tons, demonstrating high accuracy. Based on the prediction results and historical data, a weighted average method is used to calculate the comprehensive carbon asset score. For example, a forest area with a carbon sink capacity score of 85 and a carbon storage stability score of 90 has a comprehensive score of 87.5, classifying it as a Grade A carbon asset. Visualization is an effective means of intuitively displaying the assessment results. Matplotlib is used to plot curves showing changes in carbon sequestration capacity and carbon emission intensity, clearly illustrating trends. Heatmaps can display the distribution of carbon asset scores across different regions, helping policymakers quickly identify areas of key concern.
[0034] Although the present invention has been described in detail above using general descriptions and specific embodiments, it will be apparent to those skilled in the art that modifications and improvements may be made thereto. Therefore, such modifications and improvements, without departing from the spirit of the present invention, are intended to be within the scope of protection claimed herein.
Claims
1. A method for assessing carbon neutrality in an eco-park, characterized in that, The method includes: Acquire basic data on various carbon assets within the eco-park, including the carbon sink capacity of the natural ecosystem and carbon emissions from human activities; Based on carbon asset data, time series analysis is used to model the changing trends of carbon sink capacity and carbon emission intensity, and obtain a trend model. Using geographic information system technology and combining trend models, the spatial distribution of carbon storage and carbon flux is modeled to obtain a spatial distribution model; Key variables were extracted from the spatial distribution model, and multiple regression analysis was used to quantify the effects of different factors on carbon assets. Acquire carbon asset monitoring data, assess data quality, and correct outliers using data cleaning algorithms. Based on the revised carbon asset data, a random forest model is used to predict carbon sink capacity and carbon emission intensity, generating dynamic assessment results. The process of predicting carbon sink capacity and carbon emission intensity using a random forest model based on the corrected carbon asset data, and generating dynamic assessment results, also includes: Obtain carbon asset-related data and use the mean imputation method to handle missing values in the carbon asset-related data; Outliers in carbon asset-related data are identified and removed using a box plot method to obtain processed carbon asset data. The processed carbon asset data is subjected to min-max normalization to obtain a cleaned carbon asset dataset. Features related to carbon sink capacity and carbon emission intensity were extracted from the cleaned carbon asset dataset. The information gain ratio method was used to filter the features and obtain a subset of key features. A random forest regression model is constructed using a subset of key features as input. The random forest regression model has 100 trees and a minimum number of samples per node. The hyperparameters of the random forest regression model are optimized using a grid search method. The Lin regression model was trained and evaluated to obtain the mean absolute error and mean squared error. The trained random forest regression model is applied to new carbon asset data to predict carbon sink capacity and carbon emission intensity. The predicted results are compared with the actual values, and the R² value is calculated. Based on the forecast results and the dynamic changes in carbon assets, a weighted average method is used to calculate the comprehensive score of carbon assets and classify them into levels. Matplotlib is used to plot the changes in carbon sink capacity and carbon emission intensity, and a carbon asset assessment report is generated, with a heat map showing the overall score and grade distribution.
2. The carbon neutrality assessment method for eco-parks according to claim 1, characterized in that, The acquisition of basic data on various carbon assets within the eco-park, including the carbon sink capacity of the natural ecosystem and carbon emissions from human activities, includes: Data on the types and areas of natural ecosystems within the ecological park were obtained. An object-oriented classification method was used to interpret the remote sensing images of the ecological park, resulting in spatial distribution vector data of forest, grassland, and wetland natural ecosystems. For forest ecosystems, an allometric growth model is used to calculate the biomass of the forest ecosystem based on the input data of tree species diameter at breast height (DBH) and tree height. For grassland and wetland ecosystems, a regression model is established based on the statistical relationship between normalized vegetation index and biomass to estimate the biomass of grassland and wetland ecosystems. Based on the biomass estimation results and combined with the preset carbon content conversion coefficient, the carbon storage and annual carbon sink of various natural ecosystems are calculated. Obtain data on the types and scale of human activities related to energy consumption, transportation, and waste disposal within the eco-park, and establish a carbon emission inventory of human activities through field surveys and data collection; The emission factor method is used to calculate the carbon emissions of various human activities based on energy consumption and preset emission factors. By integrating carbon sink data from natural ecosystems with carbon emission data from human activities, the carbon balance within the eco-park is analyzed, yielding the results of the carbon balance analysis.
3. The carbon neutrality assessment method for eco-parks according to claim 1, characterized in that, Based on carbon asset fundamental data, time series analysis is used to model the changing trends of carbon sink capacity and carbon emission intensity, resulting in a trend model, including: Obtain basic carbon asset data, which includes carbon sink capacity data and carbon emission intensity data; Preprocessing the basic carbon asset data yields preprocessed carbon asset data. Preprocessing includes data cleaning, deduplication, and missing value imputation; The preprocessed carbon asset data is segmented according to the time dimension, and carbon sink capacity and carbon emission intensity indicators are extracted to construct a time series dataset. The ARIMA model was used to model and analyze the time series dataset to obtain a time series prediction model; The hyperparameters of the time series forecasting model are optimized by grid search and cross-validation to obtain a trend forecasting model. The trend prediction model is used to predict the future trends of carbon sink capacity and carbon emission intensity, and the prediction results are obtained within a certain time range. The prediction results are visualized to show the future changes in carbon sequestration capacity and carbon emission intensity.
4. The carbon neutrality assessment method for eco-parks according to claim 1 or 3, characterized in that, The method utilizes Geographic Information System (GIS) technology, combined with trend models, to model the spatial distribution of carbon storage and carbon flux, obtaining a spatial distribution model, including: Acquire geographic information data of the study area, including land use type and vegetation cover, and construct a geographic information database; Historical monitoring data on carbon storage and carbon flux in the study area were obtained, the data were preprocessed, outliers were identified and removed using box plot method, and missing values were filled in appropriately. Spatial correlation is established between the geographic coordinates of carbon storage and carbon flux monitoring points and geographic information data. Establish a time series prediction model based on the time-dimensional changing trends of carbon storage and carbon flux; Time series forecasting models are used to predict future changes in carbon storage and carbon flux. Spatial matching of predicted carbon storage and carbon flux changes with geographic information data; Kriging interpolation was used to interpolate the spatial distribution prediction data of carbon storage and carbon flux to obtain a continuous spatial distribution layer. Calculate the global Moran's I index for carbon storage and carbon flux, and assess its spatial autocorrelation. We fit the empirical variograms of carbon storage and carbon flux to analyze their spatial heterogeneity characteristics. A spatial distribution model of carbon storage and carbon flux was constructed and trained and optimized using the random forest algorithm; The optimized spatial distribution model was used to predict carbon storage and carbon flux in the study area. By setting different future scenarios, simulating changes in carbon storage and carbon flux under these scenarios, and assessing the carbon budget balance, we can provide a quantitative basis for formulating carbon neutrality strategies.
5. The carbon neutrality assessment method for eco-parks according to claim 1, characterized in that, The method extracts key variables from the spatial distribution model and uses multiple regression analysis to quantify the effects of different factors on carbon assets, including: Acquire multi-dimensional environmental, geographical, and climate data related to carbon assets, and construct a dataset of carbon asset influencing factors. Exploratory analysis was performed on the dataset, missing values were handled using mean imputation or KNN imputation methods, outliers were identified and processed using box plots or Z-score methods, and the data were standardized. Pearson correlation coefficient analysis was used to analyze the correlation between various influencing factors and carbon asset quantity, and key variables that are significantly correlated with carbon asset quantity were screened out. For key variables, the variance inflation factor is used to check for multicollinearity among variables. When the VIF value is greater than 10, principal component analysis is used to extract comprehensive variables. A multiple linear regression model was constructed, with carbon asset quantity as the dependent variable and key variables as independent variables. The least squares method was used to estimate the parameters of the regression model, and the t-test was used to determine the degree of influence of each key variable. Diagnostic analysis was performed on the regression model, using 10-fold cross-validation. The predictive power and generalization performance of the model were evaluated by the root mean square error and the coefficient of determination. The spatial autocorrelation of the residuals is tested by Moran's index. When the Moran's index value is greater than 0, the spatial lag model is adopted. When the Moran index is less than 0, the spatial error model is used, and the estimated values of the regression model coefficients are corrected based on the test results. Using the corrected regression model coefficients as weights, and combining them with actual data from each region, a quantitative assessment of carbon assets in each region is obtained through weighted summation, generating a spatial distribution map of carbon assets to demonstrate their spatial distribution characteristics.
6. The carbon neutrality assessment method for eco-parks according to claim 1, characterized in that, The process of acquiring carbon asset monitoring data, assessing data quality, and correcting outliers using data cleaning algorithms includes: Acquire carbon asset monitoring data and store the carbon asset monitoring data in the database; Carbon asset monitoring data is preprocessed by converting string data into numerical data and filling missing values with the mean. The preprocessed carbon asset monitoring data is assessed for quality according to the preset data quality assessment rules. The 3σ principle is used to calculate the mean and standard deviation of carbon asset monitoring data, and to determine whether each data point is within the range of mean ± 3 times the standard deviation. Data points outside the range are considered outliers. Different processing methods are adopted according to the degree of abnormality of the outliers. Extreme outliers that deviate significantly from the normal range are directly deleted. For minor outliers, use the median as a replacement. Calculate the Z-score value for each data point, mark data points with Z-score values exceeding a preset threshold as outliers, and replace them with the median; The processed carbon asset monitoring data is compared with the original data, the statistical indicators before and after data processing are calculated, the threshold parameter of Z-score normalization is adjusted, and the accuracy and recall of outlier detection and processing are optimized. The optimized data cleaning algorithm was applied to carbon asset monitoring data to obtain high-quality cleaned data.
Citation Information
Patent Citations
Plateau lake region carbon neutralization calculation method based on carbon revenue and expenditure balance analysis
CN114266003A
Carbon right asset digital acquisition method based on block chain
CN115455490A