Carbon peak reaching prediction method and system based on data analysis

By categorizing carbon emission data by industry and combining it with multidimensional correlation analysis of economic indicators, this method addresses the problem of neglecting industry differences and economic factors in existing carbon peak prediction methods, thus achieving more accurate carbon emission prediction and regional carbon peak planning.

CN121031997APending Publication Date: 2025-11-28SHANGHAI ENERGY SAVING TECH SERVICE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511557125.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing carbon peak prediction methods fail to accurately capture the carbon emission patterns of various industries, ignore industry differences, and cannot adapt to economic development and seasonal fluctuations, resulting in inaccurate prediction results.

Method used

By classifying carbon emission data by industry category, time series decomposition is performed to extract trend, seasonal and residual component features, a dynamic weight matrix of industry carbon emissions is established, and multidimensional correlation matching is performed in combination with economic development indicators to generate industry carbon emission intensity prediction curves and construct regional carbon peak prediction models.

Benefits of technology

It enables accurate prediction of carbon emissions, reflects the dynamic changes in industry characteristics and economic factors, and provides scientific carbon reduction strategies and planning guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121031997A_ABST
    Figure CN121031997A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of carbon peak reaching prediction, and discloses a carbon peak reaching prediction method and system based on data analysis. The method comprises the steps of collecting carbon emission historical data of a target area, dividing the data into a data set according to industries, and performing time sequence decomposition on the data set to extract numerical features of trend, season and residual components; establishing an industry carbon emission dynamic weight matrix based on trend and seasonal component characteristics, and performing multi-dimensional correlation matching on an economic development index data set of a target area and the matrix; generating a carbon emission intensity prediction curve of each industry according to a matching result, and finally integrating all the curves to construct a regional carbon peak prediction model. According to the method, through multi-dimensional data fusion and dynamic weight design, internal association between industry carbon emission characteristics and economic development is accurately captured, scientific prediction of regional carbon peak reaching is realized, and effective reference is provided for low-carbon development decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of carbon peak prediction technology, specifically to a carbon peak prediction method and system based on data analysis. Background Technology

[0002] Current mainstream carbon peak prediction methods have several shortcomings. Some methods adopt a holistic prediction approach, treating regional carbon emission data as a unified whole for modeling. This ignores the significant differences in energy consumption structures, production models, and emission reduction potential among different industries such as industry, transportation, and construction. Consequently, they fail to accurately capture the carbon emission patterns of each industry, resulting in significant deviations between predicted and actual results. Other methods, while considering industry classification, often use fixed weights that fail to adjust for the dynamic characteristics of carbon emissions over time. This makes them ill-suited to the long-term evolution of trend components and the cyclical fluctuations of seasonal components, and thus cannot accurately reflect the degree of influence of each component on industry carbon emissions at different times.

[0003] Carbon emissions are closely linked to economic development. Economic indicators such as economic growth rate, industrial restructuring, and technological innovation level directly or indirectly affect the scale of carbon emissions in various industries. However, most existing forecasting methods simply establish a linear correspondence between economic indicators and carbon emissions, failing to delve into the multidimensional intrinsic relationship between the two. This results in an incomplete analysis of the impact of economic factors on carbon emissions, further reducing the reliability of forecasting models. Furthermore, some methods process carbon emission data in a coarse manner, failing to perform effective component decomposition and extract key features, leading to insufficient fit of the forecast curves and an inability to provide accurate technical support for regional carbon peaking planning. These problems make existing carbon peaking forecasting methods insufficient to meet practical application needs, necessitating a more scientific and accurate forecasting solution. Summary of the Invention

[0004] The purpose of this invention is to provide a carbon peak prediction method and system based on data analysis to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides a carbon peak prediction method based on data analysis, the method comprising: Collect historical carbon emission data within the target area, and divide the historical carbon emission data according to industry categories to form an industry carbon emission dataset; The industry carbon emission dataset is decomposed into time series to extract the numerical features of trend components, seasonal components, and residual components. Based on the numerical characteristics of the trend component and the seasonal component, a dynamic weight matrix for industry carbon emissions is established. Obtain the economic development indicator dataset of the target region, and perform multi-dimensional correlation matching between the economic development indicator dataset and the industry carbon emission dynamic weight matrix. Based on the multidimensional correlation matching results, an industry carbon emission intensity prediction curve is generated. By integrating carbon emission intensity prediction curves from all industries, a regional carbon peak prediction model is constructed.

[0006] Preferably, the time-series decomposition of the industry carbon emission dataset includes: An adaptive filtering algorithm is used to smooth the industry carbon emission dataset to eliminate abnormal fluctuations in data; A frequency domain transformation was performed on the smoothed industry carbon emission dataset to separate the trend component, seasonal component, and residual component. The energy proportions of the trend component and the seasonal component are calculated and used as the basis for the weighting of the numerical features.

[0007] Preferably, establishing the industry carbon emission dynamic weight matrix includes: The weight of the long-term trend of carbon emissions in the industry is determined based on the energy proportion of the trend components. Based on the energy proportion of the seasonal components, the periodic fluctuation weight of industry carbon emissions is calculated; The long-term trend weights and periodic fluctuation weights are normalized to generate a dynamic weight matrix for industry carbon emissions.

[0008] Preferably, the acquisition of the economic development indicator dataset for the target region includes: Collect time-series data on GDP growth rate, industrial added value, total energy consumption, and population growth rate of the target region; The time series data is standardized to eliminate dimensional differences; The standardized time-series data are grouped according to industry categories to form an economic development indicator dataset.

[0009] Preferably, the step of performing multi-dimensional correlation matching between the economic development indicator dataset and the industry carbon emission dynamic weight matrix includes: Calculate the correlation coefficients between each indicator in the economic development indicator dataset and the dynamic weight matrix of industry carbon emissions; Indicators with correlation coefficients exceeding a preset threshold are selected as key influencing factors; Based on the aforementioned key influencing factors, adjust the numerical distribution of the industry's dynamic carbon emission weight matrix.

[0010] Preferably, the generation of the industry carbon emission intensity prediction curve includes: Based on the adjusted industry carbon emission dynamic weight matrix, the historical variation pattern of industry carbon emission intensity is fitted. A grey prediction algorithm is introduced to iteratively calculate the future trend of carbon emission intensity in the industry; The iterative calculation results are converted into time series curves to generate industry carbon emission intensity prediction curves.

[0011] Preferably, the construction of the regional carbon peak prediction model includes: By overlaying the carbon emission intensity prediction curves of all industries, the predicted value of the total regional carbon emissions is calculated. Identify the peak points in the predicted values ​​to determine the carbon peak time point; Based on the carbon peak time points, a regional carbon peak prediction model is generated.

[0012] Preferably, the method further includes: Real-time monitoring of actual carbon emission data in the target area, and comparison with the predicted values ​​of the carbon peak prediction model for the area; When the deviation between actual carbon emission data and predicted values ​​exceeds the preset range, the dynamic weight matrix of industry carbon emissions is readjusted.

[0013] Preferably, the readjustment of the industry carbon emission dynamic weight matrix includes: Based on the deviation between actual carbon emission data and predicted values, the numerical characteristics of the trend component and the seasonal component are corrected. Update the industry carbon emission dynamic weight matrix based on the corrected numerical characteristics; The updated industry carbon emission dynamic weight matrix is ​​input into the regional carbon peak prediction model to recalculate the carbon peak time node.

[0014] Preferably, the present invention also includes a carbon peak prediction system based on data analysis, the system including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, it implements the steps of the carbon peak prediction method based on data analysis as described above.

[0015] Compared with the prior art, the beneficial effects of the present invention are: By categorizing historical carbon emission data by industry, we can accurately distinguish the inherent differences among industries in terms of energy consumption types, production process complexity, and emission reduction potential. This avoids the distortion of patterns caused by traditional holistic forecasting methods that ignore industry heterogeneity, allowing subsequent predictive analysis to focus on the individual characteristics of each industry and making forecasts more targeted. Time-series decomposition of industry carbon emission datasets and extraction of trend, seasonal, and residual components extracts key information corresponding to different influencing factors from the messy raw data. This clearly presents the long-term evolution trend, periodic fluctuation patterns, and random disturbances of carbon emissions, providing a high-quality feature foundation for subsequent model construction, enabling models to extrapolate based on more accurate information.

[0016] A dynamic weight matrix for industry carbon emissions is established based on the numerical characteristics of trend and seasonal components. This breaks through the limitations of fixed weights in traditional forecasting methods, allowing weights to be flexibly adjusted according to the passage of time and fluctuations in industry carbon emission characteristics. This better reflects the actual situation of carbon emissions changing with economic and seasonal factors, effectively improving the rationality of weight allocation and making the quantitative analysis of industry carbon emissions more in line with objective laws. Multi-dimensional correlation matching between the economic development indicator dataset of the target region and the dynamic weight matrix for industry carbon emissions can deeply explore the complex intrinsic relationships between economic indicators such as economic growth, industrial structure, and technological progress and carbon emissions from various industries. This avoids the one-sidedness of simple linear correlations in traditional methods, fully integrating the forecasting process into the actual scenario of economic development and reducing forecast bias caused by neglecting economic correlation effects.

[0017] The industry carbon emission intensity prediction curves generated based on multidimensional correlation matching results can accurately reflect the carbon emission intensity change trends of individual industries at different development stages, providing a clear reference direction for clarifying the emission reduction priorities of each industry and formulating differentiated industry emission reduction strategies. By integrating the carbon emission intensity prediction curves of all industries to construct a regional carbon peak prediction model, comprehensive prediction coverage from the micro-level of industries to the macro-level of regions is achieved. This not only preserves the individual carbon emission characteristics of each industry but also achieves comprehensive control over the overall carbon emission trend of the region, making the final prediction results both detailed and complete, and more closely aligned with the actual progress of regional carbon peaking.

[0018] This method, through multi-stage refined data processing and multi-dimensional deep correlation analysis, constructs a regional carbon peak prediction model with enhanced adaptability and reliability, capable of more accurately presenting the dynamic trends of regional carbon emissions. Its application requires no complex hardware support, the data processing flow is clear and orderly, and it is applicable to target regions of different sizes and industrial structures, possessing broad applicability and practical value. The prediction results obtained through this method can provide clear guidance for regions to formulate scientific and reasonable carbon emission reduction plans, optimize industrial structure layout, rationally allocate emission reduction resources, and promote energy structure transformation, helping regions to more efficiently achieve carbon peak targets and promoting the comprehensive construction of a green and low-carbon development model. Attached Figure Description

[0019] Figure 1 A trend chart showing the correlation between economic indicators and total carbon emissions; Figure 2 This is a flowchart of time series decomposition and feature extraction. Figure 3 Heatmap showing the correlation coefficients between economic indicators and industry carbon emission weights; Figure 4 A flowchart for generating the industry's carbon emission intensity prediction curve. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] Please see Figure 1 This invention provides a data analysis-based method for predicting carbon peak emissions. The method includes: systematically processing carbon emission data and economic indicators within a target region to accurately predict the carbon peak time. Historical carbon emission data within the target region is collected from government statistical departments, industry reports, or monitoring platforms, and categorized according to industry sectors such as industry, energy, and transportation to form a structured industry carbon emission dataset. The categorization is based on national or international standard industry classification codes to ensure data consistency and comparability. The industry carbon emission dataset is decomposed into time series components, and statistical methods are used to extract numerical features of trend components, seasonal components, and residual components, such as the slope of the trend component, the amplitude of the seasonal component, and the variance of the residual component. Based on these numerical features, a dynamic weight matrix for industry carbon emissions is established, reflecting the degree of influence of long-term changes and cyclical fluctuations in carbon emissions from different industries. An economic development indicator dataset for the target region is obtained, including macroeconomic indicators such as GDP and industrial added value, and dimensional differences are eliminated through data cleaning and normalization. The economic development indicator dataset is then matched with the dynamic weight matrix for industry carbon emissions in a multidimensional manner, and correlation analysis or machine learning algorithms are used to identify key driving factors. The matching results are used to generate industry carbon emission intensity prediction curves, which are obtained by fitting historical data and external variables. The carbon emission intensity prediction curves from all industries are integrated, and a regional carbon peak prediction model is constructed through weighted summation or model fusion. This model outputs a predicted value of total carbon emissions over time and identifies peak points to determine the carbon peak time node.

[0022] Example 1 See Figure 2The industry carbon emission dataset is derived from historical data sets classified according to national standard industry classifications. The data collection timeframe should cover a sufficiently long period, such as at least ten consecutive years of annual or monthly data. An adaptive filtering algorithm is applied to smooth the original industry carbon emission dataset. The algorithm is based on the minimum mean square error criterion. It iterates through each data point in the dataset using a variable-length sliding window. The window length is not fixed; it is dynamically adjusted based on the local volatility of the data sequence. A longer window is used in areas of gentle data change to enhance the smoothing effect, while a shorter window is used in areas of dramatic data change to retain more detail. Internally, the algorithm defines an abnormal fluctuation threshold, which is proportional to the overall variance of the data sequence. When the deviation of a data point from the local mean within the window exceeds the threshold, the adaptive filtering algorithm marks the point as an outlier and corrects it by replacing it with the weighted average of adjacent data points within the window. The smoothing process requires iteration, with the number of iterations determined by the complexity of the data sequence, until the difference between two consecutive smoothed data sequences is less than a preset tolerance limit.

[0023] The smoothed industry carbon emission dataset enters the frequency domain transformation stage, which is implemented using the Fast Fourier Transform (FFT). The FFT transforms the carbon emission data sequence in the time domain to the frequency domain, where the data sequence is represented as a superposition of sine and cosine waves of different frequencies and amplitudes. Low-frequency components correspond to the long-term, slowly changing trend portion of the data, while high-frequency components correspond to short-term fluctuations and noise. Separating the trend component requires setting a low-frequency cutoff frequency. The cutoff frequency is determined by analyzing the power spectral density of the data sequence; the frequency corresponding to the first significant peak on the power spectral density curve is usually selected as the cutoff frequency. Retaining all frequency components below the cutoff frequency and performing an inverse FFT yields the trend component. Extracting the seasonal component relies on identifying periodic patterns in the data. By performing peak detection on the power spectral density curve, frequency peaks corresponding to known periods (such as annual or quarterly periods) are identified. These peak frequencies and their harmonic components are retained, and after an inverse FFT, the seasonal component is obtained. Subtracting the trend and seasonal components from the original data sequence leaves the residual components, which contain random fluctuations in the data that cannot be explained by trend and seasonal factors.

[0024] Calculating the energy proportions of trend and seasonal components is a crucial step in quantifying their relative importance. Energy is defined as the sum of the squares of the values ​​at each point in the data sequence. The energy of the trend component is the sum of the squares of all data points in the trend component sequence, and the energy of the seasonal component is the sum of the squares of all data points in the seasonal component sequence. The total energy of the data sequence is the sum of the squares of all data points in the original smoothed data sequence. The energy proportion of the trend component equals the trend component energy divided by the total energy of the data sequence, multiplied by 100%, and the energy proportion of the seasonal component equals the seasonal component energy divided by the total energy of the data sequence, multiplied by 100%. The energy proportion calculation results serve as the weighting basis for numerical features, including the average slope of the trend component, the amplitude of the seasonal component, and the principal period. For example, a higher energy proportion for the trend component means that the average slope of the trend component has a higher weight in subsequent analysis; a higher energy proportion for the seasonal component means that the amplitude of the seasonal component has a higher weight in subsequent analysis. The energy proportion calculation process needs to be performed separately for carbon emission data from each industry, as carbon emissions from different industries may exhibit drastically different time-series characteristics.

[0025] The entire time series decomposition process needs to be automated by a computer program. This program includes modules for data input, filtering, frequency domain transformation, component separation, and feature calculation. The data input module reads the formatted industry carbon emission dataset and verifies its integrity and consistency. The filtering module implements an adaptive filtering algorithm and provides a parameter configuration interface for setting the initial window length and anomaly detection threshold. The frequency domain transformation module calls Fast Fourier Transform (FFT) library functions, taking care to avoid spectral leakage during processing, typically using windowing functions to improve the transformation effect. The component separation module separates trend, seasonal, and residual components based on preset frequency thresholds. The feature calculation module calculates the numerical characteristics and energy proportions of each component and outputs a structured feature description file. The computational accuracy of the entire decomposition process is affected by various factors, including data sampling frequency, data sequence length, and transformation parameter settings, requiring multiple experiments to determine the optimal parameter combination. Visualizing the decomposition results helps to intuitively understand the internal structure of the carbon emission data, typically displaying the correspondence between the original data sequence and each decomposed component in an overlay chart format.

[0026] Example 2 The establishment of the industry carbon emission dynamic weight matrix begins with the analysis of the energy proportion of the trend component. The energy proportion of the trend component is derived from the time series decomposition calculation completed in the above embodiments. The long-term trend weight of industry carbon emissions is determined based on the energy proportion of the trend component. This determination process requires consideration of the historical baseline value of industry carbon emissions, which is either the total carbon emissions of the industry in a set base year or the average carbon emissions of the industry within the data time range. The formula for calculating the long-term trend weight is the product of the energy proportion of the trend component and the historical baseline value of industry carbon emissions. The product reflects the relative importance of the long-term trend of carbon emissions in a specific industry in the overall assessment. High-energy-consuming industries such as electricity and steel typically have higher energy proportions of the trend component, and their long-term trend weights are correspondingly larger. Service industries or light industries have relatively lower energy proportions of the trend component, and their long-term trend weights are also smaller. Each industry needs to independently calculate its long-term trend weight to form a comparable weight series among industries. The periodic fluctuation weight of industry carbon emissions is calculated by combining the energy proportion of the seasonal component. The energy proportion of the seasonal component also comes from the time series decomposition results. The calculation of cyclical fluctuation weights requires the introduction of an industry-specific adjustment factor. This factor is set based on the seasonal characteristics of industry production. Industries significantly affected by seasonality, such as agriculture and construction, have larger adjustment factors, while continuous-process chemical industries have smaller adjustment factors. The specific calculation method for cyclical fluctuation weights is to multiply the energy proportion of the seasonal component by the industry-specific adjustment factor, and then multiply by a standardized coefficient. The standardized coefficient is used to balance the magnitude differences in seasonal impacts between different industries, making the calculation results comparable. The cyclical fluctuation weight value reflects the degree to which an industry's carbon emissions are affected by seasonal factors; industries with larger weight values ​​indicate that their carbon emission changes exhibit a stronger cyclical pattern.

[0027] The long-term trend weights and periodic fluctuation weights need to be normalized. The purpose of normalization is to eliminate dimensional differences and ensure that all weight values ​​fall within a uniform numerical range. The normalization process uses a min-max normalization method, which linearly transforms the original weight values ​​to the closed interval [0,1]. For the normalization of long-term trend weights, the maximum and minimum values ​​of the long-term trend weights across all industries need to be identified first. For each industry's long-term trend weight, a transformation formula is applied: subtract the minimum value and divide by the difference between the maximum and minimum values. The normalization of periodic fluctuation weights uses the same method, performing a linear transformation based on the maximum and minimum values ​​of the periodic fluctuation weights across all industries. The normalized long-term trend weights and periodic fluctuation weights have the same scale and can be directly compared and combined. Generating the industry carbon emission dynamic weight matrix requires organically combining the two types of weights after normalization. The industry carbon emission dynamic weight matrix is ​​a two-dimensional data structure, where the rows represent different industry categories and the columns represent the time dimension. Each matrix element contains two weight components: a weight representing the industry's long-term trend at a specific point in time, and a weight representing cyclical fluctuations. The assignment of values ​​to matrix elements takes into account dynamic changes over time; the long-term trend weight changes slowly over time, while the cyclical fluctuation weight fluctuates according to seasonal cycles. The update frequency of the industry carbon emission dynamic weight matrix is ​​consistent with the original data collection frequency; monthly carbon emission data corresponds to a monthly updated weight matrix, and annual data corresponds to an annual updated weight matrix. The matrix is ​​stored in a structured data format to facilitate subsequent multidimensional correlation and matching operations.

[0028] The validation of the industry carbon emission dynamic weight matrix requires backtesting using historical data. The backtesting process involves correlation analysis between the weight matrix and historical carbon emission data to verify the explanatory power of the weight matrix for actual carbon emission changes. Validation indicators include the correlation coefficient between the weight matrix and carbon emission changes, and the prediction error of the weight matrix. The establishment of the industry carbon emission dynamic weight matrix provides a quantitative basis for subsequent multidimensional correlation matching. The weight values ​​in the matrix directly affect the accuracy of the carbon emission intensity prediction curve. The weight distribution characteristics of different industries also reflect the impact mechanism of industrial structure on regional carbon emissions; high-weight industries are usually the focus of carbon reduction strategies. The establishment of the industry carbon emission dynamic weight matrix realizes the transformation from raw data to quantitative analysis indicators, providing an important intermediate parameter system for carbon peak prediction. The establishment process of the industry carbon emission dynamic weight matrix requires specialized computational software support. The software modules include a weight calculation module, a normalization module, and a matrix generation module. The weight calculation module reads the energy percentage data obtained from time series decomposition and calculates the original weight values ​​in combination with basic industry information. The normalization module standardizes various weights to ensure data scale uniformity. The matrix generation module organizes the processed weights into a matrix structure according to time series and industry classification, and outputs it as a standard data file. The entire construction process requires multiple iterations and optimizations, especially the setting of industry characteristic adjustment factors, which needs to be adjusted based on actual data feedback. The final form of the industry carbon emission dynamic weight matrix is ​​a multidimensional dataset that dynamically changes over time, accurately reflecting the inherent laws and external characteristics of carbon emissions in each industry. Specifically, the construction of the industry carbon emission dynamic weight matrix is ​​based on the numerical characteristics of the trend component and seasonal component obtained from time series decomposition. The trend component represents the long-term change pattern of industry carbon emissions, while the seasonal component reflects periodic fluctuation characteristics. The matrix structure is designed in a multidimensional form, with rows corresponding to different industry categories and columns corresponding to various points in the time series. Each matrix element contains two components: a normalized long-term trend weight and a periodic fluctuation weight. As time progresses, the weight values ​​in the matrix are dynamically adjusted according to the latest data. For example, when new carbon emission data is included, the numerical characteristics of the trend and seasonal components are updated by re-decomposing the time series, thereby correcting the weight allocation and ensuring that the matrix can capture the time series changes of carbon emissions in each industry in real time. This dynamism allows the matrix to not only intrinsically reflect the long-term evolution trend and cyclical patterns of industry carbon emissions, but also, through subsequent multi-dimensional correlation and matching with economic development indicator datasets, incorporate the influence of external factors such as economic growth and industrial structure, thereby accurately presenting the intrinsic and external characteristics of industry carbon emissions. The matrix's update frequency is consistent with the original data collection cycle, such as monthly or annual updates, ensuring that the dataset continuously reflects actual changes.

[0029] The application of dynamic weighted industry carbon emission matrices is not limited to carbon peak prediction; it can also be extended to areas such as carbon emission driver analysis and industry emission reduction potential assessment. The matrix data can be visualized using a heatmap, with the horizontal axis representing time, the vertical axis representing industry category, and color intensity indicating weight, intuitively showing the spatiotemporal evolution of carbon emission impacts across industries. Establishing the dynamic weighted industry carbon emission matrix is ​​a crucial step connecting data collection and prediction models; the matrix quality directly determines the accuracy and reliability of the entire carbon peak prediction method. Each parameter setting during matrix construction must be based on rigorous data analysis and domain knowledge to ensure the scientific and rational allocation of weights.

[0030] Example 3 The construction of the economic development indicator dataset begins with the systematic collection of macroeconomic data for the target region. The collected indicators include time-series data on GDP growth rate, industrial added value, total energy consumption, and population growth rate. This data originates from the National Bureau of Statistics, local statistical yearbooks, and authoritative reports issued by industry regulatory authorities. The time frame of the data collection must completely correspond to the historical carbon emission dataset for the industry. For example, if the carbon emission data is annual data from 2000 to 2023, then the economic development indicators also require continuous data over the same time span. During the collection process, special attention must be paid to the consistency of data sources to avoid data bias caused by different statistical standards. All raw data must indicate the statistical source and unit of measurement. The raw economic indicator time-series data needs to undergo rigorous standardization to eliminate dimensional differences. The unit for GDP growth rate is percentage, the unit for industrial added value is monetary unit, the unit for total energy consumption is standard coal tons, and the unit for population growth rate is per thousand. The standardization process uses the Z-score method, calculating the arithmetic mean and standard deviation for the time series of each economic indicator. Subtracting the mean from each raw data value and dividing by the standard deviation yields a new series that conforms to a standard normal distribution. After standardization, all indicator data become dimensionless pure numerical values ​​with a mean of 0 and a standard deviation of 1, allowing indicators of different properties and dimensions to be compared and calculated within the same system. During the standardization process, it is necessary to check for outliers and review and correct data points exceeding three times the standard deviation.

[0031] The standardized time-series data is grouped according to industry categories to form a structured economic development indicator dataset. The grouping is based on the National Economic Industry Classification Standard, ensuring that each industry's carbon emission data corresponds to a specific economic development indicator. For example, carbon emission data for the industrial sector corresponds to the industrial added value indicator, carbon emission data for the construction sector corresponds to the total output value of the construction industry, and carbon emission data for the transportation sector corresponds to freight and passenger volume indicators. The grouped economic development indicator dataset is a three-dimensional data structure: the first dimension is the time series, the second dimension is the industry classification, and the third dimension is the type of economic indicator. The dataset is stored in matrix form for easy subsequent correlation analysis calculations.

[0032] The multidimensional correlation matching between the economic development indicator dataset and the industry carbon emission dynamic weight matrix was performed using correlation coefficient analysis. The correlation coefficients between each indicator in the economic development indicator dataset and the industry carbon emission dynamic weight matrix were calculated using the Pearson product-moment correlation coefficient formula.

[0033] in: This represents the correlation coefficient between the economic indicator series X and the carbon emission weight series Y. This represents the economic indicator value at the i-th time point. This represents the arithmetic mean of a time series of economic indicators. This represents the carbon emission weight value at time point i. This represents the arithmetic mean of the carbon emission weights time series, where n represents the length of the time series. This represents the summation operator. The calculated correlation coefficient ranges from -1 to 1, with positive values ​​indicating a positive correlation and negative values ​​indicating a negative correlation. The larger the absolute value, the stronger the correlation.

[0034] Indicators with correlation coefficients exceeding a preset threshold are selected as key influencing factors. This threshold is set based on historical data analysis, typically using 0.7 as the strong correlation threshold. Correlation coefficients are calculated and screened for each industry to identify the economic indicators with the most significant impact on its carbon emissions. For example, carbon emissions from energy-intensive industries often show a strong positive correlation with industrial added value and total energy consumption, while carbon emissions from the service sector may be more strongly correlated with population growth rate. The selected key influencing factors form an industry-indicator correlation mapping table, indicating the main driving factors for each industry. The numerical distribution of the industry carbon emission dynamic weight matrix is ​​adjusted based on the key influencing factors using a weighted correction method. For industries strongly correlated with key influencing factors, their weight values ​​in the weight matrix are appropriately increased, with the adjustment magnitude proportional to the correlation coefficient. During the adjustment process, the overall structure of the weight matrix must be maintained to ensure that the sum of the weights for each industry remains unchanged. The adjusted industry carbon emission dynamic weight matrix not only reflects the carbon emission characteristics of the industry itself but also incorporates the influence of economic development factors, making the weight distribution more consistent with the actual economic carbon emission relationship.

[0035] The implementation of the multidimensional correlation matching process requires specialized computational software, including modules for data reading, correlation coefficient calculation, threshold filtering, and weight adjustment. The data reading module loads a standardized economic development indicator dataset and a dynamic weight matrix for industry carbon emissions, verifying the consistency of data format and dimensions. The correlation coefficient calculation module implements the Pearson correlation coefficient algorithm to perform correlation analysis between each economic indicator and the carbon emission weights of each industry. The threshold filtering module automatically identifies key influencing factors based on preset thresholds and generates an industry-indicator correlation mapping table. The weight adjustment module optimizes the weight matrix based on the filtering results, outputting the final version of the dynamic weight matrix for industry carbon emissions. The entire matching process requires multiple iterations for optimization, especially the threshold setting, which needs to be appropriately adjusted according to the actual data characteristics to achieve the best matching effect. The multidimensional correlation matching between the economic development indicator dataset and the dynamic weight matrix for industry carbon emissions establishes a quantitative relationship between economic factors and carbon emissions. This relationship provides important input parameters for subsequent carbon emission intensity prediction. The matching results can reveal the impact mechanism of different economic development models on carbon emissions, providing data support for formulating differentiated carbon reduction schemes. The correlation coefficient matrix generated during the matching process can also be used for sensitivity analysis to identify which changes in economic indicators have the most significant impact on carbon emissions. The quality of multidimensional correlation matching directly affects the accuracy of subsequent prediction results, thus requiring strict data quality control and a robust validation mechanism. The matched dataset needs to be updated regularly to adapt to changes in economic development and industrial structure, with the update cycle consistent with the release cycle of the original data. Each update requires re-standardization, correlation coefficient calculation, and weight adjustment to ensure that the matching results always reflect the latest economic carbon emission relationships. The accumulation of historical matching results can also be used to analyze the evolution trend of economic carbon emission relationships, providing training data for long-term prediction models. The multidimensional correlation matching between the economic development indicator dataset and the industry carbon emission dynamic weight matrix is ​​a crucial bridge connecting the macroeconomic environment and industry carbon emission behavior; its scientific validity and accuracy have a decisive impact on the entire carbon peak prediction method.

[0036] See Figure 3This chart visually presents the Pearson correlation coefficient distribution between four economic development indicators—GDP growth rate, industrial added value, total energy consumption, and population growth rate—and the carbon emission weights of five major sectors: industry, electricity, construction, transportation, and other sectors. The chart shows a strong positive correlation between industrial added value and the industrial sector, reflecting the direct driving effect of industrial production activities on carbon emissions in the industrial sector; a strong positive correlation between total energy consumption and the electricity sector, reflecting the high dependence of electricity production on energy consumption; and a weak negative correlation between GDP growth rate and the transportation sector, possibly because carbon emissions in the transportation sector are influenced by multiple factors such as structural optimization and technological emission reduction, resulting in a weaker linear correlation with GDP growth. This multidimensional correlation analysis provides a quantitative basis for subsequent adjustments to the dynamic weight matrix of industry carbon emissions and for in-depth exploration of the intrinsic relationship between the economy and carbon emissions, helping to more accurately integrate economic factors into the carbon emission prediction system.

[0037] Example 4 See Figure 4 The adjusted industry carbon emission dynamic weight matrix contains information on key influencing factors of economic development, and the numerical distribution of the matrix reflects the quantitative relationship between carbon emissions and economic development indicators in each industry. Based on the adjusted industry carbon emission dynamic weight matrix, the historical variation patterns of industry carbon emission intensity are fitted using a least squares method with polynomial regression analysis. In the regression analysis, industry carbon emission intensity is used as the dependent variable, and the long-term trend weights and periodic fluctuation weights in the industry carbon emission dynamic weight matrix are used as independent variables. A goodness-of-fit index is used to evaluate the accuracy of the regression model. The fitting result yields a mathematical expression describing the relationship between industry carbon emission intensity and the weight matrix, with parameters including the regression coefficients of each weight and a constant term. Fitting historical variation patterns needs to be performed independently for each industry, with each industry having its own fitting equation. The equation form may be a linear relationship or a polynomial relationship of quadratic or higher degree. A grey prediction algorithm is introduced to iteratively calculate the future variation trend of industry carbon emission intensity. The grey prediction algorithm selected is the GM(1,1) model, i.e., a first-order univariate grey model. The basic idea of ​​the GM(1,1) model is to accumulate the original data sequence to generate a new data sequence with an exponential growth law, and then establish a first-order linear differential equation for prediction. The model construction process includes the verification and processing of the original data sequence, the accumulation and generation operation, the generation of the nearest neighbor mean, the solution of model parameters, and the establishment of the time response equation. The iterative calculation of the grey prediction algorithm is implemented by a computer program. The program sets convergence conditions to control the number of iterations, and stops the calculation when the relative error of two consecutive iterations is less than a set threshold. The grey prediction algorithm requires historical data of industry carbon emission intensity as a training set. The length of the training set affects the prediction accuracy, and generally requires at least five years of continuous data.

[0038] The iterative calculation results are converted into time series curves to generate industry carbon emission intensity prediction curves. The conversion process involves connecting the discrete predicted values ​​output by the grey prediction algorithm in chronological order, and using cubic spline interpolation to smooth and connect the curves. The industry carbon emission intensity prediction curves are plotted with time on the horizontal axis and carbon emission intensity on the vertical axis, with key feature points such as extreme points and inflection points marked on the curve. The prediction curves include an uncertainty range, which is obtained through Monte Carlo simulation and displayed as a shaded area around the curve. An independent carbon emission intensity prediction curve is generated for each industry, and the curve data is stored in a structured format, including fields such as timestamp, predicted value, and upper and lower bounds of uncertainty. The carbon emission intensity prediction curves of all industries are integrated to construct a regional carbon peak prediction model. The integration process uses a weighted superposition method, with weighting factors determined based on the historical contribution ratio of each industry to the total regional carbon emissions. The calculation of weighting factors requires historical carbon emission data for each industry, and a moving average method is used to eliminate the impact of annual fluctuations. The superposition calculation is performed point-by-point, and the predicted total regional carbon emissions at each time point are equal to the sum of the predicted carbon emission intensity values ​​of each industry at that time point multiplied by their respective industry weights. The core output of the regional carbon peak prediction model is a prediction curve of the total regional carbon emissions over time, covering a certain period in the future, such as from 2025 to 2060.

[0039] Referring to Table 1, identify the peak points in the regional carbon emission prediction curve and determine the time nodes for carbon peaking. Peak point identification uses a numerical differentiation method, calculating the first and second derivatives of the prediction curve. The point where the first derivative is zero and the second derivative is negative is the peak point. When the prediction curve has multiple extreme points, the global maximum point is selected as the carbon peak point. The carbon peaking time node is accurate to the year, corresponding to the x-coordinate value of the peak point. The final output of the regional carbon peaking prediction model includes the peak carbon emission value, the peaking time point, and an analysis of the carbon emission change trend before and after peaking.

[0040] Table 1: Weighting Allocation Table for Industry Carbon Emission Intensity Forecast Curves Industry Classification Baseline year emissions percentage (%) Weighting factors Superposition method Industrial sector 45.6 0.456 Linear weighting power sector 28.3 0.283 Linear weighting Construction Department 12.1 0.121 Linear weighting Transportation Department 9.8 0.098 Linear weighting Other departments 4.2 0.042 Linear weighting The regional carbon peak prediction model is validated through backtesting using historical data. Backtesting compares the model's predictions with actual historical data, calculating accuracy indicators such as mean absolute percentage error and root mean square error. Model sensitivity analysis examines the impact of changes in key parameters on the prediction results, particularly the sensitivity of weighting factors and grey prediction model parameters. The regional carbon peak prediction model requires regular updates, with the update cycle consistent with the update cycle of the base data, typically annually. Model outputs include data tables and visualizations. The visualizations show the overlay effect of the predicted carbon emission intensity curves for each industry and the predicted total regional carbon emissions curve, with different industries distinguished by different colors. Determining the carbon peak time point requires considering the range of uncertainty; the output should include an estimate of the time interval for peak time, not just a single point estimate. The construction of the regional carbon peak prediction model completes the scale transformation from industry-level analysis to regional-level prediction, and the model can reflect the impact of industrial structure changes on the regional carbon emission peak time. In model application, by adjusting economic development scenario parameters, the carbon peak time points under different development paths can be simulated, providing multi-scenario analysis support for carbon reduction scheme formulation. The regional carbon peak prediction model's architecture is designed to support modular expansion, allowing for the easy addition of new industry categories or new influencing factors, thereby enhancing the model's adaptability and practicality.

[0041] Example 5 The sources of actual carbon emission data include environmental monitoring station reports, corporate carbon emission accounting reports, and statistical bulletins issued by government departments. The data collection frequency matches the time resolution of the prediction model; annual prediction models correspond to annual data collection, and quarterly prediction models correspond to quarterly data collection. Monitoring data needs to cover all industry categories included in the regional carbon peak prediction model to ensure consistency between the data and the prediction model at the industry level. The data acquisition system establishes an automated data transmission channel, directly obtaining the latest carbon emission information from the data source system, reducing errors that may be introduced by human intervention. The collected actual carbon emission data undergoes format standardization processing, converting it into the same unit of measurement and data structure as the predicted values, preparing for subsequent comparative analysis. The actual carbon emission data is compared and analyzed with the predicted values ​​of the regional carbon peak prediction model. The comparison process is conducted within a unified time base and spatial range. The comparison calculation adopts a point-in-time comparison method, calculating the absolute and relative deviations between the actual and predicted values ​​for each monitoring time point. The absolute deviation is the arithmetic difference between the actual and predicted values, and the relative deviation is the percentage value obtained by dividing the absolute deviation by the predicted value and multiplying by 100%. The deviation calculation results are stored in a specially designed deviation record table, which includes fields such as timestamp, industry classification, actual value, predicted value, absolute deviation, and relative deviation. The system sets deviation warning thresholds, determined based on historical prediction accuracy and industry characteristics; thresholds are set relatively higher for highly volatile industries and relatively lower for stable industries. Deviation monitoring is a continuous process; each update of new actual carbon emission data automatically triggers a new round of comparative analysis. When the deviation between actual carbon emission data and predicted values ​​exceeds a preset range, the system initiates a model adjustment procedure. The preset range is expressed as upper and lower limits of relative deviation. Determining whether a deviation exceeds the limit requires examining the deviation trend over multiple consecutive time points to avoid unnecessary model adjustments due to single-point data anomalies. The deviation exceeding the limit judgment logic has two levels: the first level checks whether the deviation at a single time point exceeds the threshold, and the second level checks whether the average deviation over three consecutive time points continues to exceed the limit. Only when both levels of checks meet the conditions is model adjustment confirmed. This dual verification mechanism effectively avoids over-adjustment due to temporary fluctuations or data anomalies.

[0042] The core step in revising the forecasting model is to readjust the dynamic weight matrix of industry carbon emissions. The adjustment is based on the deviation pattern and magnitude between actual carbon emission data and predicted values. Revising the numerical characteristics of the trend and seasonal components requires analyzing the temporal distribution of the deviation. If the deviation exhibits a continuous unidirectional trend, the numerical characteristics of the trend component are primarily revised. If the deviation exhibits a periodic fluctuation pattern, the numerical characteristics of the seasonal component are primarily revised. The revision of the trend component's numerical characteristics is achieved by adjusting the slope and intercept of the trend line, with the adjustment magnitude proportional to the magnitude of the deviation. The revision of the seasonal component's numerical characteristics involves adjusting the amplitude and phase. Amplitude revision changes the intensity of seasonal fluctuations, while phase revision changes the timing of seasonal fluctuations. The numerical characteristic revision uses a recursive least squares method, giving greater weight to new data and gradually decreasing the weight of older data, making the revision results more reflective of the latest changes. Based on the revised numerical characteristics, the dynamic weight matrix of industry carbon emissions is updated. The update process recalculates the long-term trend weights and periodic fluctuation weights. The update of the long-term trend weights considers the revised trend component slope and the latest value of the industry carbon emission baseline, while the update of the periodic fluctuation weights combines the revised seasonal component amplitude and the latest industry characteristic adjustment factors. After the weight matrix is ​​updated, it needs to be re-normalized to ensure that the sum of all weight values ​​remains constant. The updated industry carbon emission dynamic weight matrix replaces the corresponding data in the original matrix, and the matrix version number is updated accordingly, recording the adjustment time and reason. The updated industry carbon emission dynamic weight matrix is ​​input into the regional carbon peak prediction model to recalculate the carbon peak time node. The recalculation process repeats the steps of generating the industry carbon emission intensity prediction curve and predicting the regional total carbon emissions in the above embodiments, but uses the updated weight matrix data. The recalculated carbon peak time node is compared with the original predicted value to analyze the magnitude and direction of the changes before and after the adjustment. The recalculation of the carbon peak time node needs to consider the propagation of uncertainty; the uncertainty brought about by the weight matrix adjustment will be transmitted to the final prediction result. The recalculation process generates a new prediction report, which clearly indicates that this adjustment is based on the latest monitoring data and details the adjustment content and its impact on the prediction result.

[0043] A real-time monitoring and dynamic adjustment mechanism establishes a complete feedback loop. After each model adjustment, subsequent actual carbon emission data is continuously monitored to verify the adjustment effect. If the deviation decreases significantly after adjustment, it indicates that the adjustment direction is correct. If the deviation persists, further analysis of the cause of the deviation is needed to consider whether there are new influencing factors not considered by the model. The monitoring-comparison-adjustment process forms a closed-loop control system, enabling the regional carbon peak prediction model to have continuous learning capabilities and adapt to changes in economic development and energy structure. The system records detailed logs for each adjustment, including triggering conditions, adjustment parameters, and adjustment results, providing data support for model performance evaluation and subsequent optimization. The implementation of the dynamic adjustment mechanism requires dedicated software system support, which includes a data acquisition module, a deviation analysis module, a weight adjustment module, and a model recalculation module. The data acquisition module is responsible for automatically acquiring the latest carbon emission data from multiple data sources, performing format conversion and quality checks. The deviation analysis module calculates the deviation between the actual and predicted values ​​in real time and executes the deviation exceeding limit judgment logic. The weight adjustment module automatically corrects the numerical characteristics of the trend component and seasonal component based on the deviation analysis results and updates the industry carbon emission dynamic weight matrix. The model recalculation module calls the core algorithm of the prediction model and regenerates the prediction results using the updated weight matrix. The entire system adopts a modular design, with modules exchanging data through standard interfaces, facilitating system maintenance and functional expansion. The implementation of real-time monitoring and dynamic adjustment mechanisms transforms regional carbon peak prediction from static to dynamic, enabling the prediction model to adjust its trajectory promptly based on actual carbon emission changes. This dynamic adaptability is particularly important for medium- and long-term carbon peak prediction because economic development paths are subject to significant uncertainty. By continuously incorporating the latest monitoring data, the prediction model can gradually correct biases in its initial assumptions, improving the reliability and practicality of the prediction results. The dynamic adjustment mechanism also provides a verification and calibration pathway for the core prediction model; repeated calibration based on actual data can optimize model parameter settings and enhance the model's applicability under complex real-world conditions.

[0044] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0045] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A carbon peak prediction method based on data analysis, characterized in that, include: Collect historical carbon emission data within the target area, and divide the historical carbon emission data according to industry categories to form an industry carbon emission dataset; The industry carbon emission dataset is decomposed into time series to extract the numerical features of trend components, seasonal components, and residual components. Based on the numerical characteristics of the trend component and the seasonal component, a dynamic weight matrix for industry carbon emissions is established. Obtain the economic development indicator dataset of the target region, and perform multi-dimensional correlation matching between the economic development indicator dataset and the industry carbon emission dynamic weight matrix. Based on the multidimensional correlation matching results, an industry carbon emission intensity prediction curve is generated. By integrating carbon emission intensity prediction curves from all industries, a regional carbon peak prediction model is constructed.

2. The carbon peak prediction method based on data analysis according to claim 1, characterized in that, The time-series decomposition of the industry carbon emission dataset includes: An adaptive filtering algorithm is used to smooth the industry carbon emission dataset to eliminate abnormal fluctuations in data; A frequency domain transformation was performed on the smoothed industry carbon emission dataset to separate the trend component, seasonal component, and residual component. The energy proportions of the trend component and the seasonal component are calculated and used as the basis for the weighting of the numerical features.

3. The carbon peak prediction method based on data analysis according to claim 2, characterized in that, The establishment of the industry carbon emission dynamic weight matrix includes: The weight of the long-term trend of carbon emissions in the industry is determined based on the energy proportion of the trend components. Based on the energy proportion of the seasonal components, the periodic fluctuation weight of industry carbon emissions is calculated; The long-term trend weights and periodic fluctuation weights are normalized to generate a dynamic weight matrix for industry carbon emissions.

4. The carbon peak prediction method based on data analysis according to claim 1, characterized in that, The acquisition of the economic development indicator dataset for the target region includes: Collect time-series data on GDP growth rate, industrial added value, total energy consumption, and population growth rate of the target region; The time series data is standardized to eliminate dimensional differences; The standardized time-series data are grouped according to industry categories to form an economic development indicator dataset.

5. The carbon peak prediction method based on data analysis according to claim 4, characterized in that, The step of performing multidimensional correlation matching between the economic development indicator dataset and the industry carbon emission dynamic weight matrix includes: Calculate the correlation coefficients between each indicator in the economic development indicator dataset and the dynamic weight matrix of industry carbon emissions; Indicators with correlation coefficients exceeding a preset threshold are selected as key influencing factors; Based on the aforementioned key influencing factors, adjust the numerical distribution of the industry's dynamic carbon emission weight matrix.

6. The carbon peak prediction method based on data analysis according to claim 1, characterized in that, The generated industry carbon emission intensity prediction curve includes: Based on the adjusted industry carbon emission dynamic weight matrix, the historical variation pattern of industry carbon emission intensity is fitted. A grey prediction algorithm is introduced to iteratively calculate the future trend of carbon emission intensity in the industry; The iterative calculation results are converted into time series curves to generate industry carbon emission intensity prediction curves.

7. The carbon peak prediction method based on data analysis according to claim 6, characterized in that, The construction of the regional carbon peak prediction model includes: By overlaying the carbon emission intensity prediction curves of all industries, the predicted value of the total regional carbon emissions is calculated. Identify the peak points in the predicted values ​​to determine the carbon peak time point; Based on the carbon peak time points, a regional carbon peak prediction model is generated.

8. The carbon peak prediction method based on data analysis according to claim 7, characterized in that, The method further includes: Real-time monitoring of actual carbon emission data in the target area, and comparison with the predicted values ​​of the carbon peak prediction model for the area; When the deviation between actual carbon emission data and predicted values ​​exceeds the preset range, the dynamic weight matrix of industry carbon emissions is readjusted.

9. The carbon peak prediction method based on data analysis according to claim 8, characterized in that, The readjustment of the industry carbon emission dynamic weight matrix includes: Based on the deviation between actual carbon emission data and predicted values, the numerical characteristics of the trend component and the seasonal component are corrected. Update the industry carbon emission dynamic weight matrix based on the corrected numerical characteristics; The updated industry carbon emission dynamic weight matrix is ​​input into the regional carbon peak prediction model to recalculate the carbon peak time node.

10. A carbon peak prediction system based on data analysis, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the carbon peak prediction method based on data analysis as described in any one of claims 1 to 9.