Carbon emission reduction evolution path generation method and device, computer equipment and storage medium
By quantitatively assessing the degree of missing data in multiple dimensions of carbon emission data and dynamically matching the optimal data supplementation strategy, the problem of insufficient data in carbon emission reduction path planning is solved, and high-precision and highly adaptable path generation is achieved.
Patent Information
- Application Number
- CN202510785191.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies for carbon emission reduction pathway planning suffer from insufficient data quality, leading to inaccurate predictions by mechanistic models. This affects the accuracy and operability of pathway planning. Furthermore, the lack of a dynamic adaptation mechanism results in algorithm overfitting and inaccurate predictions.
By quantitatively assessing the degree of missing data in multiple dimensions of carbon emission data, the optimal data supplementation strategy is dynamically matched, including cross-regional data migration and time-series prediction, to generate a highly adaptable carbon emission reduction evolution path.
It significantly improves the reliability of data and the accuracy of path generation under conditions of data scarcity, ensuring that the generated paths are highly accurate and adaptable to regions with different data endowments.
Smart Images

Figure CN120875898A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of carbon emission technology, specifically to a method, apparatus, computer equipment, and storage medium for generating carbon emission reduction evolution paths. Background Technology
[0002] Current regional carbon emission reduction pathway planning mainly relies on traditional mechanistic models, which are highly dependent on complete and high-quality input data. However, in practice, there are serious data shortcomings, such as fragmented energy data, missing or lagging key economic data, and insufficient spatial precision.
[0003] In related technologies, when dealing with data missing issues, there is often a lack of multi-dimensional quantitative assessment of the "degree of lack" of data in the target area. Instead, a rigid approach is adopted, using simple interpolation, extrapolation, or proportional allocation as single supplementation strategies, failing to intelligently match the optimal method based on specific missing characteristics. This results in poor data quality input to the mechanistic model, ultimately severely limiting the model's prediction accuracy and the accuracy and operability of the planned emission reduction pathways. Summary of the Invention
[0004] In view of this, the present invention provides a method, apparatus, computer equipment and storage medium for generating carbon emission reduction evolution paths, in order to solve the problem that inaccurate predictions by mechanistic models due to insufficient carbon emission data quality, which in turn affects the effectiveness of carbon emission reduction path planning.
[0005] In a first aspect, the present invention provides a method for generating a carbon emission reduction evolution path, comprising: acquiring raw carbon emission data of a target region; determining a target data deficiency coefficient of the raw carbon emission data in the target region based on the multi-dimensional characteristics of the raw carbon emission data; determining a target data supplementation strategy corresponding to the target data deficiency coefficient based on the correspondence between the data deficiency coefficient and the data supplementation strategy, and supplementing the raw carbon emission data according to the target data supplementation strategy to obtain target carbon emission data; and generating a carbon emission reduction evolution path corresponding to the target region according to the target carbon emission data.
[0006] The carbon emission reduction evolution path generation method provided in this invention significantly improves data reliability under data scarcity conditions by quantitatively assessing the multi-dimensional missingness of carbon emission data in the target area and dynamically matching the optimal data supplementation strategy, thereby generating a more accurate and adaptable carbon emission reduction evolution path.
[0007] In one optional implementation, the multi-dimensional features include data integrity, data continuity, and data resolution. Based on the multi-dimensional features of the original carbon emission data, the target data deficiency coefficient of the original carbon emission data in the target area is determined, including: obtaining the data integrity index, data continuity index, and data resolution index corresponding to the original carbon emission data; dynamically allocating the weight coefficients of the data integrity index, data continuity index, and data resolution index according to the carbon emission entity type of the target area, obtaining the first weight coefficient corresponding to the data integrity index, the second weight coefficient corresponding to the data continuity index, and the third weight coefficient corresponding to the data resolution index; and using the first weight coefficient, the second weight coefficient, and the third weight coefficient, performing a weighted summation of the data integrity index, the data continuity index, and the data resolution index to obtain the target data deficiency coefficient of the target area.
[0008] The carbon emission reduction evolution path generation method provided in this invention dynamically allocates weight coefficients for data integrity, continuity, and resolution indicators based on the type of carbon emission entities in the target area. This makes the evaluation results of data lacking coefficients more consistent with the actual entity characteristics, significantly improving the objectivity and pertinence of data quality quantitative analysis.
[0009] In one optional implementation, based on the correspondence between the data deficiency coefficient and the data supplementation strategy, a target data deficiency coefficient corresponding to the target data deficiency strategy is determined, and the original carbon emission data is supplemented according to the target data supplementation strategy to obtain the target carbon emission data. This includes: when the target data deficiency coefficient exceeds a first preset threshold, determining a first data supplementation strategy corresponding to the target data deficiency coefficient; acquiring carbon emission source data from multiple other regions, and determining target migration data that meets preset conditions from the multiple carbon emission source data; and using the first data supplementation strategy to supplement the original carbon emission data with the target migration data to obtain the target carbon emission data.
[0010] The carbon emission reduction evolution path generation method provided in this invention automatically triggers a cross-regional data migration strategy by setting a preset threshold, filters target migration data that meets preset conditions, and uses the target migration data to supplement the original data that is severely missing in the target area, significantly improving the feasibility and cross-domain adaptability of data generation in scenarios with severe data scarcity.
[0011] In one optional implementation, target migration data that meets preset conditions is determined from multiple carbon emission source data, including: obtaining a first feature vector of the target region and second feature vectors of each other region; determining the feature similarity between the target region and each other region based on the first and second feature vectors; extracting a first energy intensity of the target region from the original carbon emission data, extracting the second energy intensity of each other region from the carbon emission source data, and determining the energy intensity difference between the first energy intensity and each second energy intensity; and filtering the carbon emission source data using feature similarity and energy intensity difference to obtain the target migration data.
[0012] The carbon emission reduction evolution path generation method provided in this invention uses a dual screening mechanism of feature similarity and energy intensity difference to accurately match external data sources that are highly compatible with the carbon emission entity structure and energy consumption pattern of the target area. This ensures the consistency of the supplementary data in terms of carbon emission entity correlation and energy-driven logic, fundamentally avoiding pseudo-similar interference in cross-domain knowledge transfer.
[0013] In one optional implementation, carbon emission source data is screened using feature similarity and energy intensity difference to obtain target migration data, including: selecting regions from multiple other regions where feature similarity exceeds a similarity threshold and energy intensity difference does not exceed an intensity threshold as carbon emission data migration sources; and determining the carbon emission source data corresponding to the carbon emission data migration sources as target migration data.
[0014] The carbon emission reduction evolution path generation method provided in this invention strictly locks high-quality migration source regions that match both carbon emission entity characteristics and energy efficiency by setting a dual hard condition of a similarity threshold and an energy intensity difference threshold, thereby ensuring the core quality and scenario adaptability of cross-domain data supplementation from a mechanism perspective.
[0015] In one optional implementation, when the target data deficiency coefficient exceeds a second preset threshold, a second data supplementation strategy corresponding to the target data deficiency coefficient is determined; the original carbon emission data is preprocessed to obtain preprocessed time series data; the preprocessed time series data is truncated using a sliding time window to obtain truncated data of a preset duration; the predicted carbon emission data for the next target duration is predicted according to the truncated data using the second data supplementation strategy, and the original carbon emission data is supplemented using the predicted carbon emission data to obtain the target carbon emission data.
[0016] The carbon emission reduction evolution path generation method provided in this embodiment of the invention performs time-series slicing on the preprocessed raw carbon emission data through a preset sliding time window, and uses the extracted data to predict key carbon emission data for future periods, so as to dynamically supplement the gaps in the original data and achieve efficient self-sufficient data generation in mild data-scarce scenarios.
[0017] In one optional implementation, generating a carbon reduction evolution path corresponding to the target region based on the target carbon emission data includes: extracting scenario parameter data corresponding to each preset carbon emission scenario from the target carbon emission data; and using the scenario parameter data to simulate the carbon reduction evolution path under different preset carbon emission scenarios corresponding to the target region.
[0018] The carbon emission reduction evolution path generation method provided in this invention accurately extracts a subset of data strongly correlated with a specific preset scenario from the supplemented target carbon emission data. This ensures that the generated carbon emission reduction evolution path is entirely based on data-driven logic customized for that scenario, preserving global data consistency while achieving precise adaptation of scenario-based path generation. Simultaneously, the supplemented target carbon emission data is combined with emission reduction path deductions under different preset carbon emission scenarios, achieving full automation from data input to path generation. This method can adapt to the carbon emission reduction path deduction needs of regions with different data endowments, providing high-precision and highly adaptable carbon peaking path decision support for differentiated regional scenarios.
[0019] Secondly, the present invention provides a carbon emission reduction evolution path generation device, comprising: an acquisition module for acquiring raw carbon emission data of a target area; a determination module for determining a target data deficiency coefficient of the raw carbon emission data in the target area based on the multi-dimensional characteristics of the raw carbon emission data; a supplementation module for determining a target data supplementation strategy corresponding to the target data deficiency coefficient based on the correspondence between the data deficiency coefficient and the data supplementation strategy, and supplementing the raw carbon emission data according to the target data supplementation strategy to obtain target carbon emission data; and a generation module for generating a carbon emission reduction evolution path corresponding to the target area according to the target carbon emission data.
[0020] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the carbon emission reduction evolution path generation method of the first aspect or any corresponding embodiment described above.
[0021] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the carbon emission reduction evolution path generation method of the first aspect or any corresponding embodiment thereof.
[0022] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the carbon emission reduction evolution path generation method of the first aspect or any corresponding embodiment thereof. Attached Figure Description
[0023] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0024] Figure 1 This is a schematic flowchart of a carbon emission reduction evolution path generation method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating another method for generating carbon emission reduction evolution paths according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the process of generating a carbon emission reduction evolution path using the LEAP model according to an embodiment of the present invention; Figure 4 This is a structural block diagram of a carbon emission reduction evolution path generation device according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] Numerous and significantly diverse grassroots regions urgently require accurate carbon emission modeling and differentiated emission reduction pathway planning. However, current mainstream traditional mechanistic models heavily rely on detailed data input, and their prediction accuracy is severely limited by the widespread lack of energy metering and economic data at the grassroots level.
[0027] While hybrid approaches combining machine learning and mechanistic models ("data + mechanism") have been used to improve modeling efficiency in data-scarce scenarios, these solutions generally lack dynamic adaptation mechanisms. They fail to adjust algorithms based on the type and extent of missing data, making it difficult to select the optimal algorithm. This "one-size-fits-all" approach can lead to algorithm overfitting, causing inaccurate mechanisms and ultimately model distortion, severely impacting prediction accuracy and the accuracy and feasibility of emission reduction pathways.
[0028] In view of this, the technical solution of this invention quantifies the degree of data loss in a region by constructing a multi-dimensional dynamic evaluation mechanism (data integrity, continuity, resolution) to generate a data lack coefficient, and adaptively triggers a differentiated supplementation strategy based on the threshold of this coefficient: for regions with severe data lack, reliable cross-regional data migration is achieved through dual screening of feature similarity and energy intensity, avoiding mechanism distortion caused by forced fitting of machine learning; for regions with mild data lack, sliding window time series prediction is used to autonomously supplement local data, reducing dependence on external factors; finally, multi-scenario emission reduction paths are generated based on the supplemented full-dimensional parameter data, overcoming the problem of overfitting and prediction inaccuracy caused by the lack of dynamic algorithm adaptation in traditional hybrid models under the condition of fragmented grassroots data, and significantly improving the accuracy and operability of the paths.
[0029] According to an embodiment of the present invention, a method for generating carbon emission reduction evolution paths is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0030] This embodiment provides a method for generating carbon emission reduction evolution paths, which can be achieved using computer devices such as desktop computers and laptops. Figure 1 This is a flowchart of a carbon emission reduction evolution path generation method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Obtain raw carbon emission data for the target area.
[0031] The target region refers to the geographical unit where carbon emission reduction pathways need to be developed, such as a county, city, or district. Raw carbon emission data refers to the original, unprocessed carbon emission-related data for the target region. Specifically, this involves collecting relevant carbon emission data for the target region, including socio-economic statistics (such as GDP, population, and industrial structure), monitoring data (such as energy consumption and traffic flow of key enterprises), and publicly available carbon accounting data. This data can be obtained through government-released statistical reports, energy consumption data provided by enterprises, and data from environmental monitoring agencies, thus forming the raw carbon emission data for the target region.
[0032] Step S102: Based on the multi-dimensional characteristics of the original carbon emission data, determine the target data lack coefficient of the original carbon emission data in the target region.
[0033] Multidimensional features refer to multiple core dimensions used to assess data quality. The target data lack coefficient is a comprehensive index that quantifies the degree of data missing in a target region. Specifically, raw carbon emission data is analyzed to determine its multidimensional features, such as completeness, continuity, and resolution. The degree of data missing in the raw carbon emission data is then determined using these multidimensional features and quantified to derive the target data lack coefficient.
[0034] Step S103: Based on the correspondence between the data deficiency coefficient and the data supplementation strategy, determine the target data supplementation strategy corresponding to the target data deficiency coefficient, and supplement the original carbon emission data according to the target data supplementation strategy to obtain the target carbon emission data.
[0035] The data lack coefficient is a comprehensive index that quantifies the degree of data missing in a region. A data imputation strategy refers to a system of data repair methods selected based on the data lack coefficient. A target data imputation strategy is a specific strategy matched to the target data lack coefficient for the current target region. Target carbon emission data refers to complete and reliable input data processed by the data imputation strategy. Specifically, the relationship between the data lack coefficient and the data imputation strategy is defined by preset rules or algorithms. Based on the severity and type of missing data, an appropriate imputation strategy is selected, such as using Long Short-Term Memory (LSTM) networks or transfer learning. Based on the target data lack coefficient for the target region, an appropriate strategy is matched, and the original carbon emission data is imputed to obtain the final target carbon emission data.
[0036] Step S104: Generate the carbon emission reduction evolution path corresponding to the target region based on the target carbon emission data.
[0037] A carbon emission reduction evolution path refers to the trajectory of carbon emission changes, reflecting the long-term trend and evolution of emission reduction. Specifically, based on target carbon emission data, the carbon emission reduction evolution path can be generated by establishing mathematical models or simulation algorithms. These models consider factors such as the long-term trend of carbon emissions within a region, policy impacts, and technological advancements, generating a predicted emission reduction trajectory. By using target carbon emission data as input, combined with existing emission reduction policies and targets, the carbon emission reduction evolution path for the target region can be derived, reflecting the emission reduction effects and trends over a future period.
[0038] The carbon emission reduction evolution path generation method provided in this invention significantly improves data reliability under data scarcity conditions by quantitatively assessing the multi-dimensional missingness of carbon emission data in the target area and dynamically matching the optimal data supplementation strategy, thereby generating a more accurate and adaptable carbon emission reduction evolution path.
[0039] This embodiment provides a method for generating carbon emission reduction evolution paths, which can be achieved using computer devices such as desktop computers and laptops. Figure 2 This is a flowchart of a carbon emission reduction evolution path generation method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Obtain raw carbon emission data for the target area. For details, please refer to [link to relevant documentation]. Figure 1 Step S101 of the illustrated embodiment will not be described again here.
[0040] Step S202: Based on the multi-dimensional characteristics of the original carbon emission data, determine the target data lack coefficient of the original carbon emission data in the target region.
[0041] Specifically, the multi-dimensional features include data integrity, data continuity, and data resolution, and step S202 above includes: Step S2021: Obtain the data integrity indicators, data continuity indicators, and data resolution indicators corresponding to the original carbon emission data.
[0042] The data integrity index (C) refers to the proportion of missing core fields required for modeling (such as GDP, energy consumption, population, and other key carbon emission data). Its quantification formula is: .
[0043] The data continuity index (T) is used to assess the continuity of data over time (whether there are time gaps). The higher the percentage of missing time steps, the worse the data continuity. Its quantification formula is: .
[0044] The data resolution metric (R) is used to comprehensively evaluate the granularity of data in terms of time, space, and dimensions. The lower the data resolution, the higher the lack of granularity. Specifically, values are assigned from three perspectives: time, space, and dimensions, and then the average is calculated. Year = 0.6, Quarter = 0.5, Month = 0.4, Day = 0.3 (the shorter the interval, the higher the resolution). County-wide = 0.6, Townships = 0.4, Enterprises = 0.2 (the finer the spatial detail, the higher the resolution); Total amount = 0.6, by sector = 0.4, by energy type = 0.2 (the finer the dimension, the higher the resolution).
[0045] Data integrity, data continuity, and data resolution indicators are obtained through analysis and calculation of raw carbon emission data.
[0046] Step S2022: Based on the carbon emission entity type of the target area, dynamically allocate the weight coefficients of the data integrity index, data continuity index, and data resolution index to obtain the first weight coefficient corresponding to the data integrity index, the second weight coefficient corresponding to the data continuity index, and the third weight coefficient corresponding to the data resolution index.
[0047] Carbon emission entity type refers to the dominant industrial economic structure type of the target region, which may include heavy industry-led, service industry-led, and mixed types. First weighting coefficient ( The second weighting factor (C) refers to the weight of data integrity. The third weighting coefficient refers to the weight of data continuity (T). ) refers to the weight of the data resolution (R). Specifically, the industrial structure (industry / service sector ratio) of the target region is analyzed to determine the dominant type: heavy industry-led (industry > 50%), service sector-led (service sector > 50%), and mixed (balanced industry). The weighting coefficients of these three indicators can be dynamically allocated based on the carbon emission entity type of the target region. Different carbon emission entity types have different data requirements. For example, for counties with heavy industry, energy consumption is relatively concentrated and emission sources are few, but data is easily compromised due to a lack of corporate privacy (e.g., key enterprises' energy consumption data is not publicly available), requiring greater reliance on data integrity. For counties with service industries, emissions are relatively dispersed (e.g., transportation, commercial buildings), but high-frequency data (e.g., electricity consumption, traffic flow) is easier to obtain, requiring a focus on data resolution. Dynamic weighting enables adaptive matching of "data assessment - industry characteristics," improving model accuracy. The rules for dynamic weight allocation are shown in Table 1. Table 1
[0048] Step S2023: Using the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient, the data integrity index, the data continuity index, and the data resolution index are weighted and summed to obtain the target data lack coefficient for the target area.
[0049] Using a weighted summation method, based on the data integrity index, data continuity index, and data resolution index calculated above, and combined with the corresponding weighting coefficients (first weighting coefficient, second weighting coefficient, and third weighting coefficient), the target data lack coefficient (DLI) for the target region can be calculated. The specific formula is as follows: .
[0050] The carbon emission reduction evolution path generation method provided in this invention dynamically allocates weight coefficients for data integrity, continuity, and resolution indicators based on the type of carbon emission entities in the target area. This makes the evaluation results of data lacking coefficients more consistent with the actual entity characteristics, significantly improving the objectivity and pertinence of data quality quantitative analysis.
[0051] Step S203: Based on the correspondence between the data deficiency coefficient and the data supplementation strategy, determine the target data supplementation strategy corresponding to the target data deficiency coefficient, and supplement the original carbon emission data according to the target data supplementation strategy to obtain the target carbon emission data.
[0052] Specifically, step S203 includes: Step S2031: When the target data deficiency coefficient exceeds the first preset threshold, determine the first data supplementation strategy corresponding to the target data deficiency coefficient.
[0053] The first preset threshold refers to a pre-defined critical value for the data lack coefficient, used to characterize scenarios with severe data lack; for example, it can be 0.4. The first data imputation strategy refers to a transfer learning data imputation method designed for target data lack coefficients exceeding the first preset threshold. Specifically, when the data lack coefficient of the target region exceeds the first preset threshold, the first data imputation strategy is automatically triggered. The first data imputation strategy is defined as a cross-domain data imputation method based on transfer learning.
[0054] Step S2032: Obtain carbon emission source data from multiple other regions, and determine the target migration data that meets the preset conditions from the multiple carbon emission source data.
[0055] Carbon emission source data refers to raw carbon emission datasets collected from other counties. Target migration data refers to qualified migration source data selected through preset criteria. Specifically, carbon emission datasets from other regions are collected; these datasets may include core data related to carbon emissions, such as GDP, energy consumption, and industrial structure. Then, based on preset screening criteria, these collected data are filtered to select target migration data that meets the criteria.
[0056] In some optional implementations, target migration data that meets preset conditions is determined from multiple carbon emission source data, including: Step a1: Obtain the first feature vector of the target region and the second feature vectors of each of the other regions.
[0057] The first eigenvector represents the distribution of the industrial structure in the target region. The second eigenvector represents the distribution of the industrial structure in other regions. Specifically, industrial structure data (such as the proportion of the three major industries or the output value ratio of sub-sectors) are collected for the target region and other regions to construct the first eigenvector for the target region. and the second feature vectors of each other region. .
[0058] Step a2: Based on the first feature vector and the second feature vector, determine the feature similarity between the target region and each other region.
[0059] Feature similarity refers to the similarity of the industrial structure between the target region and other regions. Specifically, by introducing industry carbon emission weights, an improved weighted cosine similarity is used to measure the industrial similarity between the target region and other regions. The specific formula is as follows: .
[0060] Among them, weight Based on industry-specific carbon emission intensity, high-carbon-emission industries are given higher weight to strengthen sector relevance (e.g., the steel industry). The figure is 1.5 for the service industry. (0.5).
[0061] Step a3: Extract the first energy intensity of the target area from the original carbon emission data, extract the second energy intensity of each other area from the carbon emission source data, and determine the energy intensity difference between the first energy intensity and each second energy intensity.
[0062] First energy intensity refers to energy consumption per unit of GDP in the target region. Second energy intensity refers to energy consumption per unit of GDP in other regions. Energy intensity difference refers to the absolute difference in energy intensity between the target region and other regions. Specifically, the formula for calculating energy consumption per unit of GDP in the target region and each of the other regions is as follows: .
[0063] Select other regions whose energy intensity differs from the target region by less than a threshold (e.g., ±20%). For example, if the target region has an energy intensity of 1.5 tons of standard coal equivalent per 10,000 yuan of GDP, then select other regions with energy intensity in the range of 1.2-1.8 tons of standard coal equivalent per 10,000 yuan of GDP to ensure consistency in energy-driven patterns.
[0064] Step a4: Use feature similarity and energy intensity difference to filter carbon emission source data to obtain target migration data.
[0065] Carbon emission source data is screened using a dual filtering condition of feature similarity and energy intensity difference, and carbon emission data corresponding to other regions that simultaneously meet the conditions are used as target migration data.
[0066] In the above implementation, the dual screening mechanism of feature similarity and energy intensity difference is used to accurately match external data sources that are highly consistent with the carbon emission entity structure and energy consumption pattern of the target area, ensuring the consistency of the supplementary data in terms of carbon emission entity correlation and energy driving logic, and fundamentally avoiding pseudo-similar interference in cross-domain knowledge transfer.
[0067] In some alternative implementations, step a4 above includes: Step a41: From multiple other regions, select regions whose feature similarity exceeds a similarity threshold and whose energy intensity difference does not exceed an intensity threshold as carbon emission data migration sources.
[0068] Carbon emission data migration sources refer to the set of regions that simultaneously meet the criteria of feature similarity exceeding a similarity threshold and energy intensity difference not exceeding an intensity threshold. Specifically, the industry similarity between the target region and other regions is calculated based on weighted cosine similarity, and regions exceeding the similarity threshold (e.g., 0.7) are selected. The energy intensity difference between the target region and other regions is calculated, and regions with differences not exceeding an intensity threshold (e.g., ±20%) are selected. For other regions that simultaneously meet the above conditions, the Top-K regions (K≈5~10) are selected in descending order of similarity as carbon emission data migration sources, and migration weights are assigned (the higher the weight, the greater the influence of other regions on the transfer learning model). .
[0069] Step a42: The carbon emission source data corresponding to the carbon emission data migration source is determined as the target migration data.
[0070] The complete carbon emission dataset corresponding to the carbon emission data migration source is defined as the target migration data. Specifically, after the carbon emission data migration source is determined, its carbon emission source data (including historical energy series, GDP, population, and other full information) is extracted. This carbon emission source data is the target migration data used for transfer learning and is directly input into the transfer learning model for knowledge transfer.
[0071] In the above implementation, by setting a dual hard condition of similarity threshold and energy intensity difference threshold, high-quality migration source areas with both carbon emission entity characteristics and energy efficiency are strictly locked, thus ensuring the core quality and scenario adaptability of cross-domain data supplementation from a mechanism perspective.
[0072] Step S2033: Using the first data supplementation strategy, the original carbon emission data is supplemented with the target migration data to obtain the target carbon emission data.
[0073] On the selected target migration data, an LSTM model is constructed. The specific process includes: Calculate the forget gate. The forget gate determines how much old information needs to be forgotten. It works by weighting the hidden state from the previous time step and the input from the current time step, then applying the sigmoid function to calculate a value between 0 and 1. The specific formula is as follows: .
[0074] in, It is the output of the forget gate; and These are the weight matrix and the bias terms; It is to hide the previous state. and current input The result of piecing everything together. The output of the forget gate. It will be used to update the state of the memory unit. The closer to 0, the more information will be forgotten; the closer to 1, the more information will be retained.
[0075] Calculate the input gate and candidate memory. The input gate determines how much new information should be written into the memory cells; the specific formula is as follows: .
[0076] in, The output of the input gate represents the proportion of new information entering the memory unit.
[0077] Next, LSTM calculates candidate memory cells: .
[0078] in, and It consists of a new weight matrix and bias terms, and candidate memories. The output of the input gate Together they decide how much new information to write into the memory cells.
[0079] Update the memory cells. LSTM updates the memory cells using the values from the forget gate and the input gate, as shown in the following formula: .
[0080] in, It is an updated memory unit; It is the memory unit of the previous time step; This indicates that the previously remembered portion is retained; This indicates the new information to be added. In this way, LSTM can simultaneously retain long-term dependent information and add new, useful information.
[0081] Calculate the output gate and output hidden state. LSTM uses the output gate to determine the output information at the current time step and updates the hidden state. The specific formula is as follows: .
[0082] in, It is the output of the output gate, indicating how much of the information in the current memory cell will be used to generate a new hidden state.
[0083] The method for updating the hidden state is: .
[0084] in, It is the hidden state at the current time step; This is the updated memory unit state. Hidden state. As the output of the LSTM, it will also be passed to the next time step.
[0085] After constructing the LSTM model according to the above process, the input includes historical energy use sequence data and exogenous variable data (GDP growth rate, population, etc.), and the output is predicted carbon emissions and energy intensity time series data. During fine-tuning of the target region, the first layer is frozen, and only the second layer and the fully connected layer are trained. Specifically, the target region is transferred through data... Compared with raw carbon emission data The data should be mixed, while retaining a portion of the target region's samples as a validation set. For both input formats, it's crucial to ensure that the carbon emission data migration source and the exogenous variables (GDP growth rate, population, etc.) of the target region are consistent. The specific formula is as follows: .
[0086] Among them, for mixed weights The initial value is 1 (completely dependent on the carbon emission data migration source), which is gradually reduced to 0.5 during training.
[0087] Finally, the reliability of the established transfer learning model is determined using the root mean square error method.
[0088] The original carbon emission data is supplemented using the established transfer learning model to obtain the target carbon emission data.
[0089] The carbon emission reduction evolution path generation method provided in this invention automatically triggers a cross-regional data migration strategy by setting a preset threshold, filters target migration data that meets preset conditions, and uses the target migration data to supplement the original data that is severely missing in the target area, significantly improving the feasibility and cross-domain adaptability of data generation in scenarios with severe data scarcity.
[0090] Step S2034: When the target data deficiency coefficient exceeds the second preset threshold, determine the second data supplementation strategy corresponding to the target data deficiency coefficient.
[0091] The second preset threshold refers to a pre-defined critical value for the data deficiency coefficient, used to characterize mild data deficiency scenarios. The second data imputation strategy refers to an LSTM-based time series prediction method designed for mild data deficiency scenarios. Specifically, when the data deficiency coefficient of the target area exceeds the second preset threshold, the imputation strategy for mild data deficiency scenarios is automatically triggered, i.e., the LSTM-based time series modeling method.
[0092] In addition, a second data supplementation strategy can be determined when the target data lack coefficient does not exceed the first preset threshold.
[0093] Step S2035: Preprocess the raw carbon emission data to obtain preprocessed time series data; use a sliding time window to extract data from the preprocessed time series data to obtain extracted data of a preset duration.
[0094] Preprocessed time-series data refers to the normalized time-series dataset after preprocessing the original carbon emission data. A sliding time window refers to a time-series data truncation rule, using a fixed duration (e.g., 3 years) as the window and sliding it step-by-step to extract continuous subsequences as model input. For example, input: carbon emission sequence [t-2 year, t-1 year, t year]; output: predicted value for t+1 year. Truncation refers to the fixed-length subsequence cut from the preprocessed time-series data using a sliding time window. Specifically, for mildly missing data, cleaning and standardization are required. For single missing values (missing percentage < 5%), linear interpolation is used to impute them. .
[0095] If continuity is missing, it is filled by extracting the trend term, seasonal term, and residual term using Seasonal Decomposition (STL): .
[0096] All features need to be Z-score standardized to eliminate dimensional differences: .
[0097] in, and These represent the mean and standard deviation of the features, respectively. A time-series window is then constructed to extract data from the preprocessed time-series data, for example, extracting 3 years of historical data (endogenous variables: carbon emission series; exogenous variables: GDP growth rate, energy intensity, etc.) to predict carbon emissions for the next year. The input / output format is as follows: ; .
[0098] Step S2036: Using the second data supplementation strategy, predict the carbon emission data for the next target duration based on the intercepted data, and supplement the original carbon emission data with the predicted carbon emission data to obtain the target carbon emission data.
[0099] Predicted carbon emission data is future carbon emission prediction data generated by a pre-trained LSTM model based on truncated data. Specifically, the truncated data is input into the pre-trained LSTM model, which outputs future dynamic data (such as carbon emissions and energy intensity), i.e., predicted carbon emission data. The predicted carbon emission data is then used to directly replace the missing or insufficient parts of the original carbon emission data to generate complete target carbon emission data.
[0100] The carbon emission reduction evolution path generation method provided in this embodiment of the invention performs time-series slicing on the preprocessed raw carbon emission data through a preset sliding time window, and uses the extracted data to predict key carbon emission data for future periods, so as to dynamically supplement the gaps in the original data and achieve efficient self-sufficient data generation in mild data-scarce scenarios.
[0101] Step S204: Generate the carbon emission reduction evolution path corresponding to the target region based on the target carbon emission data.
[0102] Specifically, step S204 includes: Step S2041: Extract scenario parameter data corresponding to each preset carbon emission scenario from the target carbon emission data.
[0103] Preset carbon emission scenarios refer to carbon emission evolution simulation schemes under differentiated policy interventions. Scenario parameter data refers to the parameter data extracted from the target carbon emission data corresponding to each preset carbon emission scenario. Specifically, after obtaining complete and reliable target carbon emission data, specific data subsets or features matching the parameters of each predefined carbon emission scenario are extracted; these are the scenario parameter data. Each preset scenario (e.g., "policy reinforcement scenario," "technology breakthrough scenario," "baseline scenario," etc.) clearly specifies which key indicators or variables need to be focused on (e.g., emission intensity of specific industries, energy structure ratio, technological progress rate, policy constraint coefficient, etc.). The extraction process involves, based on the definitions of these preset scenarios, locating and extracting specific parameter values or datasets directly related to the core driving factors and boundary conditions of each scenario from the comprehensive data pool of target carbon emission data.
[0104] Step S2042: Using scenario parameter data, simulate the carbon emission reduction evolution path under different preset carbon emission scenarios corresponding to the target area.
[0105] After obtaining the scenario parameter data corresponding to a specific scenario, these data are used as key inputs and boundary conditions to drive the Long-range Energy Alternatives Planning System (LEAP) model, as shown in the following formula: ; .
[0106] in, and Represents total energy use and carbon emissions. They represent different economic sectors within different industries. Representing different types of energy, These represent the corresponding emission factors, sector activity levels, and sector energy intensity, respectively.
[0107] like Figure 3 As shown, this model comprehensively considers a series of factors such as population, economy, and technology, including energy supply, energy processing and conversion, and end-use energy demand. Combining input scenario parameter data, it forecasts medium- and long-term energy demand and carbon emissions for sectors such as energy, transportation, housing, and industry under different preset scenarios, and forecasts supply-demand, pollutant emissions, and carbon emissions for different energy types. It generates a trajectory showing how carbon emissions in the target region change over time under a specific preset scenario, i.e., a carbon reduction evolution path. Each path visually demonstrates the dynamic process the region undergoes to achieve its emission reduction targets under the corresponding scenario setting.
[0108] The carbon emission reduction evolution path generation method provided in this invention accurately extracts a subset of data strongly correlated with a specific preset scenario from the supplemented target carbon emission data. This ensures that the generated carbon emission reduction evolution path is entirely based on data-driven logic customized for that scenario, preserving global data consistency while achieving precise adaptation of scenario-based path generation. Simultaneously, the supplemented target carbon emission data is combined with emission reduction path deductions under different preset carbon emission scenarios, achieving full automation from data input to path generation. This method can adapt to the carbon emission reduction path deduction needs of regions with different data endowments, providing high-precision and highly adaptable carbon peaking path decision support for differentiated regional scenarios.
[0109] This embodiment also provides a carbon emission reduction evolution path device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0110] This embodiment provides a carbon emission reduction evolution path device, such as... Figure 4 As shown, it includes: Module 401 is used to acquire raw carbon emission data for the target area; Module 402 is used to determine the target data lack coefficient of the original carbon emission data in the target area based on the multi-dimensional characteristics of the original carbon emission data. The supplementation module 403 is used to determine the target data supplementation strategy corresponding to the target data deficiency coefficient based on the correspondence between the data deficiency coefficient and the data supplementation strategy, and to supplement the original carbon emission data according to the target data supplementation strategy to obtain the target carbon emission data. The generation module 404 is used to generate the carbon emission reduction evolution path corresponding to the target region based on the target carbon emission data.
[0111] In some alternative implementations, the determining module 402 includes: The first acquisition submodule is used to acquire the data integrity indicators, data continuity indicators, and data resolution indicators corresponding to the raw carbon emission data; The allocation submodule is used to dynamically allocate the weight coefficients of data integrity indicators, data continuity indicators, and data resolution indicators according to the carbon emission entity type of the target area, so as to obtain the first weight coefficient corresponding to the data integrity indicator, the second weight coefficient corresponding to the data continuity indicator, and the third weight coefficient corresponding to the data resolution indicator. The weighted summation submodule is used to perform a weighted summation of the data integrity index, data continuity index, and data resolution index using the first weight coefficient, the second weight coefficient, and the third weight coefficient to obtain the target data lack coefficient for the target area.
[0112] In some alternative implementations, the supplementation module 403 includes: The first determining submodule is used to determine the first data supplementation strategy corresponding to the target data deficiency coefficient when the target data deficiency coefficient exceeds the first preset threshold. The second acquisition submodule is used to acquire carbon emission source data from multiple other regions and determine the target migration data that meets preset conditions from the multiple carbon emission source data. The first data supplementation submodule is used to supplement the original carbon emission data with the target migration data through the first data supplementation strategy to obtain the target carbon emission data.
[0113] In some optional implementations, the second acquisition submodule includes: The acquisition unit is used to acquire the first feature vector of the target region and the second feature vectors of each of the other regions; The first determining unit is used to determine the feature similarity between the target region and each other region based on the first feature vector and the second feature vector; The second determining unit is used to extract the first energy intensity of the target area from the original carbon emission data, extract the second energy intensity of each other area from the carbon emission source data, and determine the energy intensity difference between the first energy intensity and each second energy intensity. The filtering unit is used to filter carbon emission source data based on feature similarity and energy intensity difference to obtain target migration data.
[0114] In some alternative implementations, the filtering unit includes: The filtering subunit is used to select regions from multiple other regions whose feature similarity exceeds a similarity threshold and whose energy intensity difference does not exceed an intensity threshold as carbon emission data migration sources. The sub-unit is used to identify the carbon emission source data corresponding to the carbon emission data migration source as the target migration data.
[0115] In some alternative implementations, the supplementation module 403 further includes: The second determination submodule is used to determine the second data supplementation strategy corresponding to the target data deficiency coefficient when the target data deficiency coefficient exceeds the second preset threshold. The preprocessing submodule is used to preprocess the raw carbon emission data to obtain preprocessed time-series data; The data extraction submodule is used to extract data from preprocessed time series data using a sliding time window to obtain extracted data of a preset duration. The second data supplementation submodule is used to supplement the original carbon emission data with the predicted carbon emission data for the next target duration according to the intercepted data using the second data supplementation strategy, and to obtain the target carbon emission data.
[0116] In some alternative implementations, the generation module 404 includes: The extraction submodule is used to extract scenario parameter data corresponding to each preset carbon emission scenario from the target carbon emission data; The simulation submodule is used to simulate the carbon emission reduction evolution path under different preset carbon emission scenarios corresponding to the target area using scenario parameter data.
[0117] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0118] The carbon emission reduction evolution path device in this embodiment is presented in the form of functional units. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0119] The carbon emission reduction evolution path generation device provided in this embodiment of the invention significantly improves the reliability of data under data scarcity conditions by quantitatively assessing the multi-dimensional missingness of carbon emission data in the target area and dynamically matching the optimal data supplementation strategy, thereby generating a more accurate and adaptable carbon emission reduction evolution path.
[0120] This invention also provides a computer device having the above-described features. Figure 4 The carbon emission reduction evolution path device is shown.
[0121] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 5 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 5 Take a processor 10 as an example.
[0122] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.
[0123] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.
[0124] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0125] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0126] The computer device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.
[0127] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the computer device, such as a touchscreen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touchscreen.
[0128] The computer device also includes a communication interface for communicating with other devices or communication networks.
[0129] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.
[0130] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0131] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for generating carbon emission reduction evolution pathways, characterized in that, The method includes: Obtain raw carbon emission data for the target area; Based on the multi-dimensional characteristics of the raw carbon emission data, the target data lack coefficient of the raw carbon emission data in the target region is determined; Based on the correspondence between the data deficiency coefficient and the data supplementation strategy, the target data supplementation strategy corresponding to the target data deficiency coefficient is determined, and the original carbon emission data is supplemented according to the target data supplementation strategy to obtain the target carbon emission data. Generate the carbon reduction evolution path corresponding to the target region based on the target carbon emission data.
2. The method according to claim 1, characterized in that, The multi-dimensional features include data integrity, data continuity, and data resolution; The determination of the target data deficiency coefficient of the original carbon emission data in the target region based on the multi-dimensional features of the original carbon emission data includes: Obtain the data integrity index, data continuity index, and data resolution index corresponding to the raw carbon emission data; Based on the carbon emission entity type of the target area, the weight coefficients of the data integrity index, the data continuity index, and the data resolution index are dynamically allocated to obtain the first weight coefficient corresponding to the data integrity index, the second weight coefficient corresponding to the data continuity index, and the third weight coefficient corresponding to the data resolution index. Using the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient, the data integrity index, the data continuity index, and the data resolution index are weighted and summed to obtain the target data lack coefficient for the target region.
3. The method according to claim 1 or 2, characterized in that, Based on the correspondence between the data deficiency coefficient and the data supplementation strategy, the target data deficiency coefficient is determined, and the target data supplementation strategy is applied to the original carbon emission data to obtain the target carbon emission data, including: When the target data deficiency coefficient exceeds a first preset threshold, a first data supplementation strategy corresponding to the target data deficiency coefficient is determined. Acquire carbon emission source data from multiple other regions, and determine target migration data that meets preset conditions from the multiple carbon emission source data; The target carbon emission data is obtained by supplementing the original carbon emission data with the target migration data using the first data supplementation strategy.
4. The method according to claim 3, characterized in that, Target migration data that meets preset conditions is determined from multiple carbon emission source data, including: Obtain the first feature vector of the target region and the second feature vectors of each of the other regions; Based on the first feature vector and the second feature vector, the feature similarity between the target region and each of the other regions is determined; Extract the first energy intensity of the target region from the original carbon emission data, extract the second energy intensity of each of the other regions from the carbon emission source data, and determine the energy intensity difference between the first energy intensity and each of the second energy intensities; The carbon emission source data is filtered using the feature similarity and the energy intensity difference to obtain the target migration data.
5. The method according to claim 4, characterized in that, The process of filtering the carbon emission source data using the feature similarity and the energy intensity difference to obtain the target migration data includes: From the other regions, regions whose feature similarity exceeds a similarity threshold and whose energy intensity difference does not exceed an intensity threshold are selected as carbon emission data migration sources. The carbon emission source data corresponding to the carbon emission data migration source is determined as the target migration data.
6. The method according to claim 3, characterized in that, Also includes: When the target data deficiency coefficient exceeds the second preset threshold, a second data supplementation strategy corresponding to the target data deficiency coefficient is determined; The raw carbon emission data is preprocessed to obtain preprocessed time-series data; The preprocessed time series data is truncated using a sliding time window to obtain truncated data of a preset duration. The second data supplementation strategy is used to predict carbon emission data for the next target duration based on the extracted data, and the original carbon emission data is supplemented using the predicted carbon emission data to obtain the target carbon emission data.
7. The method according to claim 1, characterized in that, The step of generating the carbon reduction evolution path corresponding to the target region based on the target carbon emission data includes: Extract scenario parameter data corresponding to each preset carbon emission scenario from the target carbon emission data; Using the scenario parameter data, the carbon emission reduction evolution path under different preset carbon emission scenarios corresponding to the target area is simulated.
8. A carbon emission reduction evolution path generation device, characterized in that, The device includes: The acquisition module is used to acquire raw carbon emission data for the target area; The determination module is used to determine the target data deficiency coefficient of the original carbon emission data in the target region based on the multi-dimensional characteristics of the original carbon emission data; The data supplementation module is used to determine the target data supplementation strategy corresponding to the target data deficiency coefficient based on the correspondence between the data deficiency coefficient and the data supplementation strategy, and to supplement the original carbon emission data according to the target data supplementation strategy to obtain the target carbon emission data. The generation module is used to generate a carbon reduction evolution path corresponding to the target region based on the target carbon emission data.
9. A computer device, characterized in that, include: A memory and a processor are interconnected, the memory stores computer instructions, and the processor executes the carbon emission reduction evolution path generation method according to any one of claims 1 to 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to execute the carbon emission reduction evolution path generation method according to any one of claims 1 to 7.