A method, device, equipment, medium and product for constructing a carbon emission prediction model
By constructing a carbon emission sample set and a candidate influencing factor set, screening and integrating key carbon emission influencing factors, and assigning weights to them, and using a neural network model for training, the problem of insufficient identification of influencing factors in the carbon emission prediction model of coal enterprises was solved, and the prediction accuracy and stability were improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENHUA SHENDONG COAL GRP
- Filing Date
- 2026-04-09
- Publication Date
- 2026-06-02
AI Technical Summary
In existing technologies, carbon emission prediction models for coal enterprises suffer from problems such as insufficient identification of key influencing factors, subjective setting of input feature weights, and insufficient model adaptability, resulting in poor prediction accuracy and stability.
By constructing a carbon emission sample set and a candidate influencing factor set, key carbon emission influencing factors are screened and fused to obtain them, and weighted accordingly. A carbon emission prediction model is then established by training a neural network model.
It improves the accuracy, stability, and engineering applicability of carbon emission prediction, and enhances the model's ability to characterize complex nonlinear relationships and adapt to data from different stages and operating conditions.
Smart Images

Figure CN122133877A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of carbon emission prediction technology, and in particular to a method, apparatus, equipment, medium and product for constructing a carbon emission prediction model. Background Technology
[0002] Accurate prediction of carbon emissions is of great significance in the carbon emission management and emission reduction decision-making of coal enterprises. Due to the complex sources of emissions from coal enterprises, numerous influencing factors, and the strong nonlinearity and uncertainty of the carbon emission process, there is an urgent need to provide a method for constructing a high-precision and highly adaptable carbon emission prediction model in order to achieve effective prediction of carbon emissions.
[0003] In related technologies, the construction scheme usually adopts the method of selecting influencing factors based on experience and combining them with conventional prediction models. However, this construction method has problems such as insufficient identification of key influencing factors, subjective setting of input feature weights, and insufficient model adaptability, resulting in poor prediction accuracy and stability of the constructed model, which is difficult to meet the actual application needs of carbon emission prediction for coal enterprises. Summary of the Invention
[0004] This disclosure provides a method, apparatus, equipment, medium, and product for constructing a carbon emission prediction model.
[0005] According to a first aspect of this disclosure, a method for constructing a carbon emission prediction model is provided, the method comprising: Based on the carbon emission activity data of the target object, a carbon emission sample set and a candidate influencing factor set are constructed; the candidate influencing factor set is used to characterize the candidate factors affecting the carbon emission changes of the target object. Based on the carbon emission sample set, the candidate influencing factor set is screened and fused to obtain key carbon emission influencing factors; The key carbon emission influencing factors are weighted to obtain the factor weights corresponding to each key carbon emission influencing factor. The key carbon emission influencing factors are weighted according to the weights of each factor to obtain a weighted input feature set. The neural network model is then trained based on the weighted input feature set and the carbon emission sample set to obtain a carbon emission prediction model.
[0006] Furthermore, the screening and fusion of the candidate influencing factor set based on the carbon emission sample set to obtain key carbon emission influencing factors includes: A data availability analysis is performed on the candidate impact factor set to obtain the impact factor set to be analyzed; wherein, the impact factor set to be analyzed consists of candidate impact factors that meet preset data integrity conditions; Based on the carbon emission sample set, a correlation analysis is performed on the set of influencing factors to be analyzed to obtain the first influencing factor and the second influencing factor. The second impact factor is fused to obtain the composite impact factor; Based on the first influencing factor and the composite influencing factor, the key carbon emission influencing factors are determined.
[0007] Further, the correlation analysis of the set of influencing factors to be analyzed based on the carbon emission sample set to obtain the first influencing factor and the second influencing factor includes: Using the carbon emissions corresponding to the carbon emission sample set as the dependent variable, the correlation between each of the influencing factors to be analyzed and the carbon emissions is calculated using the Pearson correlation coefficient. The influencing factors whose correlation satisfies the preset correlation threshold are determined as the first influencing factor; The second influence factor is obtained by identifying the influence factors that meet the preset correlation conditions from the influence factors whose correlation does not meet the preset correlation threshold.
[0008] Further, the fusion processing of the second impact factor to obtain a composite impact factor includes: The second impact factor is standardized to obtain the third impact factor; K-means cluster analysis was performed on the third influencing factor to obtain the cluster analysis results; Based on the clustering analysis results, the distance results are calculated; The distance results are normalized to obtain normalized distance results, and the composite influence factor is constructed based on the normalized distance results.
[0009] Furthermore, the weighting process for the key carbon emission influencing factors to obtain the factor weights corresponding to each key carbon emission influencing factor includes: Construct the original index sample matrix corresponding to the key carbon emission influencing factors, and normalize the original index sample matrix to obtain the normalized index matrix. A proportion matrix is constructed based on the normalized index matrix, and the entropy value corresponding to each of the key carbon emission influencing factors is determined based on the proportion matrix. The redundancy of each of the key carbon emission influencing factors is calculated based on the entropy value. The factor weights corresponding to each of the key carbon emission influencing factors are determined based on the redundancy.
[0010] Furthermore, the neural network model is a backpropagation neural network model; The step of training a neural network model based on the weighted input feature set and the carbon emission sample set to obtain a carbon emission prediction model includes: Construct the backpropagation neural network model; wherein the number of input layer nodes of the backpropagation neural network model corresponds to the number of features in the weighted input feature set; The weighted input feature set is input into the backpropagation neural network model, and the carbon emissions corresponding to the carbon emission sample set are used as the target output to train the backpropagation neural network model to obtain the carbon emission prediction model.
[0011] According to a second aspect of this disclosure, an apparatus for constructing a carbon emission prediction model is provided, the apparatus comprising: A construction module is used to construct a carbon emission sample set and a candidate influencing factor set based on the carbon emission activity data of the target object; the candidate influencing factor set is used to characterize the candidate factors affecting the carbon emission changes of the target object; The first processing module is used to screen and fuse the candidate influencing factor set based on the carbon emission sample set to obtain key carbon emission influencing factors. The second processing module is used to assign weights to the key carbon emission influencing factors to obtain the factor weights corresponding to each key carbon emission influencing factor. The training module is used to weight the key carbon emission influencing factors based on the weights of each factor to obtain a weighted input feature set, and to train the neural network model based on the weighted input feature set and the carbon emission sample set to obtain a carbon emission prediction model.
[0012] According to a third aspect of this disclosure, an electronic device is provided. The electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the method described above.
[0013] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the methods described above.
[0014] According to a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes a computer program that, when executed by a processor, implements the methods described above in this disclosure.
[0015] This disclosure provides a method, apparatus, device, medium, and product for constructing a carbon emission prediction model. First, based on the carbon emission activity data of the target object, a carbon emission sample set and a candidate influencing factor set are constructed. The candidate influencing factor set is used to characterize candidate factors affecting changes in the target object's carbon emissions. Then, based on the carbon emission sample set, the candidate influencing factor set is screened and fused to obtain key carbon emission influencing factors. Next, the key carbon emission influencing factors are weighted to obtain the factor weights corresponding to each key carbon emission influencing factor. Finally, based on the factor weights, the key carbon emission influencing factors are weighted to obtain a weighted input feature set. A neural network model is trained based on the weighted input feature set and the carbon emission sample set to obtain a carbon emission prediction model.
[0016] As described above, the embodiments of this disclosure first screen and fuse candidate influencing factor sets based on a carbon emission sample set to obtain key carbon emission influencing factors, thereby reducing the bias caused by subjective selection of input factors and improving the accuracy of the input factors in representing changes in carbon emissions. Second, the embodiments of this disclosure assign weights to the key carbon emission influencing factors to obtain the factor weights corresponding to each key carbon emission influencing factor, thereby establishing an objective weighting mechanism based on data features and improving the rationality of the model input structure. Finally, the embodiments of this disclosure train the neural network model based on a weighted input feature set, which can improve the carbon emission prediction model's ability to characterize complex nonlinear relationships and its adaptability to data from different stages and operating conditions, thereby improving the accuracy, stability, and engineering practicality of carbon emission prediction. Attached Figure Description
[0017] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0018] Figure 1 A flowchart illustrating a method for constructing a carbon emission prediction model provided as an exemplary embodiment of this disclosure; Figure 2 One of the flowcharts for a method of constructing a carbon emission prediction model provided as another exemplary embodiment of this disclosure; Figure 3 A second flowchart of a method for constructing a carbon emission prediction model provided for another exemplary embodiment of this disclosure; Figure 4 A flowchart of a method for constructing a carbon emission prediction model provided for another exemplary embodiment of this disclosure; Figure 5A schematic diagram of the topology of a BP neural network model provided in an exemplary embodiment of this disclosure; Figure 6 A schematic block diagram of the functional modules of an apparatus for constructing a carbon emission prediction model provided in an exemplary embodiment of this disclosure; Figure 7 A structural block diagram of an electronic device provided as an exemplary embodiment of this disclosure; Figure 8 A structural block diagram of a computer system provided as an exemplary embodiment of this disclosure; Figure 9 A structural block diagram of a computer program product provided for an exemplary embodiment of this disclosure. Detailed Implementation
[0019] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0020] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0021] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc., used in this disclosure are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0022] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0023] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0024] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0025] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0026] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device. It is understood that the above notification and user authorization process is merely illustrative and does not constitute a limitation on the implementation of this disclosure; other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0027] Carbon emission prediction for coal enterprises is a complex modeling problem driven by multiple coupled factors. It involves not only acquiring carbon emission data but also selecting influencing factors, identifying relationships between features, determining factor weights, and constructing a prediction model. In existing technologies, the selection of input factors generally relies on empirical judgment or literature review, lacking a systematic analysis that combines actual production processes and multidimensional driving factors. This easily leads to the omission of key factors or interference from irrelevant factors. Simultaneously, the weighting of input factors often lacks objective basis based on data characteristics, easily resulting in an unreasonable model input structure, thus affecting the accuracy and stability of the prediction results. Furthermore, while some neural network models possess a certain degree of nonlinear fitting ability, they still suffer from insufficient adaptability and limited generalization ability when faced with the diversity and uncertainty of emission data at different stages and under different operating conditions. Therefore, it is necessary to provide a carbon emission prediction model construction scheme that can balance key factor identification, objective weighting, and nonlinear modeling capabilities.
[0028] In one embodiment, such as Figure 1 As shown, a method for constructing a carbon emission prediction model is provided, including the following steps: Step 101: Based on the carbon emission activity data of the target object, construct a carbon emission sample set and a candidate influencing factor set.
[0029] Here, the implementing entity can construct a carbon emission sample set and a candidate influencing factor set based on the carbon emission activity data of the target object. The candidate influencing factor set is used to characterize the candidate factors that affect the carbon emission changes of the target object.
[0030] In one possible embodiment, the target object can be an emission entity with carbon emission activities. For example, the target object can be a coal enterprise. Specifically, a coal enterprise can be an enterprise entity that includes one or more production links such as coal mining, washing and processing, transportation, and coal-fired power generation. Carbon emission activity data can be understood as production operation data, energy consumption data, resource consumption data, transportation activity data, and statistical data reflecting the external development environment of the target object related to the carbon emission formation process. The carbon emission sample set is a set of sample data formed based on the carbon emission activity data of the target object in multiple sample periods, used to characterize the carbon emission level of the target object in each sample period. The candidate influencing factor set can be understood as a set of candidate factors initially selected from multiple dimensions that may affect the carbon emission changes of the target object, used to characterize the potential driving factors of the carbon emission changes of the target object. Furthermore, the candidate influencing factor set can be constructed based on the system boundary of the target object, the detailed carbon emission sources of each link and their accounting models, combined with measured data and regional statistical data, and can cover multiple dimensions such as energy resources, economic development, social factors, environmental factors and management efficiency.
[0031] In one possible embodiment, the target is a coal enterprise. The specific steps for the implementing entity to construct a carbon emission sample set and a candidate influencing factor set based on the target enterprise's carbon emission activity data are as follows: First, the implementing entity uses the life cycle assessment method to divide the coal enterprise's production process into stages and sets system boundaries to clarify the scope of carbon emission accounting. Specifically, the implementing entity can divide the enterprise's production process into four core links: coal mining, washing and processing, transportation, and coal-fired power generation. The system boundary is limited to the period from the start of coal mining to its consumption within the mining area or its transportation to the mine exit (including railway and highway shipping points), covering all direct or indirect carbon emissions, including the main production system (mining equipment, generator sets), auxiliary systems (ventilation systems, water supply and drainage systems), and ancillary facilities (office areas, transportation roads), but excluding carbon emissions generated from long-distance transportation outside the mining area, end-user consumption, and secondary processing.
[0032] Secondly, the implementing entities identify and classify carbon emission sources at each stage, and refine their respective carbon emission sources based on the actual production activities and characteristics of each stage of the enterprise. Specifically, carbon emission sources can be divided into direct emission sources, energy indirect emission sources, and other indirect emission sources. Direct emission sources include direct carbon dioxide emissions from energy consumption, methane and carbon dioxide escape emissions from mining or open-pit mining processes, methane and carbon dioxide escape emissions from post-mining activities, and carbon emissions from uncontrolled oxidation and spontaneous combustion of coal or waste stockpiles. Energy indirect emission sources include implicit carbon emissions from purchased electricity and heat consumption. Other indirect emission sources include water consumption required for coal mine dust removal, carbon emissions from land use changes, and carbon emissions caused by coal seam oxidation. Furthermore, based on the production activities and characteristics of each stage of the enterprise, the coal mining stage may include operations such as coal breaking, loading, and ventilation, and its carbon emission sources may include gas leakage, electricity consumption, water consumption, diesel engine transportation, and dust control systems. The washing and processing stage may involve operations such as crushing, screening, dehydration, and drying, and its carbon emission sources may include heating fuel, electric motor drive, water supply, and in-plant transportation. The transportation stage may cover the entire process of raw coal from underground to the collection and distribution station, with electric locomotives and fuel vehicles as the main emission sources. The coal-fired power generation stage may involve fuel combustion and post-treatment processes, and the emission composition may include emissions from coal combustion, coal slime, coal gangue combustion, and emissions from cooling water circulation and desulfurization processes.
[0033] Based on this, the implementing entity, through a detailed analysis of the carbon emission sources at each stage, constructs a segmented carbon emission accounting model. Specifically, the carbon emission accounting model for the coal mining stage... middle, This represents the total carbon emissions from the coal mining process. This indicates the amount of carbon emissions produced by burning raw coal. This indicates the carbon emissions generated by fuel consumption during automobile transportation. This indicates the amount of carbon emissions generated by gas release during coal mining. This indicates the amount of carbon emissions generated by consuming water. This represents the carbon emissions generated from electricity consumption, with all units expressed as tCO2e; a carbon emission accounting model for the coal washing and processing process. middle, This represents the total carbon emissions from the coal washing and processing process. This indicates the amount of carbon emissions produced by burning raw coal. This indicates the carbon emissions generated by the fuel consumed by vehicles during the coal washing and processing process. This indicates the amount of carbon emissions generated by consuming water. This represents the carbon emissions generated from electricity consumption, with all units expressed as tCO2e; a carbon emission accounting model for coal transportation. middle, This indicates the total carbon emissions during coal transportation. This represents the carbon emission factor for the distance corresponding to the i-th mode of transportation during the transportation process, expressed in tCO2e / km. This represents the transport distance of the i-th type of locomotive during the transportation process, in km; Carbon emission accounting model for coal-fired power generation. middle, This represents the total carbon emissions from the coal-fired power generation process in a self-owned power plant. This indicates the amount of carbon emissions produced by burning raw coal. This indicates the amount of carbon emissions produced by burning coal slime. This indicates the amount of carbon emissions produced by the combustion of coal gangue. This indicates the amount of carbon emissions generated from electricity consumption. This indicates the carbon emissions generated by fuel consumption during automobile transportation. This indicates the amount of carbon emissions generated by consuming water. This represents the carbon emissions generated from the desulfurization of exhaust gases, with all units expressed as tCO2e. Based on the above segmented carbon emission accounting model, the implementing entity can obtain the sample data foundation for characterizing the carbon emission levels of each sample period.
[0034] Then, based on the aforementioned system boundaries, the detailed carbon emission sources and their accounting models for each link, and combined with measured data and historical statistics, the implementing entity constructs a preliminary impact factor database. This preliminary impact factor database follows the principle of "covering major emission sources and integrating macro background variables" and covers multiple dimensions such as energy resources, economic development, social factors, environmental factors and management efficiency.
[0035] In this embodiment, as an exemplary implementation, the executing entity can initially select 16 indicators from the aforementioned multiple dimensions as candidate influencing factors. These include: energy resource indicators such as fuel consumption at each stage (X1), gas extraction volume (X2), purchased electricity consumption (X3), purchased heat consumption (X4), annual water consumption (X5), and average coal transportation distance (X6); economic development indicators such as regional GDP (X7), secondary industry proportion (X8), total energy consumption (X9), and coal share in the energy structure (X10); social factors indicators such as resident population size (X11), total electricity consumption (X12), and road density (X13); environmental factors indicators such as forest coverage rate (X14); and management efficiency indicators such as equipment utilization rate (X15) and average annual operating hours (X16). It should be understood that the aforementioned 16 indicators are merely illustrative examples. In other embodiments, the number and specific content of candidate influencing factors can be increased, decreased, or adjusted based on the type of the target object, the system boundary scope, data availability, and actual application needs. Thus, the executing entity can obtain a preliminary composition of the candidate influencing factors.
[0036] Based on the above, the implementing entity can further combine historical data from multiple sample periods to organize the preliminary composition of the aforementioned sample data foundation and candidate impact factors, forming a carbon emission sample set and a candidate impact factor set. This provides a sample foundation for subsequent data availability analysis, correlation analysis, cluster fusion, and weighted modeling of the candidate impact factor set. Thus, the implementing entity completes the construction of the carbon emission sample set and the candidate impact factor set.
[0037] Step 102: Based on the carbon emission sample set, the candidate influencing factor set is screened and fused to obtain the key carbon emission influencing factors.
[0038] Here, after constructing a carbon emission sample set and a candidate influencing factor set based on the carbon emission activity data of the target object, the implementing entity can screen and fuse the candidate influencing factor set based on the carbon emission sample set to obtain key carbon emission influencing factors. These key carbon emission influencing factors can be understood as: among the candidate influencing factors that meet data integrity requirements, those that are significantly correlated with total carbon emissions, and composite influencing factors that, although not showing a significant linear correlation, can characterize the potential coupling relationship of carbon emissions after fusion processing. Through this step, the implementing entity can further extract input factors from the initially constructed candidate influencing factors that have a more representative and discriminative ability for carbon emission prediction, providing a foundation for subsequent weighting processing and model training.
[0039] In one possible embodiment, such as Figure 2 As shown, the key carbon emission influencing factors are obtained by screening and fusing the candidate influencing factor set based on the carbon emission sample set, including the following steps: Step 1021: Perform data availability analysis on the candidate impact factor set to obtain the impact factor set to be analyzed.
[0040] Here, after constructing a carbon emission sample set and a candidate impact factor set based on the carbon emission activity data of the target object, the implementing entity can perform a data availability analysis on the candidate impact factor set to obtain the impact factor set to be analyzed. The impact factor set to be analyzed consists of candidate impact factors that meet the preset data integrity conditions. The preset data integrity conditions can be used to characterize whether the data missingness of the candidate impact factors in multiple sample periods is within an acceptable range, so as to ensure the quality and integrity of the data used in subsequent correlation analysis and fusion processing.
[0041] In one possible embodiment, the executing entity can statistically analyze the data missing rate and completeness of each indicator in the candidate impact factor set over multiple sample periods, and set a removal threshold of 20%. For candidate impact factors with a missing rate exceeding the removal threshold, the executing entity can remove them from the candidate impact factor set. For candidate impact factors with a missing rate not exceeding the removal threshold, the executing entity can retain them in the impact factor set to be analyzed.
[0042] Specifically, following the previous example, for the 16 initially selected candidate influencing factors, the implementing entity statistically analyzed the data missingness of these factors over 10 sample periods from 2011 to 2020, and set the elimination threshold at 20%. Statistical analysis showed that road density (X13) and forest coverage (X14) had data missingness for more than three consecutive years, so X13 and X14 were eliminated. The remaining 14 indicators had relatively complete data and were retained as the set of influencing factors to be analyzed for subsequent correlation analysis.
[0043] Step 1022: Based on the carbon emission sample set, perform correlation analysis on the set of influencing factors to be analyzed to obtain the first influencing factor and the second influencing factor.
[0044] Here, after the implementing entity performs data availability analysis on the candidate impact factor set to obtain the impact factor set to be analyzed, it can perform correlation analysis on the impact factor set to be analyzed based on the carbon emission sample set to obtain the first impact factor and the second impact factor. The first impact factor can be understood as the impact factor that is significantly related to carbon emissions, and the second impact factor can be understood as the impact factor that has insufficient correlation but still has value for further fusion analysis.
[0045] In one possible embodiment, a correlation analysis is performed on the set of influencing factors to be analyzed based on a carbon emission sample set to obtain a first influencing factor and a second influencing factor, including the following steps: Using the carbon emissions corresponding to the carbon emission sample set as the dependent variable, the correlation between each influencing factor to be analyzed and the carbon emissions was calculated using the Pearson correlation coefficient. The influencing factors whose correlation meets the preset correlation threshold are identified as the first influencing factors. From the influencing factors whose correlation does not meet the preset correlation threshold, the influencing factors that meet the preset correlation conditions are identified, and the second influencing factor is obtained.
[0046] Specifically, after conducting data availability analysis on the candidate influencing factor set and obtaining the set of influencing factors to be analyzed, the implementing entity first uses carbon emissions E as the dependent variable and calculates the correlation between each influencing factor to be analyzed and carbon emissions E using the Pearson correlation coefficient, setting a correlation threshold r=0.9. Then, the implementing entity determines the influencing factors to be analyzed that meet the correlation threshold as the first influencing factor. From the influencing factors to be analyzed that do not meet the correlation threshold, it further determines the influencing factors to be analyzed that meet the preset correlation conditions as the second influencing factor. The Pearson correlation coefficient is used to characterize the strength of the linear relationship between variables. The closer the absolute value of the calculated Pearson correlation coefficient is to 1, the stronger the linear relationship between the two variables; the closer it is to 0, the weaker the linear relationship. The calculation formula is as follows:
[0047] in, and Let be the i-th observation of the two variables in the sample. and These are the means of the corresponding variables, and n is the sample size.
[0048] In one possible embodiment, the executing entity uses the carbon emissions E corresponding to the carbon emission sample set as the dependent variable and calculates the correlation between each influencing factor to be analyzed and the carbon emissions using the Pearson correlation coefficient. For example, the correlation analysis results are shown in the table below:
[0049] Among them, the correlation threshold r=0.9, the absolute value of the Pearson correlation coefficient for each stage of fuel consumption X1 is 0.946, gas extraction X2 is 0.910, purchased electricity consumption X3 is 0.931, purchased heat consumption X4 is 0.828, annual water consumption X5 is 0.698, average coal transportation distance X6 is 0.914, total regional GDP X7 is 0.926, the proportion of secondary industry X8 is 0.755, total energy consumption X9 is 0.810, the proportion of coal in the energy structure X10 is 0.925, the size of the resident population X11 is 0.931, and total electricity consumption X12 is... The comprehensive utilization rate of equipment (X15) is 0.784, the average annual operating hours (X16) is 0.582, and the average annual operating hours (X16) is 0.630. Therefore, the implementing entity can identify the seven indicators (X1, X2, X3, X6, X7, X10, and X11) that are significantly related to carbon emissions as the first influencing factor. The five indicators (X4, X5, X8, X9, and X12) that do not show a significant linear correlation but can comprehensively reflect the potential coupling relationship of the carbon emission process from multiple dimensions such as energy structure, resource consumption, and industrial composition can be identified as the second influencing factor. The remaining indicators with weak correlation and that do not meet the preset correlation conditions will not be included in the subsequent fusion process.
[0050] Step 1023: Perform fusion processing on the second impact factor to obtain the composite impact factor.
[0051] Here, after the implementing entity performs correlation analysis on the set of influencing factors to be analyzed based on the carbon emission sample set to obtain the first and second influencing factors, it can perform fusion processing on the second influencing factors to obtain composite influencing factors. The composite influencing factor can be understood as a comprehensive characterizing variable formed by feature fusion of multiple second influencing factors with insufficient correlation but potential coupling relationship. It is used to reflect the overall changing trend of this type of influencing factor in the sample space. Through this step, the influencing factors with insufficient correlation but still of practical significance can be avoided from being directly discarded, thereby improving the ability of the input variables to characterize the carbon emission process.
[0052] In one possible embodiment, the second impact factor is fused to obtain a composite impact factor, including the following steps: The second impact factor was standardized to obtain the third impact factor. K-means cluster analysis was performed on the third influencing factor to obtain the cluster analysis results; Based on the cluster analysis results, the distance results are calculated; The distance results are normalized to obtain normalized distance results, and a composite influence factor is constructed based on the normalized distance results.
[0053] Specifically, after obtaining the second influencing factor, the implementing entity first performs Z-score standardization on the sample data corresponding to the second influencing factor to eliminate the influence of dimensions, thus obtaining the third influencing factor. Then, the implementing entity uses the K-means clustering method to perform cluster analysis on the third influencing factor, and determines the optimal number of clusters K=2 through silhouette coefficient analysis, thus obtaining the cluster analysis results. Afterwards, based on the cluster analysis results, the implementing entity calculates the Euclidean distance between each sample point and the center of its cluster, thus obtaining the distance results. Finally, the implementing entity performs range normalization on the distance results to obtain normalized distance results, and constructs a composite influencing factor Z1 based on the normalized distance results to characterize the overall trend of this type of indicator in the sample space.
[0054] In one possible embodiment, following the previous example, this embodiment selects X4, X5, X8, X9, and X12 as the second influencing factors and introduces K-means clustering for feature fusion to uncover their collaborative variation patterns in the sample space. Specifically, firstly, the executing entity performs Z-score standardization on the sample data of the above five indicators to eliminate the influence of dimensions. Subsequently, K-means clustering is used, and the optimal number of clusters K=2 is determined through silhouette coefficient analysis. K-means clustering achieves data aggregation by iteratively adjusting the cluster centers by minimizing the sum of squared distances between sample points within a cluster and the cluster centers. Its objective function is:
[0055] Where K is the number of clusters, C k For the k-th cluster, x i For sample points, μ k Let Z be the centroid of the k-th cluster. Based on this, the executing entity performs cluster analysis on all sample data, further calculates the Euclidean distance between each sample point and the center of its cluster, and performs range normalization on the distance results. Finally, a fusion variable Z1 is constructed to characterize the overall trend of this type of indicator in the sample space. After the above fusion process, the executing entity can obtain the fusion variable Z1. Z1 is included as a new composite input factor in the set of key carbon emission influencing factors and participates in the weighting model as one of the input indicators in the subsequent original indicator sample matrix.
[0056] Step 1024: Based on the first impact factor and the composite impact factor, determine the key carbon emission influencing factors.
[0057] Here, after the implementing entity fuses the second influencing factor to obtain the composite influencing factor, it can determine the key carbon emission influencing factors based on the first influencing factor and the composite influencing factor. The key carbon emission influencing factors include the first influencing factor that is significantly related to carbon emissions, as well as the composite influencing factor formed by the fusion of the second influencing factor. This takes into account both the direct characterization effect of highly correlated influencing factors on carbon emission changes and the comprehensive characterization effect of influencing factors with low correlation but potential coupling relationship on carbon emission changes.
[0058] In one possible embodiment, the executing entity can identify eight factors as key carbon emission influencing factors: first influencing factors X1, X2, X3, X6, X7, X10, and X11, and a composite influencing factor Z1 obtained by fusing second influencing factors X4, X5, X8, X9, and X12. Here, X1 represents fuel consumption at each stage, X2 represents gas extraction, X3 represents purchased electricity consumption, X6 represents the average coal transportation distance, X7 represents the total regional GDP, X10 represents the proportion of coal in the energy structure, X11 represents the size of the resident population, and Z1 represents the composite input factor obtained by fusing the second influencing factors. Thus, the executing entity completes the screening and fusion of the candidate influencing factor set and uses these eight key carbon emission influencing factors as the input basis for subsequent weighting processing and neural network model training.
[0059] In this embodiment, firstly, the executing entity performs data availability analysis on the candidate influencing factor set to obtain the influencing factor set to be analyzed; then, the executing entity performs correlation analysis on the influencing factor set to be analyzed based on the carbon emission sample set to obtain the first influencing factor and the second influencing factor; subsequently, the executing entity performs fusion processing on the second influencing factor to obtain the composite influencing factor; finally, the executing entity determines the key carbon emission influencing factors based on the first influencing factor and the composite influencing factor.
[0060] As described above, this embodiment, by performing data availability analysis, correlation analysis, and fusion processing on the candidate influencing factor set, can screen out the first influencing factor that is significantly related to carbon emissions while ensuring data quality and integrity. It also performs fusion processing on the second influencing factor, which has insufficient correlation but possesses comprehensive characterization value, thereby obtaining the key carbon emission influencing factors. Thus, this embodiment not only reduces the interference of irrelevant or inefficient factors on subsequent modeling but also retains information from low-correlation factors that has a potential coupling relationship with the carbon emission process, improving the accuracy of key influencing factor identification and the characterization ability of input features, providing a reliable foundation for subsequent weighting processing and the construction of carbon emission prediction models.
[0061] Step 103: Assign weights to the key carbon emission influencing factors to obtain the factor weights corresponding to each key carbon emission influencing factor.
[0062] Here, the implementing entity screens and merges the candidate influencing factor set based on the carbon emission sample set to obtain the key carbon emission influencing factors. Then, it assigns weights to these key carbon emission influencing factors to obtain the factor weights corresponding to each key carbon emission influencing factor. The factor weights can be used to characterize the importance of the corresponding key carbon emission influencing factor to the input of the carbon emission prediction model. Through this step, the weights of each key carbon emission influencing factor can be determined based on the objective distribution characteristics of the sample data, avoiding reliance on manual experience to set the weights, thereby improving the rationality of the subsequent model input structure.
[0063] In one possible embodiment, such as Figure 3 As shown, the key carbon emission influencing factors are weighted to obtain the factor weights corresponding to each key carbon emission influencing factor, including the following steps: Step 1031: Construct the original indicator sample matrix corresponding to the key carbon emission influencing factors, and normalize the original indicator sample matrix to obtain the normalized indicator matrix.
[0064] Here, the implementing entity screens and merges the candidate influencing factor set based on the carbon emission sample set to obtain the key carbon emission influencing factors. Then, it can construct the original indicator sample matrix corresponding to the key carbon emission influencing factors and normalize the original indicator sample matrix to obtain the normalized indicator matrix. The original indicator sample matrix can be used to characterize the original values of each key carbon emission influencing factor in each sample period, while the normalized indicator matrix can be used to eliminate the dimensional influence and order of magnitude differences between different indicators, so as to facilitate the subsequent calculation of the proportion matrix and entropy value.
[0065] In one possible embodiment, following the previous example, the implementing entity can construct an original indicator sample matrix from the aforementioned eight key carbon emission influencing factors, namely X1, X2, X3, X6, X7, X10, X11, and the fused composite indicator Z1. Specifically, assuming the number of samples is m and the number of indicators is n, the original indicator sample matrix can be expressed as:
[0066] in, Here, m represents the original value of the j-th indicator for the i-th sample, m is the total number of samples, and n is the total number of indicators. In this embodiment, the number of samples m=10 and the number of indicators n=8. The specific original data can be found in the table below:
[0067] Furthermore, the implementing entity can perform range normalization on the original indicator sample matrix to obtain a normalized indicator matrix. The normalization formula can be expressed as:
[0068] in, These are the standardized sample values. This represents the maximum value of the j-th indicator in all samples. This represents the minimum value of the j-th indicator among all samples.
[0069] In this embodiment, the executing entity uses the sample data from 2011 to 2020 in the table above as a basis, and takes X1 (fuel consumption), X2 (gas extraction), X3 (purchased electricity), X6 (coal transportation distance), X7 (regional GDP), X10 (coal share in energy structure), X11 (resident population size) and Z1 (composite indicator) as input indicators. Together with carbon emissions E, they form the basis for subsequent modeling. Among them, X1 to X11 and Z1 constitute the eight input indicators in the original indicator sample matrix, and E represents the carbon emissions for the corresponding sample period, which is used as the output variable in the subsequent neural network model training.
[0070] Step 1032: Construct a proportion matrix based on the normalized index matrix, and determine the entropy value corresponding to each key carbon emission influencing factor based on the proportion matrix.
[0071] Here, after the executing entity normalizes the original indicator sample matrix to obtain the normalized indicator matrix, it can construct a proportion matrix based on the normalized indicator matrix and determine the entropy value corresponding to each key carbon emission influencing factor based on the proportion matrix. The proportion matrix is used to characterize the relative proportion of each sample value under each indicator dimension, and the entropy value is used to reflect the information dispersion of the corresponding indicator.
[0072] In one possible embodiment, the executing entity can base its actions on a normalized index matrix. Calculate the scale matrix Its calculation formula can be expressed as:
[0073] in, This represents the proportion of the i-th sample under the j-th indicator.
[0074] Furthermore, the implementing entity can calculate the entropy value corresponding to each indicator based on the proportion matrix. Its calculation formula can be expressed as:
[0075] in, Let m represent the entropy value of the j-th indicator, and m represent the total number of samples.
[0076] In this embodiment, following the previous example, the executing entity can construct a proportion matrix based on the original sample data in the table above after range normalization, and further calculate the entropy values corresponding to each of the eight key carbon emission influencing factors.
[0077] Step 1033: Calculate the redundancy of each key carbon emission influencing factor based on the entropy value.
[0078] Here, after determining the entropy value corresponding to each key carbon emission influencing factor based on the proportion matrix, the implementing entity can calculate the redundancy corresponding to each key carbon emission influencing factor based on the entropy value. The redundancy can be used to characterize the amount of effective information provided by the corresponding indicator. Generally speaking, the smaller the entropy value of an indicator, the greater the information difference of the indicator, and the higher its corresponding redundancy.
[0079] In one possible implementation, the executing entity can base its actions on the entropy values corresponding to each metric. Calculate its redundancy The calculation formula can be expressed as:
[0080] in, The redundancy of the j-th indicator is represented by the redundancy value. In this embodiment, the executing entity can calculate the redundancy value of each of the aforementioned eight key carbon emission influencing factors based on their respective entropy values, for use in subsequent weight calculation.
[0081] Step 1034: Determine the factor weights corresponding to each key carbon emission influencing factor based on redundancy.
[0082] Here, after the implementing entity calculates the redundancy of each key carbon emission influencing factor based on the entropy value, it can determine the factor weight of each key carbon emission influencing factor based on the redundancy. The factor weight can be used to characterize the relative importance of the corresponding key carbon emission influencing factor in the input indicator system, thereby forming the weight basis for the subsequent weighted input feature set.
[0083] In one possible implementation, the executing entity can base its actions on the redundancy corresponding to each metric. Calculate its weight The calculation formula can be expressed as:
[0084] in, This represents the factor weight of the j-th indicator, and n represents the total number of indicators.
[0085] Specifically, in this embodiment, the executing entity can obtain the entropy weight method calculation results as shown in the table below based on the aforementioned calculation results:
[0086] Among them, X1 has an entropy of 0.9178, a redundancy of 0.0822, and a weight of 0.0788; X2 has an entropy of 0.8575, a redundancy of 0.1425, and a weight of 0.1366; X3 has an entropy of 0.8009, a redundancy of 0.1991, and a weight of 0.1909; and X6 has an entropy of 0.8918, a redundancy of 0.1082, and a weight of 0.1037. The entropy value of X7 is 0.8848, the redundancy is 0.1152, and the weight is 0.1104; the entropy value of X10 is 0.8812, the redundancy is 0.1188, and the weight is 0.1139; the entropy value of X11 is 0.9074, the redundancy is 0.0926, and the weight is 0.0888; and the entropy value of Z1 is 0.8156, the redundancy is 0.1844, and the weight is 0.1768. Therefore, the implementing entity can objectively assign weights to key carbon emission influencing factors and form a weighting system for subsequent input indicators.
[0087] In one possible embodiment, the factor weights can also support a dynamic update mechanism, that is, when the input indicator data changes with time, operating conditions or system parameters, the executing entity can adaptively adjust the factor weights corresponding to each key carbon emission influencing factor to improve the applicability of the subsequent model in different operating stages and different scenarios.
[0088] In this embodiment, firstly, the executing entity constructs an original indicator sample matrix corresponding to the key carbon emission influencing factors, and normalizes the original indicator sample matrix to obtain a normalized indicator matrix; then, the executing entity constructs a proportion matrix based on the normalized indicator matrix, and determines the entropy value corresponding to each key carbon emission influencing factor based on the proportion matrix; subsequently, the executing entity calculates the redundancy corresponding to each key carbon emission influencing factor based on the entropy value; finally, the executing entity determines the factor weight corresponding to each key carbon emission influencing factor based on the redundancy.
[0089] As described above, this embodiment, by constructing an original index sample matrix and combining entropy, redundancy, and weight calculation processes, can determine the importance of each key carbon emission influencing factor based on the objective distribution characteristics of the sample data, avoiding biases caused by subjective weighting, thereby improving the rationality and objectivity of the input index weight system and providing a reliable foundation for the subsequent construction of weighted input feature sets and the training of carbon emission prediction models.
[0090] Step 104: Based on the weights of each factor, the key carbon emission influencing factors are weighted to obtain a weighted input feature set. The neural network model is then trained based on the weighted input feature set and the carbon emission sample set to obtain a carbon emission prediction model.
[0091] Here, after assigning weights to key carbon emission influencing factors to obtain the corresponding factor weights for each key carbon emission influencing factor, the implementing entity can further weight the key carbon emission influencing factors based on these factor weights to obtain a weighted input feature set. The neural network model is then trained based on this weighted input feature set and the carbon emission sample set to obtain a carbon emission prediction model. The weighted input feature set can be understood as the set of input features formed by combining key carbon emission influencing factors with their corresponding factor weights, used to characterize the relative importance of each key carbon emission influencing factor in the model input. It should be noted that since the factor weights corresponding to each key carbon emission influencing factor have already been determined using the entropy weight method in step 103, and the weighted and optimized indicator system is used as input to construct and train the BP neural network model in step 104, the final prediction model established in this embodiment is an entropy weight-BP neural network model. The BP neural network model is used to characterize the network model ontology, and the entropy weight-BP neural network model is used to characterize the overall prediction model after incorporating the entropy weighting mechanism.
[0092] In one possible embodiment, the executing entity weights the key carbon emission influencing factors based on the weights of each factor to obtain a weighted input feature set. Specifically, the executing entity can combine the sample values of each key carbon emission influencing factor with its corresponding factor weight to obtain weighted input feature values. For example, taking the weight results given in the table in step 103, if the purchased electricity consumption X3 of a certain sample is 90, and the factor weight corresponding to X3 is 0.1909, then the weighted input feature value of this sample in the X3 dimension can be calculated as 0.1909 × 90 = 17.181. It should be understood that this embodiment is only an exemplary weighting implementation. In other embodiments, other weighting combinations adapted to the factor weights can also be used to form a weighted and optimized input indicator system. Thus, the executing entity can construct a weighted input feature set based on the weighted input feature values of all samples, serving as the input data basis for the subsequent entropy-weighted BP neural network model.
[0093] In one possible embodiment, the neural network model is a backpropagation neural network model, i.e., a BP neural network model, such as... Figure 4 As shown, a neural network model is trained based on a weighted input feature set and a carbon emission sample set to obtain a carbon emission prediction model, including the following steps: Step 1041: Construct the backpropagation neural network model.
[0094] Here, after the implementing entity performs weighted processing on the key carbon emission influencing factors based on the weights of each factor to obtain a weighted input feature set, it can construct a backpropagation neural network model. The number of input layer nodes in the backpropagation neural network model corresponds to the number of features in the weighted input feature set. The backpropagation neural network model is a multi-layer feedforward neural network model, namely a BP neural network model, which can update the model parameters through the error backpropagation mechanism, thereby gradually reducing the prediction error.
[0095] In one possible embodiment, the BP neural network model in this embodiment adopts a three-layer feedforward neural network, such as... Figure 5 As shown, Figure 5 An exemplary diagram of the BP neural network model topology is shown. The number of nodes in the input layer of the BP neural network model is set to 8, corresponding to the 8 key carbon emission influencing factors identified above. The number of nodes in the hidden layer is based on an empirical formula. The number of nodes was set to 5 based on experimental optimization, where m is the number of input nodes, n is the number of output nodes, a is a constant, and the number of output layer nodes was set to 1 to output the predicted carbon emission value. Furthermore, the hidden layer of the BP neural network model uses the ReLU activation function, and the output layer uses the linear activation function. To improve training stability, the input features were normalized using Min-Max.
[0096] Step 1042: Input the weighted input feature set into the backpropagation neural network model, and use the carbon emissions corresponding to the carbon emission sample set as the target output to train the backpropagation neural network model to obtain the carbon emission prediction model.
[0097] Here, after constructing the backpropagation neural network model, the executing entity can input the weighted input feature set into the backpropagation neural network model and use the carbon emissions corresponding to the carbon emission sample set as the target output to train the backpropagation neural network model, thereby obtaining a carbon emission prediction model. The training process may include forward propagation to calculate predicted values, calculating errors based on the loss function, and updating network parameters through the error backpropagation algorithm until model training is complete. It should be noted that since the input to the backpropagation neural network model is the weighted and optimized input index system constructed based on the entropy weighting method in step 103, the final carbon emission prediction model obtained in this embodiment is an entropy weight-BP neural network model. In other words, the backpropagation neural network model is used to represent the network model itself, while the entropy weight-BP neural network model is used to represent the overall prediction model constructed and trained based on the entropy weighting mechanism.
[0098] In one possible implementation, the backpropagation algorithm is used for weight updates during training. The loss function is mean squared error (MSE), the optimizer is Adam, the learning rate is set to 0.01, the maximum number of iterations is set to 2000, the batch size is set to 4, and an early stopping mechanism is introduced. When the loss does not improve significantly in 50 consecutive training rounds and the threshold is set to 1e-4, the training is terminated early and the optimal model weights are restored. Thus, the execution entity can complete the training of the entropy weight-BP neural network model.
[0099] In one possible embodiment, taking into account the small sample data characteristics of this embodiment, the executing entity can use the 5-fold cross-validation method to train and test the entropy weight-BP neural network model. Specifically, all samples can be randomly divided into 5 equal-sized subsets. In each iteration, one subset is selected as the validation set, and the remaining 4 subsets are used for model training. A total of 5 rounds of training and validation are carried out. Finally, the average value of the 5 rounds of validation results is used as the comprehensive performance evaluation index of the entropy weight-BP neural network model.
[0100] Furthermore, to verify the effectiveness of the entropy-weighted BP neural network model established in this embodiment, the executing entity can train and evaluate the traditional BP neural network model and the entropy-weighted BP neural network model established in this embodiment under the same 5-fold cross-validation conditions. The model performance comparison is shown in the table below:
[0101] Specifically, the mean absolute error (MAE) of the BP neural network model is 50352.78tCO2e, the root mean square error (RMSE) is 56746.29tCO2e, and the mean absolute percentage error (MAPE) is 2.49%; the mean absolute error (MAE) of the entropy weight-BP neural network model is 24211.46tCO2e, the root mean square error (RMSE) is 28820.56tCO2e, and the mean absolute percentage error (MAPE) is 1.20%. As shown in the table above, both models can fit the carbon emission data of coal enterprises well. However, the entropy weight-BP neural network model established in this embodiment outperforms the traditional BP neural network model in all three evaluation indicators (MAE, RMSE, and MAPE). This indicates that after introducing the feature weights calculated by the entropy weight method, the model's sensitivity to key influencing factors and its fitting ability are significantly improved. This result further verifies the applicability and superiority of the entropy weight-BP neural network model established in this embodiment in predicting carbon emissions from coal enterprises.
[0102] In this embodiment, firstly, the executing entity weights the key carbon emission influencing factors based on the weights of each factor to obtain a weighted input feature set; then, the executing entity constructs a backpropagation neural network model; subsequently, the executing entity inputs the weighted input feature set into the backpropagation neural network model and uses the carbon emission amount corresponding to the carbon emission sample set as the target output to train the backpropagation neural network model; finally, the executing entity obtains the entropy weight-BP neural network carbon emission prediction model.
[0103] As described above, this embodiment introduces the factor weights determined by the entropy weight method into the training process of the BP neural network model to establish an entropy weight-BP neural network model. This model can improve the sensitivity of the model to key carbon emission influencing factors and its ability to fit complex nonlinear relationships while taking into account the objective weight distribution of input factors. At the same time, combined with the 5-fold cross-validation and the model performance evaluation results shown in the table above, it can be seen that the entropy weight-BP neural network model established in this embodiment has smaller prediction errors and higher prediction accuracy than the traditional BP neural network model, thus providing more reliable model support for the analysis and dynamic prediction of carbon emission trends of target objects.
[0104] This disclosure provides a method, apparatus, device, medium, and product for constructing a carbon emission prediction model. First, based on the carbon emission activity data of the target object, a carbon emission sample set and a candidate influencing factor set are constructed. The candidate influencing factor set is used to characterize candidate factors affecting changes in the target object's carbon emissions. Then, based on the carbon emission sample set, the candidate influencing factor set is screened and fused to obtain key carbon emission influencing factors. Next, the key carbon emission influencing factors are weighted to obtain the factor weights corresponding to each key carbon emission influencing factor. Finally, based on the factor weights, the key carbon emission influencing factors are weighted to obtain a weighted input feature set. A neural network model is trained based on the weighted input feature set and the carbon emission sample set to obtain a carbon emission prediction model.
[0105] As described above, the embodiments of this disclosure first screen and fuse candidate influencing factor sets based on a carbon emission sample set to obtain key carbon emission influencing factors, thereby reducing the bias caused by subjective selection of input factors and improving the accuracy of the input factors in representing changes in carbon emissions. Second, the embodiments of this disclosure assign weights to the key carbon emission influencing factors to obtain the factor weights corresponding to each key carbon emission influencing factor, thereby establishing an objective weighting mechanism based on data features and improving the rationality of the model input structure. Finally, the embodiments of this disclosure train the neural network model based on a weighted input feature set, which can improve the carbon emission prediction model's ability to characterize complex nonlinear relationships and its adaptability to data from different stages and operating conditions, thereby improving the accuracy, stability, and engineering practicality of carbon emission prediction.
[0106] By dividing each functional module according to its corresponding function, this disclosure provides an apparatus for constructing a carbon emission prediction model. The apparatus for constructing a carbon emission prediction model can be a server or a chip applied to a server. Figure 6 A schematic block diagram of the functional modules of an apparatus for constructing a carbon emission prediction model provided as an exemplary embodiment of this disclosure. Figure 6 As shown, the apparatus for constructing this carbon emission prediction model includes: The construction module 601 is used to construct a carbon emission sample set and a candidate influencing factor set based on the carbon emission activity data of the target object; the candidate influencing factor set is used to characterize the candidate factors affecting the carbon emission changes of the target object. The first processing module 602 is used to screen and fuse the candidate influencing factor set based on the carbon emission sample set to obtain key carbon emission influencing factors. The second processing module 603 is used to assign weights to the key carbon emission influencing factors to obtain the factor weights corresponding to each key carbon emission influencing factor. The training module 604 is used to perform weighted processing on the key carbon emission influencing factors based on the weights of each factor to obtain a weighted input feature set, and to train the neural network model based on the weighted input feature set and the carbon emission sample set to obtain a carbon emission prediction model.
[0107] In one embodiment, the first processing module 602 includes: The first processing unit is used to perform data availability analysis on the candidate impact factor set to obtain the impact factor set to be analyzed; wherein, the impact factor set to be analyzed consists of candidate impact factors that meet preset data integrity conditions; The second processing unit is used to perform correlation analysis on the set of influencing factors to be analyzed based on the carbon emission sample set to obtain the first influencing factor and the second influencing factor. The fusion unit is used to fuse the second impact factor to obtain the composite impact factor; The first determining unit is used to determine the key carbon emission influencing factors based on the first influencing factor and the composite influencing factor.
[0108] In one embodiment, the first processing module 602 includes: The first calculation unit is used to calculate the correlation between each of the influencing factors to be analyzed and the carbon emissions using the carbon emissions corresponding to the carbon emission sample set as the dependent variable and the Pearson correlation coefficient. The second determining unit is used to determine the influence factors to be analyzed that satisfy the preset correlation threshold as the first influence factor. The third determining unit is used to determine the second influence factor from the influence factors to be analyzed that meet the preset correlation conditions from the influence factors whose correlation does not meet the preset correlation threshold.
[0109] In one embodiment, the first processing module 602 includes: The third processing unit is used to standardize the second impact factor to obtain the third impact factor; The fourth processing unit is used to perform K-means clustering analysis on the third influencing factor to obtain the clustering analysis results; The second calculation unit is used to calculate the distance result based on the clustering analysis result; The first construction unit is used to normalize the distance results to obtain normalized distance results, and to construct the composite influence factor based on the normalized distance results.
[0110] In one embodiment, the second processing module 603 includes: The second construction unit is used to construct the original indicator sample matrix corresponding to the key carbon emission influencing factors, and to normalize the original indicator sample matrix to obtain the normalized indicator matrix. The fourth determining unit is used to construct a proportion matrix based on the normalized index matrix, and to determine the entropy value corresponding to each of the key carbon emission influencing factors based on the proportion matrix. The third calculation unit is used to calculate the redundancy of each of the key carbon emission influencing factors based on the entropy value. The fifth determining unit is used to determine the factor weights corresponding to each of the key carbon emission influencing factors based on the redundancy.
[0111] In one embodiment, training module 604 includes: The third construction unit is used to construct the backpropagation neural network model; wherein the number of input layer nodes of the backpropagation neural network model corresponds to the number of features in the weighted input feature set; The training unit is used to input the weighted input feature set into the backpropagation neural network model, and use the carbon emissions corresponding to the carbon emission sample set as the target output to train the backpropagation neural network model to obtain the carbon emission prediction model.
[0112] Figure 7 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this disclosure. For example... Figure 7 As shown, the electronic device 700 includes at least one processor 701 and a memory 702 coupled to the processor 701. The processor 701 can perform the corresponding steps in the methods disclosed in the embodiments of this disclosure.
[0113] The processor 701 described above can also be called a central processing unit (CPU), which can be an integrated circuit chip with signal processing capabilities. Each step in the method disclosed in this embodiment can be implemented by the integrated logic circuitry in the processor 701 or by software instructions. The processor 701 can be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this embodiment can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can be located in the memory 702, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The processor 701 reads information from the memory 702 and, in conjunction with its hardware, completes the steps of the method described above.
[0114] Furthermore, various operations / processes according to this disclosure, implemented via software and / or firmware, can be transmitted from a storage medium or network to a computer system with a dedicated hardware architecture, such as... Figure 8 The computer system 800 shown is equipped with the programs that constitute the software. When various programs are installed, the computer system is able to perform various functions, including those described above. Figure 8 A block diagram of a computer system provided for an exemplary embodiment of this disclosure.
[0115] Computer system 800 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0116] like Figure 8As shown, the computer system 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 may also store various programs and data required for the operation of the computer system 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0117] Multiple components in the computer system 800 are connected to the I / O interface 805, including: an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 can be any type of device capable of inputting information into the computer system 800. The input unit 806 can receive input numerical or character information and generate key signal inputs related to user settings and / or function control of the electronic device. The output unit 807 can be any type of device capable of presenting information and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. The storage unit 808 may include, but is not limited to, a hard disk and an optical disk. The communication unit 809 allows the computer system 800 to exchange information / data with other devices via a network such as the Internet, and may include, but is not limited to, a modem, network card, infrared communication device, wireless communication transceiver, and / or chipset, such as Bluetooth™ device, WiFi device, WiMax device, cellular communication device, and / or the like.
[0118] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above. For example, in some embodiments, the methods disclosed in this disclosure can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 700 via ROM 802 and / or communication unit 809. In some embodiments, the computing unit 801 can be configured to perform the methods disclosed in this disclosure by any other suitable means (e.g., by means of firmware).
[0119] This disclosure also provides a computer-readable storage medium, wherein when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is able to perform the methods disclosed in this disclosure.
[0120] The computer-readable storage medium in this disclosure can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. The aforementioned computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specifically, the aforementioned computer-readable storage medium may include electrical connections based on one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0121] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0122] Figure 9 A computer program product 900 is provided as an exemplary embodiment of the present disclosure. The computer program product 900 includes a computer program 901, wherein the computer program 901, when executed by a processor, implements the methods disclosed in the embodiments of the present disclosure.
[0123] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer.
[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0125] The modules, components, or units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules, components, or units do not necessarily constitute a limitation on the module, component, or unit itself.
[0126] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0127] The above description is merely an embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0128] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. A method for constructing a carbon emission prediction model, characterized in that, The method includes: Based on the carbon emission activity data of the target object, a carbon emission sample set and a candidate influencing factor set are constructed; the candidate influencing factor set is used to characterize the candidate factors affecting the carbon emission changes of the target object. Based on the carbon emission sample set, the candidate influencing factor set is screened and fused to obtain key carbon emission influencing factors; The key carbon emission influencing factors are weighted to obtain the factor weights corresponding to each key carbon emission influencing factor. The key carbon emission influencing factors are weighted according to the weights of each factor to obtain a weighted input feature set. The neural network model is then trained based on the weighted input feature set and the carbon emission sample set to obtain a carbon emission prediction model.
2. The method according to claim 1, characterized in that, The process of screening and fusing the candidate influencing factor set based on the carbon emission sample set yields key carbon emission influencing factors, including: A data availability analysis is performed on the candidate impact factor set to obtain the impact factor set to be analyzed; wherein, the impact factor set to be analyzed consists of candidate impact factors that meet preset data integrity conditions; Based on the carbon emission sample set, a correlation analysis is performed on the set of influencing factors to be analyzed to obtain the first influencing factor and the second influencing factor. The second impact factor is fused to obtain the composite impact factor; Based on the first influencing factor and the composite influencing factor, the key carbon emission influencing factors are determined.
3. The method according to claim 2, characterized in that, The correlation analysis of the set of influencing factors to be analyzed based on the carbon emission sample set to obtain the first influencing factor and the second influencing factor includes: Using the carbon emissions corresponding to the carbon emission sample set as the dependent variable, the correlation between each of the influencing factors to be analyzed and the carbon emissions is calculated using the Pearson correlation coefficient. The influencing factors whose correlation satisfies the preset correlation threshold are determined as the first influencing factor; The second influence factor is obtained by identifying the influence factors that meet the preset correlation conditions from the influence factors whose correlation does not meet the preset correlation threshold.
4. The method according to claim 2, characterized in that, The process of fusing the second impact factor to obtain a composite impact factor includes: The second impact factor is standardized to obtain the third impact factor; K-means cluster analysis was performed on the third influencing factor to obtain the cluster analysis results; Based on the clustering analysis results, the distance results are calculated; The distance results are normalized to obtain normalized distance results, and the composite influence factor is constructed based on the normalized distance results.
5. The method according to claim 1, characterized in that, The weighting process for the key carbon emission influencing factors, to obtain the factor weights corresponding to each key carbon emission influencing factor, includes: Construct the original index sample matrix corresponding to the key carbon emission influencing factors, and normalize the original index sample matrix to obtain the normalized index matrix. A proportion matrix is constructed based on the normalized index matrix, and the entropy value corresponding to each of the key carbon emission influencing factors is determined based on the proportion matrix. The redundancy of each of the key carbon emission influencing factors is calculated based on the entropy value. The factor weights corresponding to each of the key carbon emission influencing factors are determined based on the redundancy.
6. The method according to claim 1, characterized in that, The neural network model is a backpropagation neural network model; The step of training a neural network model based on the weighted input feature set and the carbon emission sample set to obtain a carbon emission prediction model includes: Construct the backpropagation neural network model; wherein the number of input layer nodes of the backpropagation neural network model corresponds to the number of features in the weighted input feature set; The weighted input feature set is input into the backpropagation neural network model, and the carbon emissions corresponding to the carbon emission sample set are used as the target output to train the backpropagation neural network model to obtain the carbon emission prediction model.
7. An apparatus for constructing a carbon emission prediction model, characterized in that, include: The module is used to construct a carbon emission sample set and a candidate influencing factor set based on the carbon emission activity data of the target object. The candidate influencing factor set is used to characterize the candidate factors affecting the carbon emission changes of the target object; The first processing module is used to screen and fuse the candidate influencing factor set based on the carbon emission sample set to obtain key carbon emission influencing factors. The second processing module is used to assign weights to the key carbon emission influencing factors to obtain the factor weights corresponding to each key carbon emission influencing factor. The training module is used to weight the key carbon emission influencing factors based on the weights of each factor to obtain a weighted input feature set, and to train the neural network model based on the weighted input feature set and the carbon emission sample set to obtain a carbon emission prediction model.
8. An electronic device, characterized in that, include: At least one processor; Memory for storing the at least one processor-executable instruction; The at least one processor is configured to execute the instructions to implement the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.