A carbon emission prediction method considering master device difference metering characteristics

By using improved particle swarm optimization and random forest algorithm for data clustering, combined with deep belief network model and grey wolf optimization algorithm, the accuracy and stability problems of existing carbon emission prediction methods under complex data are solved, and high-precision carbon emission prediction and trend analysis are achieved.

CN118674092BActive Publication Date: 2025-10-17STATE GRID FUJIAN POWER ELECTRIC CO ECONOMIC RESEARCH INSTITUTE +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410663326.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-27
Publication Date
2025-10-17
Estimated Expiration
2044-05-27

AI Technical Summary

Technical Problem

Existing carbon emission prediction methods have difficulty achieving satisfactory prediction accuracy when processing complex and changeable industrial data, and fail to fully utilize advanced algorithms to automatically discover key influencing factors, resulting in insufficient data analysis accuracy and model performance.

Method used

An improved particle swarm algorithm is used for cluster analysis, combined with the random forest algorithm to identify factors affecting carbon emissions, and a deep belief network model is constructed for prediction. The model parameters are optimized by using the improved gray wolf optimization algorithm to ensure data quality and prediction accuracy.

Benefits of technology

It improves the accuracy and stability of carbon emission predictions and enhances the generalization ability of the model. It can effectively predict the carbon emission trends of individual devices and equipment clusters, and support the formulation of precise emission reduction strategies and optimization of production processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118674092B_ABST
    Figure CN118674092B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of carbon emission prediction methods considering transformer difference metering characteristics, comprising the following steps: collecting the carbon emission related basic data of different transformers, and the data is preprocessed;The carbon emission related basic data of different transformers after preprocessing is carried out cluster analysis by improved particle swarm algorithm, and the transformer data cluster of different categories is generated;Based on the influence degree of transformer carbon emission, the features of transformer data cluster are sorted, and the first g features after sorting are extracted as training features;Based on training feature extraction, the corresponding data in each transformer data cluster is extracted and constructed as training set;A transformer carbon emission prediction model is built, the model parameters are trained and optimized by training set, and the trained transformer carbon emission prediction model is obtained, and the carbon emission of transformer is predicted by the trained transformer carbon emission prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a carbon emission prediction method considering the differential metering characteristics of a main device and belongs to the technical field of carbon emission determination. BACKGROUND

[0002] With the severe challenge of global climate change, reducing carbon emissions has become the consensus of the international community, and accurate carbon emission prediction of main devices in industrial production has become the key to achieving the goal of energy saving and emission reduction. Traditional carbon emission prediction methods mostly rely on simple statistical models or basic machine learning algorithms. These methods often fail to achieve satisfactory prediction accuracy when dealing with complex and variable industrial data, especially when facing a large number of heterogeneous data sources, nonlinear relationships, and high-dimensional feature spaces. Therefore, there is an urgent need for a carbon emission prediction technology that can consider multiple factors, efficiently process big data, and have high prediction accuracy.

[0003] Prior art such as patent number CN117408380A discloses a device carbon emission prediction method and device based on big data and electronic equipment. The carbon emission of the device of the production type device group and the historical data of the infrared image of the device taken by the corresponding infrared thermal imager, and the historical data of the carbon emission of the device of the non-production type device group and the historical data of the infrared image of the device taken by the corresponding infrared thermal imager are used to train neural network models. The infrared image of the device is combined with the carbon emission of the device to predict the carbon emission. Since various factors during the use of the device will affect the carbon emission, including the temperature characteristics of the device, the slow rise of the temperature of the device during use and the accumulation of the temperature will have important influence on the carbon emission. By considering the infrared image of the device in the carbon emission prediction factors of the device, it is helpful to more accurately predict the carbon emission of the device.

[0004] However, in the above calculation scheme, only a single model is used, which is difficult to fully capture and express the complex dynamic characteristics of the carbon emission of the main device. At the same time, the data size and dimension used for model prediction are relatively low, and the data quality is not high, which affects the accuracy of subsequent analysis.

[0005] At the same time, in other prior art, when identifying the factors affecting carbon emission, it may rely too much on subjective judgment or simple statistical indicators, and fail to fully utilize advanced algorithms to automatically find key influencing factors. At the same time, the model optimization method is relatively traditional, and advanced optimization algorithms are not fully utilized to improve the performance of the prediction model. SUMMARY

[0006] In order to solve the problems existing in the prior art, the application proposes a carbon emission prediction method considering the differential metering characteristics of a transformer.

[0007] The technical scheme of the application is as follows:

[0008] In one aspect, the present application provides a carbon emission prediction method considering transformer difference measurement characteristics, comprising the following steps:

[0009] Collecting carbon emission related basic data of different transformers and preprocessing the data;

[0010] Performing cluster analysis on the preprocessed carbon emission related basic data of different transformers by an improved particle swarm algorithm to generate transformer data clusters of different categories;

[0011] Based on the random forest algorithm, the features of the transformer data clusters are sorted according to the influence degree of transformer carbon emission, the top g features after sorting are extracted as training features, and the corresponding data in each transformer data cluster is extracted based on the training features and constructed as a training set;

[0012] Constructing a transformer carbon emission prediction model, training and optimizing the model parameters through the training set, obtaining the trained transformer carbon emission prediction model, and predicting the carbon emission of the transformer through the trained transformer carbon emission prediction model.

[0013] As a preferred embodiment of the present application, the carbon emission related basic data of the transformer includes geographical position data and technical condition data of the transformer;

[0014] The geographical position data of the transformer includes installation site longitude and latitude coordinates, regional climate data, local policy and legal rule information;

[0015] The technical condition data includes manufacturing process, magnetic flux density, power loss, service life and recovery rate, cooling system information and operating conditions of the transformer core.

[0016] As a preferred embodiment of the present application, the preprocessing step of the carbon emission related basic data of the transformer is:

[0017] The box plot method is used to evaluate the outlying value measurement of the carbon emission related basic data of the transformer, and the specific steps are:

[0018] The carbon emission related basic data of the transformer is divided into four equal parts, and the maximum value, minimum value, median, first quartile and second quartile in the sample data are taken;

[0019] Based on the first quartile and the second quartile, a box plot is constructed to determine the outlying value standard Q, as shown in the following formula:

[0020]

[0021] Wherein: Q L is the first quartile; Q UQ is a second quartile; Q IQR Q is a quartile distance;

[0022] After the outlier in the data is screened by the box plot judgment outlier standard Q, a reasonable interval distribution model is constructed based on Chebyshev inequality, and the threshold interval of the data is calculated, and the specific steps are as follows:

[0023] The confidence level is defined So that at least Data samples fall within k standard deviations:

[0024]

[0025] Wherein: x represents a data sample; μ represents a data expectation; σ represents a data variance;

[0026] The target function f min (x) is constructed with the preset threshold interval length and the minimum difference of 90% proportion as the target:

[0027]

[0028] Wherein: Q(x) is all sample intervals obtained according to the preset threshold interval length; q(μ-kσ≤x≤μ+kσ) is the interval length falling into μ-kσ≤x≤μ+kσ;

[0029] Based on the target function value, the golden section method is used to optimize the value of k, and the threshold interval of the data is obtained based on the optimized value of k, the data outside the threshold interval is regarded as an abnormal value, and the abnormal value is repaired, that is, the abnormal value is set as the average value of the threshold interval.

[0030] As a preferred embodiment of the present application, the clustering analysis step of the different transformer carbon emission related basic data is:

[0031] Initialize particle swarm algorithm parameters;

[0032] Chaotic numbers are generated by chaotic mapping, the chaotic numbers are mapped into the data sample interval, the reverse position of the particle is constructed by reverse learning, and the fitness of the particle and its reverse position is compared, the position corresponding to the particle with the closest fitness to the reverse position of the particle is set as the initial position of the particle;

[0033] The particle starts searching from the initial position, iteratively updates the particle, and controls the boundary by varying the particle position through Cauchy;

[0034] The data sample is divided according to the nearest neighbor rule, and the fitness of the particle is recalculated;

[0035] The optimal position and optimal fitness of the particle individual and group before and after are compared and updated;

[0036] When the iteration stopping condition is reached, the iteration is stopped, and an optimal solution is obtained.

[0037] The optimal solution is taken as a clustering center, and the data samples are divided according to a nearest neighbor rule to obtain K clusters, each of which is a transformer data cluster.

[0038] As a preferred embodiment of the present application, the extraction step of the training features is:

[0039] The random forest algorithm constructs multiple decision trees, each of which draws part of the samples from all transformer data clusters as a training set with replacement, and randomly selects part of the features from all features of the transformer data cluster as candidate features;

[0040] Each decision tree is trained by its corresponding training set. When the decision tree is split each time, the optimal feature is selected as the split point in its candidate features based on the minimum square error criterion or the minimum Gini index criterion. When a preset stopping condition is reached, the training is completed.

[0041] The average impurity reduction of all features in all decision trees is calculated, and all features are sorted according to the size of the average impurity reduction, and the first n features are selected as the training features.

[0042] As a preferred embodiment of the present application, the transformer carbon emission prediction model is constructed based on a deep belief network model, and the deep belief network model includes multiple layers of restricted Boltzmann machines. The specific training steps are:

[0043] The feature vector of the training set is input into the first layer of the restricted Boltzmann machine for training;

[0044] When the first layer of the restricted Boltzmann machine has completed the learning of the features in the training set through training, the feature data is output, and these feature data are taken as the next layer of the restricted Boltzmann machine for continuous training;

[0045] Each subsequent layer of the restricted Boltzmann machine is trained through the above steps. When all the restricted Boltzmann machines are trained, the last layer of the restricted Boltzmann machine is taken as the output feature, and the local optimal parameters of each layer of the restricted Boltzmann machine after training are obtained.

[0046] Based on the output feature and the local optimal parameters of each layer of the restricted Boltzmann machine, the error back propagation algorithm is used to adjust the parameters of each layer of the restricted Boltzmann machine, and finally the trained deep belief network model is obtained.

[0047] As a preferred embodiment of the present application, the deep belief network model is optimized by improving the grey wolf optimization algorithm, and the specific steps are:

[0048] A set of solutions is randomly generated, each solution representing a configuration of connection weights in the deep belief network model, and the set of solutions is regarded as an original population, and an opposite individual of each original individual in the original population is generated, as shown in the following formula:

[0049] X' = L b + U b - X

[0050] Wherein: X' represents the position vector of the opposite individual; X represents the position vector of the original individual; L b , U b respectively represent the upper and lower boundaries of X;

[0051] The fitness of each original individual and its opposite individual is compared respectively, and the one with high fitness is selected as an initial population individual, and finally an initial population is obtained;

[0052] The initial population is iteratively updated by the improved grey wolf algorithm to obtain the optimal configuration of the connection weights in the deep belief network model, and a convergence factor a with a cosine variation law is used in the iteration process, as shown in the following formula:

[0053]

[0054] Wherein: a max represents the maximum value of the convergence factor; t max represents the maximum iteration number; t represents the current iteration number; and n represents a decreasing index.

[0055] On the other hand, the application also provides a carbon emission prediction system considering the difference measurement characteristics of transformers, comprising a data acquisition module, a data clustering module, a feature extraction module and a prediction module;

[0056] The data collection module is used to collect carbon emission related basic data of different transformers, and to preprocess the data;

[0057] The data clustering module is used to perform clustering analysis on the preprocessed carbon emission related basic data of different transformers by using an improved particle swarm algorithm, to generate transformer data clusters of different categories;

[0058] The feature extraction module is used to sort all features of the transformer data clusters based on the influence degree of the transformer carbon emission according to the random forest algorithm, to extract the first g features after sorting as training features, and to extract corresponding data in each transformer data cluster based on the training features and construct a training set;

[0059] The prediction module is used for constructing a transformer carbon emission prediction model, training and optimizing model parameters through a training set, obtaining a trained transformer carbon emission prediction model, and predicting the carbon emission of the transformer through the trained transformer carbon emission prediction model.

[0060] In another aspect, the application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of the embodiments of the application when executing the program.

[0061] In another aspect, the application further provides a computer readable storage medium, having a computer program stored thereon, wherein the program, when executed by a processor, implements the method according to any one of the embodiments of the application.

[0062] The application has the following beneficial effects:

[0063] 1. The application not only covers basic information such as geographic location and technical conditions in data collection and integration, but also goes deep into details such as equipment operation state and maintenance history, ensuring the comprehensiveness and accuracy of the data. By introducing advanced data cleaning and correction algorithms, the problems of data missing, errors and inconsistency are solved, ensuring the data quality. This process not only enhances the reliability of model input, but also lays a solid foundation for subsequent analysis, overcoming the shortcomings of traditional methods in data processing.

[0064] 2. The application uses an improved particle swarm algorithm to perform cluster analysis on different main devices. This innovative method not only improves the efficiency of clustering, but also enhances the stability of the clustering results, making the classification of equipment more scientific and reasonable. Subsequently, with the help of the random forest algorithm, the top five main factors affecting carbon emissions are quickly and accurately identified from numerous features. This process is highly automated, reducing the bias of human judgment. Compared with traditional feature selection based on experience or a single statistical indicator, the objectivity and accuracy of the analysis are significantly improved.

[0065] 3. The application uses a deep belief network model to make predictions. Deep belief networks can learn multi-level abstract features of data, and are particularly suitable for handling nonlinear and high-dimensional carbon emission prediction problems. Compared with traditional shallow models, the deep structure of the deep belief network model can capture more complex internal laws, improving the accuracy and generalization ability of the prediction.

[0066] 4、The application creatively introduces an improved grey wolf algorithm to adjust the connection weights of the deep belief network. As a newly emerging heuristic optimization algorithm, the grey wolf algorithm has shown great potential in the field of parameter optimization with its excellent search ability and convergence speed. Through improvement, the global search ability and local search precision of the algorithm are further improved, ensuring the optimal configuration of the model parameters and improving the overall performance of the model.

[0067] 5、Based on the above series of innovations, the carbon emission prediction model constructed by the application has reached an advanced level in the industry in terms of prediction accuracy, stability and computational efficiency. The model not only can effectively predict the carbon emissions of a single device, but also can be extended to a larger scale of device cluster, and even the carbon emission trend analysis of the entire industry. This provides a powerful decision support tool for governments, enterprises and research institutions, which helps to develop more accurate emission reduction strategies and optimize production processes, and promotes the green transformation of the economy and society. BRIEF DESCRIPTION OF DRAWINGS

[0068] Figure 1 is a flowchart of the method of the application;

[0069] Figure 2 is a flowchart of the golden section algorithm of the application;

[0070] Figure 3 is a schematic diagram of the random forest algorithm of the application. DETAILED DESCRIPTION

[0071] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.

[0072] It should be understood that the step numbers used herein are only for the convenience of description, and are not limited to the execution sequence of the steps.

[0073] It should be understood that the terms used in the specification of the application are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in the specification and the appended claims of the application, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0074] The terms "include" and "contain" indicate the presence of the described features, whole, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, whole, steps, operations, elements, components and / or sets thereof.

[0075] The term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items.

[0076] Example 1:

[0077] See also Figure 1 , a carbon emission prediction method considering the differential metering characteristics of transformers, comprising the following steps:

[0078] S101, collect basic data related to carbon emissions of different transformers and pre-process the data;

[0079] S102, performing cluster analysis on the pre-processed carbon emission-related basic data of different transformers using an improved particle swarm algorithm to generate transformer data clusters of different categories;

[0080] S103, sorting all features of the transformer data cluster according to the degree of influence on the transformer carbon emissions based on the random forest algorithm, extracting the first g features after sorting as training features, and extracting corresponding data from each transformer data cluster based on the training features and constructing them into a training set;

[0081] S104, constructing a transformer carbon emission prediction model, training and optimizing model parameters through a training set to obtain a trained transformer carbon emission prediction model, and predicting the carbon emissions of the transformer through the trained transformer carbon emission prediction model.

[0082] As a preferred implementation of this embodiment, the basic data related to carbon emissions of the transformer includes geographic location data and technical condition data of the transformer;

[0083] The geographic location data of the transformer includes: the longitude and latitude coordinates of the installation location, regional climate data (such as annual average temperature and humidity, which may affect equipment efficiency and energy consumption), and local policies and laws and regulations (such as whether it is located in an environmentally sensitive area, local emission standards, etc.);

[0084] Since the iron core is the main body of the transformer, its size greatly affects the carbon emissions of the converter transformer. Therefore, in this embodiment, the iron core is used as the main measurement factor for data collection and carbon emissions prediction. The technical condition data includes:

[0085] Manufacturing process of transformer core:

[0086] The manufacturing process of the core (such as cold rolling, hot rolling, coating technology) has a significant impact on material utilization and energy consumption, and thus affects carbon emissions. It also includes aspect ratio, thickness and material density.

[0087] Magnetic flux density:

[0088] The magnetic flux density requirement of the core in operation also affects material selection and energy consumption, and high magnetic flux density often requires higher performance materials, which may increase carbon emissions.

[0089] Power loss:

[0090] The eddy current loss and hysteresis loss of the core are related to its size, shape, material, and operating frequency, which indirectly affect the energy efficiency and carbon emissions of the device.

[0091] Service life and recycling rate:

[0092] The service life and recycling rate of the core are closely related to carbon emissions, and long service life and high recycling rate can significantly reduce the carbon footprint during the life cycle.

[0093] Cooling system information:

[0094] For devices that require cooling, the efficiency of the cooling system and the choice of cooling medium also affect the overall energy efficiency and carbon emissions.

[0095] Operating conditions: The actual working environment and operating conditions (such as load changes, working cycle) also have some impact on energy consumption and carbon emissions.

[0096] As a preferred embodiment of the present embodiment, the preprocessing step of the transformer's carbon emission-related basic data is:

[0097] Outlier detection is a hypothesis testing process, common detection methods include distribution-based, distance-based, density-based, clustering-based, number-based, dimensionality reduction-based, classification-based, prediction-based, etc. Considering the one-dimensionality of the data analyzed in this embodiment and the applicability of the method, the box plot method based on distribution detection is selected for outlier measurement evaluation;

[0098] The box plot is a commonly used method for quickly identifying outliers in sample data, and can be used for outlier judgment of sample data conforming to normal distribution, as well as for outlier judgment of sample data not conforming to normal distribution, with a wide range of applications.

[0099] Through the box plot method, the transformer's carbon emission-related basic data is measured and evaluated for outliers, and the specific steps are as follows:

[0100] After sorting the transformer's carbon emission-related basic data from small to large, divide it into four equal parts, and take the maximum value, minimum value, median (second partition point of four equal parts), first quartile (first partition point of four equal parts), and second quartile (third partition point of four equal parts) of the sample data;

[0101] Based on the first quartile and the second quartile, a box plot is constructed to determine the outlier standard Q, as shown in the following formula:

[0102]

[0103] Q L is the first quartile; Q U is the second quartile; Q IQR is the quartile distance; quartiles are not affected by outliers, and the box plot criterion is based on quartiles and quartile distances, which can accurately and stably depict the discrete distribution of data, and is conducive to data cleaning; after processing outliers, the sample data analysis result is more accurate;

[0104] After screening outliers in the data by the box plot judgment criterion Q, a reasonable interval distribution model is constructed based on Chebyshev inequality to calculate the threshold interval of the data, and the specific steps are as follows:

[0105] Define the confidence level so that there are at least data samples falling within k standard deviations:

[0106]

[0107] wherein x represents a data sample; μ represents a data expectation; and σ represents a data variance;

[0108] A target function f min (x) is constructed with the preset threshold interval length and the minimum difference of 90% proportion as the target:

[0109]

[0110] wherein Q(x) is all sample intervals obtained according to the preset threshold interval length; and q(μ-kσ≤x≤μ+kσ) is the interval length falling within μ-kσ≤x≤μ+kσ;

[0111] Referring to Figure 2 , based on the target function value, the golden section method (also can select the secant method, the step forward and backward method, the parabolic difference method and many other optimization methods) is used to optimize the value of k, and the interval containing the optimal solution is gradually reduced until the interval length is less than a given precision, and the specific steps are as follows:

[0112] 1) Given interval [a, b] and a < b and very small ε, ε > 0;

[0113] 2) Calculate k1 = a + 0.382(b-a), k2 = a + 0.618(b-a);

[0114] 3) If k2-k1 < ε, output the calculation is completed, otherwise go to the next step;

[0115] 4) If f(k2) > f(k1), let b = k2, go back to 2), otherwise let a = k1, go back to 2).

[0116] Based on the optimized k value, the threshold interval of the data is obtained, the data outside the threshold interval is regarded as an abnormal value, and the abnormal value is repaired, that is, the abnormal value is set as the average value of the threshold interval.

[0117] As a preferred embodiment of the present embodiment, the clustering analysis step of the different transformer carbon emission related basic data is:

[0118] Initialize particle swarm algorithm parameters: particle swarm size N, sample feature dimension d, cluster number K, maximum iteration number T, position boundary x min , x max , and maximum speed v max ;

[0119] Generate chaotic numbers through chaotic mapping, map the chaotic numbers to the data sample interval, construct the reverse position of the particle through reverse learning at the same time, and compare the fitness of the particle and its reverse position. The position corresponding to the particle with the closest fitness to the reverse position of the particle is set as the initial position of the particle;

[0120] The particle starts searching from the initial position, iteratively updates the particle, and controls the boundary through Cauchy mutation of the particle position;

[0121] Divide the data sample according to the nearest neighbor rule, and recalculate the fitness of the particle;

[0122] Compare and update the optimal position and optimal fitness value of the particle individual and group before and after updating;

[0123] When the iteration stopping condition is reached (the preset maximum iteration number is reached or the difference between the optimal position and optimal fitness value of the particle individual and group before and after updating reaches the preset threshold value), stop iteration and obtain the optimal solution;

[0124] Take the optimal solution as the clustering center, divide the data sample according to the nearest neighbor rule, and obtain K clusters, each cluster being a transformer data cluster.

[0125] Update the particle velocity and position through the following formula:

[0126]

[0127]

[0128] ω = ω max -(ω max -ω min ) × q ÷ T

[0129] Wherein: represents the dthdimensional velocity of particle i at the t+1thiteration, the unit is consistent with the data; c1, c2 represent learning factors, representing the cognitive degree of the particle to itself and the group; r1, r2 represent random numbers between (0, 1); represents the dthdimensional position of particle i at the t+1thiteration; P id , g id respectively represent the optimal solution of the current particle and the group; ω represents the inertia weight; q represents the current iteration number, and T is the maximum iteration number;

[0130] The chaotic number is generated by a chaotic mapping formula:

[0131]

[0132] x' = x max +x min -x

[0133] wherein: z k is the chaotic number generated for the kthtime, β takes a value between (0, 1); x represents the index data; x' represents the inverse solution of the index data x; x max , x min respectively are the upper and lower limits of the value of the index data x;

[0134] The ability of the particle to jump out of the local optimum is enhanced by the following Cauchy variation formula:

[0135] x b ' = x x (1-tan (π (u-0.5))

[0136] wherein: x b ' represents the position after variation; u represents a random number in the interval (0, 1);

[0137] The particle fitness function adopts the clustering evaluation index error sum of squares SSE:

[0138]

[0139] wherein: K represents the number of clusters; μ i represents the cluster center of cluster C i ; x represents the sample of cluster C i ; the smaller the SSE value is, the more compact the cluster is, and the better the clustering effect is.

[0140] As a preferred embodiment of the present embodiment, the extraction step of the training features is:

[0141] Referring to Figure 3The random forest algorithm constructs multiple decision trees, each of which will draw part of the samples from all transformer data clusters as a training set with replacement, and randomly select part of the features from all features of the transformer data clusters as candidate features, so as to increase the diversity of the model and reduce the correlation between the trees;

[0142] Each decision tree is trained by its corresponding training set. When the decision tree splits each time, it selects the optimal feature as the split point based on the minimum square error criterion or the minimum Gini index criterion among its candidate features. Specifically, if a feature causes a significant decrease in impurity in multiple splits, it is considered more important for classification or regression tasks. When the preset stopping condition is reached, the training is completed.

[0143] The average impurity reduction of all features in all decision trees is calculated, and the top five features are selected as training features according to the size of the average impurity reduction.

[0144] To better understand the importance of each feature, the importance score of each feature can be displayed through a bar chart. The higher the score, the greater the contribution of the feature to the model. In addition, a threshold for feature importance can be used to filter the main influencing factors, and only those features that exceed the threshold are retained.

[0145] Random Forest (RF) is one of the important machine ensemble learning algorithms. In addition to using different data-guided samples to build each tree, RF algorithm also changes the way of building classification trees or regression trees. In the RF algorithm, each node uses the best prediction in the prediction subset randomly selected at the node to split; compared with many other classifiers, such as SVM and neural network, RF algorithm has better robustness to overfitting. RF algorithm has the following advantages for regression analysis:

[0146] 1. Simple inclusion or exclusion of predictors based on data availability and user needs;

[0147] 2. May include continuous and categorical predictors, which allows, for example, to combine land use information;

[0148] 3. Relatively few model parameters that must be specified by the user;

[0149] 4. Minimal risk of overfitting;

[0150] 5. Automatic calculation of variable importance scores that evaluate the contribution of individual predictors to the final model.

[0151] The RF model is a set of decision tree classifiers {h(X,Θ kan ensemble classification model consisting of a set of decision trees {h(X; Θk), k = 1, 2,..., N}; where Θk is a random vector independently and identically distributed with the kth decision tree, and can represent the growth process of the kth decision tree, and X is the sample to be classified. k is a random vector independently and identically distributed with the kth decision tree, and can represent the growth process of the kth decision tree, and X is the sample to be classified.

[0152] When the RF model is input with a sample X to be classified, the sample X will enter all the decision trees generated by training. The decision trees will select and determine the type of data X according to the characteristics of the data sample. After all the decision trees obtain their respective classification results, the RF model performs a summary vote to predict the classification category.

[0153] As an ensemble algorithm based on decision tree algorithm, the RF model selects and extracts different training sets to train the decision trees in the algorithm during construction and training, thereby improving the differentiation between each classifier and improving the classification effect of the RF algorithm, which is superior to each decision tree in the algorithm construction. The randomness of the RF model can improve the performance of the algorithm, which is embodied as follows:

[0154] (1) Sampling

[0155] In the RF model, a plurality of training samples are randomly and with replacement extracted from the training set to form a sub-sample set, and the data amount of the sub-sample set is the same as that of the original sample set input. The training set of each tree is different and may contain repeated samples.

[0156] (2) Feature selection

[0157] The decision trees in the RF model only select part of the features when splitting. The RF model first randomly selects a part of the total available features, and then selects the optimal feature from the randomly selected features when the decision tree splits each time. Each decision tree grows as much as possible without pruning.

[0158] Therefore, the randomness of the RF model in the entire training process improves the classification accuracy of the unassociated decision trees in the model, makes the model less likely to fall into overfitting, and thereby enhances the noise resistance and generalization of the algorithm.

[0159] Whether the RF model is used for classification or regression depends on whether the classification on and regression tree (CART) is a classification tree or a regression tree.

[0160] If the CART is a classification tree, the key of the algorithm is to select the test attribute of the node and to divide the data purity. The calculation principle of the CART classification tree is the Gini index, and the smaller the Gini index, the smaller the probability of misclassified samples.

[0161] The Gini index is defined as follows:

[0162]

[0163] where p(i|t) is the probability that the test variable t belongs to class i; n is the number of classes.

[0164] When k Gini = 0, all samples belong to one class. The CART decision tree generation algorithm selects the split attribute rule according to the minimum principle of k Gini index. Assuming that attribute A in training set C divides C into C1 and C2, the k Gini index of given division C is:

[0165]

[0166] The decision tree cannot grow indefinitely, and the stopping conditions of the decision tree are:

[0167] The data volume of the node is less than a specified value;

[0168] The index is less than a threshold value;

[0169] The depth of the decision tree reaches a specified value;

[0170] All features have been used up.

[0171] If CART is a regression tree, the minimum mean square error calculation principle is adopted. That is, for a randomly divided feature, the feature and the division point corresponding to the minimum mean square error of the two data sets divided by any division point are calculated, and the sum of the mean square errors of the two data sets is minimized:

[0172]

[0173] where D1 and D2 are the divided data sets, A is an arbitrary divided feature, s is an arbitrary division point, and c1 and c2 are the sample output means of D1 and D2, respectively. Therefore, the regression model of random forest is based on the mean value of the prediction values of all decision trees when performing prediction analysis.

[0174] Based on the calculated feature importance values, random forest can sort the features, with high importance features ranked first. This step helps identify which features play a major role in predicting the target variable. Finally, the top five elements are selected as input parameters for the prediction model, which are the length-width ratio, thickness, material density, magnetic flux density, and average air temperature of the core.

[0175] As a preferred embodiment of the present embodiment, the transformer carbon emission prediction model is constructed based on a deep belief network model, the deep belief network model comprising a plurality of layers of restricted Boltzmann machines, and the specific training steps are as follows:

[0176] The feature vector of the training set is input into the first layer of the restricted Boltzmann machine for training.

[0177] When the first layer of the restricted Boltzmann machine has completed learning of the features in the training set through training, the feature data is output, and the feature data is used as input for the next layer of the restricted Boltzmann machine for continued training.

[0178] Each subsequent layer of the restricted Boltzmann machine is trained through the above steps, and when all the layers of the restricted Boltzmann machines have been trained, the last layer of the restricted Boltzmann machine is used as the output feature, and the locally optimal parameters of each layer of the restricted Boltzmann machine after training are obtained.

[0179] Based on the output feature and the locally optimal parameters of each layer of the restricted Boltzmann machine, the parameters of each layer of the restricted Boltzmann machine are adjusted through an error back propagation algorithm, and finally a trained deep belief network model is obtained.

[0180] The deep belief network model (DBN) is a derivative model of a multi-layer neural training network, and the difference lies in abstracting low-level features of model data and mining inherent distribution characteristics of data to obtain essential features of data using less data samples. At the same time, the deep belief network model well inherits the robustness of the neural network training model and has the ability to process complex functions under the condition of less data samples.

[0181] The deep belief network model mainly integrates the restricted Boltzmann machine (RBM) and the adaptive intelligent algorithm. The training idea is as follows: 1. Extract the bottom layer data feature quantity of the deep learning model as the input variable of the top layer learning of the model design, and use the mode of training from the bottom layer to the high layer of the model; 2. After training to the top layer of the model, the adaptive particle swarm algorithm is used to adaptively optimize and adjust the entire training network to ensure that the model training result can jump out of the local solution.

[0182] The RBM is composed of a visible layer v i and a hidden layer h i . The last layer is a BP network, which mainly performs weight fine-tuning from top to bottom.

[0183] In the RBM, according to the given (v, h), the energy function is:

[0184]

[0185] Where: θ = {w, a, b} is the network parameters, w is the weight between the visible layer and the hidden layer, a and b are the bias of the visible layer and the hidden layer, m and n are the number of neurons of the visible layer and the hidden layer.

[0186] According to the energy function, the following joint probability distribution function can be obtained:

[0187]

[0188] Where: Z(θ) = ∑ v,h e -E(v,h|θ) , is the normalization factor, representing the algebraic sum of all variable energy functions.

[0189] When the v state of the visible layer is determined, the activation probability of the hidden layer unit is:

[0190]

[0191] When the h state of the hidden layer is determined, the activation probability of the visible layer unit is:

[0192]

[0193] When the number of training samples is K, the parameters θ can be determined by solving the maximum likelihood function problem, and the objective function of the maximum likelihood function problem is given as follows:

[0194]

[0195] Where: maxL(θ) is obtained by the stochastic gradient method;

[0196] Through the Gibbs sampling repetition, the updating rule of the RBM parameters can be obtained as follows:

[0197] Δw ij = ε(<v i h j > data -<v i h j > recon )

[0198] Δa i = ε(<v i > data -<v i > recon )

[0199] Δb i = ε(<h j > data -<h j > recon )

[0200] where ε is the RBM learning rate, <·> data and <·> recon are the mathematical expectations of the input data and the reconstructed data, respectively.

[0201] In this embodiment, the specific training steps of the deep belief network model are as follows:

[0202] (1) The original data is input as an input layer vector into the first layer RBM to complete unsupervised training.

[0203] (2) After the first layer RBM completes feature learning on the original data, feature data is obtained, and these feature data are taken as input vectors of a new layer and input into the next layer RBM to continue unsupervised training.

[0204] (3) Steps (1) and (2) are repeatedly performed until the RBM of each layer is trained and learned to completion, the features obtained in the last layer RBM are taken as output features, and the locally optimal parameters in each layer RBM are obtained.

[0205] (4) The error back propagation algorithm is used to perform supervised fine-tuning on the RBM from top to bottom, adjust the parameters of each layer RBM, and finally obtain the global optimal parameters of the entire DBN network model. The DBN network uses a plurality of RBM units to form a basic network architecture, so that the DBN network has both unsupervised pre-training and supervised fine-tuning. Such combination of unsupervised and supervised methods not only solves the gradient dispersion problem existing in traditional methods, but also solves the problem that the network is easily trapped in local optimum.

[0206] As a preferred embodiment of the present embodiment, the deep belief network model is optimized by improving the grey wolf optimization algorithm, and the specific steps are as follows:

[0207] In the standard grey wolf optimization algorithm (GWO), α, β, δ and ω represent grey wolf individuals, where α represents the individual that makes decisions and manages the wolf pack, β and δ have lower fitness than α, and ω is a common individual. The specific behaviors of the GWO algorithm are surrounding, hunting and attacking.

[0208] (1) Surrounding behavior

[0209] The data model of the grey wolf surrounding prey can be expressed as follows:

[0210] D = | C · X p (t) - X(t) |

[0211] X(t+1) = X p (t) - A · D

[0212] where D represents the distance between the wolf pack and the prey; A = 2a r1-a, a is a convergence factor; C = 2 r2, t represents the number of iterations, X p and X represent the positions of the prey and the wolf pack respectively, r1, r2 are random quantities with a value range of [0, 1], and a has a value range of [0, 2].

[0213] (2) Hunting behavior

[0214] Assuming that a, b, d represent the global optimal solution, the second solution and the third solution of the gray wolf individual, and the optimization positioning is performed, the distances are represented as formulas.

[0215] D α = |C1 X α -X|

[0216] D β = |C2 X β -X|

[0217] D δ = |C2 X δ -X|

[0218] where D α , D β , D δ represent the approximate distances of the individuals a, b, d and the current position X, X α , X β , X δ represent the positions of the global optimal solution, the second solution and the third solution respectively; C1, C2, C3 represent random vectors with a value range of [0, 1]. The iteration steps of X are shown in the following formula:

[0219] X1 = X α -A1 (D α )

[0220] X2 = X β -A2 (D β )

[0221] X3 = X δ -A3 (D δ )

[0222]

[0223] where: X(t+1) represents the updated solution; A1, A2, A3 represent random quantities.

[0224] (3) Attack behavior

[0225] The attack is the last stage of the wolf pack hunting behavior, which can be realized by adjusting the parameter a.

[0226] If A≤1, the wolf pack approaches the prey and concentrates on attacking the prey (X * ,Y * ); otherwise, the wolf pack gradually moves away from the prey.

[0227] A set of solutions is randomly generated, each solution representing a configuration of connection weights in the deep belief network model. This set of solutions is considered as the original population. The standard GWO algorithm initializes the population positions using a random initialization method, which may result in an excessively large search range for the wolf pack, leading to a longer search time. Therefore, the opposite individual of each original individual in the original population is also generated, as shown in the following formula:

[0228] X′=L b +U b -X

[0229] where X' represents the position vector of the opposite individual; X represents the position vector of the original individual; L b , U b represent the upper and lower boundaries of X, respectively;

[0230] The fitness of each original individual and its opposite individual is compared, and the one with higher fitness is selected as the initial population individual. Finally, the initial population is obtained.

[0231] The improved grey wolf algorithm is used to iteratively update the initial population to obtain the optimal configuration of connection weights in the deep belief network model. According to the formula before improvement, the parameter A is determined by a, which linearly decreases from 2 to 0. In the actual algorithm iteration process, the search space is large at the beginning of iteration, and global search is needed, so the a decreasing speed should be slowed down, and the a decreasing speed should be accelerated in the later iteration, so as to improve the local optimization and speed up the convergence speed. Therefore, a convergence factor a with cosine variation rule is used in the iteration process, as shown in the following formula:

[0232]

[0233] where a max represents the maximum value of the convergence factor; t max represents the maximum iteration number; t represents the current iteration number; n represents the decreasing index, 0<n<1.

[0234] In this embodiment, the optimization process of the grey wolf algorithm can be expanded in detail as follows:

[0235] 1. Initialization stage

[0236] Population initialization: first, a set of solutions (i.e., initial weight matrix) is randomly generated, each solution representing a configuration of connection weights in the DBN. These solutions constitute the initial population of the grey wolf algorithm.

[0237] Role Assignment: Based on the performance of these initial solutions (weight configurations) on a predefined objective function (usually prediction error or loss function), each set of solutions is determined as Alpha (the leader wolf, the optimal solution), Beta (the second-best solution), Delta (the third-best solution), or a regular wolf.

[0238] 2. Hunting and Exploration Phase

[0239] Update Strategy: The algorithm updates each solution (weight configuration) by simulating the behavior of a wolf pack hunting prey. The update process involves calculating the direction and distance of each wolf relative to the Alpha, Beta, and Delta positions, then adjusting their positions (i.e., weight values) in the solution space based on this information.

[0240] Search Range Adjustment: As iterations progress, the search range may gradually narrow, mimicking the behavior of a wolf pack concentrating its attack once it locks onto prey, allowing for more refined exploration of the solution space.

[0241] 3. Convergence Judgment and Iteration

[0242] Iteration and Evaluation: The above update process is repeated, and the performance of each set of solutions is re-evaluated after each iteration, and the roles of Alpha, Beta, and Delta may be reassigned.

[0243] Stopping Criteria: One or more stopping conditions are set, such as reaching a maximum number of iterations, the improvement in solutions falling below a certain threshold, or finding a satisfactory solution.

[0244] 4. Application to Deep Belief Network Optimization

[0245] Connection Weight Optimization: In the context of a deep belief network model, the grey wolf algorithm optimizes the connection weights to minimize prediction error. This means that through constant iteration, GWO adjusts the weights between layers in the DBN to find a configuration that minimizes the error between predicted output and actual carbon emissions data.

[0246] Performance Evaluation: After each iteration, the model's performance is evaluated using a validation set to ensure that the optimization process not only reduces training error but also generalizes to unseen data.

[0247] 5. Results Application

[0248] Final Model Training: Once the optimal or near-optimal weight configuration is found, the DBN model is trained on the entire training set (or a larger dataset) for the final model to be used for carbon emissions prediction.

[0249] Example Two:

[0250] The application discloses a carbon emission prediction system considering transformer difference measurement characteristics, which comprises a data collection module, a data clustering module, a feature extraction module and a prediction module.

[0251] The data collection module is used for collecting carbon emission related basic data of different transformers and pre-processing the data.

[0252] The data clustering module is used for clustering analysis on the pre-processed carbon emission related basic data of different transformers by using an improved particle swarm algorithm to generate transformer data clusters of different categories.

[0253] The feature extraction module is used for sorting all features of the transformer data clusters based on the influence degree of transformer carbon emission by using a random forest algorithm, extracting the first g features after sorting as training features, extracting corresponding data in each transformer data cluster based on the training features and constructing a training set.

[0254] The prediction module is used for constructing a transformer carbon emission prediction model, training and optimizing model parameters through the training set, obtaining a trained transformer carbon emission prediction model, and predicting the carbon emission of the transformer through the trained transformer carbon emission prediction model.

[0255] The system is used for realizing the functions in the first embodiment, and details are not described herein.

[0256] Embodiment three:

[0257] The embodiment provides an electronic device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor realizes the method according to any one of the embodiments of the application when executing the program.

[0258] Embodiment four:

[0259] The embodiment provides a computer readable storage medium, which stores a computer program, and the program is executable on a processor to realize the method according to any one of the embodiments of the application.

[0260] In the embodiments of the present application, "at least one" means one or more, and "multiple" means two or more. The "and / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the cases of A alone, A and B together, and B alone. Wherein A and B can be singular or plural. The character " / " generally represents an "or" relationship between the front and rear associated objects. "At least one of the following" and the like means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.

[0261] Those of ordinary skill in the art can realize that the units and algorithm steps described in the embodiments disclosed herein can be realized in electronic hardware, computer software, and a combination of electronic hardware and computer software. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0262] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.

[0263] In several embodiments provided in the present application, any function realized in the form of a software function unit and sold or used as an independent product can be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (Read-Only Memory; hereinafter referred to as: ROM), a random access memory (Random Access Memory; hereinafter referred to as: RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0264] The above is only an embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation based on the content of the present application specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A carbon emission prediction method considering transformer differential metering characteristics, characterized in that: The following steps are involved: Collect basic data related to carbon emissions of different transformers and pre-process the data; The improved particle swarm algorithm is used to perform cluster analysis on the pre-processed carbon emission-related basic data of different transformers to generate transformer data clusters of different categories. Based on the random forest algorithm, all features of the transformer data cluster are sorted according to the degree of influence on the transformer carbon emissions. The first g features after sorting are extracted as training features. Based on the training features, the corresponding data in each transformer data cluster is extracted and constructed as a training set. Construct a transformer carbon emission prediction model, train it using a training set and optimize the model parameters to obtain a trained transformer carbon emission prediction model, and use the trained transformer carbon emission prediction model to predict the transformer's carbon emissions; The basic data related to carbon emissions of the transformer includes geographical location data and technical condition data of the transformer; The geographic location data of the transformer includes: the longitude and latitude coordinates of the installation location, regional climate data, and local policies and legal regulations; The technical condition data includes the manufacturing process of the transformer core, magnetic flux density, power loss, service life and recovery rate, cooling system information and operating conditions; The cluster analysis steps for the basic data related to carbon emissions of different transformers are as follows: Initialize the particle swarm algorithm parameters; Generate chaotic numbers through chaotic mapping, map the chaotic numbers to the data sample interval, and construct the reverse position of the particle through reverse learning. Then compare the fitness of the particle and its reverse position, and set the position corresponding to the particle with the closest fitness to its reverse position as the initial position of the particle. The particles start searching from the initial position, iteratively update the particles, and perform boundary control by using Cauchy mutation. Divide the data samples according to the nearest neighbor rule and recalculate the fitness of the particles; Compare and update the optimal positions and optimal fitness values ​​of individual particles and groups before and after; When the iteration stop condition is reached, the iteration is stopped and the optimal solution is obtained; Taking the optimal solution as the cluster center, the data samples are divided according to the nearest neighbor rule to obtain K clusters, each cluster is a transformer data cluster; The transformer carbon emission prediction model is constructed based on a deep belief network model, which includes a multi-layer restricted Boltzmann machine. The specific training steps are: Input the feature vector of the training set into the first layer of restricted Boltzmann machine for training; When the first layer of restricted Boltzmann machine completes learning the features in the training set through training, it outputs the feature data, which is used as the feature data in the next layer of restricted Boltzmann machine to continue training; Each subsequent layer of restricted Boltzmann machines is trained through the above steps. After all restricted Boltzmann machines are trained, the last layer of restricted Boltzmann machines is used as the output feature, and the local optimal parameters of each layer of restricted Boltzmann machines after training are obtained. Based on the output features and the local optimal parameters of each layer of restricted Boltzmann machine, the parameters of each layer of restricted Boltzmann machine are adjusted through the error back propagation algorithm, and finally a trained deep belief network model is obtained.

2. The carbon emission prediction method considering transformer differential metering characteristics according to claim 1 is characterized in that: The preprocessing steps for the basic data related to carbon emissions of the transformer are as follows: The box plot method is used to conduct outlier measurement assessment on the basic data related to carbon emissions of transformers. The specific steps are as follows: Divide the basic data related to carbon emissions of transformers into four equal parts, and take the maximum value, minimum value, median, first quartile, and second quartile of the sample data; Based on the first quartile and the second quartile, a box plot is constructed to determine the outlier standard Q, as shown in the following formula: Where: Q L is the first quartile; Q U is the second quartile; Q IQR is the interquartile range; After using the box plot to determine the outlier standard Q to screen out the outliers in the data, a reasonable interval distribution model is constructed based on Chebyshev's inequality to calculate the threshold interval of the data. The specific steps are as follows: Defining the confidence level So that there is at least Data samples fall within k standard deviations: Where: x represents the data sample; μ represents the data expectation; σ represents the data variance; The objective function f is constructed with the preset threshold interval length and the minimum difference of 90% ratio as the goal. min (x): Where: Q(x) is the interval of all samples obtained according to the preset threshold interval length; q(μ-kσ≤x≤μ+kσ) is the length of the interval falling into μ-kσ≤x≤μ+kσ; Based on the objective function value, the golden section method is used to optimize the k value. The threshold interval of the data is obtained based on the optimized k value. The data outside the threshold interval is regarded as outliers, and then the outliers are repaired, that is, the outliers are set to the average value of the threshold interval.

3. The carbon emission prediction method considering transformer differential metering characteristics according to claim 1 is characterized in that: The steps of extracting the training features are: The random forest algorithm constructs multiple decision trees. Each decision tree extracts some samples from all transformer data clusters with replacement as the training set, and randomly selects some features from all features of the transformer data cluster as candidate features. Each decision tree is trained using its corresponding training set. Each time the decision tree splits, the optimal feature is selected as the split point among its candidate features based on the square error minimization criterion or the Gini index minimization criterion. When the preset stopping condition is reached, the training is completed. Calculate the average impurity reduction of all features in all decision trees, sort all features according to the size of the average impurity reduction, and select the top n features as training features.

4. The carbon emission prediction method considering transformer differential metering characteristics according to claim 1 is characterized in that: The deep belief network model is optimized by improving the gray wolf optimization algorithm. The specific steps are as follows: A set of solutions is randomly generated, each solution represents a configuration of the connection weights in the deep belief network model. The set of solutions is regarded as the original population and the opponent individuals of each original individual in the original population are generated at the same time, as shown in the following formula: X′=L b +U b -X Where: X′ represents the position vector of the opposing individual; X represents the position vector of the original individual; L b 、U b Represent the upper and lower boundaries of X respectively; Compare the fitness of each original individual and its opposing individual, select the individual with higher fitness as the initial population, and finally obtain the initial population; The initial population is iteratively updated through the improved gray wolf algorithm to obtain the optimal configuration of the connection weights in the deep belief network model. During the iteration process, a convergence factor α with a cosine variation law is used, as shown in the following formula: Where: α max Indicates the maximum value of the convergence factor; t mɑx represents the maximum number of iterations; t represents the current number of iterations; and n represents the decrement index.

5. A carbon emission prediction system considering transformer differential metering characteristics, which uses a carbon emission prediction method considering transformer differential metering characteristics according to any one of claims 1 to 4, characterized in that: It includes data acquisition module, data clustering module, feature extraction module and prediction module; The data collection module is used to collect basic data related to carbon emissions of different transformers and pre-process the data; The data clustering module is used to perform cluster analysis on the pre-processed carbon emission-related basic data of different transformers through an improved particle swarm algorithm to generate transformer data clusters of different categories; The feature extraction module is used to sort all features of the transformer data cluster according to the degree of influence of the transformer carbon emissions based on the random forest algorithm, extract the first g features after sorting as training features, and extract corresponding data in each transformer data cluster based on the training features and construct them into a training set; The prediction module is used to construct a transformer carbon emission prediction model, train and optimize model parameters through a training set, obtain a trained transformer carbon emission prediction model, and predict the carbon emissions of the transformer through the trained transformer carbon emission prediction model.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Equipment carbon emission prediction method and device based on big data and electronic equipment

    CN117408380A

  • Carbon emission prediction method and device, electronic equipment and storage medium

    CN114662780A

  • Carbon emission prediction method, device and equipment based on AO algorithm

    CN115759346A