Method and system for constructing electricity price prediction large model based on federated learning

The federated learning method for building a large-scale electricity price forecasting model solves the problems of data sharing and multi-source data integration in the power industry, achieves high accuracy and generalization capabilities in electricity price forecasting, and ensures data security and privacy.

CN120744559APending Publication Date: 2025-10-03POWERCHINA RENEWABLE ENERGY CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510718644.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In existing technologies, data in the power industry is difficult to share, and federated learning cannot be used to integrate multi-source heterogeneous data, resulting in low electricity price forecasting accuracy and poor generalization ability of the forecasting model.

Method used

Through the method of building a large-scale electricity price prediction model based on federated learning, data collection of time windows is performed, window data sets are established, sample feature extraction and multi-dimensional feature clustering are performed, cluster clusters are configured, decision trees are trained using independent and joint cluster clusters, a large-scale electricity price prediction model is established, and data security and privacy are guaranteed through encryption mechanisms.

Benefits of technology

It has achieved collaborative modeling of multi-institutional data, enhanced the generalization ability and accuracy of the electricity price prediction model, and ensured the security and privacy of the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744559A_ABST
    Figure CN120744559A_ABST
Patent Text Reader

Abstract

The invention discloses an electricity price prediction large model construction method and system based on federated learning, and relates to the related technical field of electricity price prediction, and the method comprises the steps: executing the data collection of a time window in a preset time window, and building a window data set; executing sample feature extraction; performing multi-dimensional feature clustering by using the multi-dimensional feature identifier and configuring N clusters; performing cluster division based on double identification channels, and establishing an independent cluster and a combined cluster; independently training the decision tree by using the independent cluster, and executing weak classifier mapping training of the combined cluster; and establishing an electricity price prediction large model. The technical problems that in the prior art, power industry data are difficult to share, multi-source heterogeneous data cannot be integrated through federal learning, the electricity price prediction precision is low, and the prediction model generalization ability is poor are solved, and the technical effects that multi-mechanism data collaborative modeling is achieved, and the electricity price prediction model generalization ability and the electricity price prediction precision are enhanced are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field related to electricity price prediction, and specifically to a method and system for constructing a large-scale electricity price prediction model based on federated learning. Background Art

[0002] As more and more new energy power generation companies participate in electricity market transactions, the output of new energy has the characteristics of randomness and instability. The randomness of the output is transmitted to the power market as a sharp fluctuation in electricity prices. Accurate electricity price prediction is conducive to the stable operation of the power market and the effective allocation of resources. However, traditional electricity price prediction methods such as time series method and neural network prediction method cannot fully utilize multi-source data scattered in various institutions, and the prediction accuracy is greatly limited. Among them, a single data cannot fully reflect the many complex factors affecting electricity prices, and data from different institutions are difficult to effectively integrate, affecting the prediction accuracy.

[0003] Therefore, in the current relevant technologies, there are technical problems such as difficulty in sharing data in the power industry and inability to use federated learning to integrate multi-source heterogeneous data, resulting in low electricity price forecasting accuracy and poor generalization ability of the forecasting model. Summary of the Invention

[0004] This application provides a method and system for constructing a large-scale electricity price prediction model based on federated learning, thereby solving the technical problems in the existing technology that power industry data is difficult to share and multi-source heterogeneous data cannot be integrated using federated learning, resulting in low electricity price prediction accuracy and poor generalization ability of the prediction model. It achieves the technical effect of realizing collaborative modeling of multi-institutional data, enhancing the generalization ability of the electricity price prediction model and the accuracy of electricity price prediction.

[0005] The present application provides a method for constructing a large-scale electricity price prediction model based on federated learning, which includes: executing data collection of the time window within a preset time window to establish a window data set, wherein the window data set includes a first data set and a second data set, wherein the first data set is an internal data set and the second data set is a homomorphically encrypted data set; executing sample feature extraction of the window data set to establish a multidimensional feature identification of the sample; performing multidimensional feature clustering using the multidimensional feature identification of the sample, and configuring N clustering clusters according to the multidimensional feature clustering results, wherein the multidimensional feature clustering is a repeatable clustering; performing cluster division of the N clustering clusters based on dual recognition channels to establish independent clustering clusters and joint clustering clusters; after training a decision tree separately using the independent clustering clusters, performing weak classifier mapping training of the joint clustering cluster; and establishing a large-scale electricity price prediction model based on the trained decision tree and weak classifier.

[0006] In a possible implementation, the method for constructing a large model for electricity price prediction based on federated learning also performs the following processing: extracting cluster bias features based on the joint cluster cluster mapped by the weak classifier; configuring the global weight dynamic update layer of the mapped weak classifier using the cluster bias features; and establishing the large model for electricity price prediction based on the decision tree, the weak classifier, and the global weight dynamic update layer.

[0007] In a possible implementation, the method for constructing a large model for electricity price prediction based on federated learning also performs the following processing: performing multidimensional sample analysis of the weak classifier, and establishing a first basic weight based on the multidimensional sample analysis results, the multidimensional sample analysis including sample quantity analysis, sample quality analysis, and sample coverage analysis; performing prediction state analysis of the weak classifier, and establishing a second basic weight; constructing a matching mapping weight library based on the cluster bias characteristics, and establishing dynamic weights; and configuring a global weight dynamic update layer using the first basic weight, the second basic weight, and the dynamic weight.

[0008] In a possible implementation, the method for constructing a large electricity price prediction model based on federated learning also performs the following processing: encrypting the model parameters of the large electricity price prediction model and uploading them to a central server; receiving global feedback from the central server, wherein the global feedback is global feedback established by the central server after aggregating the model parameters through a global aggregation algorithm after receiving multiple locally uploaded encrypted model parameters; and updating the parameters of the large electricity price prediction model according to the global feedback.

[0009] In a possible implementation, the method for constructing a large-scale electricity price prediction model based on federated learning also performs the following processing: after selecting one or more features in the multidimensional feature identification as core features, performing clustering analysis centered on the core features; after any sample is clustered, creating a copy sample of the corresponding sample to continue clustering processing; when clustering is completed, performing attraction analysis between clusters, and performing cluster merging based on the attraction analysis results to configure N clusters.

[0010] In a possible implementation, the method for constructing a large-scale electricity price prediction model based on federated learning further performs the following processing: establishing a cluster number threshold discrimination sub-channel, using the cluster number threshold discrimination sub-channel to discriminate the size of each cluster within N clusters, and establishing a first discrimination trust; establishing a cluster feature alienation threshold discrimination sub-channel, using the cluster feature alienation threshold discrimination sub-channel to discriminate the degree of feature alienation of each cluster within N clusters, and establishing a second discrimination trust; completing cluster division according to the first discrimination trust and the second discrimination trust.

[0011] In a possible implementation, the method for constructing a large-scale electricity price prediction model based on federated learning also performs the following processing: in the sample feature extraction of the window data set, the extracted features include time features, power demand features, weather environment features, power supply features, market features, economic features, and geographic space features.

[0012] In a possible implementation, the method for constructing a large-scale electricity price prediction model based on federated learning also performs the following processing: after sorting the window data set in time series, a time linkage window is established; when performing sample feature extraction, an electricity price lag feature is established based on the time linkage window, and the electricity price lag feature is added to the multidimensional feature identifier.

[0013] In a possible implementation, the method for constructing a large-scale electricity price prediction model based on federated learning also performs the following processing: establishing an evaluation indicator set; using the evaluation indicator set to perform model testing of the large-scale electricity price prediction model and establishing model testing feedback; and using the model testing feedback to perform focused optimization of the large-scale electricity price prediction model.

[0014] The present application also provides a system for building a large-scale electricity price prediction model based on federated learning, which includes: a window data set establishment module, which is used to perform data collection of a time window within a preset time window and establish a window data set, wherein the window data set includes a first data set and a second data set, wherein the first data set is an internal data set and the second data set is a homomorphically encrypted data set; a sample feature extraction module, which is used to perform sample feature extraction of the window data set and establish a multidimensional feature identification of the sample; a multidimensional feature clustering module, which is used to perform multidimensional feature clustering using the multidimensional feature identification of the sample and configure N cluster clusters according to the multidimensional feature clustering results, wherein the multidimensional feature clustering is a repeatable clustering; a cluster partitioning module, which is used to perform cluster partitioning of N cluster clusters based on dual recognition channels and establish independent cluster clusters and joint cluster clusters; a mapping training module, which is used to perform weak classifier mapping training of the joint cluster cluster after training the decision tree separately using the independent cluster cluster; an electricity price prediction large-scale model establishment module, which is used to establish an electricity price prediction large-scale model based on the trained decision tree and weak classifier.

[0015] The proposed method and system for constructing a large-scale electricity price prediction model based on federated learning in this application aims to collect data from a preset time window to establish a window dataset; extract sample features; cluster features using multidimensional feature identification and configure N clusters; perform cluster division based on dual recognition channels to establish independent clusters and joint clusters; train decision trees using independent clusters and perform weak classifier mapping training for the joint clusters; and establish a large-scale electricity price prediction model. This solves the existing technical problems of difficulty in sharing power industry data and inability to integrate multi-source heterogeneous data using federated learning, resulting in low electricity price prediction accuracy and poor generalization of the prediction model. It achieves the technical effect of enabling collaborative modeling of multi-institutional data, enhancing the generalization ability of the electricity price prediction model, and improving the accuracy of electricity price prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments of the present disclosure are briefly introduced below. Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in precise order. Instead, various steps may be processed in reverse order or simultaneously as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.

[0017] Figure 1 A flow chart of the method for constructing a large-scale electricity price prediction model based on federated learning provided in an embodiment of the present application.

[0018] Figure 2 A schematic diagram of the system structure for building a large-scale electricity price prediction model based on federated learning provided in an embodiment of the present application.

[0019] Explanation of the reference numerals: window data set establishment module 10 , sample feature extraction module 20 , multi-dimensional feature clustering module 30 , clustering cluster division module 40 , mapping training module 50 , electricity price prediction large model establishment module 60 . DETAILED DESCRIPTION

[0020] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below.

[0021] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0022] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict, and the terms “first\second” involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. The terms “including” and “having” and any variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used herein are for the purpose of describing the embodiments of this application only.

[0023] The present application embodiment provides a method for constructing a large-scale electricity price prediction model based on federated learning, such as Figure 1 As shown, the method includes:

[0024] Preferably, federated learning can coordinate data usage and machine learning modeling. Through parameter exchange under an encryption mechanism, it can not only fully tap the value of data from various institutions, but also ensure the security and privacy of the data. Building a large electricity price forecasting model based on federated learning can integrate the internal data of power companies and relevant data from external institutions such as meteorology, comprehensively capture various factors affecting electricity prices, and achieve more accurate electricity price forecasts.

[0025] Step S100: within a preset time window, execute data collection of the time window to establish a window data set, wherein the window data set includes a first data set and a second data set, wherein the first data set is an internal data set and the second data set is a homomorphically encrypted data set.

[0026] Preferably, a specific time period is used as a preset time window, such as 1 day or 1 week, to limit the scope of data collection, and data collection of the time window is performed, that is, within the preset time window, multiple types of electricity price forecast-related data are collected from different data sources to establish a window data set, including a first data set (i.e., an internal data set) and a second data set (i.e., a homomorphically encrypted data set). Specifically, the internal data set refers to the management data generated within the power enterprise or institution, which is usually stored in the enterprise's private database and may include but is not limited to historical electricity price data, grid load data, generator set operating status data, meteorological data (such as temperature, humidity, etc.); the homomorphically encrypted data set is data involving privacy or sensitive information obtained from an external data source, which needs to be homomorphically encrypted before use. It may include user electricity consumption behavior data, market transaction data, etc., which usually comes from external cooperative enterprises and contains user privacy information or corporate secrets; wherein, homomorphic encryption processing allows calculations to be performed directly on encrypted data without decryption first. The calculation result after decryption is the same as the result of the same calculation on the original data. For electricity price prediction based on federated learning, the use of homomorphic encryption can ensure that external data participates in model training without being decrypted, thereby protecting the privacy of the data.

[0027] Step S200 , performing sample feature extraction of the window data set, and establishing a multi-dimensional feature identifier of the sample.

[0028] Step S200 further includes performing sample feature extraction of the window data set, where the extracted features include time features, power demand features, weather environment features, power supply features, market features, economic features, and geographic space features.

[0029] Preferably, sample feature extraction is performed on the window data set, including extracting time features, power demand features, weather and environmental features, power supply features, market features, economic features, and geographic space features. Specifically, time features reflect the impact of the time dimension on electricity prices, including periodic patterns, time period characteristics, etc., including peak periods (such as peak electricity consumption in the morning and evening on weekdays), valley periods (such as nighttime), flat periods, and other time series associations, as well as time intervals of historical data for the same period (such as "electricity prices for the same period in the previous three days"), special time nodes (such as electricity price policy adjustment days). Power demand features directly reflect the demand intensity and changing trends at the power consumption end, including real-time demand indicators such as total power consumption within the time window, load peaks, load valleys, and load curve shapes (such as whether there is a steep rise / fall trend), demand-side behaviors of different user types such as industry, commerce, and residents, and fluctuations in electric vehicle charging loads and air conditioning loads, as well as demand forecast data.

[0030] Preferably, weather and environmental characteristics indirectly affect electricity prices by influencing users' electricity usage habits (such as cooling / heating demand) or renewable energy output (such as photovoltaic and wind power). They include basic meteorological data such as temperature, humidity, wind speed, precipitation, and sunshine duration, extreme weather events (whether there are abnormal weather such as heavy rain, high temperature, cold wave, and its duration and intensity) and seasonal environment (such as high temperature in summer leading to a surge in air conditioning load and increased heating demand in winter). Power supply characteristics characterize the supply capacity and structure of the power production end, directly affecting the balance of supply and demand in electricity prices. They include the proportion of power generation of different energy types, the available capacity of energy storage facilities, the operating status of the unit (the start and stop plan of the generator unit, maintenance status, maximum / minimum output limit) and transmission network constraints (such as the transmission capacity and congestion of key transmission lines, and cross-regional power dispatch data). Market characteristics reflect the impact of electricity market trading rules on electricity prices, including market trading data (real-time market trading volume within the time window, transaction price range, and number of bidding entities) and market supply and demand signals (reserve capacity rate, that is, the ratio of available power generation capacity to maximum load; and status indicators of supply shortage or oversupply).

[0031] Preferably, economic characteristics represent the linkage between macroeconomic indicators and electricity prices, including macroeconomic indicators (GDP growth rate, industrial added value, manufacturing PMI index), industry electricity consumption correlation and electricity price elasticity coefficient (the sensitivity of electricity price changes to electricity consumption). Geographical spatial characteristics refer to the differentiated impact of power grid structure, resource distribution and user density in different geographical locations on electricity prices, including regional power grid attributes, spatial distance and transmission cost (the distance between the power source and the load center, the transmission line loss rate) and distributed energy distribution. Finally, a multidimensional feature identification of the sample is established, that is, each sample (such as electricity price data in a certain time window) is converted into a set of multidimensional feature vectors based on the extracted multiple features, so that the federated learning model can capture the law of electricity price fluctuations from multiple angles such as time, space, market, environment, and economy.

[0032] Further step S200 also includes step S210, after sorting the window data set in time sequence, establishing a time linkage window; step S220, when performing sample feature extraction, establishing an electricity price lag feature based on the time linkage window, and adding the electricity price lag feature to the multidimensional feature identifier.

[0033] Preferably, the samples in the window data set are arranged in chronological order (such as by minutes, hours, and days) to form a continuous time series, and then a time linkage window is established, that is, taking the current prediction moment as the benchmark, a fixed-length historical time period is traced back to form a sliding window containing past-present electricity price data. By constructing a temporally continuous window, the law of electricity price and feature evolution over time, such as intraday fluctuations and daily trends, is captured; the electricity price lag feature refers to the electricity price data at a historical moment, which is used as the input feature of the current sample. When extracting sample features, the number of historical periods that need to be traced back is determined within the time linkage window, such as a lag of 1 hour, 3 hours, 1 day, etc., and the electricity price value at the corresponding historical moment is extracted from the window as a new feature dimension and added to the multidimensional feature identifier, so that the multidimensional feature identifier of each sample contains historical electricity price information, so as to directly learn the autocorrelation of the electricity price sequence and improve the prediction accuracy of trend and cyclical fluctuations.

[0034] Step S300 : performing multidimensional feature clustering using the multidimensional feature identifiers of the samples, and configuring N clusters according to the multidimensional feature clustering results, wherein the multidimensional feature clustering is repeatable clustering.

[0035] Step S300 further includes step S310, after selecting one or more features in the multidimensional feature identification as core features, performing cluster analysis centered on the core features; step S320, after any sample is clustered, creating a copy sample of the corresponding sample to continue clustering processing; step S330, when clustering is completed, performing attraction analysis between clusters, and performing cluster merging according to the attraction analysis results to configure N clusters.

[0036] Preferably, one or more features that have the most significant impact on electricity prices are selected from the multidimensional feature identifiers of the sample (such as time, load, weather, electricity price lag value, etc.) as core features. For example, the Pearson coefficient is used to identify features that are highly correlated with electricity price tags, such as load peak and real-time electricity price lag value, and then a clustering analysis centered on the core features is performed, that is, clustering algorithms such as K-means and DBSCAN are used to cluster samples with similar core features into one category based on the numerical distance or density distribution of the core features to obtain preliminary clusters.

[0037] Preferably, when a sample is classified into a cluster, a copy of the sample is generated to continue to participate in subsequent clustering iterations. For example, based on the load peak core feature, sample A is classified into the high load cluster and a copy of sample A is generated. 1 , then based on the core characteristics of weather temperature, sample A 1It may be classified into the low-temperature cluster, thereby recording the multiple attribution relationships of sample A under different core features, and then forming the multi-cluster association attributes of the sample, reflecting the complexity of factors affecting electricity prices. For example, high load may be accompanied by high or low temperature at the same time, corresponding to different electricity price adjustment mechanisms; avoiding the limitations of a single clustering result, allowing the same original data to be competed by different clusters in different iterations, and improving the flexibility of clustering; for example, a sample belongs to the high-load cluster in the load peak dimension, but may belong to the low-temperature (high heating demand) cluster in the weather temperature dimension. Duplicating the sample can enable it to be evaluated multiple times in clusters with different feature combinations.

[0038] Preferably, after clustering is completed, an attraction analysis between clusters is performed, that is, the similarity or correlation between different clusters is evaluated to determine whether adjacent clusters need to be merged into a larger cluster, such as calculating the Euclidean distance of the cluster center in the multidimensional feature space, whether the mean or distribution of electricity prices of different clusters are close, etc.; when the attraction between clusters exceeds a threshold (such as the distance is less than a set value, the electricity price distribution overlap rate is >80%), they are merged into a new cluster, such as merging the high load-high temperature cluster and the high load-low temperature cluster into a high load cluster, whose core feature load peak is the same, and the electricity prices are driven by supply and demand tensions, and only temperature needs to be retained as a subdivision feature within the cluster; through repeated iterative attraction analysis and merging, the optimal number of clusters N is finally determined, so that the samples within each cluster are highly homogeneous in core features and electricity price patterns, and there are significant differences between clusters, where N is a positive integer, such as N=3, which helps to build a large electricity price prediction model that takes into account both privacy protection and prediction accuracy.

[0039] Step S400 : performing cluster division of N clusters based on the dual recognition channels, and establishing independent clusters and joint clusters.

[0040] Step S400 further includes step S410, establishing a cluster number threshold discrimination sub-channel, using the cluster number threshold discrimination sub-channel to perform size discrimination of each cluster within N clusters, and establishing a first discrimination confidence level; step S420, establishing a cluster feature alienation threshold discrimination sub-channel, using the cluster feature alienation threshold discrimination sub-channel to perform feature alienation degree discrimination of each cluster within N clusters, and establishing a second discrimination confidence level; step S430, completing cluster division according to the first discrimination confidence level and the second discrimination confidence level.

[0041] Preferably, a dual recognition channel is used to perform cluster division of N clusters, that is, the quality of the clustering results is evaluated from the two dimensions of scale rationality and feature consistency, and a quantifiable discrimination confidence is constructed, wherein the dual recognition channel includes a cluster number threshold discrimination sub-channel and a cluster feature alienation threshold discrimination sub-channel. Specifically, the cluster number threshold discrimination sub-channel is used to perform size discrimination of each cluster within the N clusters, that is, to evaluate whether the number of samples in each cluster is reasonable, to avoid noise clusters with too few samples or mixed clusters with too many samples, to set a minimum sample threshold, that is, to set the minimum number of samples for each cluster (such as 100). If the sample size of a cluster is lower than this value, it may be an invalid cluster formed by an outlier or noise; to set a maximum sample threshold, that is, to set the maximum number of samples for each cluster (such as 20% of the total samples) to prevent a single cluster from containing too many heterogeneous samples, resulting in feature dilution. Then the first discrimination confidence is calculated. If the number of cluster samples is between the minimum sample threshold and the maximum sample threshold, the first discrimination confidence is 1; if the number of cluster samples is less than the minimum sample threshold, the first discrimination confidence is the ratio of the number of cluster samples to the minimum sample threshold; if the number of cluster samples is greater than the maximum sample threshold, the first discrimination confidence is the ratio of the maximum sample threshold to the number of cluster samples; among them, the closer the first discrimination confidence is to 1, the more reasonable the cluster size is; when the confidence is lower than 0.5, the cluster may need to be re-clustered or merged.

[0042] Preferably, the cluster feature alienation threshold discrimination subchannel is used to discriminate the degree of feature alienation of each cluster within the N clusters, that is, to evaluate the degree of dispersion of samples within the cluster on key features to ensure that samples within the same cluster have high feature consistency. For example, the variance of samples within the cluster on core features (such as load peak and electricity price lag value) is calculated. The larger the variance, the more serious the feature alienation. Then, the cluster feature alienation threshold is set, that is, the maximum allowable variance of each feature, such as the load peak variance ≤ 100MW2. If the threshold is exceeded, the heterogeneity of the samples within the cluster is considered too high. Then, the second discrimination trust is calculated, that is, by multiplying the trust factors of each feature, a comprehensive assessment of the consistency of the features within the cluster is made. If the variance of a feature far exceeds the threshold, the overall trust will be significantly reduced.

[0043] Preferably, cluster division is completed based on the first discrimination trust and the second discrimination trust. Specifically, the weights of the first discrimination trust and the second discrimination trust are set according to historical data and business needs, and the weighted fusion is performed to obtain a comprehensive trust. If the comprehensive trust is ≥0.8, it is directly retained as an independent cluster; if the comprehensive trust is <0.5, it is split into subclusters or noise samples are discarded; if 0.5≤comprehensive trust <0.8, the cluster is optimized. For example, if the first trust is low (abnormal scale), it is merged with the adjacent cluster; if the second trust is low (feature alienation), the samples in the cluster are re-clustered. Finally, independent clusters and joint clusters are obtained, where the joint cluster indicates that a certain type of electricity price fluctuation requires a multi-regional collaborative explanation, such as price linkage caused by cross-regional electricity transactions, and the sample features within the cluster are highly complementary, such as the load data in area A and the new energy output data in area B jointly affect the electricity price. Through the cluster division mechanism of dual identification channels, the clustering results are intelligently classified into independent and joint modes, thereby maximizing the value of multi-source heterogeneous data and ensuring the accuracy and completeness of electricity price forecasts under the dual constraints of data silos and privacy protection in the power industry.

[0044] Step S500 : After training the decision tree using the independent clusters, perform weak classifier mapping training of the joint clusters.

[0045] Preferably, independent clusters are used to train decision trees separately. Specifically, local features (such as load peak, temperature, and electricity price lag value) are extracted from independent clusters, and irrelevant or redundant features are filtered out. Then, the time series features are differentially transformed and Fourier decomposed to extract periodic patterns; distance calculation and regional coding are performed on geographic spatial features; random forests or gradient boosting trees are selected to improve the model's noise resistance and generalization capabilities, and then feature importance evaluation is performed, for example, through Gini impurity or information gain, to calculate the weight of the impact of each feature on electricity prices; pruning strategies (such as minimum sample number, maximum tree depth) are used to prevent overfitting; multiple sub-decision trees are trained separately and integrated into a decision tree for predicting local electricity price fluctuations (such as specific scenarios such as industrial areas and residential areas).

[0046] Preferably, weak classifier mapping training of the joint clustering cluster is performed. Specifically, the federated averaging method or the secure multi-party computing (MPC) framework is adopted to realize cross-domain parameter exchange, and differential privacy (ε=0.5, δ=1e-5) is used to add noise to the gradient to protect the privacy of the original data; then a lightweight neural network (such as a multi-layer perceptron MLP) or logistic regression is selected as a weak classifier to map the multidimensional features of the joint clustering cluster (such as cross-regional load, renewable energy output) to a low-dimensional space. Each participant uses local data to train a weak classifier, outputs gradient updates, and then performs federated aggregation, that is, the gradients of each participant are aggregated through a central server, the global weak classifier parameters are updated, and then iterative optimization is performed until the model converges. The heterogeneous features of different participants (such as the load in area A and the weather in area B) are mapped to a unified feature space through principal component analysis (PCA) to generate a final weak classifier for capturing the cross-regional electricity price linkage law, such as the impact of inter-regional power trading on electricity prices.

[0047] Step S600: Establishing a large electricity price prediction model based on the trained decision tree and weak classifier.

[0048] Step S600 further includes step S610, extracting cluster bias features based on the joint cluster cluster mapped by the weak classifier; step S620, configuring the global weight dynamic update layer of the mapped weak classifier using the cluster bias features; step S630, establishing the electricity price prediction model based on the decision tree, the weak classifier, and the global weight dynamic update layer.

[0049] Preferably, features that can reflect the fluctuation pattern of electricity prices of specific clusters, i.e., cluster bias features, are extracted from the joint clusters mapped by the weak classifiers. For example, fluctuations in features such as load peaks and equipment start and stop times affect electricity prices, and features such as photovoltaic / wind power output forecast errors and sudden weather changes are strongly correlated with electricity prices. Specifically, the features are extracted based on feature importance analysis, such as using the SHAP value to calculate the contribution of each feature to the output of the weak classifier, and screening features with a contribution ≥ 0.1 as bias features; performing intra-cluster distribution difference detection, such as calculating the distribution entropy of features in each joint cluster. If the entropy value of a feature in a specific cluster is significantly lower than the global entropy (such as the difference > 2 times the standard deviation), it is regarded as a bias feature of the cluster; and finally, the cluster bias features of the joint clusters are obtained.

[0050] Preferably, the weight ratio of the independent decision tree and the joint weak classifier is dynamically adjusted according to the feature distribution of the real-time input data, the cluster to which the input sample belongs is identified, and the bias feature of the corresponding cluster is called to perform weight calculation. Finally, the independent decision trees, joint weak classifiers, and the global weight dynamic update layer are deeply integrated. Specifically, the independent decision tree ensemble is responsible for predicting electricity price components dominated by local factors (such as industrial load fluctuations and local policy impacts); the joint weak classifier network captures the impact of cross-regional synergistic factors (such as cross-regional power trading and fluctuations in renewable energy output) on electricity prices. The leaf node outputs of the decision trees (i.e., local electricity price patterns) are then encoded as vectors and spliced ​​with the intermediate layer features of the weak classifiers. For example, the decision tree output of the electricity price forecast for high-load and high-temperature scenarios is fused with the cross-regional transmission price fluctuation characteristics output by the weak classifier. The global weight dynamic update layer is then used to dynamically adjust the contribution weights of the independent decision trees and the joint weak classifier based on the feature distribution of the input samples. For example, the weight of the load-related decision tree is increased during peak hours, and the weight of the renewable energy output prediction model is increased during off-peak hours. Alternatively, the influence of the joint weak classifier is enhanced in areas with high renewable energy penetration. Ultimately, a large electricity price prediction model is obtained that can capture local electricity price characteristics while reflecting global market laws, ensuring the accuracy and stability of electricity price predictions and a balance between privacy protection and efficiency.

[0051] Furthermore, step S620 also includes step S621, performing multidimensional sample analysis of the weak classifier, and establishing a first basic weight based on the multidimensional sample analysis results, the multidimensional sample analysis including sample quantity analysis, sample quality analysis, and sample coverage analysis; step S622, performing prediction state analysis of the weak classifier, and establishing a second basic weight; step S623, constructing a matching mapping weight library based on the cluster bias characteristics, and establishing a dynamic weight; step S624, configuring a global weight dynamic update layer using the first basic weight, the second basic weight, and the dynamic weight.

[0052] Preferably, a multi-dimensional sample analysis of the weak classifier is performed, including sample size analysis, sample quality analysis, and sample coverage analysis, to quantify the credibility of the weak classifier input samples and reflect the data's ability to support model predictions. Specifically, sample size analysis refers to calculating the global proportion of sample sizes in each region / time period, and assigning higher basic weights to categories with sufficient sample sizes (such as weekday load data); sample quality analysis refers to constructing a quality score (0 to 1 point) based on missing rate, outlier ratio, and feature correlation, and filtering out interference from low-quality data, such as abnormal load data caused by sensor failure; sample coverage analysis refers to calculating the coverage of the sample feature space (such as the proportion of convex hull volume), and assigning higher weights to samples that cover scarce features (such as new energy output data under extremely high temperatures); the three analysis results are weighted fused to obtain the first basic weight.

[0053] Preferably, the prediction status analysis of the weak classifier is performed, including time series stability, trend consistency and abnormal response capability evaluation, to evaluate the prediction reliability of the weak classifier in real time and reflect the current health status of the model. Specifically, time series stability is to calculate the standard deviation of the recent prediction error, such as the fluctuation of MAE in the past 7 days; trend consistency is to judge the degree of fit between the predicted trend and the actual trend, such as the accuracy of the prediction direction for 3 consecutive time periods; abnormal response capability is the rationality of the model prediction deviating from the historical pattern in the event of sudden change in electricity price (such as policy adjustment), and is scored by expert rules (0 to 1 points); the analysis results of these three dimensions are weighted and fused to obtain the second basic weight.

[0054] Preferably, a matching mapping weight library is constructed based on the cluster bias characteristics, and dynamic weights are established. This allows the model preference to be dynamically adjusted based on the business attributes of the cluster (such as regional energy structure and user type). Specifically, key business attributes are parsed from the cluster labels, with an emphasis on load correlation characteristics, with weights tilted toward the decision tree; and with an emphasis on weather and output forecast characteristics, with weights tilted toward the weak classifier. A mapping relationship between business attributes and weights is established to obtain a matching mapping weight library, and then the dynamic weights are calculated and determined. Finally, the first basic weight, the second basic weight, and the dynamic weight are used to configure the global weight dynamic update layer, so that the large-scale electricity price forecast model maintains high accuracy and robustness in a complex electricity market environment.

[0055] Furthermore, step S600 also includes step S640, encrypting the model parameters of the electricity price prediction model and uploading them to the central server; step S650, receiving global feedback from the central server, the global feedback is the global feedback established by the central server after aggregating the model parameters through a global aggregation algorithm after receiving multiple locally uploaded encrypted model parameters; step S660, updating the parameters of the electricity price prediction model according to the global feedback.

[0056] Preferably, the model parameters of the large electricity price prediction model are encrypted, which may include encrypting the tree structure parameters of the independent decision tree (such as split nodes, leaf node values), the neural network weights of the joint weak classifiers, and the fusion parameters of the global weight dynamic update layer. The RSA homomorphic encryption algorithm is used to allow the central server to perform addition and multiplication operations (such as aggregate summation) on the ciphertext without decrypting the original parameters. After each participant (such as a regional power grid company) completes the model training locally, the parameters are encrypted and compressed, and the encrypted parameters are uploaded to the central server through a secure channel (such as TLS1.3). The SHA-256 hash is used for integrity verification during the transmission process.

[0057] Preferably, the central server aggregates the encrypted model parameters of multiple participants into global model parameters to simulate the effect of virtual centralized training while avoiding contact with the original data. Specifically, the central server receives the encrypted parameters of each participant, performs weighted summation based on the proportion of the participant's sample size based on the federal averaging method, and returns the aggregated global parameters to each participant; then performs secure multi-party computing, such as through secret sharing, to decompose the parameters into multiple shares, each participant only holds a part of the share, and the central server aggregates the shares through linear algebra operations without exposing the complete parameters throughout the process; finally, global feedback is generated, including the global aggregated parameters.

[0058] Preferably, the parameters of the large electricity price prediction model are updated according to the global feedback. Specifically, after receiving the global feedback, the participants use the local private key to decrypt, obtain the global aggregate parameters, and calculate the difference between the local model parameters and the global parameters (i.e., gradient update); then regularization is applied to the global update to avoid overfitting cross-domain data, and finally the updated model performance is evaluated using the local validation set. If the performance is improved by ≥3%, the global update is accepted; if the performance degrades, the rollback mechanism is activated, the parameters are restored to the previous version, and the abnormal code is fed back to the central server; thereby, collaborative modeling of multiple institutions is achieved under the premise of data privacy compliance, thereby ensuring the security, efficiency and accuracy of electricity price forecasts.

[0059] Furthermore, step S600 also includes step S670, establishing an evaluation index set; step S680, using the evaluation index set to perform model testing on the electricity price prediction model and establish model testing feedback; step S690, using the model testing feedback to perform focus optimization on the electricity price prediction model.

[0060] Preferably, an evaluation indicator set is composed of prediction accuracy indicators, trend consistency indicators, stability indicators and business adaptation indicators, and the evaluation indicator set is used to conduct multi-scenario testing on the large electricity price prediction model to expose potential problems and generate model test feedback. Specifically, the error indicator is calculated for a single time point or sample to locate the specific prediction error (such as the electricity price is seriously overestimated at a certain moment); statistical indicators are calculated according to time windows (such as daily, weekly, and monthly) to analyze the performance of the model at different time granularities; compared with traditional models (such as ARIMA, LSTM) or benchmark models (such as historical mean prediction) to verify the gain of the large model; and then generate model test feedback, which may include comparison of the specific values ​​of each indicator with the threshold, such as whether the RMSE is lower than the business allowable range; error distribution characteristics, such as positive errors are concentrated in peak hours and negative errors are concentrated in trough hours; weak scenario identification results, such as the significant increase in prediction deviation under extreme weather conditions; model efficiency bottlenecks, such as the decrease in inference speed in complex clustering scenarios. Finally, based on the feedback from model testing, the large electricity price prediction model is optimized in a targeted manner. For example, additional data is collected or synthesized for small sample or low-quality sample scenarios. If the prediction error of an independent cluster cluster is consistently high, the cluster is re-divided, the number of clusters is increased / decreased, and the feature weights of the dual recognition channels are adjusted. According to the cluster cluster bias characteristics (such as industrial user clusters are sensitive to policies, and residential clusters are sensitive to weather), the parameters in the matching mapping weight library are adjusted to strengthen the influencing factors of key features. For cluster clusters or scenarios with large errors, incremental training or transfer learning is performed to avoid local noise interference in the global model. Ultimately, the accuracy, robustness and business adaptability of the large electricity price prediction model are improved.

[0061] In the above, refer to Figure 1 The method for constructing a large-scale electricity price prediction model based on federated learning according to an embodiment of the present invention is described in detail. Figure 2 The present invention describes a system for constructing a large-scale model for electricity price prediction based on federated learning according to an embodiment of the present invention.

[0062] The large-scale electricity price forecasting model construction system based on federated learning according to the embodiment of the present invention is used to solve the technical problems existing in the prior art, such as the difficulty in sharing data in the power industry and the inability to use federated learning to integrate multi-source heterogeneous data, resulting in low electricity price forecasting accuracy and poor generalization ability of the forecasting model. It achieves the technical effect of realizing collaborative modeling of multi-institutional data, enhancing the generalization ability of the electricity price forecasting model and the accuracy of electricity price forecasting. Figure 2 As shown, the electricity price prediction large model construction system based on federated learning includes: a window data set establishment module 10, a sample feature extraction module 20, a multi-dimensional feature clustering module 30, a clustering cluster division module 40, a mapping training module 50, and an electricity price prediction large model establishment module 60.

[0063] A window data set establishment module 10 is used to perform data collection of a time window within a preset time window and establish a window data set, wherein the window data set includes a first data set and a second data set, wherein the first data set is an internal data set and the second data set is a homomorphically encrypted data set; a sample feature extraction module 20 is used to perform sample feature extraction of the window data set and establish a multidimensional feature identification of the sample; a multidimensional feature clustering module 30 is used to perform multidimensional feature clustering using the multidimensional feature identification of the sample, and configure N cluster clusters according to the multidimensional feature clustering results, wherein the multidimensional feature clustering is a repeatable clustering; a cluster partitioning module 40 is used to perform cluster partitioning of N cluster clusters based on dual recognition channels, and establish independent cluster clusters and joint cluster clusters; a mapping training module 50 is used to perform weak classifier mapping training of the joint cluster cluster after training the decision tree separately using the independent cluster cluster; a large electricity price prediction model establishment module 60 is used to establish a large electricity price prediction model based on the trained decision tree and weak classifier.

[0064] The specific configuration of the electricity price prediction large model establishment module 60 will be described in detail below. The electricity price prediction large model establishment module 60 further includes: extracting cluster bias features based on the joint clusters mapped by the weak classifiers; configuring the global weight dynamic update layer of the mapped weak classifiers using the cluster bias features; and establishing the electricity price prediction large model based on the decision tree, the weak classifiers, and the global weight dynamic update layer.

[0065] The specific configuration of the electricity price prediction large model establishment module 60 will be described in detail below. The electricity price prediction large model establishment module 60 further includes: performing multidimensional sample analysis of the weak classifier, establishing a first basic weight based on the multidimensional sample analysis results, and the multidimensional sample analysis includes sample quantity analysis, sample quality analysis, and sample coverage analysis; performing prediction state analysis of the weak classifier to establish a second basic weight; constructing a matching mapping weight library based on the cluster bias characteristics and establishing dynamic weights; and configuring a global weight dynamic update layer using the first basic weight, the second basic weight, and the dynamic weight.

[0066] The specific configuration of the electricity price prediction model building module 60 will be described in detail below. The electricity price prediction model building module 60 further includes: encrypting the model parameters of the electricity price prediction model and uploading them to a central server; receiving global feedback from the central server, wherein the central server aggregates the model parameters using a global aggregation algorithm after receiving multiple locally uploaded encrypted model parameters; and updating the parameters of the electricity price prediction model based on the global feedback.

[0067] The specific configuration of the multidimensional feature clustering module 30 will be described in detail below. The multidimensional feature clustering module 30 further includes: after selecting one or more features from the multidimensional feature identifier as core features, performing a cluster analysis centered on the core features; after any sample is clustered, creating a copy of the corresponding sample to continue the clustering process; after clustering is completed, performing an attraction analysis between clusters, and merging clusters based on the attraction analysis results to configure N clusters.

[0068] The specific configuration of the clustering module 40 will be described in detail below. The clustering module 40 further includes: establishing a cluster quantity threshold discrimination sub-channel, using the cluster quantity threshold discrimination sub-channel to discriminate the size of each cluster within the N clusters, and establishing a first discrimination confidence level; establishing a cluster feature alienation threshold discrimination sub-channel, using the cluster feature alienation threshold discrimination sub-channel to discriminate the degree of feature alienation of each cluster within the N clusters, and establishing a second discrimination confidence level; and completing clustering based on the first discrimination confidence level and the second discrimination confidence level.

[0069] The specific configuration of the sample feature extraction module 20 will be described in detail below. The sample feature extraction module 20 further includes: in the execution of the sample feature extraction of the window data set, the extracted features include time features, power demand features, weather environment features, power supply features, market features, economic features, and geographic space features.

[0070] The specific configuration of the sample feature extraction module 20 will be described in detail below. The sample feature extraction module 20 further includes: after sorting the window dataset in time sequence, establishing a time linkage window; when performing sample feature extraction, establishing an electricity price hysteresis feature based on the time linkage window, and adding the electricity price hysteresis feature to the multidimensional feature identifier.

[0071] The specific configuration of the electricity price prediction model building module 60 will be described in detail below. The electricity price prediction model building module 60 further includes: establishing an evaluation index set; using the evaluation index set to perform model testing on the electricity price prediction model, establishing model testing feedback; and using the model testing feedback to optimize the electricity price prediction model.

[0072] The system for constructing a large-scale electricity price prediction model based on federated learning provided in an embodiment of the present invention can execute the method for constructing a large-scale electricity price prediction model based on federated learning provided in any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution method.

[0073] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, any number of different modules may be used and run on the user terminal and / or server, and the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the present invention.

[0074] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.

Claims

1. A method for constructing a large-scale electricity price prediction model based on federated learning, characterized in that: The method comprises: Within a preset time window, perform data collection of the time window to establish a window data set, wherein the window data set includes a first data set and a second data set, wherein the first data set is an internal data set and the second data set is a homomorphically encrypted data set; Execute sample feature extraction of the window data set and establish a multi-dimensional feature identification of the sample; Performing multidimensional feature clustering using the multidimensional feature identifiers of the samples, and configuring N clusters according to the multidimensional feature clustering results, wherein the multidimensional feature clustering is repeatable clustering; Based on the dual recognition channels, N clusters are divided into clusters, and independent clusters and joint clusters are established; After training the decision tree separately using the independent clusters, performing weak classifier mapping training of the joint clusters; A large electricity price prediction model is established based on the trained decision tree and weak classifier.

2. The method for constructing a large-scale electricity price prediction model based on federated learning according to claim 1, characterized in that: The large electricity price prediction model is established based on the trained decision tree and weak classifier, including: Extracting cluster bias features according to the joint clusters mapped by the weak classifiers; A global weight dynamic update layer of a mapping weak classifier is configured using the cluster bias feature; The electricity price prediction model is established based on the decision tree, the weak classifier, and the global weight dynamic update layer.

3. The method for constructing a large-scale electricity price prediction model based on federated learning according to claim 2, wherein: The global weight dynamic update layer of the weak classifier configured by using the cluster bias feature mapping includes: Performing a multidimensional sample analysis of the weak classifier, and establishing a first basic weight according to the multidimensional sample analysis results, the multidimensional sample analysis including sample quantity analysis, sample quality analysis, and sample coverage analysis; Performing a prediction state analysis of the weak classifier to establish a second basic weight; Constructing a matching mapping weight library based on the cluster bias characteristics and establishing dynamic weights; The global weight dynamic update layer is configured using the first basic weight, the second basic weight, and the dynamic weight.

4. The method for constructing a large-scale electricity price prediction model based on federated learning according to claim 1, wherein: After the large electricity price prediction model is established based on the trained decision tree and weak classifier, the following steps are included: Encrypting the model parameters of the large electricity price prediction model and uploading them to the central server; Receiving global feedback from the central server, wherein the global feedback is global feedback established by the central server after aggregating the model parameters using a global aggregation algorithm after receiving multiple locally uploaded encrypted model parameters; Parameters of the electricity price prediction model are updated according to the global feedback.

5. The method for constructing a large-scale electricity price prediction model based on federated learning according to claim 1, wherein: The multidimensional feature identification of the sample is used to perform multidimensional feature clustering, Configure N clusters based on the multi-dimensional feature clustering results, including: After selecting one or more features in the multidimensional feature identifier as core features, performing cluster analysis centered on the core features; After any sample is clustered, a copy of the corresponding sample is created to continue the clustering process; After clustering is completed, an attraction analysis is performed between clusters, and clusters are merged according to the attraction analysis results to configure N clusters.

6. The method for constructing a large-scale electricity price prediction model based on federated learning according to claim 5, characterized in that: The cluster division of N clusters based on the dual recognition channels and the establishment of independent clusters and joint clusters include: Establishing a cluster number threshold discrimination sub-channel, using the cluster number threshold discrimination sub-channel to discriminate the size of each cluster in the N clusters, and establishing a first discrimination confidence level; Establishing a cluster feature alienation threshold discrimination sub-channel, using the cluster feature alienation threshold discrimination sub-channel to discriminate the degree of feature alienation of each cluster in the N clusters, and establishing a second discrimination confidence level; Cluster division is completed according to the first discrimination confidence and the second discrimination confidence.

7. The method for constructing a large-scale electricity price prediction model based on federated learning according to claim 1, wherein: In the sample feature extraction of the window data set, the extracted features include time features, power demand features, weather environment features, power supply features, market features, economic features, and geographic space features.

8. The method for constructing a large-scale electricity price prediction model based on federated learning according to claim 1, wherein: The performing of sample feature extraction of the window data set further includes: After sorting the window data set in time sequence, a time linkage window is established; When performing sample feature extraction, an electricity price hysteresis feature is established based on the time linkage window, and the electricity price hysteresis feature is added to the multidimensional feature identifier.

9. The method for constructing a large-scale electricity price prediction model based on federated learning according to claim 1, wherein: The method of establishing a large electricity price prediction model based on the trained decision tree and weak classifier further includes: Establish a set of evaluation metrics; Using the evaluation index set to perform model testing on the large electricity price prediction model, and establishing model testing feedback; The model test feedback is used to perform focused optimization of the large electricity price prediction model.

10. A large-scale model construction system for electricity price prediction based on federated learning, characterized by: The system is used to implement the method for constructing a large-scale electricity price prediction model based on federated learning according to any one of claims 1 to 9, and the system includes: A window data set establishment module is used to perform time window data acquisition within a preset time window and establish a window data set, wherein the window data set includes a first data set and a second data set, wherein the first data set is an internal data set and the second data set is a homomorphically encrypted data set; A sample feature extraction module, configured to extract sample features from the window data set and establish a multi-dimensional feature identifier for the sample; A multidimensional feature clustering module, configured to perform multidimensional feature clustering using the multidimensional feature identifiers of the samples, and configure N clusters according to the multidimensional feature clustering results, wherein the multidimensional feature clustering is a repeatable clustering; Clustering cluster partitioning module, used to perform cluster partitioning of N clusters based on dual recognition channels, and establish independent clusters and joint clusters; A mapping training module, configured to perform mapping training of weak classifiers of the joint clusters after training the decision trees separately using the independent clusters; The electricity price prediction large model establishment module is used to establish the electricity price prediction large model based on the trained decision tree and weak classifier.

Citation Information

Cited By

  • Short-term electricity price prediction method and device and computer readable storage medium

    CN121212488A

  • Self-adaptive hierarchical distributed learning method and system for industrial production line data isomerism

    CN121543033A

  • An Adaptive Hierarchical Distributed Learning Method and System for Heterogeneous Data in Industrial Production Lines

    CN121543033B

  • Agent-assisted master-slave federal evolution feature selection model construction method

    CN121682188A

  • A Method for Constructing a Master-Slave Federated Evolutionary Feature Selection Model with Agent Assistance

    CN121682188B