Game mechanism based scenario clustering method for planning of integrated energy system of hydrogen and electricity

By introducing a game theory mechanism to improve the K-means clustering algorithm and combining it with the detection of local anomalies in extreme scenarios, the problem of balancing extreme and conventional scenarios in the electrothermal-hydrogen integrated energy system is solved, generating highly representative planning scenarios and improving clustering accuracy and efficiency.

CN121479372BActive Publication Date: 2026-05-08ELECTRIC POWER RES INST OF STATE GRID ZHEJIANG ELECTRIC POWER COMAPNY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ELECTRIC POWER RES INST OF STATE GRID ZHEJIANG ELECTRIC POWER COMAPNY
Filing Date
2026-01-07
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Traditional clustering methods are difficult to effectively balance extreme and normal scenarios in integrated electrothermal hydrogen energy systems, resulting in low clustering accuracy, insufficient diversity, and low computational efficiency.

Method used

An improved K-means clustering method based on game theory is adopted, combined with the local anomaly factor algorithm, to simulate the game interaction between extreme and normal scenarios. The cluster center selection and anomaly handling strategies are dynamically optimized. The game equilibrium solution is found through collaborative optimization of two parameters, generating a highly representative planning scenario.

Benefits of technology

It improves clustering accuracy and scene diversity, enhances computational efficiency, achieves a quantitative balance between clustering accuracy in normal scenarios and integrity preservation in extreme scenarios, and provides accurate scene support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121479372B_ABST
    Figure CN121479372B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of data clustering, and particularly relates to a method for clustering planning scenarios of an electric-thermal-hydrogen integrated energy system based on a game mechanism. In view of the fact that the clustering method used by the existing electric-thermal-hydrogen integrated energy system cannot simultaneously well handle extreme scenarios and regular scenarios, the method comprises the following steps: preparing and preprocessing multi-dimensional scenario data; applying a local outlier factor algorithm to a multi-dimensional scenario feature matrix to identify extreme daily scenarios; introducing a game mechanism based on game theory to improve K-means clustering, simulating the game interaction between extreme scenarios and regular scenarios, quantifying the dynamic balance relationship between extreme scenarios and regular scenarios, establishing a global optimization interaction logic of the local outlier factor algorithm and the K-means clustering algorithm, finding a game equilibrium solution through double-parameter combination collaborative optimization, and generating a representative planning scenario set containing regular and extreme scenarios. The present application realizes a quantitative balance between the clustering accuracy of regular scenarios and the integrity of extreme scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of integrated energy system planning data clustering technology, specifically involving a clustering method for integrated energy system planning scenarios based on game theory mechanisms. Background Technology

[0002] With the rapid development of renewable energy and the deepening of energy transition, integrated electric-thermal-hydrogen energy systems have become an important way to achieve efficient and low-carbon energy supply because they can integrate renewable energy sources such as photovoltaics and wind turbines, as well as various energy demands such as electricity load, heat load, and hydrogen load. However, due to the intermittency and volatility of renewable energy sources and the complexity of multidimensional loads, the planning of integrated electric-thermal-hydrogen energy systems faces significant challenges.

[0003] In long-term planning, clustering and reduction processing is required for the 8760 hours of multidimensional time-series data throughout the year. Directly using all the original data for optimization calculations would lead to the "curse of dimensionality," resulting in high computational costs and low efficiency. In particular, in the multidimensional scenario of the electrothermal hydrogen system, the diversity of system operation data and the existence of abnormal scenarios (such as extreme weather or load anomalies) make it difficult for traditional clustering methods to generate highly representative planning scenarios. Traditional clustering methods usually treat extreme scenarios as noise and remove them or perform simple smoothing, resulting in insufficient reliability of planning results when dealing with extreme operating conditions, and failing to achieve a good balance between computational efficiency and scenario representativeness.

[0004] Traditional K-means clustering is widely used in energy system scenario analysis, but it has the following limitations: First, the algorithm lacks an active mechanism for identifying extreme scenarios, indiscriminately including extreme scenarios (such as abnormal wind and solar power output due to extreme weather) in the same clustering category as normal scenarios. This approach forces extreme values ​​to participate in the construction of normal clusters, causing cluster centers to shift towards extreme values, thus reducing the clustering accuracy of typical scenarios. Second, although the optimal number of clusters K can be determined through methods such as the elbow method and silhouette coefficient, traditional K-means still focuses on the compactness of normal data when determining the number of clusters, easily overlooking extreme scenarios. This clustering strategy, dominated by normal data, systematically misses extreme operating conditions, resulting in a lack of diversity in the generated planning scenarios and posing reliability risks when dealing with abnormal operating conditions.

[0005] While some existing improvement methods enhance the performance of the K-means algorithm by introducing anomaly detection or optimizing initial point selection, they still fail to address the core challenge of balancing extreme and normal scenarios. For example, anomaly detection methods based on Local Outlier Factor (LOF) can identify outliers, but they are difficult to seamlessly integrate with the K-means clustering process to form a global optimization logic. More critically, existing methods generally lack an effective quantitative balancing mechanism when dealing with the interaction between extreme and normal scenarios, making it impossible to reasonably control the ratio of extreme scenarios to typical normal clustering scenarios: deliberately increasing the proportion of extreme scenarios can improve scenario diversity to cover more system anomalies, but it will also increase data processing volume and scenario complexity, leading to increased computational load and reduced efficiency in subsequent planning of integrated electrothermal-hydrogen energy systems; conversely, focusing on retaining normal scenarios to simplify calculations and reduce load will result in the loss of key system anomalies due to excessive compression of the proportion of extreme scenarios, making the clustering results unable to fully reflect the fluctuation risks and extreme operating conditions in the operation of electrothermal-hydrogen systems, thus limiting its application effectiveness in the planning of integrated electrothermal-hydrogen energy systems. Summary of the Invention

[0006] This invention addresses the shortcomings of existing clustering methods used in integrated electrothermal-hydrogen energy systems, which fail to adequately handle both extreme and normal scenarios simultaneously. It provides a game-theoretic clustering method for planning scenarios in integrated electrothermal-hydrogen energy systems. By simulating the strategic interactions between extreme and normal scenarios through a game mechanism, it dynamically optimizes cluster center selection and abnormal scenario handling strategies, effectively improving clustering accuracy and scenario diversity, and enhancing computational efficiency. This provides precise scenario support for the optimized planning of integrated electrothermal-hydrogen energy systems.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: a clustering method for planning scenarios of an integrated electrothermal-hydrogen energy system based on a game-theory mechanism, wherein the clustering method for planning scenarios of an integrated electrothermal-hydrogen energy system based on a game-theory mechanism includes the following steps:

[0008] S1. Preparation and preprocessing of historical time-series multi-dimensional scenario data of the electrothermal hydrogen integrated energy system;

[0009] S2. Apply the local anomaly factor algorithm to the feature matrix obtained by multi-dimensional scene preprocessing to identify extreme day scenes, and store the extreme day scenes in the extreme scene library, while the rest are stored in the regular scene library.

[0010] S3. Improve K-means clustering by adopting a game theory-based game mechanism, simulate the game interaction between extreme and normal scenarios, quantify the dynamic balance relationship between extreme and normal scenarios, establish the global optimization interaction logic between the local anomaly factor algorithm and the K-means clustering algorithm, find the game equilibrium solution through dual-parameter combination collaborative optimization, and finally generate a representative planning scenario set containing normal and extreme scenarios.

[0011] As an improvement, in S1, the data preparation and preprocessing process includes:

[0012] S1.1 Data Acquisition and Cleaning: Collect M-dimensional time series data of the electrothermal hydrogen integrated energy system with a sequence duration of T, form a data matrix, and denote it as the original input dataset. Fill and clean the data, and calculate the statistical characteristics of the data, where T = N days × 24 hours.

[0013] S1.2, Data Segmentation by Day: The cleaned and filled dataset is divided into N days of daily scene data in 24-hour units, and independent daily scenes are constructed. Each daily scene is a 24-row × M-column matrix. The daily scenes are flattened into n-dimensional vectors to generate daily vector matrices and form a set of daily vectors for subsequent reconstruction error calculation.

[0014] S1.3 Extract statistical features from the scene matrix of the d-th day to capture the operational characteristics of each dimension. Calculate one or more of the following features for each dimension: mean, standard deviation, peak-to-valley difference, and quantile.

[0015] S1.4 Integrate the above features by dimension to form the feature vector of the d-th day scene, construct the feature matrix, and standardize the feature matrix. The standardized feature matrix is ​​used as the input for the subsequent local anomaly factor method for extreme scene detection.

[0016] As an improvement, in S2, the process of identifying extreme day scenes includes:

[0017] S2.1 Apply the local anomaly factor method to the feature matrix to identify extreme scene points by calculating the deviation between the local reachability density and the nearest neighbor density of each day's scene;

[0018] S2.2 Dynamic parameter optimization: Traverse the outlier ratio threshold and evaluate the impact of the outlier ratio threshold on the clustering quality.

[0019] As an improvement, the process in S2.1 includes:

[0020] S2.1.1 Define reachability distance, which is used to measure the reachability between samples;

[0021] S2.1.2 Calculate the local reachability density based on the distance between the sample and its nearest neighbor, reflecting the density level around the sample;

[0022] S2.1.3. Quantify the degree of outlier by comparing the local reachability density of the sample with its nearest neighbors;

[0023] S2.1.4 Set the abnormal ratio threshold to a dynamic adjustment value and determine the criteria for judging extreme days;

[0024] S2.1.5 Output the extreme scenario index and the normal scenario index, representing extreme days and typical days, respectively.

[0025] As an improvement, in S2.2, the clustering quality indicators include the regular day reconstruction error and the silhouette coefficient. The lower the regular day reconstruction error and the higher the silhouette coefficient, the better the compactness and accuracy of the regular day clustering. The number of extreme days and the number of regular days for each test are recorded.

[0026] As an improvement, in S3, the construction process of the game-theory-based game mechanism, S3.1, includes:

[0027] S3.1.1, Treat the clustering party in the normal scenario and the retention party in the extreme scenario as game participants;

[0028] S3.1.2 Quantitative Game Mechanism Payoff Function: Define the positive payoff of the clustering party in the normal scenario as the clustering profile coefficient, the higher the value, the better the compactness of the clustering on the normal day; define the positive payoff of the retention party in the extreme scenario as the coverage rate on the extreme day.

[0029] S3.1.3 Strategies for Constructing the Game Mechanism: The strategies for the regular scenario clustering side include: using the K-means clustering algorithm to group the regular day data after extreme days are filtered, and controlling the compression degree of the regular day representative scenario by setting the number of clusters; the strategies for the extreme scenario retention side include: using the local anomaly factor algorithm to identify extreme days, controlling the filtering ratio of extreme days by setting an anomaly ratio threshold, not performing clustering on the identified extreme days, and directly retaining the original data as the representative scenario.

[0030] As an improvement, in S3, the process S3.2 of improving the K-means clustering algorithm based on the game theory mechanism includes:

[0031] S3.2.1 Input and output of the improved algorithm: The preprocessed standard daily scene features, daily scene vector set, extreme scene index set, normal scene index set, and dynamically optimized anomaly ratio threshold are used as input; the normal daily clustering label set, normal daily cluster center set, and normal scene clustering quality evaluation results are used as output. The normal scene clustering quality evaluation includes normal daily reconstruction error and contour coefficient.

[0032] S3.2.2 Data partitioning based on scene classification results: Extract extreme scene vector subsets and regular scene vector subsets from the daily scene vector set. The extreme scene vector subset is directly used as the representative vector of extreme scenes, and the regular scene vector subset is used as the core processing object of K-means clustering in this step. Based on the statistical results of the number of regular days and extreme days, determine the initial range of the number of clusters for K-means clustering on regular days.

[0033] S3.2.3, Optimization of conventional daily K-means clustering to adapt to game objectives.

[0034] As an improvement, in S3.2.3, the conventional daily K-means clustering optimization process adapted to the game objective includes:

[0035] Construct a clustering objective function. The clustering objective function is to minimize the sum of the squares of the Euclidean distances between the regular day and the corresponding cluster center. The smaller the objective function value, the more concentrated the scene features within the cluster.

[0036] K-means iterative optimization includes cluster center initialization, cluster affiliation assignment, and cluster center update;

[0037] Quality verification: The clustering effect under the current number of clusters is evaluated by using the contour coefficient and the regular daily reconstruction error. If the set target requirements are not met, the number of clusters is adjusted and the iteration is restarted.

[0038] As an improvement, in S3, the construction process of the global optimization interaction logic between the local anomaly factor algorithm and the K-means clustering algorithm, S3.3, includes:

[0039] S3.3.1 Construct a global optimization objective function with clustering quality, global error, and scene diversity as its core;

[0040] S3.3.2, Two-parameter combination collaborative optimization: By traversing the pre-set range of two parameters, the abnormality ratio threshold and the number of clusters, the global objective function value corresponding to each combination is calculated, and the optimal parameter pair is selected.

[0041] As an improvement, in S3, representative scene sets and corresponding visualization curves are generated based on the clustering results. S3.4 includes:

[0042] S3.4.1 Constructing a global representative scene set: including extraction of representative scenes for regular scenes, extraction of representative scenes for extreme scenes, and integration and annotation of the global scene set;

[0043] S3.4.2. Take the scenes in the global scene set as a unit of day, and output the normal day and extreme day corresponding to each dimension separately, and present them in a visual image.

[0044] The beneficial effects of the game-theoretic clustering method for planning scenarios of integrated electrothermal-hydrogen energy systems of this invention are as follows: By introducing a game mechanism to improve the traditional K-means clustering algorithm, and combining it with the detection of extreme scenarios by local anomalies, the method optimizes clustering for multidimensional energy data to generate highly representative planning scenarios; by simulating the strategic interaction between extreme and normal scenarios through the game mechanism, the method dynamically optimizes the selection of cluster centers and the handling strategy for abnormal scenarios, effectively improving clustering accuracy and scenario diversity, and enhancing computational efficiency. It achieves a quantitative balance between clustering accuracy in normal scenarios and the preservation of integrity in extreme scenarios, effectively solving the problems of strong subjectivity in manually setting algorithm parameters and difficulty in balancing clustering accuracy, extreme scenario coverage, and computational efficiency; and providing precise scenario support for the optimized planning of integrated electrothermal-hydrogen energy systems. Attached Figure Description

[0045] Figure 1 This is an overall flowchart of the clustering method for planning scenarios of an integrated electrothermal hydrogen energy system based on a game-theoretic mechanism, according to an embodiment of the present invention.

[0046] Figure 2 This is a flowchart of step S3 of the game-theoretic mechanism-based clustering method for planning scenarios of an integrated electrothermal hydrogen energy system according to an embodiment of the present invention.

[0047] Figure 3 This is a multi-dimensional time-series scene curve of the electrothermal hydrogen integrated energy system obtained after preprocessing in step S1 according to an embodiment of the present invention.

[0048] Figure 4 This is a distribution map of conventional and extreme scenarios obtained by the game-theoretic mechanism-based clustering method for planning scenarios of an integrated electrothermal hydrogen energy system according to an embodiment of the present invention.

[0049] Figure 5 This is a conventional scenario curve obtained by the game-theoretic mechanism-based clustering method for planning scenarios of an integrated electrothermal hydrogen energy system according to an embodiment of the present invention.

[0050] Figure 6 and Figure 7 This is an extreme scenario curve obtained by the game-theoretic mechanism-based clustering method for planning scenarios of an integrated electrothermal hydrogen energy system according to an embodiment of the present invention. Detailed Implementation

[0051] The technical solutions of the embodiments of the present invention will be explained and described below. However, the following embodiments are only preferred embodiments of the present invention and not all of them. Other embodiments obtained by those skilled in the art based on the embodiments in the implementation methods without creative effort are all within the protection scope of the present invention.

[0052] See Figure 1The game-theoretic clustering method for planning scenarios of integrated electrothermal-hydrogen energy systems according to embodiments of the present invention includes the following steps:

[0053] S1. Preparation and preprocessing of historical time-series multi-dimensional scenario data of the electrothermal hydrogen integrated energy system;

[0054] S2. Apply the local anomaly factor algorithm to the feature matrix obtained by multi-dimensional scene preprocessing to identify extreme day scenes, and store the extreme day scenes in the extreme scene library, while the rest are stored in the regular scene library.

[0055] S3. Improve K-means clustering by adopting a game theory-based game mechanism, simulate the game interaction between extreme and normal scenarios, quantify the dynamic balance relationship between extreme and normal scenarios, establish the global optimization interaction logic between the local anomaly factor algorithm and the K-means clustering algorithm, find the game equilibrium solution through dual-parameter combination collaborative optimization, and finally generate a representative planning scenario set containing normal and extreme scenarios.

[0056] In step S1, historical time-series scenario data of the electrothermal hydrogen integrated energy system is acquired, cleaned, segmented, and feature-extracted to prepare data for subsequent clustering. Step S1 includes steps S1.1 to S1.3.

[0057] Step S1.1: Data Acquisition and Cleaning. Historical time-series multi-dimensional scenario data of the integrated electric-thermal-hydrogen energy system is collected. This data consists of M-dimensional time-series data with a sequence duration of T (N days × 24 hours), covering data on photovoltaic output, wind turbine output, electrical load, thermal load, and hydrogen load. This data forms a data matrix, denoted as the original input dataset, and is represented as follows:

[0058]

[0059] in, Indicates the first t Hour i Dimensional observations, T Indicates the duration of the sequence. M This represents the dimension of the time-series scenario; for missing data, a forward imputation method is used, filling the missing values ​​with the values ​​from the previous time step, i.e. For long-term missing values ​​that cannot be repaired by forward padding, they are filled with 0 values. Calculate the statistical properties of the data, including the mean, standard deviation, minimum, and maximum values, and verify them to ensure data quality.

[0060] Step S1.2: Split the data by day. The cleaned and populated dataset... The daily scene data is divided into N days based on 24-hour intervals to construct independent daily scenes. Each daily scene is a 24-row × M-column matrix (M=5 in this embodiment), containing data such as photovoltaic output, wind turbine output, electrical load, heat load, and hydrogen load for that day, denoted as […]. Flatten each day's scene into an n-dimensional vector to generate a daily vector matrix, i.e. The vector dimension is 24 hours × M dimensions, thus forming a daily vector set. This is used for subsequent reconstruction error calculation.

[0061] Step S1.3: Feature extraction. For the first... d individual scenes Extract statistical features to capture the operational characteristics of each dimension. i Calculate the following features:

[0062] (1) Mean: This reflects the daily average level of this dimension;

[0063] (2) Standard deviation: To measure the degree of intraday volatility;

[0064] (3) Peak-to-valley difference: This characterizes the intraday extreme value differences;

[0065] (4) Quantiles: including the 25th percentile , median 75th percentile This reflects the characteristics of data distribution.

[0066] The above features are integrated according to dimensions to form the feature vector of the d-th day scene. And construct the feature matrix .

[0067] To eliminate the influence of different dimensions of features (such as the difference in the numerical range of the mean and standard deviation), the feature matrix is ​​standardized using the z-score standardization formula: ,in This is the mean vector of all daily scene features. The standard deviation vector is the standardized feature matrix. This serves as input for subsequent local anomaly factor methods in extreme scenario detection, ensuring the accuracy of subsequent cluster analysis.

[0068] In step S2, the Local Anomaly Factor (LOF) method is used to identify extreme daily scenarios (such as abnormal loads or abnormal weather), which are then distinguished from regular daily scenarios and added to the extreme scenario library. Step S2 includes steps S2.1 and S2.2.

[0069] Step S2.1: In the standardized feature matrix The LOF method is applied to identify extreme scene points by calculating the deviation between the local reachability density and the nearest neighbor density for each day's scene. Specifically:

[0070] (1) Define reachability distance to measure the reachability between samples. For day scene d and sample o, the reachability distance is: ,in Let be the Euclidean distance between samples a and b. Let be the set of k nearest neighbors of 0. This definition ensures that the reachability distance of nearest neighbor samples is not affected by their own density.

[0071] (2) Calculate the local reachability density (lrd) based on its distance to its nearest neighbor, reflecting the density level around the sample. The local reachability density of day scene d is: Lrd is the reciprocal of the average reachable distance of k-nearest neighbors; the higher the density, the larger the Lrd value.

[0072] (3) Local Outlier Factor (LOF) quantifies the degree of outlierness by comparing the local reachability density of a sample with that of its nearest neighbors: .when When the value is lower than the density of its nearest neighbors, it indicates that the density of d is lower than that of its nearest neighbors. The larger the value, the stronger the outlier, which is an anomaly (extreme scenario). When d is close to the density of its nearest neighbors, it indicates that it is a normal sample.

[0073] (4) Set the threshold for the abnormal proportion The value is dynamically adjusted, ranging from 0.01 to 0.05, to determine the criteria for extreme days. After sorting the LOF values ​​for all daily scenarios, the cumulative percentage is taken. The sample is used as an extreme day (i.e. ,in (The LOF value corresponds to the quantile), while the rest are regular days;

[0074] Output extreme scenario index and regular scene index , representing extreme days and normal days, respectively.

[0075] Step S2.2: Dynamic parameter optimization. Iterate through the abnormal proportion thresholds. To assess its impact on clustering quality, quality indicators include routine daily reconstruction error. And the profile coefficient S, its calculation formula is as follows:

[0076]

[0077]

[0078] in, This is the original vector for a regular day. Its corresponding cluster center; The average distance between a sample and other samples in the same cluster. The average distance between a sample and its nearest heterogeneous sample. S The higher the value, The lower the value, the better the compactness and accuracy of the regular daily clustering.

[0079] Record the extreme number of days m and the normal number of days n for each test to support subsequent optimization of the game mechanism.

[0080] Step S3: Introduce a game theory-based mechanism to improve K-means clustering, simulate the game interaction between extreme and normal scenarios, quantify the dynamic equilibrium relationship between extreme and normal scenarios, establish the global optimization interaction logic between the local anomaly factor algorithm and the K-means clustering algorithm, and find the game equilibrium solution through collaborative optimization of two parameters, generating a representative planning scenario set containing both normal and extreme scenarios. Step S3 includes steps S3.1 to S3.4, such as... Figure 2 As shown.

[0081] Step S3.1, Game Mechanism Design. Based on game theory, a game mechanism is designed, treating extreme daily scenarios and normal daily scenarios as two participants in the game. This includes the following steps.

[0082] Step S3.1.1: Define the game participants. The game participants in this mechanism are two co-optimizing decision-makers:

[0083] Clustering for Common Scenarios: Indexed by Common Scenarios The corresponding data composition and core function is to cluster and compress the majority of regular daily data, replacing a large number of similar scenarios with a small number of cluster centers. The goal is to maximize cluster compactness and reduce the total number of representative scenarios through K-means clustering.

[0084] Extreme scenario preserver: Indexed by extreme scenarios The corresponding data composition focuses on identifying and retaining extreme days with unique characteristics. The goal is to maintain independent representativeness, ensure that the original features of extreme scenarios are not absorbed and smoothed by cluster centers, maximize the fidelity of abnormal patterns and global scenario coverage, and avoid the decline in the integrity of the representative set due to the loss of extreme information.

[0085] Step S3.1.2: Quantification of Game Payoff Function. This step focuses on the global representative scenario quality, defining the payoff composition and quantification formulas for both sides, making the payoffs calculable and comparable.

[0086] Benefits of clustering in typical scenarios: The positive benefit is the cluster profile coefficient S, a higher value indicates better compactness of clustering on a typical day; the cost is the reconstruction error on a typical day. The root mean square error is used as a metric; the smaller the value, the more complete the information retained by the clustering, and the higher the benefit.

[0087] Benefits for the extreme scenario retainer: The positive benefit is the extreme daily coverage rate C, calculated using the following formula: ,in To correctly identify the number of extreme days, This represents the total number of actual extreme days; the cost item represents the error increment caused by misjudgment on extreme days. Its calculation formula is ,in The total reconstruction error, including extreme days, is calculated using the following formula: . A smaller value indicates better retention on extreme days and higher returns. The overall return can be expressed as: ,in These are the weighting coefficients. .

[0088] Step S3.1.3: Constructing the game strategy set. This step sets the strategy space and parameter constraints for both sides to ensure that the strategies are adjustable and adaptable to clustering requirements.

[0089] The strategies for clustering regular scenarios include: using the K-means clustering algorithm to group regular daily data, controlling the compression degree of the regular daily representative scenario by setting the number of clusters k, and only performing clustering on regular daily samples after extreme day screening in step S2 to avoid interference from extreme samples to the cluster centers.

[0090] The strategy for preserving extreme scenarios includes: using the LOF algorithm described in step S2 to identify extreme days, and setting an anomaly ratio threshold. Control the screening ratio of extreme days, and do not perform clustering on the identified extreme days, directly retaining the original data as representative scenarios.

[0091] The strategies of the two modules are subject to multi-dimensional mutual constraints:

[0092] 1. Abnormal proportion threshold The constraint effect: Taking too high a value increases the positive return C of the extreme day retention side, but the reduction of the sample size on regular days will lead to insufficient sample representativeness of the regular day clusters, decreased cluster center stability, and ultimately a decrease in the cluster profile coefficient S, resulting in poorer cluster compactness. When the value is too low, the sufficient sample size of regular days increases S, but it may misclassify some extreme days as regular days. Their unique features are smoothed by clustering, leading to an increase in the misclassification error of extreme days. An upward movement indicates insufficient completeness of the set.

[0093] 2. The constraint effect of the number of clusters k: When the value of k increases, the positive return S of the regular day clustering method can be improved in the early stage by subdividing similar scenarios. However, if k is too high, it will lead to redundancy in the number of clusters and may split similar regular days into different clusters, which will actually reduce S. If the value of k is too low, it will merge regular days with significant differences into the same cluster, increase the data dispersion within the cluster, and increase the reconstruction error of regular days. The increased efficiency makes it difficult to accurately distinguish between different routine operating scenarios, and the compression efficiency target of routine daily clustering is difficult to achieve.

[0094] 3. Indirect constraints with k: Adjustment This will change the regular daily sample size, thus affecting the optimal value of k, and together determining the efficiency and completeness of the global representative set. For example, when the regular daily sample size decreases, k needs to be reduced accordingly to ensure sufficient samples within the cluster.

[0095] Step S3.2: Improve the K-means clustering algorithm based on game theory mechanism. The specific steps are as follows:

[0096] Step S3.2.1: Improve the input and output definitions of the algorithm.

[0097] Input: Standard day scene feature matrix after preprocessing in step S1 Daily scene vector set The extreme scenario index set output by step S2.1 Common scenario index set The abnormal proportion threshold determined after dynamic optimization in step S2.2 .

[0098] Output: Set of regular daily clustering labels ; Regular solar cluster center set ,in , The center of the k-th cluster; Clustering quality assessment results for typical scenarios, including daily reconstruction error. And the profile coefficient S.

[0099] Step S3.2.2: Data partitioning based on scene classification results. This step, based on the accurate separation of extreme and normal scenes completed in step S2, performs targeted partitioning of the input daily scene vector set V, providing a targeted data subset for subsequent clustering.

[0100] Scene vector subset extraction: based on and Extract extreme scenario vector subsets from V respectively Typical scenario vector subset .in, Directly used as a representative vector for extreme scenarios, This is the core processing object of K-means clustering in this step.

[0101] Clustering basic parameters are determined as follows: Based on the statistical results of the number of regular days (n) and the number of extreme days (m) recorded in step S2.2, the initial range of the number of clusters (k) for the K-means clustering of regular days is determined. For example, when n = 350 (regular days account for approximately 96%), based on the requirement to minimize the reconstruction error of regular days in step S2.2, the initial range of k is set to... This avoids the confusion of scene features due to too few clusters, and the data redundancy due to too many clusters.

[0102] Step S3.2.3: Optimization of Conventional Daily K-means Clustering to Adapt to Game Objectives. This step is guided by the core objective of clustering in conventional scenarios within the game mechanism, and optimizes the clustering based on the daily K-means. The specific process for performing K-means clustering optimization is as follows:

[0103] Clustering objective function construction:

[0104]

[0105] in, Let i be the set of daily scenes for the i-th regular scene cluster. Let i be the center vector of the i-th cluster. The objective function value is the squared Euclidean distance between the regular day d and the corresponding cluster center. The smaller the objective function value, the more concentrated the scene features within the cluster.

[0106] K-means iterative optimization:

[0107] Cluster center initialization: using the distance-maximization sampling method from The initial cluster center is selected to avoid local optima caused by random initialization and to ensure that the initial center covers common scenarios with different features;

[0108] Cluster affiliation assignment: Calculate each regular day vector The Euclidean distance to all initial cluster centers will Assign to the nearest cluster and generate initial cluster labels;

[0109] Cluster center update: Based on the current scene set of each cluster, recalculate the cluster center vector. Repeat the cluster affiliation assignment and cluster center update process until the change in the cluster center vector is less than the convergence threshold (e.g., ...). (t is the number of iterations) or reaching the maximum number of iterations (e.g., 100 times).

[0110] Clustering quality verification: The silhouette coefficient S from step S2.2 was compared with the routine daily reconstruction error. Evaluate the clustering effect at the current number of clusters k. If the set target criteria are not met, adjust the number of clusters k and iterate again.

[0111] Step S3.3: Optimize the clustering process through a game theory mechanism to finally determine the game equilibrium solution, which includes the following sub-steps.

[0112] Step S3.3.1: Construction of the global optimization objective function. Based on the profit demands of both sides in the game, and combining the quantitative indicators described in Steps S2.2 and S3.2, a global optimization objective function is constructed with clustering quality, global error, and scenario diversity as its core:

[0113]

[0114] in, The weights are used to balance the clustering quality S with the global error. .

[0115] Step S3.3.2: Two-parameter combination collaborative optimization. This involves iterating through pre-set parameters... Calculate the global objective function value for each parameter range combination and select the optimal parameter pair. For example, when the two-parameter combination takes the value of... When, the objective function If the minimum value is obtained, then For an equilibrium solution, the clustering process for this parameter pair is the optimal process.

[0116] At this point, the conventional daily clustering method achieves S-maximization and Minimize, extreme day retain the square to maximize C and Minimize the Pareto optimality of the benefits for both parties, meaning it cannot be achieved by adjusting the balance alone. Alternatively, k can improve the benefit of one party without harming the other party, and the global representative set can simultaneously satisfy the goals of "efficient clustering in normal scenarios" and "complete preservation of features in extreme scenarios".

[0117] Step S3.4: Generate a representative scene set and corresponding visualization curves based on the clustering results, which specifically includes the following sub-steps:

[0118] Step S3.4.1: Construction of a globally representative scenario set. Based on the parameter constraints of the game equilibrium solution, and integrating the clustering features of regular scenarios with the uniqueness of extreme scenarios, a globally representative scenario set with a clear structure and complete features is constructed:

[0119] Extracting representative data from typical scenarios: calling the corresponding... The set of regular solar cluster centers Each cluster center Let be the 24-hour multi-dimensional representative vector of the i-th type of regular scenario.

[0120] Extreme scenario representative extraction: call the corresponding extreme daily vector subset This vector can be directly used as a representative vector for extreme scenarios without additional compression, ensuring that the original operational features of extreme scenarios are fully preserved.

[0121] Global scene set integration and annotation: merging and To form a global representative scene set And create an indexed association table for each representative scenario that links to the original data's natural date.

[0122] Step S3.4.2: Generation of multi-dimensional visualization curves. The scenes in the global scene set are divided into daily units, and the corresponding regular and extreme days for each dimension are output separately and presented as visualized images.

[0123] For a multi-dimensional time-series scenario of an integrated electrothermal-hydrogen energy system in a certain year, the preprocessed curve is as follows: Figure 3 As shown, Figure 3 In the diagram, the horizontal axis represents time (in hours), and the vertical axis represents load (which can be considered unitless after standardization). The distributions of normal and extreme scenarios obtained using the method of this embodiment are shown below. Figure 4 As shown, the majority of the days are regular, totaling 346, while the number of extreme days is 19.

[0124] The curves obtained using the method of this embodiment for a typical scenario are as follows: Figure 5 As shown. The curves of the extreme scenarios obtained using the method of this embodiment are as follows. Figure 6 and Figure 7 As shown.

[0125] Furthermore, to verify the effectiveness of the game-theoretic mechanism-based clustering method for planning scenarios of integrated electrothermal hydrogen energy systems proposed in this embodiment, the method of this embodiment is compared with the isolated forest method and the percentile quantile method. Four indicators are selected for evaluation: global root mean square error (RMSE), profile coefficient, extreme scenario coverage, and number of extreme scenarios.

[0126] Global RMSE: The root mean square error of the entire dataset reconstructed using representative scenes. It is calculated as the square root of the average of the MSE (mean-square error) of each daily vector and the nearest representative scene. This metric measures the representativeness and accuracy of the overall clustering method; a lower value indicates a smaller reconstruction error and that the representative scenes better cover and approximate the original data.

[0127] Silhouette coefficient: Calculated only for regular daily clustering, it measures intra-cluster compactness and inter-cluster separation. This metric evaluates the quality of regular daily clustering; a higher value indicates higher intra-cluster similarity and greater inter-cluster dissimilarity.

[0128] Extreme scenario coverage: The proportion of detected extreme days to the total number of days, quantifying the coverage of extreme events and evaluating the sensitivity of the method to anomalies and extreme scenarios.

[0129] The metrics calculated by the three algorithms are shown in Table 1 below.

[0130] Table 1 Comparison of Results

[0131]

[0132] As shown in Table 1, the game-theoretic clustering method for integrated energy system planning scenarios based on electrothermal hydrogen proposed in this embodiment demonstrates superior performance compared to other comparative algorithms in balancing conventional and extreme scenarios. The method proposed in this patent maintains a silhouette coefficient of 0.145 while increasing the coverage of extreme scenarios to 5.2%; in contrast, the isolated forest and percentile quantile methods only cover 3.0% and 1.6% of extreme scenarios, respectively, clearly favoring conventional scenarios and neglecting the representativeness of extreme scenarios. This embodiment's method, while improving extreme scenario coverage, still maintains the lowest global RMSE of 87.30, proving that it achieves a balance between the clustering quality of conventional scenarios and the completeness of extreme scenario coverage through a game-theoretic mechanism, providing a scenario set that is both typical and covers extreme scenarios for integrated energy system planning.

[0133] The beneficial effects of the game-theory-hydrogen integrated energy system planning scenario clustering method proposed in this embodiment are as follows: By introducing a game-theory-based mechanism, extreme daily scenarios and normal daily scenarios are regarded as non-cooperative game participants. This not only overcomes the problem that traditional methods are unable to effectively distinguish between extreme and normal scenarios, resulting in the generated scenario set not covering extreme scenarios, but also establishes a global optimization interaction logic between the Local Anomaly Factor (LOF) method and the K-means clustering algorithm, making up for the lack of anomaly detection and the lack of a global closed-loop mechanism in the clustering algorithm. By using the anomaly ratio threshold parameter of LOF and the number of clusters in K-means as game constraints, the strategic interaction between extreme and normal scenarios is simulated. This ensures that normal scenarios minimize intra-cluster reconstruction errors through K-means clustering while retaining extreme scenarios as independent representative scenarios, preventing them from being absorbed by normal clustering. By simulating the strategic interaction between extreme and normal scenarios through the game mechanism, the cluster center selection and anomaly scenario handling strategies are dynamically optimized, effectively improving clustering accuracy and scenario diversity, and enhancing computational efficiency, providing accurate scenario support for the optimized planning of the electric-thermal-hydrogen integrated energy system.

[0134] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Those skilled in the art should understand that the present invention includes, but is not limited to, the content described in the above specific embodiments. Any modifications that do not depart from the functional and structural principles of the present invention will be included within the scope of the claims.

Claims

1. A clustering method for planning scenarios of an integrated electrothermal-hydrogen energy system based on a game theory mechanism, characterized by: The game-theoretic-based clustering method for planning scenarios of integrated electrothermal-hydrogen energy systems includes the following steps: S1. Preparation and preprocessing of historical time-series multi-dimensional scenario data of the electrothermal hydrogen integrated energy system; S2. Apply the local anomaly factor algorithm to the feature matrix obtained by multi-dimensional scene preprocessing to identify extreme day scenes, and store the extreme day scenes in the extreme scene library, while the rest are stored in the regular scene library. S3. Improve K-means clustering by adopting a game theory-based game mechanism, simulate the game interaction between extreme and normal scenarios, quantify the dynamic balance relationship between extreme and normal scenarios, establish the global optimization interaction logic between the local anomaly factor algorithm and the K-means clustering algorithm, find the game equilibrium solution through dual-parameter combination collaborative optimization, and finally generate a representative planning scenario set containing normal and extreme scenarios. In S3, the construction process of the game-theory-based game mechanism, S3.1, includes: S3.1.1, Treat the clustering party in the normal scenario and the retention party in the extreme scenario as game participants; S3.1.2 Quantitative Game Mechanism Payoff Function: Define the positive payoff of the clustering party in the normal scenario as the clustering profile coefficient, the higher the value, the better the compactness of the clustering on the normal day; define the positive payoff of the retention party in the extreme scenario as the coverage rate on the extreme day. S3.1.3, Constructing the strategy set for the game mechanism: The strategy of the regular scenario clustering side includes: using the K-means clustering algorithm to group the regular day data after extreme days are filtered, and controlling the compression degree of the regular day representative scenario by setting the number of clusters; The strategy of the extreme scenario retention side includes: using the local anomaly factor algorithm to identify extreme days, controlling the screening ratio of extreme days by setting an anomaly ratio threshold, and not performing clustering on the identified extreme days, directly retaining the original data as the representative scenario; In S3, the process S3.2 of improving the K-means clustering algorithm based on the game theory mechanism includes: S3.2.1 Input and output of the improved algorithm: The preprocessed standard daily scene features, daily scene vector set, extreme scene index set, normal scene index set, and dynamically optimized anomaly ratio threshold are used as input; the normal daily clustering label set, normal daily cluster center set, and normal scene clustering quality evaluation results are used as output. The normal scene clustering quality evaluation includes normal daily reconstruction error and contour coefficient. S3.2.2 Data partitioning based on scene classification results: Extract extreme scene vector subsets and regular scene vector subsets from the daily scene vector set. The extreme scene vector subset is directly used as the representative vector of extreme scenes, and the regular scene vector subset is used as the core processing object of K-means clustering in this step. Based on the statistical results of the number of regular days and extreme days, determine the initial range of the number of clusters for K-means clustering on regular days. S3.2.3, Optimization of conventional daily K-means clustering to adapt to game objectives; In S3.2.3, the conventional daily K-means clustering optimization process for adapting the game objective includes: Construct a clustering objective function. The clustering objective function is to minimize the sum of the squares of the Euclidean distances between the regular day and the corresponding cluster center. The smaller the objective function value, the more concentrated the scene features within the cluster. K-means iterative optimization includes cluster center initialization, cluster affiliation assignment, and cluster center update; Quality verification: The clustering effect under the current number of clusters is evaluated by using the silhouette coefficient and the daily reconstruction error. If the set target requirements are not met, the number of clusters is adjusted and the iteration is repeated. In S3, the construction process of the global optimization interaction logic between the local anomaly factor algorithm and the K-means clustering algorithm, S3.3, includes: S3.3.1 Construct a global optimization objective function with clustering quality, global error, and scene diversity as its core; S3.3.2, Two-parameter combination collaborative optimization: By traversing the pre-set range of two parameters, the abnormality ratio threshold and the number of clusters, the global objective function value corresponding to each combination is calculated, and the optimal parameter pair is selected.

2. The clustering method for planning scenarios of an integrated electrothermal-hydrogen energy system based on a game-theory mechanism as described in claim 1, characterized in that: In S1, the data preparation and preprocessing process includes: S1.1 Data Acquisition and Cleaning: Collect M-dimensional time series data of the electrothermal hydrogen integrated energy system with a sequence duration of T, form a data matrix, and denote it as the original input dataset. Fill and clean the data, and calculate the statistical characteristics of the data, where T = N days × 24 hours. S1.2, Data Segmentation by Day: The cleaned and filled dataset is divided into N days of daily scene data in 24-hour units, and independent daily scenes are constructed. Each daily scene is a 24-row × M-column matrix. The daily scenes are flattened into n-dimensional vectors to generate daily vector matrices and form a set of daily vectors for subsequent reconstruction error calculation. S1.3 Extract statistical features from the scene matrix of the d-th day to capture the operational characteristics of each dimension. Calculate one or more of the following features for each dimension: mean, standard deviation, peak-to-valley difference, and quantile. S1.4 Integrate the above features by dimension to form the feature vector of the d-th day scene, construct the feature matrix, and standardize the feature matrix. The standardized feature matrix is ​​used as the input for the subsequent local anomaly factor method for extreme scene detection.

3. The clustering method for planning scenarios of an integrated electrothermal hydrogen energy system based on a game-theory mechanism as described in claim 1, characterized in that: In S2, the process of identifying extreme day scenes includes: S2.1 Apply the local anomaly factor method to the feature matrix to identify extreme scene points by calculating the deviation between the local reachability density and the nearest neighbor density of each day's scene; S2.2 Dynamic parameter optimization: Traverse the outlier ratio threshold and evaluate the impact of the outlier ratio threshold on the clustering quality.

4. The clustering method for planning scenarios of an integrated electrothermal-hydrogen energy system based on a game-theory mechanism as described in claim 3, characterized in that: The process in S2.1 includes: S2.1.1 Define reachability distance, which is used to measure the reachability between samples; S2.1.2 Calculate the local reachability density based on the distance between the sample and its nearest neighbor, reflecting the density level around the sample; S2.1.

3. Quantify the degree of outlier by comparing the local reachability density of the sample with its nearest neighbors; S2.1.4 Set the abnormal ratio threshold to a dynamic adjustment value and determine the criteria for judging extreme days; S2.1.5 Output the extreme scenario index and the normal scenario index, representing extreme days and typical days, respectively.

5. The clustering method for planning scenarios of an integrated electrothermal-hydrogen energy system based on a game-theory mechanism as described in claim 4, characterized in that: In S2.2, the clustering quality indicators include the regular day reconstruction error and the silhouette coefficient. The lower the regular day reconstruction error and the higher the silhouette coefficient, the better the compactness and accuracy of the regular day clustering. Record the number of extreme days and the number of regular days for each test.

6. The clustering method for planning scenarios of an integrated electrothermal-hydrogen energy system based on a game-theory mechanism as described in claim 1, characterized in that: In S3, representative scene sets and corresponding visualization curves are generated based on the clustering results. S3.4 includes: S3.4.1 Constructing a global representative scene set: including extraction of representative scenes for regular scenes, extraction of representative scenes for extreme scenes, and integration and annotation of the global scene set; S3.4.

2. Take the scenes in the global scene set as a unit of day, and output the normal day and extreme day corresponding to each dimension separately, and present them in a visual image.

Citation Information

Patent Citations

  • Power consumption data abnormal value detection method and system

    CN118070191A

  • Game optimization operation method and device based on multiple energy storage operation subjects

    CN118690950A