Comprehensive energy ems source and load collaborative optimization method based on multi-dimensional feature k-means clustering

CN122840801APending Publication Date: 2026-09-29POTEVIO TELECOMM CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611065715.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-17
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

现有相关技术主要集中于单一维度的场景聚类,例如仅针对荷侧负荷曲线进行聚类,或仅关注源侧分布式电源出力特征,典型现有技术可参考中国发明专利CN114863768A《一种基于K-means聚类的微电网源荷协同优化方法》,该专利仅围绕负荷侧的用电特征进行聚类,未考虑源侧、网侧及环境侧的耦合影响

Benefits of technology

1.聚类精度显著提升,场景划分更精准:本专利通过构建四维耦合多维特征集,结合动态权重机制,解决了现有技术特征单一、权重固定的问题,能够更全面、精准地刻画综合能源系统的运行状态;同时,DBSCAN辅助初始化避免了传统K-means的局部最优问题,聚类结果稳定性显著提升,场景识别更贴合系统实际运行状态,为源荷协同优化提供了可靠的场景依据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840801A_ABST
    Figure CN122840801A_ABST
Patent Text Reader

Abstract

This invention, titled "A Source-Load Co-optimization Method for Integrated Energy Management Systems Based on Multi-Dimensional Feature K-means Clustering," belongs to the interdisciplinary field of power system and integrated energy management. The technical problems it addresses are the limitations of existing technologies, such as single feature dimensions, poor stability of clustering algorithms, lack of dynamic adaptability, and unreasonable weight allocation. The key technical solution involves constructing a four-dimensional coupled feature set of sources, grid, load, and ring, employing an improved K-means clustering algorithm, and combining it with DBSCAN algorithm initialization and a dynamic weighting mechanism to accurately classify integrated energy system operation scenarios, match corresponding source-load co-optimization strategies, and dynamically update the model and strategies through a closed-loop iterative mechanism. Ultimately, this achieves economical, efficient, and low-carbon operation of the integrated energy system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of power system and integrated energy management, specifically to the direction of smart energy management and optimized dispatch. Background Technology

[0002] Currently, Integrated Energy Systems (IES) have become the core carrier for achieving "dual carbon" goals and building new power systems. They integrate multiple energy forms such as electricity, heat, cooling, and gas, and are characterized by multi-source complementarity, diverse load sides, and network coupling. The Energy Management System (EMS), as the "brain" of the IES, is responsible for the coordination of source-side power output scheduling and load-side load regulation. Its optimization effect directly determines the system's economic efficiency, high efficiency, and low-carbon operation.

[0003] Among existing technologies both domestically and internationally, source-load co-optimization methods based on clustering algorithms have been widely applied. K-means clustering, in particular, has become the mainstream technology for scene segmentation and source-load matching due to its simple principle and high computational efficiency. Existing related technologies mainly focus on single-dimensional scene clustering, such as clustering only the load-side load curve or only focusing on the output characteristics of distributed power sources on the source side. A typical existing technology can be found in Chinese invention patent CN114863768A, "A Microgrid Source-Load Co-optimization Method Based on K-means Clustering," which only clusters based on the electricity consumption characteristics of the load side and does not consider the coupling effects of the source side, grid side, and environment side.

[0004] Research and analysis have revealed that existing integrated energy system (EMS) source-load co-optimization technologies based on K-means clustering have the following core drawbacks and shortcomings, which seriously affect the optimization effect and stability of the system and cannot meet the operational requirements of complex integrated energy systems: 1. Limited feature dimension and incomplete characterization: Existing technologies mostly focus on features of only one dimension on the source side or the load side, ignoring key coupling factors such as network side (line loss, network topology) and environment side (light, temperature, wind speed). This results in clustering scenarios failing to truly reflect the actual operating status of the integrated energy system, thus affecting the accuracy of source-load collaborative optimization.

[0005] 2. Poor stability of clustering algorithms: The initial cluster centers of the traditional K-means clustering algorithm are randomly selected, which is prone to getting trapped in local optima. This leads to large fluctuations in clustering results. The same set of running data may produce different scenario partitioning results, which cannot provide a stable and reliable scheduling basis for EMS.

[0006] 3. Lack of dynamic adaptive capability: Existing technologies mostly adopt fixed clustering models and optimization strategies, without considering the concept drift during system operation. When environmental conditions, load characteristics, and source-side output change, the clustering model cannot be updated in real time, the robustness of the optimization strategy is poor, and the optimization effect decreases significantly after long-term operation.

[0007] 4. Unreasonable weight allocation: In existing technologies, if multi-feature clustering is involved, a fixed weight allocation method is often used. The weights are not dynamically adjusted according to the fluctuation of different features and their impact on system optimization, which weakens the role of key features and results in insufficient clustering accuracy. Summary of the Invention

[0008] The purpose of this invention is to provide: A source-load coordinated optimization method for integrated energy systems (EMS) based on multi-dimensional feature K-means clustering is proposed. The core idea is to construct a four-dimensional coupled feature set of source, grid, load, and ring, adopt an improved K-means clustering algorithm, and combine the DBSCAN algorithm initialization and dynamic weight mechanism to achieve accurate division of integrated energy system operation scenarios, match corresponding source-load coordinated optimization strategies, and realize dynamic updates of models and strategies through a closed-loop iterative mechanism, ultimately achieving economical, efficient, and low-carbon operation of integrated energy systems.

[0009] Terminology Explanation: Unless otherwise defined, all technical terms in this document have the same meanings as commonly understood by one of ordinary skill in the art to which the subject matter of the claims pertains. Unless otherwise stated, all patents, patent inventions, and publications cited in this document are incorporated herein by reference in their entirety. If multiple definitions exist for terms in this document, the definitions in this chapter shall prevail.

[0010] It should be understood that the above brief description and the following detailed description are exemplary and for illustrative purposes only, and do not limit the subject matter of the invention in any way. In this invention, the singular is used in conjunction with the plural unless otherwise specifically stated. It should also be noted that, unless otherwise stated, the use of “or” or “or” means “and / or”. Furthermore, the use of the term “comprising” and other forms such as “including,” “containing,” and “contains” are not limiting.

[0011] Unless otherwise stated, existing or conventional methods within the scope of the art shall be used.

[0012] The term "K-means clustering" used in this article refers to a classic unsupervised partitioning clustering algorithm. Its core principle is to iteratively divide the dataset into K user-specified clusters, with the goal of minimizing the within-cluster squared error, making the features of data in the same cluster as similar as possible, and the differences between different clusters as obvious as possible. In the scenario of integrated energy EMS source-load collaborative optimization, K-means clustering mainly undertakes the functions of data dimensionality reduction and feature classification.

[0013] The term "source-load coordinated optimization" used in this article is a core technology for integrated energy systems to achieve supply and demand balance, reduce costs and increase efficiency. It mainly addresses the system stability and economic issues caused by fluctuations in output on the source side (energy production end) and load on the load side (user demand end). By coordinating and scheduling resources on both sides, it achieves precise matching of energy supply and demand.

[0014] The term "Integrated Energy System (IES)" used in this article refers to an integrated system for production, supply, and sales in the field of new energy. Its core is to improve energy utilization efficiency through multi-energy complementarity and support sustainable energy development under the "dual carbon" goal. It is supported by advanced physical information technology and innovative management models. In the planning, construction, and operation stages, it organically coordinates and optimizes the production, transmission, conversion, storage, and consumption of various energy sources. It integrates multiple heterogeneous energy subsystems such as electricity, natural gas, heat, and renewable energy to form an integrated energy production, supply, and sales system. Its core connotation is multi-energy complementarity and coordinated optimization.

[0015] The term "Energy Management System (EMS)" used in this article refers to an energy management system based on information technology. Its core function is to optimize energy use through real-time monitoring and analysis, thereby reducing energy costs and improving energy efficiency. It is currently widely used in power grid dispatching, integrated energy, and energy storage. The EMS adopts a layered distributed architecture, mainly comprising four layers: a data acquisition layer (sensors, smart meters), a data transmission layer (4G / IoT platform), a data storage layer (time-series database), and an application layer (analysis, monitoring, and optimization modules). It supports unified management of multiple energy sources, including electricity, gas, and water.

[0016] The term "DBSCAN (Density-Based Spatial Clustering of Applications with Noise)" used in this paper refers to density-based spatial clustering applications with noise. It is a typical density-based unsupervised clustering algorithm, which differs significantly from partition-based K-means clustering. It is suitable for discovering clusters of arbitrary shapes and can automatically identify noise outliers.

[0017] The term "scene drift judgment" used in this paper is a technical means used in machine learning modeling to detect changes in data distribution over time, which leads to a decline in the performance of the original model. In a cluster-based source-load collaborative optimization system, it is used to identify the feature distribution shift that occurs in the operation scenario of integrated energy source and load, and to ensure the long-term effectiveness of the optimization model.

[0018] The term “storage SOC” used in this article refers to the state of charge of an energy storage battery. It is the most critical monitoring parameter in an electrochemical energy storage system, indicating the percentage of the energy storage battery’s current remaining capacity relative to its rated capacity. It is used to intuitively reflect the remaining capacity of the energy storage battery.

[0019] In a first aspect, the present invention provides: The integrated energy EMS source-load collaborative optimization method based on multidimensional feature K-means clustering includes the following steps: Step S1, Data Acquisition and Preprocessing: Input the raw operating data of the integrated energy EMS, collect four-dimensional raw data covering the source side, load side, grid side, and environment side, and perform data preprocessing on the four-dimensional raw data; Step S2, Multidimensional Feature Set Construction and Normalization: Based on the preprocessed four-dimensional raw data, core dimensions are selected, redundant features are removed, and a multidimensional feature set of source, network, load, and ring four-dimensional coupling is constructed. Step S3: Improved K-means clustering initialization: The DBSCAN algorithm is used to identify the core feature points in the multidimensional feature set. These core feature points are then clustered, and the mean vector is calculated as the initial center of the K-means cluster to determine the optimal number of clusters K. The dynamic weight vector is calculated using the analytic hierarchy process (AHP) combined with the variance and coefficient of variation of the feature data. W ; Step S4, Improve K-means clustering scenario partitioning: Based on the dynamic weight vector W Calculate the weighted Euclidean distance from each feature point to the initial center, assign each feature point to the category of the nearest cluster center, forming K initial clusters; calculate the weighted mean of all feature points in each cluster as the new cluster center, and iteratively calculate the final clustering result; based on the clustering result, identify the system operation scenario; Step S5, Source-Load Co-optimization Strategy Matching and Execution: A source-load co-optimization strategy library is preset. Based on the current system operation scenario identified in step S4, the corresponding optimization strategy is matched from the strategy library. The energy management system (EMS) issues scheduling instructions to the source-side, load-side, and grid-side devices to perform corresponding operations. Step S6, Optimization effect monitoring and scene drift judgment: Real-time monitoring of key operating indicators of the optimized system to generate an optimization effect report; Calculation of the weighted Euclidean distance between the real-time feature vector and each cluster center to determine whether the scene drift condition is met; Step S7, Iterative Update or Continuous Operation: If the judgment result is "drift", return to steps S2-S6 to re-execute the clustering process; if the judgment result is "no drift", maintain the current clustering model and optimization strategy, continue to execute the source-load co-optimization operation in step S5, and continuously monitor the system operation data to enter the next round of monitoring and judgment cycle.

[0020] Based on further solutions to the technical problems of the present invention, or simultaneous solutions to multiple technical problems, the preferred solution in the technical solution provided in the first aspect of the present invention includes: The first preferred option: In step S3, the method for calculating the dynamic weights includes: determining the initial weights using the analytic hierarchy process (AHP) and then dynamically correcting them, with the specific formula as follows: ; in, W i represents the final dynamic weight of the i-th feature dimension, which is dimensionless and ranges from [0,1]. The sum of the weights of all dimensions is 1. This represents the initial subjective weight of the i-th feature dimension determined by the Analytic Hierarchy Process (AHP), which is dimensionless and ranges from [0,1]. α This represents the weight adjustment coefficient, which is dimensionless and ranges from [0.1, 0.5]. It is used to adjust the influence of the variance coefficient on the weights. The coefficient of variation (COP) represents the variance of the sample data for the i-th feature dimension. It is dimensionless and reflects the degree of fluctuation in the feature data for that dimension. The calculation formula is: ; in, Let be the standard deviation of the feature data in the i-th dimension. Let be the mean of the feature data in the i-th dimension. The larger the value, the greater the feature fluctuation, and the greater the impact on the clustering results; therefore, the weight should be increased accordingly.

[0021] The second preferred option: In step S4, the method for calculating the weighted Euclidean distance from the feature points to each initial cluster center includes: ; in, d ij This represents the weighted Euclidean distance from the i-th feature point to the j-th cluster center. x ik This represents the feature value of the k-th dimension of the i-th feature point. c jk This represents the feature value of the k-th dimension of the j-th cluster center. Wk This represents the dynamic weight of the k-th dimension.

[0022] Preferably, in step S4, the termination condition for the iterative calculation of cluster centers is that the change in cluster centers is less than a preset threshold, or the number of iterations reaches a preset maximum value.

[0023] Preferably, in step S4, the typical scenarios identified based on the clustering results include high energy consumption peak scenarios, stable operation scenarios, and high renewable energy output scenarios. The characteristics of the high-energy-consuming peak-hour scenario are: the total load is more than 1.2 times the average load in the current analysis period, the proportion of interruptible load is lower than the preset demand response threshold, the line loss rate is higher than the average, and the output of renewable energy is lower than the average. The characteristics of the stable operation scenario are: the total load is between 0.8 and 1.2 times the average load in the current analysis period, the load fluctuation coefficient is lower than the preset fluctuation threshold, the line loss rate is near the average, and the output of renewable energy is stable; The characteristics of the high renewable energy output scenario are: the output of photovoltaic / wind turbines is more than 1.3 times higher than the average output in the current analysis period, the total load is lower than the average, the energy storage SOC is higher than the preset high threshold, and the line transmission efficiency is higher than the average.

[0024] The third preferred option: In step S5, the EMS scheduling instructions include source-side equipment output instructions, load-side load control instructions, and grid-side line adjustment instructions.

[0025] Fourth preferred option: In step S6, the scene drift condition includes calculating the weighted Euclidean distance between the real-time feature vector and each cluster center. If continuous t Each data collection moment, t If the value is ≥10, it is considered a preset warning value. If the distance between the real-time feature vector and the cluster center of the scene it belongs to is greater than the preset drift threshold, then scene drift is determined to have occurred; otherwise, it is determined to have no scene drift.

[0026] The fifth preferred option: The four-dimensional raw data includes source side, load side, network side, and environment side.

[0027] Preferably, in step S1, the core dimension includes: Source-side power output characteristics: normalized value of actual photovoltaic power output, normalized value of actual wind turbine power output, and normalized value of energy storage SOC; Load characteristics on the load side: normalized total load, proportion of interruptible load, load fluctuation coefficient; Network topology characteristics: normalized line loss, node voltage deviation rate, and network transmission efficiency; Environmental impact characteristics: normalized values ​​of light intensity, normalized values ​​of ambient temperature, and normalized values ​​of wind speed.

[0028] The sixth preferred option: In step S2, the method for judging redundant features includes: using the Pearson correlation coefficient method to calculate the correlation coefficient between any two features. If the absolute value of the correlation coefficient is greater than or equal to 0.8, it is judged as a redundant feature.

[0029] The present invention has at least the following beneficial effects: 1. Significantly improved clustering accuracy and more precise scene segmentation: This patent solves the problems of single features and fixed weights in existing technologies by constructing a four-dimensional coupled multi-dimensional feature set and combining it with a dynamic weight mechanism. This enables a more comprehensive and accurate characterization of the operating status of the integrated energy system. At the same time, DBSCAN-assisted initialization avoids the local optima problem of traditional K-means, significantly improves the stability of clustering results, and makes scene recognition more consistent with the actual operating status of the system, providing a reliable scene basis for source-load collaborative optimization.

[0030] 2. Significant system optimization results and improved energy utilization efficiency: This patent achieves coordinated control of source-side output and load-side load through precise scenario segmentation and strategy matching, effectively improving the renewable energy absorption capacity, reducing line losses, and significantly improving the economy and efficiency of system operation; combined with a closed-loop iteration mechanism, the optimization effect can be maintained for a long time, avoiding optimization failure caused by system state drift.

[0031] 3. The algorithm is robust and adaptable to complex operating scenarios: The closed-loop iteration mechanism of this patent can detect scenario drift in real time and dynamically update the clustering model and optimization strategy. It is adapted to the operating characteristics of integrated energy systems with "multi-source coupling and variable states". Compared with existing fixed model algorithms, the robustness is significantly improved. It can adapt to complex scenarios such as extreme weather, load changes, and source-side power output fluctuations, ensuring long-term stable operation of the system.

[0032] 4. The algorithm is highly scalable and adaptable to various application scenarios: The multi-dimensional feature set of this patent can be flexibly adjusted according to different application scenarios, the number of clusters K can be expanded according to actual needs, and the optimization strategy library can be further enriched to adapt to the optimization needs of different types of integrated energy systems, making it highly versatile.

[0033] 5. Highly efficient algorithm and easy to implement in engineering: The improved K-means clustering algorithm in this patent improves clustering accuracy without significantly increasing the amount of computation, and its computational efficiency meets the requirements of real-time online optimization; the algorithm steps are clear and the logic is rigorous, and it can be modularly deployed through programming, making it easy to interface with existing integrated energy EMS systems and highly practical for engineering applications. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 It is a comprehensive energy EMS source-load co-optimization method based on multidimensional feature K-means clustering. Detailed Implementation

[0036] The following non-limiting embodiments are intended to enable those skilled in the art to gain a more comprehensive understanding of the present invention, but do not limit the invention in any way. The following content is merely an exemplary description of the scope of protection claimed by the present invention, and those skilled in the art can make various changes and modifications to the present invention based on the disclosed content, and such changes should also fall within the scope of protection claimed by the present invention.

[0037] Example This embodiment provides a comprehensive energy EMS source-load collaborative optimization method based on multi-dimensional feature K-means clustering, including the following steps: Step S1: Data Acquisition and Preprocessing 1. Input: Raw operating data from the source side, grid side, load side, and environment side of the integrated energy system. The data comes from the system's built-in sensors, data acquisition terminals, and meteorological monitoring equipment. The acquisition frequency is 15 minutes / time, and the acquisition cycle is 72 consecutive hours (to ensure the integrity and representativeness of the data).

[0038] 2. Perform the operation: (1) Data Acquisition: Clearly define the specific data types to be acquired, ensuring coverage of four-dimensional features, specifically including: ① Source-side data: Rated power, actual output power, and ramp rate of distributed power sources (photovoltaics, wind turbines); SOC (State of Charge) value and charging / discharging power of energy storage systems; ② Load-side data: Real-time power consumption, load curves, and response potential (proportion of interruptible load and amount of transferable load) for industrial, commercial, and residential loads. ③ Network-side data: line transmission power, line loss, node voltage, network topology parameters (line resistance, reactance); ④ Environmental data: light intensity, ambient temperature, wind speed, humidity, and weather forecast level (level 1-5, with level 1 being the best and level 5 being extreme weather).

[0039] (2) Data preprocessing: For the collected raw data, the following processing operations are performed to avoid abnormal data affecting the clustering and optimization results: ① Missing value handling: Linear interpolation is used to fill in missing data. For data with missing data at 3 or more consecutive collection points, all dimensions of data in that time period are removed to avoid excessive interpolation error; ② Outlier handling: using 3 s The criteria identify outlier data (data exceeding the mean ± 3 standard deviations is considered outlier) and replace the outlier data with the mean of the data in that dimension. ③ Data Standardization: The min-max normalization method is used to map all data to the [0,1] interval to eliminate the influence of units. The normalization formula is: Where x is the original data, The minimum value of the data in this dimension. The maximum value of the data in this dimension. This is the normalized data.

[0040] 3. Output: Preprocessed source, network, load, and ring four-dimensional raw data in matrix format, with each row representing a collection time and each column representing a data dimension, without missing data, anomalies, or dimensions.

[0041] Step S2: Construction and Normalization of Multidimensional Feature Set 1. Input: The preprocessed raw data output from step 1.

[0042] 2. Perform the operation: (1) Multidimensional feature selection and construction: Based on the preprocessed raw data, core features are selected, redundant features are removed, and a multidimensional feature set containing 12 core dimensions is constructed to ensure that the features can comprehensively and accurately depict the operating status of the integrated energy system. The specific features are as follows: ① Source-side power output characteristics (3 dimensions): normalized value of actual photovoltaic power output, normalized value of actual wind turbine power output, and normalized value of energy storage SOC; ② Load characteristics (3 dimensions): normalized total load, proportion of interruptible load, load fluctuation coefficient (the ratio of the difference between the maximum and minimum hourly load to the mean). ③ Network topology characteristics (3 dimensions): normalized line loss, node voltage deviation rate, and network transmission efficiency; ④ Environmental impact characteristics (3 dimensions): normalized value of light intensity, normalized value of ambient temperature, and normalized value of wind speed.

[0043] (2) Redundant Feature Removal: The Pearson correlation coefficient method is used to calculate the correlation coefficient between any two features. If the absolute value of the correlation coefficient is greater than or equal to 0.8, it is determined to be a redundant feature. Features with a small impact on system optimization are removed (the weights are initially determined by the analytic hierarchy process, and features with lower weights are removed). Finally, 12 core dimensional features are retained to form a multidimensional feature vector X = [x1, x2, ..., x 12 ], where x i (i=1,2,...,12) represents the feature value of the i-th dimension.

[0044] (3) Feature normalization: The constructed multidimensional feature set is normalized again by min-max (consistent with the normalization method in step 1) to ensure that all features are on the same order of magnitude and to avoid the excessive influence of a certain dimension feature on the clustering results.

[0045] 3. Output: A standardized multidimensional feature set (12-dimensional feature vector matrix) for subsequent cluster analysis.

[0046] Step S3: Improved K-means clustering initialization (DBSCAN assistance + dynamic weight assignment) 1. Input: The standardized multidimensional feature set output from step 2.

[0047] 2. Execution Operation: The core of this step is to address the shortcomings of traditional K-means clustering, such as random initial centers and unreasonable weight allocation. It is executed in two sub-steps: (1) DBSCAN algorithm preprocessing and initial cluster center determination: ① DBSCAN Algorithm Definition: DBSCAN = [Density-Based Spatial Clustering of Applications with Noise] = [Density-based spatial clustering application algorithm], used to identify core points, boundary points and noise points in a multidimensional feature set. After removing noise points, core points are selected as the initial centers of K-means clustering, avoiding the local optimum problem caused by random initial centers.

[0048] ② Algorithm parameter settings: Set the neighborhood radius e =0.5 (dimensionless, calibrated according to the distribution characteristics of the multidimensional feature set), minimum number of points MinPts=5 (i.e., a core point has at least 5 feature points in its neighborhood), and the multidimensional feature set is clustered using the DBSCAN algorithm to identify all core feature points.

[0049] ③ Initial cluster center selection: Cluster the identified core feature points, calculate the mean vector of each core point, and use the mean vector as the initial center of K-means clustering; at the same time, determine the optimal number of clusters K based on the number of core points in the cluster (K=3 in this patent, corresponding to "high energy consumption peak time scenario", "stable operation scenario" and "high renewable energy output scenario", which can be dynamically adjusted according to the actual system scenario).

[0050] (2) Dynamic weighting mechanism: ① Core objective: To dynamically allocate weights based on the fluctuation of each feature and its impact on the collaborative optimization of system source and load, balance the contribution of each feature to the clustering results, and improve clustering accuracy.

[0051] ② Weight Calculation Formula: Initial weights are determined using the Analytic Hierarchy Process (AHP), and dynamically adjusted based on the variance and coefficient of variation of the feature data. The specific formula is as follows:

[0052] ③ Formula parameter description: in, W i represents the final dynamic weight of the i-th feature dimension, which is dimensionless and ranges from [0,1]. The sum of the weights of all dimensions is 1. This represents the initial subjective weight of the i-th feature dimension determined by the Analytic Hierarchy Process (AHP), which is dimensionless and ranges from [0,1]. α This represents the weight adjustment coefficient, which is dimensionless and ranges from [0.1, 0.5]. It is used to adjust the influence of the variance coefficient on the weights. The coefficient of variation (COP) represents the variance of the sample data for the i-th feature dimension. It is dimensionless and reflects the degree of fluctuation in the feature data for that dimension. The calculation formula is: ; in, Let be the standard deviation of the feature data in the i-th dimension. Let be the mean of the feature data in the i-th dimension. The larger the value, the greater the feature fluctuation, and the greater the impact on the clustering results; therefore, the weight should be increased accordingly.

[0053] ④ Weight normalization: Normalize the calculated dynamic weights of each dimension. W i Normalization is performed to ensure that the sum of the weights of all dimensions is 1, resulting in the final weight vector. W = [ W 1, W 2, ..., W 12 ].

[0054] 3. Output: Initial cluster centers of K-means clustering, optimal number of clusters K, and final dynamic weight vectors for each feature dimension. W .

[0055] Step S4: Improve the partitioning of K-means clustering scenarios 1. Input: Standardized multidimensional feature set output from step 2, initial cluster centers output from step 3, optimal number of clusters K, dynamic weight vector W .

[0056] 2. Perform the operation: (1) Improve K-means clustering operation: use dynamic weight vector W By incorporating the K-means clustering algorithm, the weighted Euclidean distance from each feature point to each initial cluster center is calculated using the following formula:

[0057] in, d ij The weighted Euclidean distance from the i-th feature point to the j-th cluster center is... x ik Let be the feature value of the k-th dimension of the i-th feature point. c jk Let k be the feature value of the j-th cluster center. W k represents the dynamic weight of the k-th dimension.

[0058] (2) Feature point classification: Each feature point is assigned to the category of the nearest cluster center to form K initial clusters.

[0059] (3) Cluster center iterative update: Calculate the weighted mean of all feature points in each cluster, use the mean as the new cluster center, repeat the steps of “calculating weighted Euclidean distance → feature point classification → updating cluster center” until the change in cluster center is less than the preset threshold (the preset threshold is 0.001, dimensionless), or the number of iterations reaches the preset maximum value (the preset maximum value is 100 times), stop the iteration, and obtain the final clustering result.

[0060] (4) Scene identification: Based on the final clustering results, the system operation scenario corresponding to each cluster is defined. Combining the operation characteristics of the integrated energy system, the three clusters correspond to the following three typical scenarios: ① Cluster 1 (High Energy Consumption Peak Hour Scenario): Total load is higher than 1.2 times the average load in the current analysis period, the proportion of interruptible load is lower than the preset demand response threshold, the line loss rate is higher than the average, and the output of renewable energy is lower than the average; ② Cluster 2 (Stable Operation Scenario): The total load is between 0.8 and 1.2 times the average load during the current analysis period, the load fluctuation coefficient is lower than the preset fluctuation threshold, the line loss rate is near the average, and the output of renewable energy is stable; ③ Cluster 3 (High Renewable Energy Output Scenarios): Photovoltaic / wind turbine output is more than 1.3 times higher than the average output in the current analysis period, total load is lower than the average, energy storage SOC is higher than the preset high threshold, and line transmission efficiency is higher than the average.

[0061] 3. Output: Clustering results of integrated energy system operation scenarios (scenario category corresponding to each feature point) and final cluster centers.

[0062] Step S5: Source-Load Collaborative Optimization Strategy Matching and Execution 1. Input: Scene clustering results output in step 4, real-time operation data of the integrated energy system (consistent with the data type collected in step 1), and a preset source-load collaborative optimization strategy library.

[0063] 2. Perform the operation: (1) Pre-set optimization strategy library: Based on the operational requirements of the integrated energy system, three core optimization strategies are preset, corresponding to three clustering scenarios. The strategy library can be expanded according to actual application scenarios, as follows: ① Optimization strategy for high-energy-consuming peak-hour scenarios: With the goal of "peak shaving and valley filling, and reducing electricity purchase costs", specific measures include: instructing photovoltaic and wind turbine systems to generate full capacity for power consumption; scheduling energy storage systems to shave peaks with maximum discharge power; issuing demand response instructions to postpone the operation of high-energy-consuming interruptible loads; and optimizing line transmission paths to reduce line losses.

[0064] ② Stable operation scenario optimization strategy: With the goal of "maintaining system stability and improving energy utilization efficiency", specific measures include: adjusting the output of distributed power sources to match real-time load demand; controlling the energy storage system to be in float charging state and maintaining the SOC between 50% and 70%; monitoring line voltage and losses to ensure stable system operation.

[0065] ③ Optimization strategy for high renewable energy output scenarios: With the goal of "improving the renewable energy absorption rate and reducing carbon emissions", specific measures include: instructing energy storage systems to be fully charged to store excess renewable energy; guiding flexible loads to stagger peak consumption; and if there is excess renewable energy output, it can be transmitted back to the grid.

[0066] (2) Strategy matching and execution: Based on the current system operation scenario identified in step S4, the corresponding optimization strategy is matched from the strategy library, and the scheduling command is issued through the integrated energy EMS. The source-side, load-side, and grid-side devices are instructed to perform corresponding operations to achieve source-load collaborative optimization.

[0067] 3. Output: EMS dispatching instructions (including source-side equipment output instructions, load-side load control instructions, and grid-side line adjustment instructions), and optimized system operating parameters (such as load curves, source-side output curves, and line loss data).

[0068] Step S6: Optimize effect monitoring and scene drift judgment 1. Input: Optimized system operating parameters output from step S5, final cluster centers output from step 4, and preset drift judgment threshold.

[0069] 2. Perform the operation: (1) Optimization effect monitoring: Real-time monitoring of key operating indicators of the optimized system, including system operating cost, renewable energy absorption rate, line loss rate, and carbon emissions, compared with the indicators before optimization, to evaluate the optimization effect and generate an optimization effect report.

[0070] (2) Scene drift judgment: Monitor the deviation between the real-time multidimensional feature data and the final cluster center output in step 4, and calculate the weighted Euclidean distance between the real-time feature vector and each cluster center (the calculation formula is the same as in step 4). If the distance between the real-time feature vector and the cluster center of the scene it belongs to is greater than the preset drift threshold (the preset threshold is 0.3, dimensionless) for 10 consecutive collection times (i.e. 2.5 hours), it is judged that scene drift has occurred (i.e. the system operating state has changed significantly and the original clustering model can no longer accurately describe the system state); otherwise, it is judged that there is no scene drift.

[0071] 3. Output: Optimization effect report, scene drift judgment result ("drift" or "no drift").

[0072] Step S7: Iterative update or continuous operation 1. Input: Scene drift judgment result output in step S6.

[0073] 2. Perform the operation: (1) If the judgment result is “drift”: trigger the model re-clustering process, return to step 2, rebuild the multi-dimensional feature set (based on the latest real-time data), and execute the subsequent steps in sequence to update the clustering model, initial cluster center, dynamic weight and optimization strategy, so as to realize the online iterative update of the model and strategy.

[0074] (2) If the judgment result is “no drift”: maintain the current clustering model and optimization strategy, continue to execute the source load co-optimization operation in step S5, and continuously monitor the system operation data to enter the next round of monitoring and judgment cycle.

[0075] 3. Output: No explicit output (continuously running) or updated clustering model and optimization strategy (iterative update status).

[0076] Finally, it should be noted that the above content is only used to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Simple modifications or equivalent substitutions made by those skilled in the art to the technical solution of the present invention do not depart from the essence and scope of the technical solution of the present invention.

Claims

1. A comprehensive energy EMS source-load collaborative optimization method based on multidimensional feature K-means clustering, characterized in that, Includes the following steps: Step S1, Data Acquisition and Preprocessing: Input the raw operating data of the integrated energy EMS, collect four-dimensional raw data, and perform data preprocessing on the four-dimensional raw data; Step S2, Multidimensional Feature Set Construction and Normalization: Based on the preprocessed four-dimensional raw data, core dimensions are selected, redundant features are removed, and a four-dimensional coupled multidimensional feature set is constructed. Step S3: Improved K-means clustering initialization: The DBSCAN algorithm is used to identify the core feature points in the multidimensional feature set. These core feature points are then clustered, and the mean vector is calculated as the initial center of the K-means cluster to determine the optimal number of clusters K. The dynamic weight vector is calculated using the analytic hierarchy process (AHP) combined with the variance and coefficient of variation of the feature data. W ; Step S4, Improve K-means clustering scenario partitioning: Based on the dynamic weight vector W Calculate the weighted Euclidean distance from the feature points to each initial center, and assign the feature points to the category of the nearest cluster center to form K initial clusters; Calculate the weighted mean of all feature points within each cluster as the new cluster center and iteratively calculate the final clustering result; based on the clustering result, identify the system operation scenario; Step S5, Source-Load Co-optimization Strategy Matching and Execution: A source-load co-optimization strategy library is preset. Based on the current system operation scenario identified in step S4, the corresponding optimization strategy is matched from the strategy library, and the scheduling command is issued through the energy management system (EMS) to execute the corresponding operation. Step S6, Optimization effect monitoring and scene drift judgment: Monitor the operation indicators of the optimized system in real time and obtain the optimization effect report; calculate the weighted Euclidean distance between the real-time feature vector and each cluster center to determine whether the scene drift condition is met; Step S7, Iterative Update or Continuous Operation: If the judgment result is "drift", return to steps S2-S6 to re-execute the clustering process; if the judgment result is "no drift", maintain the current clustering model and optimization strategy, continue to execute the source-load co-optimization operation in step S5, and continuously monitor the system operation data to enter the next round of monitoring and judgment cycle.

2. The K-means clustering-based comprehensive energy EMS source-load co-optimization method according to claim 1, characterized in that: In step S3, the method for calculating the dynamic weights includes: determining the initial weights using the analytic hierarchy process (AHP) and then dynamically correcting them, as shown in the following formula: ; in, W i represents the final dynamic weight of the i-th feature dimension, which is dimensionless and ranges from [0,1]. The sum of the weights of all dimensions is 1. This represents the initial subjective weight of the i-th feature dimension determined by the analytic hierarchy process. It is dimensionless and ranges from [0,1]. α This represents the weight adjustment coefficient, which is dimensionless and ranges from [0.1, 0.5]. It is used to adjust the influence of the variance coefficient on the weights. The coefficient of variation (COP) represents the variance of the sample data for the i-th feature dimension. It is dimensionless and reflects the degree of fluctuation in the feature data for that dimension. The calculation formula is: ; in, Let be the standard deviation of the feature data in the i-th dimension. Let be the mean of the feature data in the i-th dimension. The larger the value, the greater the feature fluctuation, and the greater the impact on the clustering results; therefore, the weight should be increased accordingly.

3. The comprehensive energy EMS source-load collaborative optimization method based on K-means clustering according to claim 1, characterized in that: In step S4, the method for calculating the weighted Euclidean distance from the feature points to each initial cluster center includes: ; in, d ij This represents the weighted Euclidean distance from the i-th feature point to the j-th cluster center. x ik This represents the feature value of the k-th dimension of the i-th feature point. c jk This represents the feature value of the k-th dimension of the j-th cluster center. W k This represents the dynamic weight of the k-th dimension.

4. The K-means clustering-based comprehensive energy EMS source-load co-optimization method according to claim 1 or 3, characterized in that: In step S4, the termination condition for the iterative calculation of cluster centers is that the change in cluster centers is less than a preset threshold, or the number of iterations reaches a preset maximum value.

5. The comprehensive energy EMS source-load collaborative optimization method based on K-means clustering according to claim 1, characterized in that: In step S4, the typical scenarios identified based on the clustering results include high energy consumption peak scenarios, stable operation scenarios, and high renewable energy output scenarios. The characteristics of the high-energy-consuming peak-hour scenario are: the total load is more than 1.2 times the average load in the current analysis period, the proportion of interruptible load is lower than the preset demand response threshold, the line loss rate is higher than the average, and the output of renewable energy is lower than the average. The characteristics of the stable operation scenario are: the total load is between 0.8 and 1.2 times the average load in the current analysis period, the load fluctuation coefficient is lower than the preset fluctuation threshold, the line loss rate is near the average, and the output of renewable energy is stable; The characteristics of the high renewable energy output scenario are: the output of photovoltaic / wind turbines is more than 1.3 times higher than the average output in the current analysis period, the total load is lower than the average, the energy storage SOC is higher than the preset high threshold, and the line transmission efficiency is higher than the average.

6. The comprehensive energy EMS source-load collaborative optimization method based on K-means clustering according to claim 1, characterized in that: In step S5, the EMS dispatching instructions include source-side equipment output instructions, load-side load control instructions, and grid-side line adjustment instructions.

7. The K-means clustering-based comprehensive energy EMS source-load co-optimization method according to claim 1, characterized in that: In step S6, the scene drift condition includes calculating the weighted Euclidean distance between the real-time feature vector and each cluster center. If continuous t Each data collection moment, t If the distance between the real-time feature vector and the cluster center of the scene it belongs to is greater than the preset warning value, then scene drift is determined to have occurred. Otherwise, it is determined as no scene drift.

8. The comprehensive energy EMS source-load co-optimization method based on K-means clustering according to claim 1, characterized in that: The four-dimensional raw data includes source side, load side, network side, and environment side.

9. The comprehensive energy EMS source-load co-optimization method based on K-means clustering according to claim 8, characterized in that: In step S1, the core dimensions include: Source-side power output characteristics: normalized value of actual photovoltaic power output, normalized value of actual wind turbine power output, and normalized value of energy storage SOC; Load characteristics on the load side: normalized total load, proportion of interruptible load, load fluctuation coefficient; Network topology characteristics: normalized line loss, node voltage deviation rate, and network transmission efficiency; Environmental impact characteristics: normalized values ​​of light intensity, normalized values ​​of ambient temperature, and normalized values ​​of wind speed.

10. The comprehensive energy EMS source-load co-optimization method based on K-means clustering according to claim 1, characterized in that: In step S2, the method for determining redundant features includes: using the Pearson correlation coefficient method to calculate the correlation coefficient between any two features; if the absolute value of the correlation coefficient is greater than or equal to 0.8, it is determined to be a redundant feature.

Citation Information

Patent Citations

  • Coriolis force measurement and qualitative verification experiment instrument

    CN114863768A