New energy consumption multi-view scene clustering method, system and device and storage medium
By extracting multi-dimensional features from the power grid operation data and combining knowledge graphs for feature screening, and using multi-view analysis and clustering algorithms for scene division, the problems of insufficient single-view analysis and insufficient scientific feature selection in the existing technology are solved, and a comprehensive and scientific scenario division and accuracy of clustering results are improved for power grid operation data.
Patent Information
- Application Number
- CN202510291462.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The existing technology has problems such as insufficient single-view analysis and insufficient scientific feature selection in the analysis of power grid operation data, which leads to one-sided results of scenario division and cannot fully reflect the operating characteristics of power grid.
The multi-view scenario clustering method of new energy consumption is adopted. Multi-dimensional features are extracted from the power grid operation data, combined with knowledge graphs and domain knowledge to perform feature screening and data view division, clustering algorithms are used in each view, and finally using cluster evaluation indicators to verify and optimize the clustering results.
It has achieved a comprehensive and scientific scenario division of power grid operation data, improved the accuracy and reliability of clustering results, and provided reliable support for new energy consumption optimization and grid operation decisions.
Smart Images

Figure CN120217026A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of new energy consumption, and particularly relates to a multi-view scenario clustering method, system, device and storage medium for new energy consumption. Background Art
[0002] With the large-scale access of new energy power generation to the power system, the randomness, volatility and intermittency of new energy such as wind power and photovoltaic power bring unprecedented challenges to the safe and stable operation of the power grid. The unpredictability and instability of new energy power generation not only increase the complexity of power grid operation, but also pose higher requirements for the dispatching and operation of the power system. During the operation of the power grid, different operating conditions (such as seasons, weather, date types, etc.) will significantly affect the output characteristics of new energy and the consumption capacity of the power grid. The interaction of these factors causes the load demand of the power grid and the power generation of new energy to show high volatility in time and space. Therefore, how to effectively analyze and mine complex power grid operation data, identify key scenarios and reveal scenario characteristics has become an important research direction in the current power system field.
[0003] As an unsupervised learning method, the clustering algorithm plays an important role in data analysis and mining. It divides data into several clusters according to the similarity or difference between samples, helping to reveal the internal structure and laws of the data. The clustering algorithm is widely used in the power system field. Through clustering analysis, different load patterns (such as peak load, valley load, etc.) can be identified, which helps to optimize power grid dispatching and improve the load capacity of the power grid.
[0004] However, traditional clustering algorithms have the following limitations in power system scenario analysis:
[0005] (1) Insufficiency of single-view analysis: Many traditional methods usually only focus on operation data of a single dimension, while ignoring the combined effect of other influencing factors such as weather, resulting in one-sided scenario division results and being unable to comprehensively reflect the operation characteristics of the power grid.
[0006] (2) Lack of scientific feature selection: In scenario analysis, the selection of features directly determines the quality of the clustering result. In traditional methods, features are usually selected based on experience or simple statistical indicators, and fail to fully combine domain knowledge for scientific screening, which may lead to the interference of redundant features or the omission of key features, thereby affecting the accuracy of the clustering result.
[0007] To address the above problems, in recent years, some researchers have proposed multi-view analysis and knowledge-driven methods. Multi-view analysis can partition power grid operation data from different perspectives and reveal data characteristics in different dimensions. Compared with single-view analysis, multi-view methods can comprehensively consider different types of data and mine more potential information. Considering multiple factors can better reflect the operation characteristics of the power grid in different scenarios, and thus achieve more accurate scenario partitioning. As a tool that can integrate domain knowledge and data information, a knowledge graph can provide a theoretical basis for feature selection and scenario partitioning. By constructing a knowledge graph in the field of power systems, various power grid operation data, equipment information, meteorological conditions, etc. can be organically associated to provide guidance for feature screening. Combining multi-view analysis with a knowledge graph can enhance the effect of traditional clustering algorithms. For example, through the domain knowledge provided by the knowledge graph, features can be more scientifically selected in multi-view analysis, improving the accuracy and reliability of clustering results. At the same time, the application of clustering algorithms can discover potential operation patterns from different views, providing more accurate decision-making support for power grid dispatching and optimization. However, although the research on the combination of multi-view analysis and knowledge graphs has gradually received attention, the research on applying these methods in combination with clustering algorithms to new energy consumption scenario partitioning is still relatively limited. Summary of the Invention
[0008] The purpose of the present invention is to provide a new energy consumption multi-view scenario clustering method, system, device, and storage medium for the above-mentioned problems in the existing technology. By combining domain knowledge for feature selection and using multi-view analysis and clustering algorithms to achieve scientific scenario partitioning, it provides reliable support for new energy consumption optimization and power grid operation decision-making, and can comprehensively integrate multi-dimensional features and domain knowledge, thereby enhancing the comprehensiveness of scenario analysis and the accuracy of clustering results.
[0009] To achieve the above purpose, the present invention has the following technical solutions:
[0010] In the first aspect, a new energy consumption multi-view scenario clustering method is provided, including:
[0011] Extract multi-dimensional features from power grid operation data;
[0012] Combine a knowledge graph and domain knowledge to screen the multi-dimensional features, and perform data view partitioning based on the screened features;
[0013] Perform clustering operations on each partitioned data view respectively, and display the scenario features under different data views through the clustering results;
[0014] Use clustering evaluation indicators to verify the rationality of the clustering results and optimize the clustering operations;
[0015] Analyze the scenario features in the data view corresponding to the verified clustering results, and screen out typical scenarios and typical days.
[0016] As a preferred solution, the extraction of multi-dimensional features from the power grid operation data includes:
[0017] Extract any one or more of the daily load rate, daily peak-valley difference rate, peak-period load rate, valley-period load rate, maximum load occurrence time, and minimum load occurrence time to reflect the load changes at different times of the day;
[0018] Extract any one or more of the average output, power generation time ratio, peak value, coefficient of variation, average fluctuation, and maximum output occurrence time to reflect the photovoltaic output level and the intraday trend changes;
[0019] Extract any one or more of the average output size, peak-valley difference, coefficient of variation, average fluctuation, maximum output occurrence time, and minimum output occurrence time to reflect the wind power output level and the intraday trend changes;
[0020] Extract any one or more of the average output size, daily peak-valley difference rate, coefficient of variation, average fluctuation, maximum output occurrence time, and minimum output occurrence time to reflect the value size of the tie line and the intraday trend changes;
[0021] Extract temperature data from the weather website as temperature features;
[0022] Combine the features extracted from the time series corresponding to the load, wind power, photovoltaic, and tie line, and the temperature features.
[0023] As a preferred solution, the feature screening of the multi-dimensional features by combining the knowledge graph and domain knowledge, and the data view division based on the screened features include:
[0024] According to the attributes and relationships of the entities in the dispatching knowledge graph, screen the features related to the target power grid optimization task, and extract the key elements affecting the dispatching optimization;
[0025] Combine domain knowledge, and through the analysis of the physical meaning and engineering feasibility of the features by experts, eliminate invalid or incorrect features, and at the same time supplement or correct the features according to expert suggestions;
[0026] By calculating the correlation and redundancy between features, screen out the core features;
[0027] According to the needs of power grid operation analysis, divide two views of season and date type; according to the seasonal fluctuation law, extract the views of the four seasons, and divide the data by season; divide the working day and holiday views according to the type of date.
[0028] As a preferred solution, the clustering operation is implemented using the K-means clustering algorithm. The K-means clustering algorithm is used to perform clustering operations on each divided data view respectively, and the scene features under different data views are presented through the clustering results, including:
[0029] Input the features corresponding to the dates in different seasons, and standardize all features;
[0030] Determine the number of scenes required for clustering of load, photovoltaic, wind power, and tie lines in each season respectively, and randomly select a specified number of sample points as the initial clustering centers;
[0031] Calculate the distances between the samples in different seasons and the clustering centers within the corresponding seasons respectively, and assign the samples to the clustering clusters with the closest distances; for each clustering cluster, calculate the mean value of all samples within the cluster as the new clustering center, and replace the old clustering center with the new clustering center; repeat the assignment and update steps until the positions of the clustering centers no longer change or the maximum number of iterations is reached;
[0032] Divide the scenes corresponding to load, photovoltaic, wind power, and tie lines in the seasonal view according to the clustering results;
[0033] Input the features corresponding to the dates of weekdays and rest days, and standardize all features;
[0034] Determine the number of scenes required for clustering of load, photovoltaic, wind power, and tie lines on weekdays and rest days respectively, and randomly select a specified number of sample points as the initial clustering centers;
[0035] Calculate the distances between the samples corresponding to weekdays and rest days and the clustering centers respectively, as well as the distances between the samples corresponding to seasons and their respective clustering centers, and assign the samples to the clustering clusters with the closest distances; for each clustering cluster, calculate the mean value of all samples within the cluster as the new clustering center, and replace the old clustering center with the new clustering center; repeat the assignment and update steps until the positions of the clustering centers no longer change or the maximum number of iterations is reached;
[0036] Divide the scenes corresponding to load, photovoltaic, wind power, and tie lines on weekdays and rest days according to the clustering results.
[0037] As a preferred solution, the steps of using clustering evaluation indicators to verify the rationality of the clustering results and optimize the clustering operation include:
[0038] Calculate the silhouette coefficient corresponding to each clustering. Measure the quality of the clustering results through the silhouette coefficient, and evaluate the compactness of each data point within the cluster and the separation from other clusters; evaluate the overall effect of the clustering by calculating the average silhouette coefficient of each clustering;
[0039] Calculate the Calinski-Harabasz CH index corresponding to each clustering. The CH index evaluates the compactness and separation of clustering by calculating the ratio of the distance between clusters to the distance within clusters. The larger the value of the CH index, the better the clustering effect, which is manifested as a higher degree of separation between clusters and a stronger degree of compactness within clusters. Combine the CH indices of multiple clusterings and select the clustering result with the best separation effect.
[0040] Calculate the Davies-Bouldin DB index corresponding to each clustering. The DB index measures the degree of compactness of the data within clusters and the degree of separation between clusters. The smaller the value of the DB index, the better the clustering effect and the stronger the separation between clusters. By calculating and comparing the DB indices of each clustering, verify the rationality of the clustering results, optimize the parameters of the clustering algorithm, or select the best number of clusters.
[0041] As a preferred solution, analyze the scenario features under the data view corresponding to the clustering results that pass the verification, and screen out typical scenarios and typical days, including:
[0042] Visualize all the curves corresponding to each scenario of seasonal load, photovoltaic, wind power, and tie line, analyze the basic features of the corresponding curves under each scenario, and obtain the corresponding scenario descriptions.
[0043] Analyze the scenarios in different seasons and analyze the reasons corresponding to the scenario features in combination with the seasonal characteristics.
[0044] Visualize all the curves corresponding to the scenarios of load, photovoltaic, wind power, and tie line on weekdays and rest days, analyze the basic features of the corresponding curves under each scenario, and obtain the corresponding scenario descriptions.
[0045] Analyze the scenarios of weekdays and rest days and analyze the reasons corresponding to the scenario features in combination with the electricity consumption characteristics.
[0046] Calculate the probabilities of the scenarios corresponding to load, photovoltaic, wind power, and tie line under each view.
[0047] Select the scenario with the highest probability corresponding to each season as the typical scenario for the corresponding season.
[0048] Select the scenarios with the highest probabilities corresponding to weekdays and rest days as the typical scenarios for weekdays and rest days.
[0049] Select the date closest to the scenario clustering center among all the dates covered by the typical scenarios as the typical day.
[0050] In the second aspect, provide a new energy consumption multi-view scenario clustering system, including:
[0051] A multi-dimensional feature extraction module for extracting multi-dimensional features from grid operation data.
[0052] A feature screening and view division module, which is used to screen features of the multi-dimensional features by combining a knowledge graph and domain knowledge, and divide data views according to the screened features;
[0053] A clustering module, which is used to perform clustering operations on each divided data view respectively, and display the scenario features under different data views through the clustering results;
[0054] A clustering result evaluation module, which is used to verify the rationality of the clustering results and optimize the clustering operations by using clustering evaluation indicators;
[0055] A scenario analysis and screening module, which is used to analyze the scenario features under the data view corresponding to the clustering results that pass the verification, and screen out typical scenarios and typical days.
[0056] As a preferred solution, the multi-dimensional feature extraction module extracts multi-dimensional features from power grid operation data, including:
[0057] Extracting any one or more of daily load rate, daily peak-valley difference rate, peak-period load rate, valley-period load rate, maximum load occurrence time, and minimum load occurrence time to reflect the load changes at different times of the day;
[0058] Extracting any one or more of average output, power generation time ratio, peak value, coefficient of variation, average fluctuation, and maximum output occurrence time to reflect the photovoltaic output level and intra-day trend changes;
[0059] Extracting any one or more of average output size, peak-valley difference, coefficient of variation, average fluctuation, maximum output occurrence time, and minimum output occurrence time to reflect the wind power output level and intra-day trend changes;
[0060] Extracting any one or more of average output size, daily peak-valley difference rate, coefficient of variation, average fluctuation, maximum output occurrence time, and minimum output occurrence time to reflect the value size of the tie line and intra-day trend changes;
[0061] Extracting temperature data from a weather website as temperature features;
[0062] Combining the features extracted from the time series corresponding to load, wind power, photovoltaic, and tie line and temperature features.
[0063] As a preferred solution, the feature screening and view division module combines a knowledge graph and domain knowledge to screen features of the multi-dimensional features, and divides data views according to the screened features, including:
[0064] According to the attributes and relationships of entities in the dispatching knowledge graph, screen the features related to the target power grid optimization task, and extract the key factors affecting dispatching optimization;
[0065] Combined with domain knowledge, through the analysis of the physical meaning and engineering feasibility of features by experts, eliminate invalid or incorrect features, and at the same time supplement or correct features according to expert suggestions;
[0066] By calculating the correlation and redundancy between features, screen out the core features;
[0067] According to the needs of power grid operation analysis, divide into two views of season and date type; according to the seasonal fluctuation law, extract the views of the four seasons, divide the data by season; divide the views of working days and holidays according to the type of date.
[0068] As a preferred solution, the clustering module uses the K-means clustering algorithm to perform clustering operations on each divided data view respectively, and displays the scenario features under different data views through the clustering results, including:
[0069] Input the features corresponding to the dates in different seasons, and standardize all features;
[0070] Determine the number of scenarios required for clustering of load, photovoltaic, wind power, and tie lines in each season respectively, and randomly select a specified number of sample points as the initial clustering centers;
[0071] Calculate the distances between samples in different seasons and the clustering centers within the corresponding seasons respectively, and assign the samples to the clustering clusters with the closest distances; for each clustering cluster, calculate the mean value of all samples within the cluster as the new clustering center, and replace the old clustering center with the new clustering center; repeat the assignment and update steps until the positions of the clustering centers no longer change or reach the maximum number of iterations;
[0072] Divide the scenarios corresponding to load, photovoltaic, wind power, and tie lines in the season view according to the clustering results;
[0073] Input the features corresponding to the dates of working days and rest days, and standardize all features;
[0074] Determine the number of scenarios required for clustering of load, photovoltaic, wind power, and tie lines on working days and rest days respectively, and randomly select a specified number of sample points as the initial clustering centers;
[0075] Calculate the distances between the samples corresponding to weekdays and rest days, as well as the samples corresponding to seasons and their respective cluster centers, and assign the samples to the cluster with the closest distance; for each cluster, calculate the mean of all samples within the cluster as the new cluster center, and replace the old cluster center with the new one; repeat the assignment and update steps until the positions of the cluster centers no longer change or the maximum number of iterations is reached;
[0076] Divide the scenarios corresponding to the load, photovoltaic, wind power, and tie lines on weekdays and rest days according to the clustering results.
[0077] As a preferred solution, the clustering result evaluation module calculates the silhouette coefficient corresponding to each clustering, measures the quality of the clustering result through the silhouette coefficient, evaluates the compactness of each data point within the cluster and the separation from other clusters; by calculating the average silhouette coefficient of each clustering, evaluate the overall effect of the clustering;
[0078] Calculate the Calinski-Harabasz CH index corresponding to each clustering. The CH index evaluates the compactness and separation of the clustering by calculating the ratio of the inter-cluster distance to the intra-cluster distance; the larger the value of the CH index, the better the clustering effect, manifested as higher inter-cluster separation and stronger intra-cluster compactness; combine the CH indices of multiple clusterings to select the clustering result with the best separation effect;
[0079] Calculate the Davies-Bouldin DB index corresponding to each clustering. The DB index measures the tightness of the data within the cluster and the separation between clusters; the smaller the value of the DB index, the better the clustering effect and the stronger the separation between clusters; by calculating and comparing the DB indices of each clustering, verify the rationality of the clustering result, optimize the clustering algorithm parameters or select the best number of clusters.
[0080] As a preferred solution, the scenario analysis and screening module visualizes all the curves under each seasonal load, photovoltaic, wind power, and tie line corresponding scenario, analyzes the basic characteristics of the corresponding curves under each scenario, and obtains the corresponding scenario description;
[0081] Analyze the scenarios under different seasons, and analyze the reasons corresponding to the scenario characteristics in combination with the seasonal characteristics;
[0082] Visualize all the curves under the scenarios corresponding to the load, photovoltaic, wind power, and tie lines on weekdays and rest days, analyze the basic characteristics of the corresponding curves under each scenario, and obtain the corresponding scenario description;
[0083] Analyze the scenarios on weekdays and rest days, and analyze the reasons corresponding to the scenario characteristics in combination with the electricity consumption characteristics;
[0084] Calculate the probabilities of the scenarios corresponding to the load, photovoltaic, wind power, and tie lines in each view;
[0085] Select the scenario with the highest probability for each season as the typical scenario for the corresponding season;
[0086] Select the scenarios with the highest probabilities for weekdays and rest days as the typical scenarios for weekdays and rest days;
[0087] Select the date closest to the scenario clustering center among all the dates covered by the typical scenario as the typical day.
[0088] In a third aspect, there is provided an electronic device, including a processor and a memory, where the processor is configured to execute a computer program stored in the memory to implement the new energy consumption multi-view scenario clustering method described above.
[0089] In a fourth aspect, there is provided a computer-readable storage medium storing at least one instruction, where when the at least one instruction is executed by a processor, the new energy consumption multi-view scenario clustering method described above is implemented.
[0090] Compared with the prior art, the first aspect of the present invention has at least the following beneficial effects:
[0091] With the rapid development of new energy power generation, the penetration rate of new energy sources such as wind power and photovoltaic power in the power grid has been continuously increasing. However, due to the randomness, volatility, and intermittency of new energy output, it has brought significant challenges to the operation and dispatching of the power grid. Especially in the operation of the power grid, different seasons, date types (such as weekdays and rest days), and equipment operating states have important impacts on the consumption capacity of new energy and the security of the power grid. Therefore, how to comprehensively and scientifically divide operation scenarios and mine key features based on power grid operation data has become an important research direction in the field of power systems. Most of the existing operation scenario division methods analyze based on single-view data, lacking the ability to integrate and mine multi-view and multi-angle data features. At the same time, the scientificity and accuracy of scenario division are also restricted by feature selection methods and clustering algorithms. The present invention deeply analyzes the relevant data of load, photovoltaic power generation, wind power generation, and tie lines in power grid operation data, and can extract key feature information of various different dimensions from it, including time series characteristics, fluctuation trend change rules, etc., providing support for in-depth information mining. By combining knowledge graph technology with domain knowledge, comprehensively utilizing the relevance of graph-structured data and the professional experience of domain experts, the features are screened and optimized to identify core features with significant value. At the same time, according to the semantic relationship, functional relevance, and application scenario requirements of the features, the data views are reasonably divided. On the basis of multi-view analysis, clustering operations are respectively performed on each divided data view, and the scenario features under different data views are presented through the clustering results, fully mining the data features and potential scenario patterns under each view. The scenario differences and distribution laws of multi-dimensional features under different views are presented through the clustering results, providing data support for subsequent scenario analysis. Using scientific clustering evaluation indicators, the scenario division results of each view are comprehensively evaluated to ensure the rationality, accuracy, and scientificity of scenario division, thereby enhancing the model's expression ability for multi-view complex features. Analyze the scenario features under the data view corresponding to the clustering results passed by verification, and screen out the most representative and practically applicable typical scenarios and typical days from them. These typical scenarios and typical days can intuitively reflect the key features and provide important basis for subsequent decision support, pattern recognition, and optimization analysis.
[0092] It can be understood that the beneficial effects of the second to fourth aspects above can refer to the relevant descriptions in the first aspect above, and will not be elaborated here. Brief Description of the Drawings
[0093] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0094] Figure 1 Flow chart of the new - energy accommodation multi - view scenario clustering method according to an embodiment of the present invention;
[0095] Figure 2 Block diagram of the structure of the new - energy accommodation multi - view scenario clustering system according to an embodiment of the present invention;
[0096] Figure 3 Schematic diagram of the physical structure of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0097] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures, technologies, etc. are set forth in order to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well - known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of the present application.
[0098] Please refer to Figure 1 , an embodiment of the present invention provides a new - energy accommodation multi - view scenario clustering method, including the following steps:
[0099] S1. Extract multi - dimensional features from the power grid operation data;
[0100] S2. Combine the knowledge graph and domain knowledge to perform feature screening on the multi - dimensional features, and perform data view division based on the screened features;
[0101] S3. Perform clustering operations on each divided data view respectively, and display the scenario features under different data views through the clustering results;
[0102] S4. Use clustering evaluation indicators to verify the rationality of the clustering results and optimize the clustering operations;
[0103] S5. Analyze the scenario features under the data views corresponding to the clustering results that pass the verification, and screen out typical scenarios and typical days.
[0104] In a possible implementation manner, in step S1, in the operation data of a certain regional power grid during a set time period, in - depth analysis is performed on the relevant data of load, photovoltaic power generation, wind power generation, and tie lines, and various key feature information of different dimensions is extracted therefrom, including time - series characteristics, fluctuation trend change rules, etc., to provide support for in - depth information mining.
[0105] Specifically, for the characteristics of regional load changes, features such as daily load rate, daily peak-valley difference rate, peak-period load rate, valley-period load rate, time of occurrence of maximum load, and time of occurrence of minimum load are extracted to reflect the load changes at different times of the day.
[0106] For the characteristics of large uncertainty and volatility in photovoltaic data output, features such as average output, power generation time ratio, peak value, coefficient of variation, average fluctuation, and time of occurrence of maximum output are extracted to reflect the photovoltaic output level and intra-day changes.
[0107] For the characteristics of intermittency, randomness, and volatility in wind power data, features such as average output magnitude, peak-valley difference, coefficient of variation, average fluctuation, time of occurrence of maximum output, and time of occurrence of minimum output are extracted to reflect the wind power output level and intra-day trend changes.
[0108] For the functional characteristics of tie-line data, which can achieve coordination and interaction between micro-power sources and the large power grid, features such as average output magnitude, daily peak-valley difference rate, coefficient of variation, average fluctuation, time of occurrence of maximum output, and time of occurrence of minimum output are extracted to reflect the value of tie-line readings and intra-day trend changes.
[0109] Temperature data for the corresponding region are extracted from weather websites as temperature features.
[0110] The features extracted from the time series corresponding to load, wind power, photovoltaic, and tie-line, as well as the temperature features, are combined and stored in the feature library.
[0111] In a possible implementation, in step S2, by combining knowledge graph technology and domain knowledge, the relevance of graph-structured data and the professional experience of domain experts are comprehensively utilized to screen and optimize the features, and the core features with significant value are identified. At the same time, according to the semantic relationships, functional correlations, and application scenario requirements of the features, the data view is reasonably divided.
[0112] Specifically, according to the attributes and relationships of entities in the dispatching knowledge graph, features closely related to the target power grid optimization task are screened, and the key elements that may affect dispatching optimization are extracted.
[0113] Combined with domain knowledge, experts verify whether the screened features are reasonable. By analyzing the physical meaning and engineering feasibility of the features by experts, potentially invalid or incorrect features are eliminated. At the same time, the feature set is supplemented or corrected according to expert suggestions to ensure that the final features have both theoretical basis and meet the actual power grid operation requirements.
[0114] Redundant or low-correlation features are removed. By calculating the correlation between features (such as Pearson correlation coefficient) and the redundancy between features, the core features with large amounts of information are screened.
[0115] According to the needs of power grid operation analysis, two types of views, namely seasonal and date type views, are divided. First, according to the seasonal fluctuation law, views of the four seasons of spring, summer, autumn, and winter are extracted, and the data is divided according to seasons. Second, the working day and holiday views are divided according to the type of date.
[0116] In a possible implementation manner, in step S3, based on multi-view analysis, the K-means clustering algorithm is applied to each view respectively for clustering operations, fully mining the data characteristics and potential scenario patterns under each view. The scene differences and the distribution laws of multi-dimensional features under different views are presented through the clustering results, providing data support for subsequent scene analysis.
[0117] Specifically, the features corresponding to the dates in different seasons are input, and all features are standardized.
[0118] The number of scenarios required for clustering of load, photovoltaic, wind power, and tie lines in each season is determined respectively, and a specified number of sample points are randomly selected as the initial clustering centers.
[0119] The distances between the samples in different seasons and the clustering centers in the same season are calculated respectively, and the samples are assigned to the clustering clusters with the closest distances. For each clustering cluster, the mean value of all samples within the cluster is calculated as the new clustering center, and the new clustering center is used to replace the old clustering center; the assignment and update steps are repeated until the positions of the clustering centers no longer change or the maximum number of iterations is reached.
[0120] According to the clustering results, the scenarios corresponding to load, photovoltaic, wind power, and tie lines in the seasonal view are divided.
[0121] The features corresponding to the dates of working days and rest days are input, and all features are standardized.
[0122] The number of scenarios required for clustering of load, photovoltaic, wind power, and tie lines for working days and rest days is determined. A specified number of sample points are randomly selected as the initial clustering centers.
[0123] The distances between the samples corresponding to working days and rest days and the samples corresponding to seasons and their respective clustering centers are calculated respectively, and the samples are assigned to the clustering clusters with the closest distances. For each clustering cluster, the mean value of all samples within the cluster is calculated as the new clustering center, and the new clustering center is used to replace the old clustering center. The assignment and update steps are repeated until the positions of the clustering centers no longer change or the maximum number of iterations is reached.
[0124] According to the clustering results, the scenarios corresponding to load, photovoltaic, wind power, and tie lines for working days and rest days are divided.
[0125] In a possible implementation, step S4 uses scientific clustering evaluation metrics to comprehensively evaluate the scene partitioning results of each view, ensuring the rationality, accuracy, and scientific nature of the scene partitioning, thereby enhancing the model's ability to express complex multi-view features.
[0126] Specifically, calculate the silhouette coefficient corresponding to each clustering. The silhouette coefficient is used to measure the quality of the clustering result and evaluate the compactness of each data point within its cluster and its separation from other clusters. The value range of the silhouette coefficient is [-1, 1]. The closer the value is to 1, the better the clustering effect. 0 indicates that the data point is near the cluster boundary, and a negative value indicates that the data may be misallocated to other clusters. Evaluate the overall clustering effect by calculating the average silhouette coefficient of each clustering.
[0127] Calculate the Calinski-Harabasz (CH) index corresponding to each clustering. The CH index evaluates the compactness and separation of the clustering by calculating the ratio of the inter-cluster distance to the intra-cluster distance; the larger the value of the CH index, the better the clustering effect, manifested as a higher inter-cluster separation and a stronger intra-cluster compactness; combine the CH indices of multiple clusterings and select the clustering result with the best separation effect.
[0128] Calculate the Davies-Bouldin (DB) index corresponding to each clustering. The DB index measures the compactness of the data within the cluster and the separation between clusters; the smaller the value of the DB index, the better the clustering effect and the stronger the separation between clusters; by calculating and comparing the DB indices of each clustering, the rationality of the clustering result can be further verified, the parameters of the clustering algorithm can be optimized, or the best number of clusters can be selected.
[0129] In a possible implementation, step S5 based on the clustering results of different views, conducts a detailed analysis of the scene features under each view, and screens out the typical scenarios and typical days that are the most representative and have practical application significance. These typical scenarios and typical days can intuitively reflect the key features and provide important basis for subsequent decision support, pattern recognition, and optimization analysis. This includes analyzing the scenarios under the season type view, analyzing the scenarios under the date type view, and screening out the typical scenarios and typical days with representativeness and practical significance.
[0130] Specifically, visualize all the curves corresponding to each scenario of each seasonal load, photovoltaic, wind power, and tie line, analyze the basic features of the corresponding curves under each scenario, and obtain the corresponding scenario description;
[0131] Analyze the scenarios under different seasons, and combine the seasonal characteristics of the region to analyze the reasons corresponding to the scenario features;
[0132] Visualize all the curves corresponding to the load, photovoltaic power, wind power, and tie lines under working days and rest days, analyze the basic characteristics of the corresponding curves in each scenario, and obtain the corresponding scenario descriptions;
[0133] Analyze the scenarios of working days and rest days, and analyze the reasons corresponding to the scenario characteristics in combination with the electricity consumption characteristics of the region;
[0134] Calculate the probabilities of the load, photovoltaic power, wind power, and tie lines corresponding to the scenarios appearing in each view;
[0135] Select the scenario with the highest probability for each season as the typical scenario for the corresponding season;
[0136] Select the scenarios with the highest probabilities for working days and rest days as the typical scenarios for working days and rest days;
[0137] Select the date closest to the scenario clustering center among all the dates covered by the typical scenarios as the typical day.
[0138] In summary, the new energy consumption multi-view scenario clustering method in the embodiments of the present invention extracts features from the operation data of a certain regional power grid, obtains multi-dimensional key information, combines the knowledge graph and domain knowledge for feature selection and view division, uses the K-means clustering algorithm to perform clustering in different views respectively, mines the scenario features under multiple views, evaluates the scenario division results through clustering metrics to ensure the scientificity and accuracy of the division, and finally conducts descriptive analysis on each scenario to screen out representative and practically significant typical scenarios and typical days.
[0139] Please refer to Figure 2 , another embodiment of the present invention also proposes a new energy consumption multi-view scenario clustering system, including:
[0140] A multi-dimensional feature extraction module 210, configured to extract multi-dimensional features from the power grid operation data;
[0141] A feature screening and view division module 220, configured to perform feature screening on the multi-dimensional features in combination with the knowledge graph and domain knowledge, and perform data view division based on the screened features;
[0142] A clustering module 230, configured to perform clustering operations on each divided data view respectively, and display the scenario features under different data views through the clustering results;
[0143] A clustering result evaluation module 240, configured to use clustering evaluation metrics to verify the rationality of the clustering results and optimize the clustering operations;
[0144] A scenario analysis and screening module 250, which is used to analyze the scenario features in the data view corresponding to the verified clustering results and screen out typical scenarios and typical days.
[0145] In a possible implementation manner, the multi-dimensional feature extraction module 210 extracts multi-dimensional features from the power grid operation data, including:
[0146] Extract any one or more of the daily load rate, daily peak-valley difference rate, peak-period load rate, valley-period load rate, maximum load occurrence time, and minimum load occurrence time to reflect the load changes at different times of the day;
[0147] Extract any one or more of the average output, power generation time ratio, peak value, coefficient of variation, average fluctuation, and maximum output occurrence time to reflect the photovoltaic output level and the intraday trend changes;
[0148] Extract any one or more of the average output size, peak-valley difference, coefficient of variation, average fluctuation, maximum output occurrence time, and minimum output occurrence time to reflect the wind power output level and the intraday trend changes;
[0149] Extract any one or more of the average output size, daily peak-valley difference rate, coefficient of variation, average fluctuation, maximum output occurrence time, and minimum output occurrence time to reflect the value size of the tie line and the intraday trend changes;
[0150] Extract temperature data from a weather website as temperature features;
[0151] Combine the features extracted from the time series corresponding to the load, wind power, photovoltaic, and tie line and the temperature features.
[0152] In a possible implementation manner, the feature screening and view division module 220 combines the knowledge graph and domain knowledge to perform feature screening on the multi-dimensional features, and performs data view division based on the screened features, including:
[0153] According to the attributes and relationships of entities in the dispatching knowledge graph, screen out the features related to the target power grid optimization task and extract the key elements affecting dispatching optimization;
[0154] Combined with domain knowledge, through the analysis of the physical meaning and engineering feasibility of the features by experts, eliminate invalid or incorrect features, and at the same time supplement or correct the features according to expert suggestions;
[0155] By calculating the correlation between features and the redundancy between features, screen out the core features;
[0156] According to the needs of power grid operation analysis, two types of views, namely seasonal and date type views, are divided; according to the seasonal fluctuation law, views of four seasons are extracted, and the data is divided according to seasons; views of working days and holidays are divided according to the type of date.
[0157] In a possible implementation manner, the clustering module 230 uses the K-means clustering algorithm to perform clustering operations on each divided data view respectively, and presents the scene characteristics under different data views through the clustering results, including:
[0158] Input the characteristics of the dates corresponding to different seasons, and standardize all the characteristics;
[0159] Respectively determine the number of scenes required for clustering of load, photovoltaic, wind power and tie lines in each season, and randomly select a specified number of sample points as the initial clustering centers;
[0160] Respectively calculate the distances between the samples in different seasons and the clustering centers within the corresponding seasons, and assign the samples to the clustering clusters with the closest distances; for each clustering cluster, calculate the mean value of all the samples within the cluster as the new clustering center, and replace the old clustering center with the new clustering center; repeat the assignment and update steps until the positions of the clustering centers no longer change or reach the maximum number of iterations;
[0161] Divide the scenes corresponding to load, photovoltaic, wind power and tie lines in the seasonal view according to the clustering results;
[0162] Input the characteristics of the dates corresponding to working days and rest days, and standardize all the characteristics;
[0163] Determine the number of scenes required for clustering of load, photovoltaic, wind power and tie lines for working days and rest days, and randomly select a specified number of sample points as the initial clustering centers;
[0164] Respectively calculate the distances between the samples corresponding to working days and rest days and the samples corresponding to seasons and their respective clustering centers, and assign the samples to the clustering clusters with the closest distances; for each clustering cluster, calculate the mean value of all the samples within the cluster as the new clustering center, and replace the old clustering center with the new clustering center; repeat the assignment and update steps until the positions of the clustering centers no longer change or reach the maximum number of iterations;
[0165] Divide the scenes corresponding to load, photovoltaic, wind power and tie lines corresponding to working days and rest days according to the clustering results.
[0166] In a possible implementation manner, the clustering result evaluation module 240 calculates the silhouette coefficient corresponding to each clustering, measures the quality of the clustering result through the silhouette coefficient, and evaluates the compactness of each data point within the cluster and the separation from other clusters; by calculating the average silhouette coefficient of each clustering, the overall effect of the clustering is evaluated;
[0167] Calculate the Calinski-Harabasz CH index corresponding to each clustering. The CH index evaluates the compactness and separation of clustering by calculating the ratio of the distance between clusters to the distance within clusters. The larger the value of the CH index, the better the clustering effect, which is manifested as higher separation between clusters and stronger compactness within clusters. Combine the CH indices of multiple clusterings and select the clustering result with the best separation effect.
[0168] Calculate the Davies-Bouldin DB index corresponding to each clustering. The DB index measures the degree of compactness of the data within clusters and the degree of separation between clusters. The smaller the value of the DB index, the better the clustering effect and the stronger the separation between clusters. By calculating and comparing the DB indices of each clustering, verify the rationality of the clustering results, optimize the parameters of the clustering algorithm or select the optimal number of clusters.
[0169] In a possible implementation, the scenario analysis and screening module 250 visualizes all the curves corresponding to each seasonal load, photovoltaic, wind power, and tie line scenario, analyzes the basic characteristics of the corresponding curves in each scenario, and obtains the corresponding scenario description.
[0170] Analyze the scenarios in different seasons and analyze the reasons corresponding to the scenario characteristics in combination with the seasonal characteristics.
[0171] Visualize all the curves corresponding to the scenarios of load, photovoltaic, wind power, and tie line on weekdays and rest days, analyze the basic characteristics of the corresponding curves in each scenario, and obtain the corresponding scenario description.
[0172] Analyze the scenarios on weekdays and rest days and analyze the reasons corresponding to the scenario characteristics in combination with the electricity consumption characteristics.
[0173] Calculate the probabilities of the scenarios corresponding to load, photovoltaic, wind power, and tie line in each view.
[0174] Select the scenario with the highest probability corresponding to each season as the typical scenario for the corresponding season.
[0175] Select the scenarios with the highest probabilities corresponding to weekdays and rest days as the typical scenarios for weekdays and rest days.
[0176] Select the date closest to the scenario clustering center among all the dates covered by the typical scenarios as the typical day.
[0177] Figure 3 Illustrate a schematic diagram of the physical structure of an electronic device, such as Figure 3As shown in the figure, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communications interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 can call the logical instructions in the memory 330 to execute the new energy consumption multi-view scenario clustering method. This method includes extracting multi-dimensional features from the power grid operation data; combining the knowledge graph and domain knowledge to perform feature screening on the multi-dimensional features, and dividing the data views according to the screened features; performing clustering operations on each divided data view respectively, and presenting the scenario features under different data views through the clustering results; using clustering evaluation indicators to verify the rationality of the clustering results and optimize the clustering operations; analyzing the scenario features under the data views corresponding to the clustering results that pass the verification, and screening out typical scenarios and typical days.
[0178] In addition, when the logical instructions in the above-mentioned memory 330 can be implemented in the form of software function units and sold or used as an independent product, they can be stored in a computer-readable storage medium.
[0179] Another embodiment of the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the new energy consumption multi-view scenario clustering method provided in the above embodiment.
[0180] An embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores at least one instruction. When the at least one instruction is executed by a processor, the new energy consumption multi-view scenario clustering method described above is implemented.
[0181] The computer program includes computer program code, which may be in the form of source code, object code, executable files, or some intermediate form, etc. The computer-readable storage medium may include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory, random access memory, electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals. For the sake of convenience of description, only the parts related to the embodiments of the present invention are shown above. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present invention. This computer-readable storage medium is non-transitory and can be stored in a storage device formed by various electronic devices, and can implement the execution process recorded in the method of the embodiments of the present invention.
[0182] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0183] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0184] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0185] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to generate a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or multiple processes and / or blocks Figure 1 one process or multiple processes and / or blocks Figure 1 or steps for implementing the functions specified in one block or multiple blocks.
[0186] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A multi-view scene clustering method for new energy consumption, characterized in that: include: Extract multi-dimensional features from power grid operation data; The multi-dimensional features are screened by combining the knowledge graph and the domain knowledge, and the data views are divided according to the screened features; Perform clustering operations on each divided data view, and use the clustering results to show the scene characteristics under different data views; Using clustering evaluation indicators to verify the rationality of the clustering results and optimize the clustering operation; Analyze the scene features under the data view corresponding to the verified clustering results, and screen out typical scenes and typical days.
2. The multi-view scene clustering method for new energy consumption according to claim 1 is characterized in that: The multi-dimensional features extracted from the power grid operation data include: Extract any one or more of the daily load rate, daily peak-to-valley difference rate, peak load rate, valley load rate, maximum load occurrence time, and minimum load occurrence time to reflect load changes at different times throughout the day; Extract any one or more of the average output, generation time ratio, peak value, coefficient of variation, average fluctuation and maximum output occurrence time to reflect the photovoltaic output level and intraday trend changes; Extract any one or more of the average output, peak-to-valley difference, coefficient of variation, average fluctuation, maximum output time and minimum output time to reflect the wind power output level and intraday trend changes; Extract any one or more of the average output, daily peak-to-valley difference, coefficient of variation, average fluctuation, maximum output time and minimum output time to reflect the value of the tie line and the intraday trend change; Extract temperature data from the weather website as temperature features; The features extracted from the time series corresponding to load, wind power, photovoltaics and tie lines are combined with the temperature features.
3. The multi-view scene clustering method for new energy consumption according to claim 1 is characterized in that: The combining of the knowledge graph and the domain knowledge to perform feature screening on the multi-dimensional features, and dividing the data views according to the screened features includes: According to the attributes and relationships of entities in the dispatch knowledge graph, the features related to the target power grid optimization task are screened, and the key factors affecting dispatch optimization are extracted; Combined with domain knowledge, experts analyze the physical meaning and engineering feasibility of features to eliminate invalid or erroneous features, and supplement or modify features based on expert suggestions; By calculating the correlation and redundancy between features, the core features are screened out; According to the needs of power grid operation analysis, two views are divided into season and date type; according to the law of seasonal fluctuations, four seasons' views are extracted and the data is divided by season; according to the type of date, the weekday and holiday views are divided.
4. The multi-view scene clustering method for new energy consumption according to claim 1 is characterized in that: The clustering operation is implemented using the K-means clustering algorithm. The K-means clustering algorithm is used to perform clustering operations on each divided data view. The clustering results show the scene characteristics under different data views, including: Input the features of dates corresponding to different seasons and standardize all features; Determine the number of scenarios required for clustering loads, photovoltaics, wind power and tie lines in each season, and randomly select a specified number of sample points as the initial cluster centers; Calculate the distance between samples in different seasons and the cluster center in the corresponding season respectively, and assign the samples to the cluster with the closest distance; for each cluster, calculate the mean of all samples in the cluster as the new cluster center, and replace the old cluster center with the new cluster center; repeat the assignment and update steps until the position of the cluster center no longer changes or the maximum number of iterations is reached; According to the clustering results, the scenarios corresponding to load, photovoltaic, wind power and tie lines in the seasonal view are divided; Input the features of the corresponding dates of working days and rest days, and standardize all the features; Determine the number of scenarios required for clustering loads, photovoltaics, wind power, and tie lines on weekdays and weekends, and randomly select a specified number of sample points as initial cluster centers; Calculate the distances between samples corresponding to working days and rest days, as well as samples corresponding to seasons, and their respective cluster centers, and assign the samples to the clusters with the closest distances; for each cluster, calculate the mean of all samples in the cluster as the new cluster center, and replace the old cluster center with the new cluster center; repeat the assignment and update steps until the position of the cluster center no longer changes or the maximum number of iterations is reached; According to the clustering results, the loads corresponding to working days and weekends, photovoltaic power, wind power and interconnection lines are divided into scenarios.
5. The multi-view scene clustering method for new energy consumption according to claim 1 is characterized in that: The step of using the clustering evaluation index to verify the rationality of the clustering result and optimizing the clustering operation includes: Calculate the silhouette coefficient corresponding to each clustering, measure the quality of the clustering results by the silhouette coefficient, and evaluate the compactness of each data point in the cluster and the degree of separation from other clusters; evaluate the overall effect of clustering by calculating the average silhouette coefficient of each clustering; The Kalinsky-Harabas CH index corresponding to each clustering was calculated. The CH index evaluates the compactness and separation of clusters by calculating the ratio of the distance between clusters to the distance within clusters. The larger the value of the CH index, the better the clustering effect, which is manifested as higher separation between clusters and stronger compactness within clusters. The CH index of multiple clusterings was combined to select the clustering result with the best separation effect. Calculate the Davis-Baulding DB index corresponding to each clustering. The DB index measures the compactness of the data within the cluster and the separation between clusters. The smaller the value of the DB index, the better the clustering effect and the stronger the separation between clusters. By calculating and comparing the DB index of each clustering, the rationality of the clustering results can be verified, the clustering algorithm parameters can be optimized, or the optimal number of clusters can be selected.
6. The multi-view scene clustering method for new energy consumption according to claim 1 is characterized in that: The scene features under the data view corresponding to the verified clustering results are analyzed to screen out typical scenes and typical days, including: Visualize all curves corresponding to each scenario of seasonal load, photovoltaic, wind power and tie line, analyze the basic characteristics of the corresponding curves in each scenario, and obtain the corresponding scenario description; Analyze scenes in different seasons and analyze the causes of scene characteristics based on seasonal characteristics; Visualize all curves corresponding to the scenarios of workday and weekend load, photovoltaic, wind power and tie lines, analyze the basic characteristics of the corresponding curves in each scenario, and obtain the corresponding scenario description; Analyze the scenarios on weekdays and weekends, and analyze the causes corresponding to the scenario characteristics based on the electricity consumption characteristics; Calculate the probability of occurrence of corresponding scenarios for load, photovoltaic, wind power and tie lines in each view; Select the scene with the highest probability for each season as the typical scene for the corresponding season; Select the scenarios with the highest probability for workdays and weekends as typical scenarios for workdays and weekends; The date closest to the scene cluster center among all the dates covered by the typical scenes is selected as the typical day.
7. A multi-view scene clustering system for new energy consumption, characterized in that: include: A multi-dimensional feature extraction module is used to extract multi-dimensional features from power grid operation data; A feature screening and view division module is used to screen the multi-dimensional features by combining the knowledge graph and domain knowledge, and to divide the data view according to the screened features; The clustering module is used to perform clustering operations on each divided data view and display the scene characteristics under different data views through the clustering results; A clustering result evaluation module is used to verify the rationality of the clustering result and optimize the clustering operation by using clustering evaluation indicators; The scenario analysis and screening module is used to analyze the scenario characteristics under the data view corresponding to the verified clustering results, and screen out typical scenarios and typical days.
8. The multi-view scene clustering system for new energy consumption according to claim 7 is characterized in that: The multi-dimensional feature extraction module extracts multi-dimensional features from the power grid operation data, including: Extract any one or more of the daily load rate, daily peak-to-valley difference rate, peak load rate, valley load rate, maximum load occurrence time, and minimum load occurrence time to reflect load changes at different times throughout the day; Extract any one or more of the average output, generation time ratio, peak value, coefficient of variation, average fluctuation and maximum output occurrence time to reflect the photovoltaic output level and intraday trend changes; Extract any one or more of the average output, peak-to-valley difference, coefficient of variation, average fluctuation, maximum output time and minimum output time to reflect the wind power output level and intraday trend changes; Extract any one or more of the average output, daily peak-to-valley difference, coefficient of variation, average fluctuation, maximum output time and minimum output time to reflect the value of the tie line and the intraday trend change; Extract temperature data from the weather website as temperature features; The features extracted from the time series corresponding to load, wind power, photovoltaics and tie lines are combined with the temperature features.
9. The new energy consumption multi-view scene clustering system according to claim 7, characterized in that: The feature screening and view division module combines the knowledge graph and domain knowledge to screen the multi-dimensional features, and divides the data view according to the screened features, including: According to the attributes and relationships of entities in the dispatch knowledge graph, the features related to the target power grid optimization task are screened, and the key factors affecting dispatch optimization are extracted; Combined with domain knowledge, experts analyze the physical meaning and engineering feasibility of features to eliminate invalid or erroneous features, and supplement or modify features based on expert suggestions; By calculating the correlation and redundancy between features, the core features are screened out; According to the needs of power grid operation analysis, two views are divided into season and date type; according to the law of seasonal fluctuations, four seasons' views are extracted and the data is divided by season; according to the type of date, the weekday and holiday views are divided.
10. The new energy consumption multi-view scene clustering system according to claim 7, characterized in that: The clustering module uses the K-means clustering algorithm to perform clustering operations on each divided data view, and displays the scene features under different data views through the clustering results, including: Input the features of dates corresponding to different seasons and standardize all features; Determine the number of scenarios required for clustering loads, photovoltaics, wind power and tie lines in each season, and randomly select a specified number of sample points as the initial cluster centers; Calculate the distance between samples in different seasons and the cluster center in the corresponding season respectively, and assign the samples to the cluster with the closest distance; for each cluster, calculate the mean of all samples in the cluster as the new cluster center, and replace the old cluster center with the new cluster center; repeat the assignment and update steps until the position of the cluster center no longer changes or the maximum number of iterations is reached; According to the clustering results, the scenarios corresponding to load, photovoltaic, wind power and tie lines in the seasonal view are divided; Input the features of the corresponding dates of working days and rest days, and standardize all the features; Determine the number of scenarios required for clustering loads, photovoltaics, wind power, and tie lines on weekdays and weekends, and randomly select a specified number of sample points as initial cluster centers; Calculate the distances between samples corresponding to working days and rest days, as well as samples corresponding to seasons, and their respective cluster centers, and assign the samples to the clusters with the closest distances; for each cluster, calculate the mean of all samples in the cluster as the new cluster center, and replace the old cluster center with the new cluster center; repeat the assignment and update steps until the position of the cluster center no longer changes or the maximum number of iterations is reached; According to the clustering results, the loads corresponding to working days and weekends, photovoltaic power, wind power and interconnection lines are divided into scenarios.
11. The new energy consumption multi-view scene clustering system according to claim 7, characterized in that: The clustering result evaluation module calculates the silhouette coefficient corresponding to each clustering, measures the quality of the clustering result by the silhouette coefficient, evaluates the compactness of each data point within the cluster and the separation from other clusters; and evaluates the overall effect of clustering by calculating the average silhouette coefficient of each clustering; The Kalinsky-Harabas CH index corresponding to each clustering was calculated. The CH index evaluates the compactness and separation of clusters by calculating the ratio of the distance between clusters to the distance within clusters. The larger the value of the CH index, the better the clustering effect, which is manifested as higher separation between clusters and stronger compactness within clusters. The CH index of multiple clusterings was combined to select the clustering result with the best separation effect. Calculate the Davis-Baulding DB index corresponding to each clustering. The DB index measures the compactness of the data within the cluster and the separation between clusters. The smaller the value of the DB index, the better the clustering effect and the stronger the separation between clusters. By calculating and comparing the DB index of each clustering, the rationality of the clustering results can be verified, the clustering algorithm parameters can be optimized, or the optimal number of clusters can be selected.
12. The new energy consumption multi-view scene clustering system according to claim 7, characterized in that: The scenario analysis and screening module visualizes all curves corresponding to each scenario of seasonal load, photovoltaic, wind power and tie line, analyzes the basic characteristics of the corresponding curves in each scenario, and obtains the corresponding scenario description; Analyze scenes in different seasons and analyze the causes of scene characteristics based on seasonal characteristics; Visualize all curves corresponding to the scenarios of workday and weekend load, photovoltaic, wind power and tie lines, analyze the basic characteristics of the corresponding curves in each scenario, and obtain the corresponding scenario description; Analyze the scenarios on weekdays and weekends, and analyze the causes corresponding to the scenario characteristics based on the electricity consumption characteristics; Calculate the probability of occurrence of corresponding scenarios for load, photovoltaic, wind power and tie lines in each view; Select the scene with the highest probability for each season as the typical scene for the corresponding season; Select the scenarios with the highest probability for workdays and weekends as typical scenarios for workdays and weekends; The date closest to the scene cluster center among all the dates covered by the typical scenes is selected as the typical day.
13. An electronic device, characterized in that: It comprises a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the multi-view scene clustering method for new energy consumption as described in any one of claims 1 to 6.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by the processor, the multi-view scene clustering method for new energy consumption according to any one of claims 1 to 6 is implemented.