A method and device for constructing a standard form of subdivided industry user load

By cleaning, standardizing, and clustering user load data from specific industries, standard load patterns for different scenarios in these industries are constructed. This addresses the shortcomings of existing load clustering methods and enables more accurate load analysis and management support.

CN117370761BActive Publication Date: 2026-07-24STATE GRID FUJIAN ELECTRIC POWER CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
STATE GRID FUJIAN ELECTRIC POWER CO LTD
Filing Date
2023-11-03
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing user load clustering methods have shortcomings in terms of scope and influencing factors, resulting in insufficient accuracy in constructing user load standard forms and failing to meet the high requirements of power systems.

Method used

By acquiring historical user load data from specific industries, cleaning and standardizing the data, initial clustering is performed to obtain typical load curves. Then, secondary clustering is performed based on industry type to construct standard load patterns for specific industries under different scenarios.

Benefits of technology

It enables more accurate and reliable load standard form analysis, better adapting to the characteristics and needs of different industries, and providing support for power system management and planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117370761B_ABST
    Figure CN117370761B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for constructing standard forms of user load in a subdivided industry, comprising: obtaining a user historical load data set in a subdivided industry; sequentially performing cleaning and normalization processing on the user historical load data set to obtain a normalized data set; traversing the user historical load data in the normalized data set, performing preliminary clustering on target user historical load data obtained in the traversal, and obtaining a typical load curve set of a target user in different scenes; and performing secondary clustering on the typical load curve set according to an industry type to obtain load standard forms of different subdivided industries in different scenes. By inducting and analyzing standard forms of industry user load characteristics in different scenes, the load standard forms of the subdivided industry are constructed, so that the characteristics and demands of different industries can be better adapted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power load analysis technology, and in particular to a method and apparatus for constructing standard forms of user load in subdivided industries. Background Technology

[0002] In traditional power system operation and control, power loads exhibit strong periodicity and randomness. In-depth analysis of power load characteristics helps to understand user load patterns, which is crucial for scheduling power generation and optimizing resource allocation. Therefore, power load characteristic analysis and load data mining play an indispensable role in the stable and economical operation of power systems.

[0003] With the advancement of smart grid construction, the power system has gradually accumulated a large amount of data resources. However, the increased frequency of data collection, the rapid growth in data volume, and the diversification of data types have placed higher demands on data processing and information mining. Currently, existing user load clustering methods have certain shortcomings in terms of scope and consideration of influencing factors. Firstly, existing user load clustering methods typically perform cluster analysis on users from one or a few industries. Secondly, fluctuations in user electricity load curves are affected by various factors such as seasons, holidays, and weather, while current methods for establishing standard user load patterns lack analysis of load influencing factors, thus limiting the analysis and mining of historical data and reducing the accuracy of user load standard pattern construction. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method and apparatus for constructing standard load patterns for users in subdivided industries, so as to achieve more accurate and reliable load standard pattern analysis and provide strong support for power system management and planning.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0006] A method for constructing a standard form of user load in a segmented industry includes:

[0007] Obtain historical user load datasets for specific industry segments;

[0008] The user historical load dataset is cleaned and normalized sequentially to obtain a normalized dataset;

[0009] Traverse the historical load data of users in the standardized dataset, perform preliminary clustering on the historical load data of the target users, and obtain a set of typical load curves of the target users in different scenarios.

[0010] The typical load curve set is clustered twice according to industry type to obtain the standard load form of different sub-industries under different scenarios.

[0011] To solve the above-mentioned technical problems, another technical solution adopted by the present invention is as follows:

[0012] A device for constructing a standard form of user load in a specific industry includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method for constructing a standard form of user load in a specific industry as described above.

[0013] The beneficial effects of this invention are as follows: After obtaining the historical load data of users in subdivided industries, the historical load data of users is cleaned and standardized to form unified standardized data of different types of historical load data. Then, the historical load data of users is initially clustered to extract the typical load curves of the users in different scenarios. Then, the typical load curves of users are clustered again according to the industry type to obtain the standard form of different load characteristics in subdivided industries. This enables more accurate and reliable load standard form analysis. Thus, by summarizing and analyzing the standard form of industry user load characteristics in different scenarios, the standard form of subdivided industry load can be constructed, which can better adapt to the characteristics and needs of different industries. Attached Figure Description

[0014] Figure 1 This is a flowchart illustrating the steps of a method for constructing a standard form of user load in a specific industry, as described in an embodiment of the present invention.

[0015] Figure 2 This is a flowchart illustrating the multi-dimensional scenario division process in a method for constructing a standard form of user load in a specific industry, as described in this embodiment of the invention.

[0016] Figure 3 This is another step in the method for constructing a standard form of user load in a subdivided industry according to an embodiment of the present invention;

[0017] Figure 4 This is a schematic diagram of the load standard form of residential load in scenario 12 in a method for constructing a subdivided industry user load standard form in an embodiment of the present invention.

[0018] Figure 5 This is a schematic diagram of the load standard form of residential load in scenario 21 in a method for constructing a subdivided industry user load standard form in an embodiment of the present invention.

[0019] Figure 6 This is a schematic diagram of the load standard form of residential load in scenario 22 in a method for constructing a subdivided industry user load standard form in an embodiment of the present invention;

[0020] Figure 7This is a schematic diagram of a device for constructing a standard form of user load in a specific industry according to an embodiment of the present invention. Detailed Implementation

[0021] To explain in detail the technical content, objectives, and effects of the present invention, the following description is provided in conjunction with the embodiments and accompanying drawings.

[0022] Please refer to Figure 1 A method for constructing a standard form of user load in a segmented industry, comprising:

[0023] Obtain historical user load datasets for specific industry segments;

[0024] The user historical load dataset is cleaned and normalized sequentially to obtain a normalized dataset;

[0025] Traverse the historical load data of users in the standardized dataset, perform preliminary clustering on the historical load data of the target users, and obtain a set of typical load curves of the target users in different scenarios.

[0026] The typical load curve set is clustered twice according to industry type to obtain the standard load form of different sub-industries under different scenarios.

[0027] As described above, the beneficial effects of this invention are as follows: After obtaining a user historical load dataset by acquiring user historical load data under specific industries, the user historical load dataset is cleaned and standardized to form unified standardized data for different types of user historical load data. Subsequently, the user historical load data is initially clustered to extract the typical load curves of the user under different scenarios. Then, the typical load curves of the user are clustered again according to the industry type to obtain the standard form of different load characteristics under specific industries, thereby achieving more accurate and reliable load standard form analysis. Thus, by summarizing and analyzing the standard form of industry user load characteristics under different scenarios, a standard form of load for specific industries can be constructed, which can better adapt to the characteristics and needs of different industries.

[0028] Furthermore, the acquisition of user historical workload datasets under specific industry segments includes:

[0029] The user's historical load data is acquired at a preset collection frequency within a preset time span;

[0030] The user's historical workload data is classified and labeled according to the user's industry type to obtain the user's historical workload dataset.

[0031] As described above, the user's historical load data is obtained based on a preset time span and a preset collection frequency, thereby enabling the selection of different time spans and preset collection frequencies for data collection according to different data volume requirements, thus meeting different accuracy requirements.

[0032] Furthermore, the step of sequentially cleaning and normalizing the user historical load dataset to obtain a normalized dataset includes:

[0033] The user's historical workload dataset is cleaned using a data cleaning algorithm to obtain cleaned data;

[0034] The missing and abnormal data in the cleaned data are filled in by the completion algorithm to obtain the completed data;

[0035] The completed data is standardized using a standardization algorithm to obtain the standardized dataset.

[0036] As described above, cleaning the user's historical workload dataset can avoid the problem of biased results and calculation errors caused by directly using the original data for clustering when there are outliers in the original data; and further, by using a completion algorithm to fill in missing and outlier data, the integrity of the data can be ensured.

[0037] Furthermore, the preliminary clustering of the traversed historical load data of the target users to obtain a set of typical load curves for the target users in different scenarios includes:

[0038] The target user's historical load data is initially clustered using the first clustering algorithm to obtain different clusters;

[0039] By analyzing the correlation between the load curves corresponding to the clusters and different scenarios, typical load curves under different scenarios are obtained.

[0040] As described above, the user's historical load dataset is initially clustered using the first clustering algorithm, and the load curves corresponding to the clusters are analyzed based on different scenarios to obtain the impact of different scenarios on the electricity load, thereby achieving accurate identification of the user's electricity consumption pattern.

[0041] Furthermore, the secondary clustering of the typical load curve set based on industry type to obtain the standard load patterns of different sub-industries under different scenarios includes:

[0042] The typical load curve set is clustered using a second clustering algorithm to obtain the clustering results;

[0043] The clustering results are iteratively optimized based on the clustering effectiveness evaluation index until the error value of the clustering results is less than the preset error value.

[0044] As described above, when clustering a typical load curve set based on the second clustering algorithm, iterative optimization of the clustering results through clustering effectiveness evaluation indicators can optimize the clustering results and improve clustering accuracy.

[0045] Furthermore, the second clustering algorithm clusters the set of typical load curves by:

[0046] One of the typical load curves is randomly selected from the set of typical load curves as the initial cluster center;

[0047] Calculate the shortest distance between each typical load curve in the typical load curve set and the initial cluster center;

[0048] Calculate the probability that each of the typical load curves in the typical load curve set will be selected as the next cluster center, and select the next cluster center according to the roulette wheel method until all cluster center points are selected.

[0049] K-means clustering is performed based on all the cluster centers to obtain the clustering results.

[0050] As described above, clustering based on DBSCAN (Density-Based Spatial Clustering of Applications with Noise) can effectively identify outliers in the dataset, thereby enabling the detection of uncommon typical load curves and the extraction of daily load patterns under typical scenarios.

[0051] Further, the step of iteratively optimizing the clustering results based on a clustering effectiveness evaluation index until the error value of the clustering results is less than a preset error value includes:

[0052] The effect of clustering pairs is evaluated using the silhouette coefficient S(i):

[0053]

[0054]

[0055] Where a(i) represents the cohesion of the sample points, and a(i) is calculated as follows:

[0056]

[0057] j represents other sample points within the same class as sample i, and distance represents the distance to j. Therefore, the smaller a(i) is, the closer the classes are. b(i) is calculated in a similar way to a(i), obtaining multiple values ​​{b1(i), b2(i), b3(i), ... bj} by traversing other clusters. m (i)} select the minimum value as the final result; where the value range of S(i) is [-1,1]. If the silhouette coefficient of a sample approaches 1, it means that the sample has been well clustered; if it approaches 0, it means that the sample is close to the cluster boundary and the clustering attribute is not strong; if it approaches -1, it means that the sample has been misclassified.

[0058] As described above, the best clustering effect can be achieved by iteratively optimizing the clustering results using the silhouette coefficient.

[0059] Furthermore, the obtained load standard forms for different sub-sectors under different scenarios also include:

[0060] Obtain a new user load dataset, and derive a new load curve set based on the new user load dataset;

[0061] Calculate the similarity between the new load curve and the standard load shape to obtain a similarity value;

[0062] Determine whether the similarity value is greater than the similarity threshold. If so, merge the new load curve with the standard load shape.

[0063] As described above, after clustering based on historical user load data, obtaining more electricity consumption data for the industry allows for iterative updates to the previous clustering results using newly acquired user load data, enabling continuous updates to the industry load standard form.

[0064] Further, the calculation of the similarity between the new load curve and the standard load shape to obtain the similarity value includes:

[0065] The similarity value is obtained by calculating the similarity between the new load curve and the standard load shape using cosine similarity and Pearson correlation coefficient.

[0066] As described above, calculating the similarity between the new load curve and the standard load shape based on cosine similarity and Pearson correlation coefficient can improve the accuracy and reliability of subdivided industry load clustering through iterative optimization and continuous updates.

[0067] Another embodiment of the present invention provides a device for constructing a standard form of user load in a segmented industry, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the various steps of the method for constructing a standard form of user load in a segmented industry as described above.

[0068] The method and apparatus for constructing standard load patterns for segmented industries provided by this invention achieve more accurate and reliable standard load patterns through the analysis and clustering of load characteristics in segmented industries, providing strong support for power system management and planning. The following detailed implementation methods illustrate this:

[0069] Example 1

[0070] Please refer to Figure 1 A method for constructing a standard form of user load in a segmented industry, comprising:

[0071] S1. Obtain the historical user load dataset for specific industry segments, specifically:

[0072] The historical load data of users is acquired within a preset time span at a preset collection frequency. For example, if the preset time span is at least one year and the collection frequency is one point every 5 minutes or 15 minutes, then data is collected at the above-mentioned collection frequency within the one-year time span. The time span and preset collection frequency can be set according to specific needs. The historical load data of users includes data on various types of user loads under subdivided industries, their regions, electricity consumption categories, industry classifications, etc. Specifically, the user load data is the user's active power data, i.e., the user's electricity consumption data, which is used for cluster analysis of user electricity consumption as needed. The regions can be divided according to cities, counties, and towns, as different regions... The electricity consumption characteristics of different regions will vary. A specific region can be designated for analysis of the standard electricity consumption curves for users in that region, improving the accuracy of data analysis. Electricity consumption categories include urban residential electricity consumption, large industrial electricity consumption, non-industrial electricity consumption, rural residential electricity consumption, general industrial electricity consumption, school electricity consumption, and commercial electricity consumption, which can be used to analyze the electricity consumption characteristics of users in different categories. Industry classifications include major categories of the national economic classification, such as agriculture, forestry, animal husbandry, and fisheries. Cluster analysis can be performed on each industry based on its type to obtain standard curves for different industries. After acquiring the data, the historical load data of the users is classified and labeled according to the industry type to which the users belong, resulting in the historical load dataset of the users.

[0073] S2. The user historical load dataset is cleaned and normalized sequentially to obtain a normalized dataset, including the following steps:

[0074] S21. The user historical load dataset is cleaned using a data cleaning algorithm to obtain cleaned data. In this embodiment, the Local Outlier Factor (LOF) algorithm is used to detect anomalies in the user historical load dataset, deleting data with obvious errors or incompleteness. The LOF detection algorithm is highly sensitive to outliers; the LOF value of normal nodes is approximately 1, while the LOF value of outliers is much greater than 1. The LOF calculation steps are as follows:

[0075] S211. Calculate the distance between each node and its nearest neighbor, i.e., the distance K. dist (p);

[0076] S212. Calculate the K-neighborhood N of each node. k (p):

[0077] N k (p)={q∈N / {p}dist(p,q)≤K dist (p)};

[0078] Where dist(p,q) is the spatial distance between the p-th object and the q-th object in the data;

[0079] S213. Determine the local reachability distance D of each node. reach (p,q):

[0080] D reach (p,q)=max{K dist (q),dist(p,q)};

[0081] S214. Calculate the local reachability density ρ of each node. Irdk (p):

[0082]

[0083] Where O represents the K-domain N k(P) Any object in;

[0084] S215. Solve for the local anomaly factors of each object:

[0085]

[0086] S22. The missing and abnormal data in the cleaned data are filled in by the completion algorithm to obtain the completed data; in this embodiment, the K nearest neighbor method is used to fill in the missing and abnormal data.

[0087] S23. The completed data is standardized using a standardization algorithm to obtain the normalized dataset; for example, maximum-minimum standardization is used, and the formula for maximum-minimum standardization is as follows:

[0088]

[0089] Where, x max To complete the maximum value in the data, x min To complete the minimum value in the data; numerical normalization can improve the efficiency of the algorithm. In this embodiment, the main purpose of the clustering algorithm is to classify the user's electricity consumption pattern and discover the hidden load pattern information. Therefore, this embodiment uses maximum and minimum value normalization to remove the order of magnitude limitation, map the load to between 0 and 1, pay more attention to the load trend, and intuitively observe the user's electricity consumption pattern.

[0090] S3. Traverse the historical load data of users in the specified dataset, and perform preliminary clustering on the traversed historical load data of target users to obtain a set of typical load curves for target users under different scenarios. The factors influencing user electricity consumption behavior mainly include seasonal factors, holiday factors, and weather factors. The seasonal periodicity of user load refers to the similarity in the monthly load curves of similar months, while the monthly load curves of distant months show significant differences in total load consumption and peak load. For residential users, seasonal periodicity affects their energy consumption patterns; for example, the use of air conditioning in summer and winter may lead to an increase in peak load. Simultaneously, seasonal periodicity also affects the production plans and energy demand of industrial users; seasonal changes in demand may lead to higher or lower peak loads in specific seasons. Similarly, holiday factors and weather factors also affect the load patterns of different industry types. Specifically, taking residential users as an example:

[0091] S31. The historical load data of the target user that has been traversed is initially clustered using the first clustering algorithm to obtain different clusters;

[0092] Please refer to Figure 2For example, K-means clustering analysis can be performed on the load curves of typical users by monthly load. The elbow rule can be used to determine the value of K, and the load curves of residential users can be divided into several different seasonal scenarios. Similar scenarios are merged from the clusters. By analyzing the relationship between months and seasons, the impact of seasonal factors on residential users can be identified. For example, it can be divided into four different seasonal scenarios: spring, summer, autumn, and winter, or three different seasonal scenarios: spring and autumn, summer, and winter, which can be classified according to specific needs. K-means clustering of the load curves under the major seasonal scenarios by daily load can obtain smaller scenarios under the influence of holiday factors (including weekdays, rest days, and special holidays) and weather factors (including sunny days, cloudy days, rainy days, snowy days, and severe weather), which constitute the subdivided scenarios of the load curves of residential users. As shown in Table 1, the load of residential users is divided into 36 subdivided scenarios. At the same time, in order to avoid the impact of severe weather and special power consumption on the load pattern of users, typical load patterns should be extracted from these load curves. Preliminary DBSCAN clustering should be performed on the load curves of each user to extract the typical load patterns of users.

[0093] Table 1. Residential User Load Breakdown Scenarios

[0094]

[0095] In this embodiment, DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is used as the first clustering algorithm for clustering, which includes the following steps:

[0096] S311. For a given dataset D, first give the neighborhood parameters Eps and MinPts;

[0097] S312. Access any unread object point and determine whether it is a core object based on Eps and MinPts. If it is not a core object, it is a boundary point or a noise point. If it is a core object, find all the sample sets that the core object can reach by density, which is a cluster.

[0098] S313. Then visit other unread core objects to find a sample set with achievable density. This will result in another cluster. Continue running until all core objects have been visited.

[0099] S32. Analyze the correlation between the load curves corresponding to the clusters and different scenarios to obtain the typical load curves under different scenarios; that is, after DBSCAN clustering is completed, extract the typical load curves of different types of users under different scenarios, including scenarios under different seasons, scenarios under holidays and weekdays and scenarios under different weather conditions, and remove the atypical load curves of users (which are outliers in the clustering).

[0100] S4. Perform secondary clustering on the typical load curve set according to industry type to obtain the standard load patterns of different sub-industries under different scenarios; for example, including industry types such as agriculture, forestry, animal husbandry, and fishery, that is, obtain the standard load patterns of industry types such as agriculture, forestry, animal husbandry, and fishery under different scenarios; such as Figures 4-6 The figures shown are schematic diagrams illustrating the standard load patterns of residential loads in scenarios 12, 21, and 22 in this embodiment.

[0101] Example 2

[0102] The difference between this embodiment and Embodiment 1 is that it defines the specific steps of the secondary clustering.

[0103] S41. The typical load curve set is clustered using a second clustering algorithm to obtain the clustering results. In this embodiment, the K-means++ clustering algorithm is used to perform secondary clustering on the typical load curve set after the scene division. K-means++ optimizes the method of randomly selecting initial cluster centers in K-means, so that the distance between the initial cluster centers is as large as possible, thereby accelerating the convergence speed of the algorithm. The calculation method of K-means++ is as follows:

[0104] S411. Randomly select one of the typical load curves from the set of typical load curves as the initial cluster center; that is, randomly select a sample from the dataset as the initial cluster center C1;

[0105] S412. Calculate the shortest distance between each typical load curve in the typical load curve set and the initial cluster center; calculate the shortest distance D between each sample and the existing cluster center. x ;

[0106] S413. Calculate the probability that each of the typical load curves in the typical load curve set is selected as the next cluster center, i.e.:

[0107]

[0108] Then, the roulette wheel method is used to select the next cluster center, and so on, until all cluster centers have been selected;

[0109] S414. Perform K-means clustering based on all the cluster centers to obtain the clustering results;

[0110] S42. Iteratively optimize the clustering results according to the clustering effectiveness evaluation index until the error value of the clustering results is less than a preset error value, specifically:

[0111] The silhouette coefficient S(i) is used to evaluate the effectiveness of clustering pairs. The formula for the silhouette coefficient is as follows:

[0112]

[0113] Where a(i) represents the cohesion of the sample points, and a(i) is calculated as follows:

[0114]

[0115] j represents other sample points within the same class as sample i, and distance represents the distance to j. Therefore, the smaller a(i) is, the closer the classes are. b(i) is calculated in a similar way to a(i), obtaining multiple values ​​{b1(i), b2(i), b3(i), ... bj} by traversing other clusters. m The smallest value is chosen from (i)} as the final result, therefore S(i) can be expressed as:

[0116]

[0117] Wherein, the value range of S(i) is [-1,1]. If the silhouette coefficient of a sample is close to 1, it means that the sample has been well clustered; if it is close to 0, it means that the sample is close to the cluster boundary and the clustering attribute is not strong; if it is close to -1, it means that the sample has been misclassified and should be assigned to the other cluster closest to it. The best K is selected based on the silhouette coefficient calculated according to different K values, and the previous clustering result is updated using the current clustering result to finally obtain the optimal distance result.

[0118] Example 3

[0119] The difference between this embodiment and Embodiment 1 or 2 is that it limits the processing method for new data;

[0120] Please refer to Figure 3 Sa1. Obtain a new user load dataset and obtain a new load curve set based on the new user load dataset; at the same time, perform data preprocessing on the obtained new user load data, including data cleaning, outlier removal, and missing value filling, to ensure data quality and availability; the data preprocessing method is the same as that in Example 1.

[0121] Sa2. Calculate the similarity between the new load curve and the standard load shape to obtain a similarity value; in an optional embodiment, the similarity between the new load curve and the standard load shape is calculated using cosine similarity and Pearson correlation coefficient to obtain the similarity value, specifically:

[0122] The formula for calculating cosine similarity is as follows:

[0123]

[0124] Where X ik and X jk These represent two load curves respectively;

[0125] The formula for calculating the Pearson correlation coefficient is as follows:

[0126]

[0127] in and They are x i and x j The sample mean;

[0128] The formula for calculating the overall similarity is as follows:

[0129] S(x i ,x j )=α1·cos(x i ,x j )+α2·r(x i ,x j );

[0130] Where α1 and α2 represent the weighting coefficients of cosine similarity and Pearson correlation coefficient, respectively; α1 + α2 = 1, S(x i ,x j The value range of ) is [-1, 1]. The closer the comprehensive similarity value is to 1, the higher the similarity between the new load curve and the existing standard load curve.

[0131] Sa3. Determine whether the similarity value is greater than the similarity threshold. If yes, merge the new load curve with the standard load shape; otherwise, discard the new load curve.

[0132] By continuously acquiring more historical data and using comprehensive similarity calculation methods, the standard load patterns of subdivided industries can be continuously improved to make them more accurate and adaptable to actual conditions. For example, once a user's electricity consumption curve is known, to determine which subdivided industry the user belongs to, the comprehensive similarity between the curve and the standard load patterns of each subdivided industry can be calculated. The industry with the highest comprehensive similarity is the industry type to which the user belongs.

[0133] Example 4

[0134] Please refer to Figure 7 A device for constructing a standard form of user load in a segmented industry includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the various steps in the method for constructing a standard form of user load in a segmented industry as described in Embodiments 1, 2, and 3.

[0135] In summary, this invention provides a method and apparatus for constructing standard load patterns for segmented industries. Based on typical load curves within a segmented industry, and considering the different characteristics of industry users, it further summarizes and analyzes the standard patterns of industry user load characteristics under different scenarios. It proposes a method for constructing standard load patterns for segmented industries based on massive historical data mining, solving the problem that existing load clustering methods lack the ability to cluster loads for segmented industries in multi-dimensional scenarios. This allows the standard load patterns to better adapt to the characteristics and needs of different industries, providing more accurate user load clustering results. Furthermore, iterative optimization and updates continuously improve the accuracy and reliability of segmented industry load clustering. Through the method proposed in this invention, the standard load patterns can better adapt to the characteristics and needs of different industries; this is of great significance for fields such as power system planning, load forecasting, and energy management.

[0136] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent modifications made based on the content of the present invention's specification and drawings, or direct or indirect applications in related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for constructing a standard form of user load in a segmented industry, characterized in that, include: Obtain historical user load datasets for specific industry segments; The user historical load dataset is cleaned and normalized sequentially to obtain a normalized dataset; Traverse the historical load data of users in the standardized dataset, perform preliminary clustering on the historical load data of the target users, and obtain a set of typical load curves of the target users in different scenarios. The typical load curve set is clustered twice according to industry type to obtain the standard load form of different sub-industries under different scenarios. The obtained load standard forms for different sub-sectors under different scenarios also include: Obtain a new user load dataset, and derive a new load curve set based on the new user load dataset; Calculate the similarity between the new load curve and the standard load shape to obtain a similarity value; Determine whether the similarity value is greater than the similarity threshold. If so, merge the new load curve with the standard load shape. The calculation of the similarity between the new load curve and the standard load shape, to obtain the similarity value, includes: The similarity value is obtained by calculating the similarity between the new load curve and the standard load shape using cosine similarity and Pearson correlation coefficient.

2. The method for constructing a standard form of user load in a segmented industry according to claim 1, characterized in that, The acquisition of historical user workload datasets under specific industry segments includes: The user's historical load data is acquired at a preset collection frequency within a preset time span; The user's historical workload data is classified and labeled according to the user's industry type to obtain the user's historical workload dataset.

3. The method for constructing a standard form of user load in a segmented industry according to claim 1, characterized in that, The process of cleaning and normalizing the user historical load dataset sequentially to obtain a normalized dataset includes: The user's historical workload dataset is cleaned using a data cleaning algorithm to obtain cleaned data; The missing and abnormal data in the cleaned data are filled in by the completion algorithm to obtain the completed data; The completed data is standardized using a standardization algorithm to obtain the standardized dataset.

4. The method for constructing a standard form of user load in a segmented industry according to claim 1, characterized in that, The preliminary clustering of the traversed historical load data of the target users to obtain a set of typical load curves for the target users in different scenarios includes: The target user's historical load data is initially clustered using the first clustering algorithm to obtain different clusters; By analyzing the correlation between the load curves corresponding to the clusters and different scenarios, typical load curves under different scenarios are obtained.

5. The method for constructing a standard form of user load in a segmented industry according to claim 1, characterized in that, The secondary clustering of the typical load curve set based on industry type yields the standard load patterns of different sub-industries under different scenarios, including: The typical load curve set is clustered using a second clustering algorithm to obtain the clustering results; The clustering results are iteratively optimized based on the clustering effectiveness evaluation index until the error value of the clustering results is less than the preset error value.

6. The method for constructing a standard form of user load in a segmented industry according to claim 5, characterized in that, The second clustering algorithm clusters the typical load curve set as follows: One of the typical load curves is randomly selected from the set of typical load curves as the initial cluster center; Calculate the shortest distance between each typical load curve in the typical load curve set and the initial cluster center; Calculate the probability that each of the typical load curves in the typical load curve set will be selected as the next cluster center, and select the next cluster center according to the roulette wheel method until all cluster center points are selected. K-means clustering is performed based on all the cluster centers to obtain the clustering results.

7. The method for constructing a standard form of user load in a segmented industry according to claim 5, characterized in that, The step of iteratively optimizing the clustering results based on the clustering effectiveness evaluation index until the error value of the clustering results is less than a preset error value includes: The effect of clustering pairs is evaluated using the silhouette coefficient S(i): ; ; Where a(i) represents the cohesion of the sample points, and a(i) is calculated as follows: ; j represents other sample points within the same class as sample i, and distance represents the distance to j. Therefore, the smaller a(i) is, the closer the classes are. b(i) is calculated in a similar way to a(i), obtaining multiple values ​​{b1(i), b2(i), b3(i), ... bj} by traversing other clusters. m (i)} select the minimum value as the final result; where the value range of S(i) is [-1,1]. If the silhouette coefficient of a sample approaches 1, it means that the sample has been well clustered; if it approaches 0, it means that the sample is close to the cluster boundary and the clustering attribute is not strong; if it approaches -1, it means that the sample has been misclassified.

8. A device for constructing a standard form of user load in a specific industry, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements each step of the method for constructing a standard form of user load in a segmented industry as described in any one of claims 1-7.