Industrial electricity data classification method and system based on association clustering analysis
By performing spatiotemporal interpolation and feature extraction of electricity consumption data within industrial parks, combined with density peak clustering and association clustering analysis, the problem of inaccurate electricity consumption data classification in traditional methods is solved, achieving more precise electricity consumption management and resource optimization.
Patent Information
- Application Number
- CN202511117005.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-08-11
AI Technical Summary
Traditional industrial electricity consumption data analysis methods rely on simple classification and clustering methods, which cannot fully explore the potential relationships and complex features in the electricity consumption data, resulting in inaccurate classification results.
By acquiring electricity consumption data from industrial parks, performing interpolation to fill in missing values related to spatiotemporal distribution, extracting electricity consumption feature vectors, calculating spatial distance and local distribution density to perform density peak clustering, and mining associated cluster terms, the accurate classification of industrial electricity consumption data can be achieved.
It improves the completeness and accuracy of electricity consumption data, reveals the peak, off-peak and stable periods of electricity consumption patterns, assesses equipment energy efficiency and cyclical electricity consumption patterns in production, provides more accurate electricity scheduling strategies and energy efficiency sharing opportunities, and improves the accuracy of classification results.
Smart Images

Figure CN120632597B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of correlation prediction technology, and in particular to a method and system for classifying industrial electricity consumption data based on correlation clustering analysis. Background Technology
[0002] In recent years, with the booming development of big data technology and the continuous advancement of data mining algorithms, data classification methods based on association rules and cluster analysis have gradually become a research hotspot. Association rule analysis, by mining the inherent relationships between data, can discover the mutual influences and usage patterns among different electrical devices, providing a more refined basis for industrial electricity management. Cluster analysis, on the other hand, can identify groups with different electricity consumption behaviors by clustering electricity consumption data, helping to identify potential abnormal electricity consumption patterns or equipment failures. However, traditional industrial electricity consumption data analysis methods mostly rely on simple statistical classification and clustering methods, such as K-means clustering or time series-based trend analysis. While these methods provide an understanding of industrial electricity demand to some extent, their over-reliance on fixed rules or simple models often fails to fully explore the potential relationships and complex features in the electricity consumption data, resulting in inaccurate classification results. Summary of the Invention
[0003] Therefore, it is necessary for the present invention to provide a method and system for classifying industrial electricity consumption data based on association clustering analysis, in order to solve at least one of the above-mentioned technical problems.
[0004] To achieve the above objectives, a classification method for industrial electricity consumption data based on association clustering analysis is proposed, comprising the following steps:
[0005] Step S1: Obtain the electricity consumption data of each industry in the industrial park, and perform missing interpolation based on the spatiotemporal distribution of the electricity consumption data of each industry to obtain the interpolated electricity consumption data of each industry.
[0006] Step S2: Perform electricity consumption feature vector analysis on the interpolated data of electricity consumption for each industry to obtain the industry electricity consumption feature vector corresponding to the electricity load distribution characteristics, peak-valley-normal period electricity consumption ratio characteristics, equipment start-up and shutdown electricity consumption characteristics, and production cycle electricity consumption characteristics of each industry.
[0007] Step S3: Obtain the spatial distance and local distribution density between the electricity consumption feature vectors of each industry through the electricity consumption feature vectors of each industry, and perform density peak clustering on the electricity consumption data of each industry based on the spatial distance and local distribution density between the electricity consumption feature vectors of each industry to obtain the electricity consumption data clusters of each industry.
[0008] Step S4: Perform association and aggregation mining analysis on the electricity consumption data clusters of each industry to obtain the industry electricity consumption association aggregation item category set; classify the industry electricity consumption data corresponding to each industry based on the industry electricity consumption association aggregation item category set to obtain the industry electricity consumption data classification results.
[0009] Furthermore, step S1 includes the following steps:
[0010] Step S11: Obtain the electricity consumption data of each industry in the industrial park, including the electricity load data, electricity time series data and equipment operation data of each industry;
[0011] Step S12: Perform anomaly identification and removal on the industrial electricity consumption data corresponding to each industry, calculate the moving average and standard deviation of the industrial electricity consumption data, determine the data fluctuation threshold based on the moving average and standard deviation, and remove abnormal data points that exceed the data fluctuation threshold based on the data fluctuation threshold to obtain the industrial electricity consumption anomaly removal data corresponding to each industry.
[0012] Step S13: Perform spatiotemporal dimension distribution analysis on the electricity consumption anomaly removal data corresponding to each industry to obtain the spatiotemporal dimension distribution of electricity consumption data for each industry;
[0013] Step S14: Based on the spatiotemporal dimension distribution of electricity consumption data for each industry and combined with the time series autoregressive model and spatial correlation, perform spatiotemporal distribution correlation prediction on the electricity consumption anomaly removal data for each industry to generate missing electricity consumption data for each industry.
[0014] Step S15: Based on the abnormal electricity consumption data of each industry, perform missing interpolation filling on the corresponding missing electricity consumption data to obtain the interpolated electricity consumption data of each industry.
[0015] Furthermore, step S2 includes the following steps:
[0016] Step S21: Plot the electricity load curves for each industry to generate a sequence of electricity load curves for each industry; perform electricity load distribution statistics on the sequence of electricity load curves for each industry to obtain the electricity load distribution characteristics for each industry, including the electricity load distribution frequency, the electricity load distribution amplitude, and the electricity load distribution phase.
[0017] Step S22: Obtain the corresponding peak electricity consumption period, valley electricity consumption period, and level electricity consumption period through the electricity consumption time series data corresponding to each industry, and calculate the electricity consumption ratio between the electricity consumption time series data in the corresponding period based on the peak electricity consumption period, valley electricity consumption period, and level electricity consumption period to obtain the peak-valley-normal period electricity consumption ratio characteristics of each industry.
[0018] Step S23: Based on the equipment operation data corresponding to each industry, perform equipment start-up and shutdown characteristic analysis on the power load data corresponding to each industry to obtain the equipment start-up and shutdown power consumption characteristics corresponding to each industry;
[0019] Step S24: Evaluate the production cycle fluctuations of the electricity consumption time series data corresponding to each industry to obtain the electricity consumption characteristics of the production cycle for each industry;
[0020] Step S25: Combine the electricity load distribution characteristics, peak-valley-normal period electricity consumption ratio characteristics, equipment start-up and shutdown electricity consumption characteristics, and production cycle electricity consumption characteristics of each industry to construct the corresponding industry electricity consumption feature vector.
[0021] Furthermore, step S23 includes the following steps:
[0022] Step S231: Obtain the corresponding equipment start-up and shutdown periods through the equipment operation data of each industry;
[0023] Step S232: Perform equipment operation frequency statistics on the equipment operation data corresponding to each industry to obtain the equipment operation frequency corresponding to each industry;
[0024] Step S233: Based on the start-up and shutdown periods of equipment in each industry, evaluate the power consumption attenuation of equipment start-up and shutdown for each industry to obtain the power consumption attenuation efficiency of equipment start-up and shutdown for each industry.
[0025] Step S234: Based on the power consumption attenuation efficiency of equipment start-up and shutdown for each industry, conduct an evaluation and analysis of the impact characteristics of equipment start-up and shutdown on power load for each industry, so as to evaluate and analyze the distribution characteristics of the impact of equipment start-up and shutdown on power load and obtain the power consumption characteristics of equipment start-up and shutdown for each industry.
[0026] Furthermore, step S24 includes the following steps:
[0027] Step S241: Obtain the production stage cycle corresponding to each industry;
[0028] Step S242: Divide the electricity consumption time series data of each industry into production cycle electricity consumption based on the production stage cycle of each industry, so as to obtain the electricity consumption time series data of each industry under different production cycles;
[0029] Step S243: Perform a gradient analysis of the electricity consumption distribution between different production cycles for the electricity consumption time series data of each industry under different production cycles, so as to obtain the electricity consumption distribution gradient between different production cycles for each industry.
[0030] Step S244: Estimate the electricity consumption fluctuation of each industry between different production cycles to obtain the electricity consumption fluctuation coefficient of each industry for the production cycle.
[0031] Step S245: Based on the electricity consumption fluctuation coefficient of the production cycle corresponding to each industry, perform production cycle fluctuation characteristic analysis on the corresponding electricity consumption time series data to obtain the electricity consumption characteristics of the production cycle corresponding to each industry.
[0032] Furthermore, step S3 includes the following steps:
[0033] Step S31: Obtain the spatial distance between the electricity consumption feature vectors of each industry through the industry electricity consumption feature vectors corresponding to each industry;
[0034] Step S32: Based on the spatial distance between the corresponding industrial electricity consumption feature vectors, perform spatial distribution analysis on the corresponding industrial electricity consumption feature vectors to obtain the vector spatial distribution between the corresponding industrial electricity consumption feature vectors;
[0035] Step S33: Calculate the spatial kernel density of the vector space distribution corresponding to the electricity consumption feature vectors of each industry to obtain the magnitude of the spatial kernel density corresponding to the electricity consumption feature vectors of each industry;
[0036] Step S34: Based on the spatial kernel density between the electricity consumption feature vectors of each industry, perform spatial local distribution estimation between the corresponding industry electricity consumption feature vectors to obtain the local distribution density between the electricity consumption feature vectors of each industry.
[0037] Step S35: Based on the spatial distance and local distribution density between the electricity consumption feature vectors of each industry, perform density peak clustering on the electricity consumption data of each industry to obtain the electricity consumption data clusters of each industry.
[0038] Furthermore, step S35 includes the following steps:
[0039] The corresponding industrial electricity consumption cluster centers are determined based on the spatial distance and local distribution density between the characteristic vectors of electricity consumption of each industry.
[0040] Electricity consumption data distribution density is statistically analyzed for each industry to obtain the electricity consumption data distribution density for each industry.
[0041] Based on the distribution density of electricity consumption data corresponding to each industry and combined with the industry electricity consumption cluster center, the industry electricity consumption data corresponding to each industry is clustered by density peak. Based on the density peak reachability between each electricity consumption data distribution density and the industry electricity consumption cluster center, the corresponding industry electricity consumption data is divided into different electricity density peak clusters to obtain each industry electricity consumption data cluster.
[0042] Furthermore, step S4 includes the following steps:
[0043] Step S41: Discretize the industrial electricity consumption data within each industrial electricity consumption data cluster to obtain each industrial electricity consumption discretization cluster;
[0044] Step S42: Perform association clustering mining on each industrial electricity consumption discretization cluster to obtain an industrial electricity consumption association clustering dataset;
[0045] Step S43: Obtain the support and confidence of each associated cluster item through the industrial electricity consumption associated cluster item dataset;
[0046] Step S44: Based on the support and confidence of each associated cluster item, classify and filter the associated cluster items in the industrial electricity consumption associated cluster item dataset to obtain the industrial electricity consumption associated cluster item class set;
[0047] Step S45: Classify the industrial electricity consumption data corresponding to each industry based on the industrial electricity consumption association cluster category set to obtain the industrial electricity consumption data classification results.
[0048] Furthermore, step S45 includes the following steps:
[0049] Step S451: Based on the industrial electricity consumption association cluster set, perform industrial electricity consumption association classification on the industrial electricity consumption data corresponding to each industry to obtain the initial classification results of industrial electricity consumption;
[0050] Step S452: Dynamically update and monitor the initial classification results of industrial electricity consumption to collect corresponding new electricity consumption data characteristics in real time;
[0051] Step S453: Calculate the similarity between the new electricity consumption data features and the existing industrial electricity consumption data clusters. If the similarity is lower than a preset threshold, calculate the distance between the new electricity consumption data features and the corresponding industrial electricity consumption cluster centers to obtain the distance value between the new electricity consumption features and the cluster centers.
[0052] Step S454: Determine the corresponding temporary electricity consumption cluster based on the distance between the new electricity consumption characteristics and the cluster center, and reclassify and adjust the initial classification results of industrial electricity consumption based on the temporary electricity consumption cluster to obtain the classification results of industrial electricity consumption data.
[0053] Furthermore, the present invention also provides an industrial electricity consumption data classification system based on association clustering analysis, used to execute the industrial electricity consumption data classification method based on association clustering analysis as described above. The industrial electricity consumption data classification system based on association clustering analysis includes:
[0054] The industrial electricity consumption filling module is used to obtain the industrial electricity consumption data corresponding to each industry in the industrial park, and to perform missing interpolation filling based on the spatiotemporal distribution of the industrial electricity consumption data corresponding to each industry to obtain the industrial electricity consumption interpolation filling data corresponding to each industry.
[0055] The electricity consumption feature vector analysis module is used to perform electricity consumption feature vector analysis on the interpolated data of electricity consumption for each industry, so as to obtain the industry electricity consumption feature vector corresponding to the electricity load distribution characteristics, peak-valley-normal period electricity consumption ratio characteristics, equipment start-up and shutdown electricity consumption characteristics, and production cycle electricity consumption characteristics of each industry.
[0056] The density peak clustering module is used to obtain the spatial distance and local distribution density between the electricity consumption feature vectors of each industry through the electricity consumption feature vectors of each industry, and to perform density peak clustering on the electricity consumption data of each industry based on the spatial distance and local distribution density between the electricity consumption feature vectors of each industry to obtain the electricity consumption data clusters of each industry.
[0057] The association-aggregated electricity consumption classification module is used to perform association-aggregated mining analysis on electricity consumption data clusters of various industries to obtain industry electricity consumption association-aggregated item categories; based on the industry electricity consumption association-aggregated item categories, the corresponding industry electricity consumption data of each industry is classified to obtain the industry electricity consumption data classification results.
[0058] The beneficial effects of this invention are:
[0059] 1. The industrial electricity consumption data classification method based on association clustering analysis proposed in this invention has the following advantages over existing technologies: it fills in the data gaps in the spatiotemporal distribution by acquiring electricity consumption data of various industries within an industrial park, and ensures the integrity and accuracy of the data. Since electricity consumption data may have time-series gaps or sensor malfunctions during the acquisition process, handling missing values is crucial. Traditional data gap handling methods, such as mean imputation and linear interpolation, often ignore the spatiotemporal correlation of the data, resulting in inaccurate interpolation results. This step, however, uses a spatiotemporal distribution-based interpolation method, which considers the time series and spatial distribution characteristics of the data, and can more accurately fill in missing data. Especially in a diversified environment like an industrial park, where each industry's electricity consumption pattern has different spatiotemporal characteristics, spatiotemporal distribution-related interpolation can optimize the filling process in a wider range of spatiotemporal dimensions, thereby improving data quality. The filled data is more continuous and consistent, helping to avoid bias caused by data gaps in subsequent analysis and ensuring the representativeness and reliability of industrial electricity consumption data. Secondly, by extracting features from the interpolated and filled industrial electricity consumption data, feature vectors that comprehensively reflect the electricity consumption patterns of each industry are generated. These features not only involve basic electricity consumption data, but also cover the temporal distribution of electricity load, peak-valley differences, equipment start-up and shutdown behavior, and electricity consumption characteristics during the production cycle. Among them, the electricity consumption ratio characteristics of peak-valley-normal periods help to reveal the electricity consumption patterns of each industry during peak, off-peak, and stable periods, thus providing a basis for optimizing electricity scheduling and energy-saving measures. Equipment start-up and shutdown electricity consumption characteristics can reflect the changes in electricity consumption when equipment is started and stopped, thereby assessing the energy efficiency and operating status of equipment. Production cycle electricity consumption characteristics help to identify industries with obvious production cycle electricity consumption patterns, assisting in the formulation of more accurate production scheduling strategies. Through the comprehensive analysis of these feature vectors, the electricity consumption behavior of each industry can be fully grasped, potential energy-saving and scheduling optimization opportunities can be explored, and basic data support can be provided for the next step of cluster analysis and classification. Then, by calculating the spatial distance and local distribution density between the characteristic vectors of industrial electricity consumption, industries with similar electricity consumption patterns are clustered together to form electricity consumption data clusters. The calculation of spatial distance helps to determine the similarity of electricity consumption characteristics between different industries, and further fully explores the potential correlations and complex features in the electricity consumption data. Through the density peak clustering method, naturally clustered industrial groups can be efficiently identified. It can not only find the similarities between industries, but also eliminate noisy data and outliers, making the clustering results more stable and reliable. The advantage of density peak clustering is that it can automatically select appropriate cluster centers based on the local distribution of data, avoiding the difficulty in selecting cluster centers and human interference that occurs in the traditional K-means method. This process helps to classify industries with similar electricity consumption patterns into the same category, effectively supporting subsequent analysis and decision-making.Finally, the main task is to mine potential correlation patterns from the segmented electricity consumption data clusters to gain a deeper understanding of the commonalities and differences in electricity consumption across industries. Association clustering mining, by discovering correlations between different industries' electricity consumption, can reveal the inter-industry correlations. For example, it can identify which industries have similar electricity consumption behaviors within the same time period, and which industries may have synergistic or complementary relationships. This analysis can provide industrial parks with more multi-dimensional electricity consumption optimization strategies, such as cross-industry collaborative scheduling and energy efficiency sharing. Based on these association clusters, industrial electricity consumption data can be accurately classified, grouping industries with similar electricity consumption characteristics into one category. This provides a basis for the optimal allocation of power resources. For example, for industries with stable electricity demand, centralized scheduling can be considered, while for industries with large fluctuations in electricity consumption, smoothing can be achieved through energy storage systems. This step not only improves the accuracy of industrial electricity consumption management but also enhances the precision of the classification results.
[0060] 2. The industrial electricity consumption data classification system based on association clustering analysis proposed in this invention is composed of an industrial electricity consumption filling module, an electricity consumption feature vector analysis module, a density peak clustering module, and an association clustering electricity consumption classification module. It can realize any industrial electricity consumption data classification method based on association clustering analysis described in this invention. It is used to combine the operations between the computer programs running on each module to realize the industrial electricity consumption data classification method based on association clustering analysis. The internal structure of the system cooperates with each other, which can greatly reduce repetitive work and manpower input, and can quickly and effectively provide a more accurate and efficient industrial electricity consumption data classification process based on association clustering analysis, thereby simplifying the operation process of the industrial electricity consumption data classification system based on association clustering analysis. Attached Figure Description
[0061] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0062] Figure 1 This is a flowchart illustrating the steps of the industrial electricity consumption data classification method based on association clustering analysis of the present invention. Detailed Implementation
[0063] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0064] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0065] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0066] To achieve the above objectives, please refer to Figure 1 The diagram shown illustrates the steps of the industrial electricity consumption data classification method based on association clustering analysis according to the present invention. In this example, the industrial electricity consumption data classification method based on association clustering analysis includes the following steps:
[0067] Step S1: Obtain the electricity consumption data of each industry in the industrial park, and perform missing interpolation based on the spatiotemporal distribution of the electricity consumption data of each industry to obtain the interpolated electricity consumption data of each industry.
[0068] In this embodiment of the invention, electricity consumption data for various industries is collected in real time within the industrial park using smart meters and equipment sensors. The smart meters collect power load data such as active power, reactive power, current, and voltage at 15-minute intervals. The equipment sensors record equipment operation data such as start / stop status, operating speed, and operating temperature every 10 seconds, along with corresponding timestamps. For example, in an electronics manufacturing industry, the smart meter collects an active power of 45kW between 9:00 and 9:15 AM on a certain day, while the equipment sensor records that the placement machine is operating normally at 2000 rpm during that time. If the electricity load data for a textile industry between 2:00 and 2:15 PM is missing, a spatiotemporal distribution correlation method is used. The missing interpolation imputation method, in terms of time dimension, refers to the distribution of electricity consumption data for the same period in the past week for this industry, and finds that the average active power is 38kW; in terms of spatial dimension, it compares the electricity consumption data of similar textile industries in the park during the same period, and the active power of three surrounding enterprises during the same period is 36kW, 39kW and 37kW respectively. The weighted average of the time and spatial data is taken (time dimension weight 0.6, spatial dimension weight 0.4), that is, (38×0.6+(36+39+37)÷3×0.4)=37.6kW. 37.6kW is filled into the missing period. The missing values of electricity consumption data for all industries in the park are processed in this way, and finally the corresponding interpolated electricity consumption data for each industry is obtained.
[0069] Step S2: Perform electricity consumption feature vector analysis on the interpolated data of electricity consumption for each industry to obtain the industry electricity consumption feature vector corresponding to the electricity load distribution characteristics, peak-valley-normal period electricity consumption ratio characteristics, equipment start-up and shutdown electricity consumption characteristics, and production cycle electricity consumption characteristics of each industry.
[0070] In this embodiment of the invention, using electricity consumption interpolation data from a certain mechanical processing industry as an example, electricity consumption feature vector analysis is performed. When plotting the electricity load curve, time is used as the horizontal axis and active power at 15-minute intervals is used as the vertical axis. Connecting the data points forms a curve. By analyzing the curve, the peak electricity consumption periods are determined to be 8:00-12:00 and 16:00-20:00, and the off-peak period is 0:00-6:00. The frequency, amplitude, and phase of the electricity load distribution are calculated. For example, the peak electricity load distribution frequency accounts for 30% of the total statistical period, the load amplitude is the difference of 65kW between the maximum value of 80kW and the minimum value of 15kW, and the phase is characterized by peaks occurring at fixed times each day, according to the local power grid peak-valley-flat electricity price period (peak period 8:00). From 12:00 and 17:00-21:00, during off-peak hours from 23:00 to 7:00 the next day, and the rest of the normal hours, the peak electricity consumption of this industry is calculated to be 45%, the off-peak hours 15%, and the normal hours 40%. The correlation between equipment operation data and electricity consumption data is analyzed. When a CNC machine tool is started, the power load instantly rises from 5kW in standby mode to 30kW, thus determining the power consumption characteristics of equipment start-up and shutdown. Combined with production plan data, the production cycle of this industry includes three stages: raw material processing, parts manufacturing, and assembly. The average electricity consumption and fluctuation of each stage are calculated to obtain the power consumption characteristics of the production cycle. These characteristics are integrated to form the power consumption characteristic vector of this mechanical processing industry. Other industries in the park also generate their own characteristic vectors according to this process.
[0071] Step S3: Obtain the spatial distance and local distribution density between the electricity consumption feature vectors of each industry through the electricity consumption feature vectors of each industry, and perform density peak clustering on the electricity consumption data of each industry based on the spatial distance and local distribution density between the electricity consumption feature vectors of each industry to obtain the electricity consumption data clusters of each industry.
[0072] In this embodiment of the invention, the spatial distance between an automotive parts manufacturing industry and an electrical equipment manufacturing industry is calculated using their respective electricity consumption characteristic vectors as examples. Assume the electricity consumption characteristic vector for the automotive parts manufacturing industry is [0.28, 70kW, 6h, 0.58, 0.13, 0.29, 22kW, 16kW, 24kW·h, 3.8kW·h, 0.24], and the electricity consumption characteristic vector for the electrical equipment manufacturing industry is [0.32, 75kW, 6.5h, 0.62, 0.11, 0.31, 25kW, 18kW, 26kW·h, 4.2kW·h, 0.26]. Using the Euclidean distance formula, the square root of the sum of squares of the differences in corresponding dimensions is used to obtain the spatial distance value. When calculating the local distribution density, a distance threshold of 0.15 is set (determined based on all spatial distance statistics). The number of other industries whose distance to the characteristic vector of a certain industry's electricity consumption is less than this threshold is counted. For example, if there are 7 industries around a certain food processing industry that meet the conditions, then its local distribution density is 7. Based on the spatial distance and local distribution density, the density peak clustering algorithm is used, with a density threshold of 0.6 and a minimum sample size of 5. The local distribution density of a certain new materials industry is 8, and it is far away from other high-density points, so it is determined as the cluster center. The electricity consumption data of each industry are divided into different industry electricity consumption data clusters according to the density peak reachability with the cluster center. For example, the electronic information industry is divided into one cluster, and the machinery manufacturing industry is divided into another cluster.
[0073] Step S4: Perform association and aggregation mining analysis on the electricity consumption data clusters of each industry to obtain the industry electricity consumption association aggregation item category set; classify the industry electricity consumption data corresponding to each industry based on the industry electricity consumption association aggregation item category set to obtain the industry electricity consumption data classification results.
[0074] In this embodiment of the invention, association aggregation mining analysis is performed on a data cluster of electricity consumption from a high-energy-consuming industry. Using the Apriori algorithm, a candidate set of association aggregation terms is formed by combining discrete ranges of electricity load (e.g., 100kW-150kW) and discrete ranges of equipment start-up and shutdown frequency (e.g., 3-5 times per day). The set is defined as "electricity load between 100kW-150kW and equipment start-up and shutdown 3-5 times per day." All industrial electricity consumption data within this cluster are traversed, and the frequency of this candidate set is counted. Assuming it appears 30 times out of 200 data points, other candidate sets are generated and counted to obtain the association aggregation term dataset. The support and confidence of each association aggregation term are calculated. The support of this association aggregation term is 30 ÷ 200 = 0.15. For confidence calculation, the number of records with electricity load between 100kW-150kW is first counted (80 records), of which 30 records have equipment start-up and shutdown 3-5 times per day. The confidence is 30 ÷ 80 = 0.375. The support is set... With a threshold of 0.1 and a confidence threshold of 0.3, relevant clusters that meet the criteria are selected to form a cluster of industrial electricity consumption clusters. Taking a precision instrument manufacturing industry as an example, its electricity consumption data is an electricity load of 120kW and equipment start-up and shutdown 4 times a day, which matches the cluster of "electricity load between 100kW and 150kW and equipment start-up and shutdown 3-5 times a day". Based on the category to which this cluster belongs (high energy consumption and frequent equipment start-up and shutdown), the industry is classified into the corresponding category. The electricity consumption data of all industries in the park are matched and classified, and finally the complete classification results of industrial electricity consumption data are obtained.
[0075] Furthermore, step S1 includes the following steps:
[0076] Step S11: Obtain the electricity consumption data of each industry in the industrial park, including the electricity load data, electricity time series data and equipment operation data of each industry;
[0077] In this embodiment of the invention, within an industrial park, smart meters collect electricity load data from various industries at 15-minute intervals. This data includes parameters such as active power, reactive power, current, and voltage. For example, a smart meter at an electronics manufacturing company collects data from 9:00 to 9:15, showing an active power of 50kW, a reactive power of 15kvar, a current of 100A, and a voltage of 380V. The electricity consumption time series data records the electricity consumption at each collection point, accurate to the minute, ensuring the data's continuity over time. Equipment operation data is acquired through sensors installed on production equipment. For CNC machine tools… Vibration sensors, temperature sensors, and speed sensors are installed on key parts such as the spindle and feed axis, and data is collected every 10 seconds. For assembly line equipment, photoelectric sensors and current sensors are deployed at locations such as conveyor belts and motors to monitor the equipment's start-up and shutdown status and operating current in real time. This data is transmitted to the park's data center via industrial Ethernet and stored in association with power load data and time series data to form a complete set of industrial power consumption data. For example, the equipment operation data of an automotive parts processing workshop includes the spindle speed of a CNC machine tool at 9:00, which is 2000 r / min, the temperature is 45℃, and the vibration amplitude is 0.05 mm / s, corresponding one-to-one with the power consumption data at the same time.
[0078] Step S12: Perform anomaly identification and removal on the industrial electricity consumption data corresponding to each industry, calculate the moving average and standard deviation of the industrial electricity consumption data, determine the data fluctuation threshold based on the moving average and standard deviation, and remove abnormal data points that exceed the data fluctuation threshold based on the data fluctuation threshold to obtain the industrial electricity consumption anomaly removal data corresponding to each industry.
[0079] In this embodiment of the invention, when identifying and removing anomalies in industrial electricity consumption data, a method based on moving average and standard deviation is adopted. Taking the active power data of a textile enterprise as an example, the moving average window is set to 12 time intervals (i.e., 3 hours). The moving average value at each time point is calculated. For example, in the period from 10:00 to 10:15, the power data of the first 12 intervals are 40kW, 42kW, 45kW, etc., and the moving average value for this period is calculated to be 43kW. At the same time, the standard deviation of these 12 data points is calculated, assumed to be 2kW. The 3σ principle is used to determine the data fluctuation threshold, which is the moving average ± 3 times the standard deviation. In this example, the lower limit of the threshold is 43 - 3 × 2 = 37 kW, and the upper limit is 43 + 3 × 2 = 49 kW. If the active power collected at a certain moment is 55 kW, which exceeds the upper limit of the threshold, it is judged as an abnormal data point and removed. Numerical data in the power load data, power time series data, and equipment operation data of all industries are processed in this way to finally obtain the corresponding industry power consumption anomaly removal data, ensuring the accuracy and reliability of the data.
[0080] Step S13: Perform spatiotemporal dimension distribution analysis on the electricity consumption anomaly removal data corresponding to each industry to obtain the spatiotemporal dimension distribution of electricity consumption data for each industry;
[0081] In this embodiment of the invention, when performing spatiotemporal distribution analysis on the abnormal electricity consumption data of industries, in terms of time, a day is divided into 96 15-minute time periods, and a month is divided into 30 days × 96 time periods. The average electricity load and average equipment operating status of each industry are statistically analyzed for each time period. For example, analyzing the data of a food company over a month, it is found that its average active power during the 8:00-9:00 time period is 35kW, and the equipment operating status shows that the production line start-up rate is 80% during this time period. In terms of space, the industries are divided into regions according to their geographical location in the park. For example, the park is divided into three regions: A, B, and C. The total electricity load and equipment operating status of all industries in each region during the same time period are statistically analyzed. For example, the average total electricity load of region A during weekdays is 500kW, and 70% of the equipment in the region is in operation. In this way, the detailed distribution of electricity consumption data of each industry in time and space is obtained, forming a spatiotemporal distribution map of electricity consumption data of each industry, which intuitively presents the changing patterns of industrial electricity consumption in different times and spaces.
[0082] Step S14: Based on the spatiotemporal dimension distribution of electricity consumption data for each industry and combined with the time series autoregressive model and spatial correlation, perform spatiotemporal distribution correlation prediction on the electricity consumption anomaly removal data for each industry to generate missing electricity consumption data for each industry.
[0083] In this embodiment of the invention, based on the spatiotemporal dimension distribution, an autoregressive model of time series and spatial correlation are used to predict the abnormal removal data of industrial electricity consumption. Taking a certain mechanical processing enterprise as an example, in terms of time series prediction, the autoregressive model AR (3) is used. Based on the active power data of the enterprise in the past three time intervals (45 minutes) (such as the previous three intervals being 48kW, 50kW, and 52kW respectively), the predicted power value of the next time interval (15 minutes later) is calculated by the model to be 54kW. In terms of spatial correlation analysis, the electricity consumption data of other similar mechanical processing enterprises in the area where the enterprise is located are observed. If the power growth trend of the three surrounding enterprises in the same time period is relatively obvious and has a strong correlation with the historical data of the enterprise (correlation coefficient reaches 0.8), the predicted value is corrected by combining the electricity consumption data of these enterprises. Assuming that the data of the surrounding enterprises show that the average power growth rate is larger, the predicted value is adjusted to 56kW. Such spatiotemporal distribution correlation prediction is performed on the electricity consumption data of all industries, and finally the predicted value of the missing electricity consumption data corresponding to each industry is generated, providing a basis for data filling.
[0084] Step S15: Based on the abnormal electricity consumption data of each industry, perform missing interpolation filling on the corresponding missing electricity consumption data to obtain the interpolated electricity consumption data of each industry.
[0085] In this embodiment of the invention, missing electricity consumption data is filled by interpolation based on industrial electricity consumption anomaly removal data. Taking a clothing company as an example, if there is missing active power data in the period from 14:00 to 14:15, linear interpolation is used to fill the missing data. Based on two adjacent time points (power of 38kW in the period from 13:45 to 13:60 and power of 42kW in the period from 14:15 to 14:30), the missing value is filled using the formula: Missing value = Previous value + The formula (subsequent value - previous value) × (missing time point - previous time point) / (subsequent time point - previous time point) is used to calculate the filled value as 38 + (42 - 38) × (14:15 - 13:45) / (14:30 - 13:45) = 40kW. For missing data caused by prediction, the same method is used, combined with the anomaly removal data of surrounding time points for filling. If a certain predicted value does not match the trend of the preceding and following data, the electricity consumption data of other similar industries in the same area are further referenced for adjustment. In this way, all missing electricity consumption data of each industry are interpolated and filled, and finally complete industrial electricity consumption interpolation and filling data are obtained, providing a complete data foundation for subsequent industrial electricity consumption data classification.
[0086] Furthermore, step S2 includes the following steps:
[0087] Step S21: Plot the electricity load curves for each industry to generate a sequence of electricity load curves for each industry; perform electricity load distribution statistics on the sequence of electricity load curves for each industry to obtain the electricity load distribution characteristics for each industry, including the electricity load distribution frequency, the electricity load distribution amplitude, and the electricity load distribution phase.
[0088] In this embodiment of the invention, when plotting the electricity load curve for each industry, time is used as the horizontal axis and the electricity load value at 15-minute intervals is used as the vertical axis. For example, in an electronics manufacturing industry, smart meters collect active power data every 15 minutes within a 24-hour period, resulting in 96 data points. The active power value for the period from 0:00 to 0:15 is marked at the position corresponding to 0:00 on the horizontal axis. For example, if the power during this period is 30kW, a point is marked at the corresponding position. In the same way, all data points for the remaining periods are marked on the coordinate system, and then these points are connected sequentially with a smooth broken line to generate the electricity load curve for the industry. This method is used to plot the electricity load curves for all industries within the park. When conducting electricity load distribution statistics, the range of electricity load values is divided into several intervals, such as 10kW per interval, from 0-10kW, 10-20kW, and so on. The frequency of each industry's electricity load data appearing in each interval is counted, and the ratio of these frequency occurrences to the total number of data points is calculated to obtain the electricity load distribution frequency. Taking a certain machinery processing industry as an example, within the statistical period, its electricity load data appeared 20 times in the 30-40kW interval, with a total of 100 data points. The frequency of the power load distribution in this interval is 20 ÷ 100 = 0.2. The amplitude of the power load distribution is the difference between the maximum and minimum power load values. If the maximum power load of this industry is 80 kW and the minimum is 15 kW, the distribution amplitude is 80 - 15 = 65 kW. The phase of the power load distribution is determined by analyzing the fluctuation period of the load curve. Observe the time interval from one peak to the next peak of the curve. If the power load curve of a certain industry has a peak every 8 hours, then its distribution phase is 8 hours. Thus, the power load distribution characteristics of each industry are obtained.
[0089] Step S22: Obtain the corresponding peak electricity consumption period, valley electricity consumption period, and level electricity consumption period through the electricity consumption time series data corresponding to each industry, and calculate the electricity consumption ratio between the electricity consumption time series data in the corresponding period based on the peak electricity consumption period, valley electricity consumption period, and level electricity consumption period to obtain the peak-valley-normal period electricity consumption ratio characteristics of each industry.
[0090] In this embodiment of the invention, when obtaining peak, valley, and normal periods through electricity consumption time series data corresponding to various industries, the peak, valley, and normal electricity price time period division standard of the local power grid is used. For example, the peak period is 8:00-12:00 and 17:00-21:00, the valley period is 23:00-7:00 the next day, and the remaining time period is the normal period. Taking a food processing industry as an example, all data points within the peak period (e.g., 8:00-12:00) are selected from its electricity consumption time series data, and the electricity consumption of these data points is added together to obtain the total electricity consumption during the peak period. Similarly, the total electricity consumption during the valley and normal periods is calculated. Assuming that the total electricity consumption during the peak period of this industry in a day is 2... The peak electricity consumption is 200 kWh, the off-peak electricity consumption is 50 kWh, and the normal electricity consumption is 150 kWh. The total electricity consumption is 200 + 50 + 150 = 400 kWh. The electricity consumption percentage for each period is calculated as follows: peak period electricity consumption percentage is 200 ÷ 400 = 0.5, or 50%; off-peak period electricity consumption percentage is 50 ÷ 400 = 0.125, or 12.5%; and normal period electricity consumption percentage is 150 ÷ 400 = 0.375, or 37.5%. This filtering and calculation is performed on the electricity consumption time series data of all industries in the park to obtain the peak-valley-normal period electricity consumption percentage characteristics of each industry, clearly showing the differences in the electricity consumption ratio of different industries at different times.
[0091] Step S23: Based on the equipment operation data corresponding to each industry, perform equipment start-up and shutdown characteristic analysis on the power load data corresponding to each industry to obtain the equipment start-up and shutdown power consumption characteristics corresponding to each industry;
[0092] In this embodiment of the invention, when analyzing the start-up and shutdown characteristics of electrical load data based on equipment operation data, the start-up and shutdown status in the equipment operation data is matched with the timestamp of the electrical load data on a per-equipment basis. For example, the equipment operation data of a certain injection molding machine shows that it starts at 9:00. During the period from 8:45 to 8:59 before the start-up, its corresponding electrical load data is stable at 5kW. At the moment of start-up (9:00-9:01), the electrical load rapidly rises to 30kW, and then gradually stabilizes to 25kW during the period from 9:02 to 9:15. The analysis shows that the change in electrical load during the start-up process is 30-5=25kW, and the average load increase rate during the start-up time is (30-5). ÷1=25kW / min. When the equipment stops, if the injection molding machine stops running at 17:00, the power load during the period from 16:45 to 16:59 before stopping is 25kW, and it drops to 8kW at the moment of stopping (17:00-17:01). The power load change during the stopping process is 25-8=17kW. This analysis is performed on the start-up and shutdown process of all equipment in the industry. The impact weight of different equipment start-up and shutdown on the overall power load is calculated, and the power consumption characteristics of equipment start-up and shutdown in the industry are summarized. For example, if there are many large equipment in a certain industry, the impact of equipment start-up on the power load is obvious. Its power consumption characteristics of equipment start-up and shutdown are manifested as a sharp increase in load during startup and a rapid decrease in load during shutdown.
[0093] Step S24: Evaluate the production cycle fluctuations of the electricity consumption time series data corresponding to each industry to obtain the electricity consumption characteristics of the production cycle for each industry;
[0094] In this embodiment of the invention, by evaluating the production cycle fluctuations of electricity consumption time series data corresponding to various industries, and based on the acquired production stage cycles of each industry (such as the stamping and welding stages in the automobile manufacturing industry), the electricity consumption time series data is divided according to the production stage. Taking a furniture manufacturing industry as an example, its production cycle includes three stages: wood processing, painting, and assembly. The electricity consumption time series data for the wood processing stage includes all data from 9:00 to 13:00, the painting stage includes data from 14:00 to 17:00, and the assembly stage includes data from 8:00 to 12:00 the next day. The average value and standard deviation of the electricity consumption data for each production stage are calculated. Analyzing the fluctuations in electricity consumption data, the average electricity consumption during the wood processing stage was 25 kW·h, with a standard deviation of 3 kW·h, indicating relatively stable electricity consumption during this stage. The average electricity consumption during the coating stage was 18 kW·h, with a standard deviation of 5 kW·h, showing relatively large fluctuations. Comparing the differences in electricity consumption between different production stages, the rate of change in electricity consumption between adjacent stages was calculated. For example, the rate of change in electricity consumption from wood processing to coating stage was (25-18)÷25=0.28. Through a comprehensive analysis of electricity consumption data at different stages of the production cycle of each industry, the electricity consumption characteristics of each industry's corresponding production cycle were obtained, such as the large fluctuations in electricity consumption in the early stages of production for some industries, which tend to stabilize later.
[0095] Step S25: Combine the electricity load distribution characteristics, peak-valley-normal period electricity consumption ratio characteristics, equipment start-up and shutdown electricity consumption characteristics, and production cycle electricity consumption characteristics of each industry to construct the corresponding industry electricity consumption feature vector.
[0096] In this embodiment of the invention, an industrial electricity consumption feature vector is constructed by merging the electricity load distribution characteristics, peak-valley-normal period electricity consumption ratio characteristics, equipment start-up and shutdown electricity consumption characteristics, and production cycle electricity consumption characteristics corresponding to each industry. The electricity load distribution characteristics are presented in the form of numerical values of distribution frequency, amplitude, and phase. For example, the electricity load distribution characteristics of a certain industry are [0.2 (frequency in the 30-40kW range), 65kW (distribution amplitude), 8h (distribution phase)]; the peak-valley-normal period electricity consumption ratio characteristics are represented as [0.5 (peak period ratio), 0.125 (valley period ratio), 0.375 (normal period ratio)]; the equipment start-up and shutdown electricity consumption characteristics record key data such as the change amplitude and rate of electricity load when the equipment starts up and stops; the production cycle electricity consumption characteristics include information such as the average value, standard deviation, and inter-stage electricity consumption change rate of each production stage. These feature data are arranged in a unified order to form a complete industrial electricity consumption feature vector. For example, the electricity consumption characteristic vector of a certain industry is [0.2, 65kW, 8h, 0.5, 0.125, 0.375, equipment start-up load variation, equipment stop-load variation, average electricity consumption in the wood processing stage, standard deviation of electricity consumption in the coating stage, rate of change of electricity consumption in the assembly and coating stages, etc.]. In this way, the complex electricity consumption characteristics of various industries are transformed into a structured vector form, providing a standardized data foundation for subsequent classification of industrial electricity consumption data based on association and aggregation analysis.
[0097] Furthermore, step S23 includes the following steps:
[0098] Step S231: Obtain the corresponding equipment start-up and shutdown periods through the equipment operation data of each industry;
[0099] In this embodiment of the invention, when obtaining the equipment start-stop period through the equipment operation data corresponding to each industry, the system relies on the real-time collection of status signals from equipment sensors. Within the industrial park, each production piece of equipment is equipped with a status monitoring sensor. For example, CNC machine tools have a start-stop status sensing device connected to their main circuit control module, and assembly line equipment has a start-stop signal acquisition device deployed in its motor control circuit. These sensors continuously record equipment operation status data at 1-second intervals. Taking a certain automotive parts manufacturing industry as an example, the operation data of a stamping machine in its workshop includes equipment status codes, where "0" represents a stopped state and "1" represents a running state. During the production process on a certain day, data is collected from the equipment's operating data... Records show that the status code of the stamping machine changed from "0" to "1" at 8:02 AM, which marks the start-up period. At 5:48 PM, the status code changed back to "0", indicating the stop-up period. All production equipment in this industry, such as welding robots and lathes, are identified in this manner, and this data is organized and stored in order of equipment number and date / time, forming a complete record table of equipment start-up and stop times. Other industries, such as electronic assembly and food processing, also use the same method to accurately extract the start-up and stop times of each piece of equipment from its operating data, providing foundational data for subsequent analysis.
[0100] Step S232: Perform equipment operation frequency statistics on the equipment operation data corresponding to each industry to obtain the equipment operation frequency corresponding to each industry;
[0101] In this embodiment of the invention, when statistically analyzing the equipment operation frequency of equipment corresponding to various industries, a day is used as the time unit, and the 24 hours of a day are divided into 1440 one-minute time intervals. The operating status of the equipment in each interval is statistically analyzed. Taking a textile industry as an example, there are 50 textile machines in its workshop. In the statistics of a certain day, the status of each textile machine in each one-minute interval is checked sequentially from the equipment operation data. If a textile machine is running between 8:00 AM and 8:01 AM, its running count for that interval is incremented by 1. After the statistics are completed, the running counts for each machine throughout the day are summed, and then divided by the total number of intervals for the day, 1440, to obtain the daily running frequency of the machine. Assuming that one textile machine has a running count of 1000 on a given day, its running frequency is 1000 ÷ 1440 ≈ 0.69 (meaning that the machine is running for approximately 69% of the day). This calculation is performed on all 50 textile machines in the industry, and the running frequencies of all machines are summed and averaged to obtain the equipment running frequency for the textile industry. The same operation is applied to other industries in the park, such as the machinery processing industry to count the running frequency of various machine tools, and the chemical industry to count the running frequency of reaction vessels, pumps, etc. Finally, the equipment running frequency data corresponding to each industry is obtained, reflecting the operational activity level of equipment in different industries.
[0102] Step S233: Based on the start-up and shutdown periods of equipment in each industry, evaluate the power consumption attenuation of equipment start-up and shutdown for each industry to obtain the power consumption attenuation efficiency of equipment start-up and shutdown for each industry.
[0103] In this embodiment of the invention, when evaluating the power consumption attenuation during equipment startup and shutdown based on the equipment's operating frequency during startup and shutdown periods, each startup and shutdown process is used as the analysis unit. Taking an SMT placement production line in an electronics manufacturing industry as an example, this production line includes multiple placement machines, reflow soldering machines, and other equipment. Records from the equipment startup and shutdown periods show that a placement machine starts at 9:00 AM, initially in standby mode, with the smart meter recording its active power as 3kW. At startup, the power rapidly increases to 25kW within 10 seconds, then gradually stabilizes at 20kW after 30 seconds. It stops operating at 5:00 PM, with a power of 20kW before stopping. During the stopping process, the power decreases to 4kW within 15 seconds. To calculate the power consumption attenuation efficiency during the startup process, the increase in power from the stable operating value to the moment of startup is first calculated (25-20=5kW), then divided by the stable operating power (5÷20=0.25), resulting in a power consumption attenuation efficiency of 0.25. This indicates that the power consumption at startup increased by 25% compared to stable operation. The calculation method for the power consumption attenuation efficiency during shutdown is similar. The power reduction from the stable operating value to the power consumption after shutdown (20-4=16kW) is calculated and divided by the stable operating power (16÷20=0.8), resulting in a power consumption attenuation efficiency of 0.8 during shutdown, meaning that the power consumption during shutdown is reduced by 80%. This calculation is performed for every startup and shutdown process of all equipment in the industry, and a weighted average is calculated according to equipment type and usage frequency to obtain the power consumption attenuation efficiency of equipment startup and shutdown in the industry. Other industries also use the same method to calculate the corresponding power consumption attenuation efficiency of equipment startup and shutdown based on the startup and shutdown periods and operating data of their respective equipment, quantifying the impact of equipment startup and shutdown processes on power consumption.
[0104] Step S234: Based on the power consumption attenuation efficiency of equipment start-up and shutdown for each industry, conduct an evaluation and analysis of the impact characteristics of equipment start-up and shutdown on power load for each industry, so as to evaluate and analyze the distribution characteristics of the impact of equipment start-up and shutdown on power load and obtain the power consumption characteristics of equipment start-up and shutdown for each industry.
[0105] In this embodiment of the invention, when evaluating and analyzing the impact characteristics of equipment start-up and shutdown on power load data based on the power attenuation efficiency of equipment start-up and shutdown, the power load curves of various industries are combined. Taking a food processing industry as an example, its power load curve shows a significant upward trend in load between 8:00 and 9:00 every day. By comparing the equipment start-up and shutdown time records, it is found that a large number of food processing equipment, such as dough mixers and baking ovens, are started up during this period. According to the previously calculated power attenuation efficiency of equipment start-up and shutdown in this industry, the power attenuation efficiency during the start-up process is 0.3, that is, the power during equipment start-up is on average 30% higher than that during stable operation. Analyzing the increase in power load during this period, assuming that the power load at 8:00 is 50kW and the power load at 9:00 reaches 70kW, an increase of 20kW, and combining the equipment start-up situation, the increase in power load caused by equipment start-up is estimated to be about 20×0.3=6kW, accounting for 30% of the total increase, thereby determining the influence weight of equipment start-up on the change in power load during this period. Similarly, the impact of equipment shutdown periods on electricity load is analyzed. For example, between 5 PM and 6 PM, equipment gradually stops operating. Based on the power consumption attenuation efficiency of 0.7 during the shutdown process, the reduction in electricity load caused by equipment shutdown is estimated. This analysis is performed on the electricity load data of this industry at different times of the day, and the distribution characteristic curve of the impact of equipment start-up and shutdown on electricity load is plotted. This clearly shows the degree and trend of the impact of equipment start-up and shutdown on electricity load at different times. The same method is used to evaluate and analyze other industries in the park. Finally, the electricity consumption characteristics of equipment start-up and shutdown for each industry are obtained, including detailed information such as the main time periods, the degree of impact, and the trend of impact of equipment start-up and shutdown on electricity load. This provides an important basis for the classification of industrial electricity consumption data based on association clustering analysis.
[0106] Furthermore, step S24 includes the following steps:
[0107] Step S241: Obtain the production stage cycle corresponding to each industry;
[0108] In this embodiment of the invention, when obtaining the production stage cycle corresponding to each industry, key information is extracted from the industry ERP system and equipment operation logs. Taking the automobile manufacturing industry in a power grid as an example, the production plan in the ERP system clearly marks the start and end times of production stages such as stamping, welding, painting, and final assembly. For example, in a certain production task, the stamping stage starts at 8:00 on March 1 and ends at 17:00 on March 3; the welding stage starts at 9:00 on March 4 and ends at 18:00 on March 6. At the same time, the equipment operation log records in detail the switching time of the operating status of each production equipment in different stages, further verifying the accuracy of the production stage cycle. For the food processing industry, the production stages include raw material pretreatment, processing and production, packaging, etc. By reviewing production work orders and equipment operation records, it can be seen that in the production of a certain batch of food, the raw material pretreatment stage lasted from 7:00 AM to 9:00 AM, the processing stage started at 9:30 AM and ended at 3:00 PM, and the packaging stage lasted from 3:30 PM to 5:00 PM. For all industries in the industrial park, the corresponding production stage cycles were extracted and organized from the ERP system and equipment operation logs in this way, providing basic data support for subsequent analysis.
[0109] Step S242: Divide the electricity consumption time series data of each industry into production cycle electricity consumption based on the production stage cycle of each industry, so as to obtain the electricity consumption time series data of each industry under different production cycles;
[0110] In this embodiment of the invention, the electricity consumption of the production cycle is divided based on the production stage cycle of the electricity consumption time series data. Taking the electronic assembly industry as an example, its smart meters collect electricity consumption data every 15 minutes to form an electricity consumption time series. If the production stage cycle of a certain production task is: the component placement stage from 8:00 to 12:00 and the assembly and testing stage from 13:00 to 17:00, then all the electricity consumption data from 8:00 to 12:00 is selected from the electricity consumption time series data as the electricity consumption time series data of the component placement stage; the data from 13:00 to 17:00 is selected as the electricity consumption time series data of the assembly and testing stage. For the textile industry, a production cycle includes stages such as spinning, weaving, and dyeing. Assuming the spinning stage occurs from 9:00 to 13:00 on a certain day, the weaving stage from 14:00 to 17:00, and the dyeing stage from 8:00 to 12:00 the next day, the electricity consumption data for the corresponding time periods is extracted according to the time intervals and divided into the corresponding production cycle stages. In this way, the electricity consumption time series data of each industry can be accurately divided, and the electricity consumption time series data corresponding to each industry under different production cycles can be obtained, showing that the electricity consumption data is closely related to the production stage.
[0111] Step S243: Perform a gradient analysis of the electricity consumption distribution between different production cycles for the electricity consumption time series data of each industry under different production cycles, so as to obtain the electricity consumption distribution gradient between different production cycles for each industry.
[0112] In this embodiment of the invention, by performing a gradient analysis of electricity consumption distribution data for various industries under different production cycles, taking the furniture manufacturing industry as an example, its production cycle includes three stages: wood processing, coating, and assembly. The average value of electricity consumption data for the wood processing stage is calculated. Assuming that 16 15-minute electricity consumption data points are collected in this stage, with values of 20 kW·h, 22 kW·h, etc., the calculated average value is 23 kW·h. The average value of 12 data points for the coating stage is 18 kW·h. The average value of 10 data points for the assembly stage is 15 kW·h. The gradient of electricity consumption distribution between adjacent production cycles is calculated. The gradient of electricity consumption from the wood processing stage to the coating stage is (23-18) / 23≈0.22, indicating that the average electricity consumption in the coating stage is about 22% lower than that in the wood processing stage. The gradient of electricity consumption from the coating stage to the assembly stage is (18-15) / 18≈0.17. Through this calculation method, the electricity consumption data between different production cycles of various industries are compared and analyzed to obtain the gradient of electricity consumption distribution between different production cycles of various industries, intuitively presenting the changing trend of electricity consumption during the production process.
[0113] Step S244: Estimate the electricity consumption fluctuation of each industry between different production cycles to obtain the electricity consumption fluctuation coefficient of each industry for the production cycle.
[0114] In this embodiment of the invention, the electricity consumption fluctuation is estimated by analyzing the electricity consumption distribution gradient between different production cycles of various industries. Taking the chemical industry as an example, its production cycle is divided into three stages: raw material reaction, product separation, and purification. The electricity consumption distribution gradient is 0.3 from the raw material reaction to the product separation stage and 0.2 from the product separation to the purification stage. When calculating the estimated value of electricity consumption fluctuation, the absolute values of the electricity consumption gradient of each stage are added together, i.e., 0.3 + 0.2 = 0.5. Considering the difference in the duration of the production cycle, the estimated value is corrected. Assuming that the duration of the raw material reaction stage is 6 hours, the product separation stage is 4 hours, and the purification stage is 5 hours, the total duration is 15 hours. Dividing the estimated value of electricity consumption fluctuation by the total duration, the electricity consumption fluctuation coefficient of the production cycle is obtained as 0.5 / 15 ≈ 0.033. According to this method, the electricity consumption distribution gradient between different production cycles of various industries is comprehensively calculated, taking into full account the gradient magnitude and cycle duration, and finally the electricity consumption fluctuation coefficient of the corresponding production cycle of each industry is obtained, quantifying the degree of electricity consumption fluctuation within the production cycle.
[0115] Step S245: Based on the electricity consumption fluctuation coefficient of the production cycle corresponding to each industry, perform production cycle fluctuation characteristic analysis on the corresponding electricity consumption time series data to obtain the electricity consumption characteristics of the production cycle corresponding to each industry.
[0116] In this embodiment of the invention, the fluctuation characteristics of the production cycle are analyzed by using the electricity consumption fluctuation coefficient based on the electricity consumption time series data. Taking the garment manufacturing industry as an example, its production cycle electricity consumption fluctuation coefficient is 0.04, indicating that the electricity consumption fluctuation during the production process is relatively small. Further analysis of its electricity consumption time series data shows that during the cutting stage, electricity consumption is relatively stable, fluctuating between 18-22 kW·h; during the sewing stage, electricity consumption increases slightly, but the fluctuation is not large, remaining at 20-25 kW·h; during the ironing and packaging stage, electricity consumption decreases, fluctuating between 15-18 kW·h. By combining the trends of electricity consumption fluctuation coefficient and electricity consumption time series data, the electricity consumption characteristics of the industry's production cycle are summarized: overall electricity consumption fluctuation is stable, and the changes in electricity consumption at each production stage are relatively mild, without significant fluctuations. For other industries, such as the machinery processing industry, if its production cycle electricity consumption fluctuation coefficient is 0.08, the analysis shows that its electricity consumption is higher and fluctuates more in the rough processing stage, while electricity consumption decreases in the fine processing stage but still fluctuates to some extent. Through detailed analysis of each industry, the corresponding production cycle electricity consumption characteristics of each industry are finally obtained, providing an important basis for the classification of industrial electricity consumption data and electricity management.
[0117] Furthermore, step S3 includes the following steps:
[0118] Step S31: Obtain the spatial distance between the electricity consumption feature vectors of each industry through the industry electricity consumption feature vectors corresponding to each industry;
[0119] In this embodiment of the invention, when obtaining spatial distance through the electricity consumption characteristic vectors corresponding to each industry, the Euclidean distance formula is used for calculation. Assuming that the industrial park contains an electronics manufacturing industry and an auto parts manufacturing industry, the electricity consumption characteristic vector of the electronics manufacturing industry is [0.25, 60kW, 5h, 0.55, 0.12, 0.33, 18kW, 12kW, 20kW·h, 3kW·h, 0.22], and the electricity consumption characteristic vector of the auto parts manufacturing industry is [0.35, 75kW, 6h, 0.6, 0.1, 0.3, 22kW, 15kW, 25kW·h, 4kW·h, 0.25]. For these two vectors, the values of their corresponding dimensions are substituted into the Euclidean distance formula: ,in For vector dimensions, and These are the two vectors, respectively. The values of each dimension are calculated. For example, the difference in the first dimension is 0.25 - 0.35 = -0.1, which is squared to get 0.01; the difference in the second dimension is 60 - 75 = -15, which is squared to get 225, and so on. The sum of the squares of the differences in all dimensions is then taken as the square root to obtain the spatial distance between the two industry electricity consumption feature vectors. The above calculation is performed on the electricity consumption feature vectors of all industries in the park pairwise to construct a complete distance matrix. For example, if there are 10 industries in the park, a 10×10 matrix is formed. The value in the i-th row and j-th column of the matrix is the spatial distance between the electricity consumption feature vectors of the i-th industry and the j-th industry, thus clearly showing the spatial distance relationship between the electricity consumption feature vectors of each industry.
[0120] Step S32: Based on the spatial distance between the corresponding industrial electricity consumption feature vectors, perform spatial distribution analysis on the corresponding industrial electricity consumption feature vectors to obtain the vector spatial distribution between the corresponding industrial electricity consumption feature vectors;
[0121] In this embodiment of the invention, when performing spatial distribution analysis on industrial electricity consumption feature vectors based on spatial distance, a two-dimensional scatter plot is used for visualization (if the vector dimension exceeds two dimensions, principal component analysis is used for dimensionality reduction before visualization). Each industrial electricity consumption feature vector is considered as a point in space. Taking the electronics manufacturing industry and the machinery processing industry as examples, the corresponding positions are determined in the graph based on the spatial distance between their electricity consumption feature vectors. If the coordinates of the point in the electronics manufacturing industry are (x1, y1) and the coordinates of the point in the machinery processing industry are (x2, y2), the distance between the two points is proportional to the previously calculated spatial distance. In a scatter plot, observing the distribution of various industry points reveals that if multiple industry points are clustered in a certain area, such as food processing and beverage production, where light industry points are concentrated in the upper left corner of the plot, it indicates that the electricity consumption characteristic vectors of these industries are spatially close and have a certain degree of similarity. On the other hand, heavy industry points such as steel manufacturing and chemicals are distributed at the other end, which are far away from the light industry points, indicating that their electricity consumption characteristics are quite different. By analyzing the entire scatter plot, the dense and sparse areas of industry points, as well as the relative positional relationships between different industry clusters, are marked. Finally, the vector space distribution of the electricity consumption characteristic vectors of each industry is obtained, intuitively showing the spatial distribution pattern of industry electricity consumption characteristics.
[0122] Step S33: Calculate the spatial kernel density of the vector space distribution corresponding to the electricity consumption feature vectors of each industry to obtain the magnitude of the spatial kernel density corresponding to the electricity consumption feature vectors of each industry;
[0123] In this embodiment of the invention, when calculating the spatial kernel density of the vector space distribution, a kernel density estimation method is used. Taking the garment manufacturing industry as an example, a circular region with a radius of r is defined centered on its vector space location (the value of r is determined based on the overall distribution of spatial distances, such as taking 1 / 5 of the average of all spatial distances). The number of other industry electricity consumption feature vectors falling into this circular region is counted. Assuming that a total of 8 industry points fall into this region, the kernel density calculation formula is used: ,in The number of industrial sites located within the region. This is the bandwidth (which can be adjusted according to the actual situation, and is generally related to the radius r). The kernel function is used (the commonly used Gaussian kernel function is adopted). Let be the point whose density is to be calculated (i.e., the vector position of the garment manufacturing industry). To determine the vector positions of each industry point within the region, relevant values are substituted into the formula to calculate the spatial kernel density of the garment manufacturing industry. This calculation is performed on the locations of the electricity consumption feature vectors of all industries within the park, such as calculating the kernel density values of the corresponding locations of the electronics manufacturing industry and the automobile manufacturing industry. Finally, the spatial kernel density between the electricity consumption feature vectors of each industry is obtained, quantifying the density of each industry's distribution in the vector space.
[0124] Step S34: Based on the spatial kernel density between the electricity consumption feature vectors of each industry, perform spatial local distribution estimation between the corresponding industry electricity consumption feature vectors to obtain the local distribution density between the electricity consumption feature vectors of each industry.
[0125] In this embodiment of the invention, when estimating the local spatial distribution based on the spatial kernel density, a smaller local region (such as a circular region with a radius half the radius of the aforementioned circular region) is set as the center of each industry's electricity consumption feature vector. Taking the furniture manufacturing industry as an example, after setting a local region around it, the number of other industry electricity consumption feature vectors and their spatial kernel density values are counted within this region. Assuming there are 5 industry points within this region, with spatial kernel density values of 0.2, 0.3, 0.25, 0.35, and 0.28 respectively, the weighted average of the spatial kernel density values of these industry points is calculated (weighted average). The weight is determined based on the spatial distance between the industrial points and the furniture manufacturing industry. The closer the distance, the higher the weight (e.g., the weight of an industrial point with a distance of d1 is 1 / d1). The kernel density value of each industrial point is multiplied by its corresponding weight, summed, and then divided by the total weight to obtain the average spatial kernel density value of the local area. This value is the local distribution density of the furniture manufacturing industry. This calculation is performed on all industries in the park, such as calculating the local distribution density of the electronic information industry and the biomedical industry. Finally, the local distribution density corresponding to the electricity consumption characteristic vector of each industry is obtained, which accurately reflects the distribution density of each industry within its local spatial range.
[0126] Step S35: Based on the spatial distance and local distribution density between the electricity consumption feature vectors of each industry, perform density peak clustering on the electricity consumption data of each industry to obtain the electricity consumption data clusters of each industry.
[0127] In this embodiment of the invention, when performing density peak clustering based on spatial distance and local distribution density, cluster centers are first determined. Industrial electricity consumption feature vectors with relatively high local distribution density and large spatial distances from other high-density points are selected as initial cluster centers. For example, if the local distribution density of a new materials industry is 0.4, and there are no other industrial points with higher local distribution density within a certain range around it (e.g., the spatial distance is greater than 1.5 times the average of all spatial distances), then the electricity consumption feature vector of the new materials industry is determined as a cluster center. Taking the plastic products industry as an example, the density peak reachability with each cluster center is calculated. Density peak reachability comprehensively considers spatial distance and local distribution density differences. If the spatial distance between the plastic products industry and a certain cluster center (e.g., the electricity consumption feature vector of the metal processing industry) is 0.12, the local distribution density of the plastic products industry is 0.3, and the local distribution density of the metal processing industry is... With a density of 0.35, the density peak reachability value of the plastic products industry to the cluster center is calculated using specific calculation rules (such as weighted summation of the reciprocal of spatial distance and the difference in local distribution density, with weights set according to actual conditions, such as a weight of 0.6 for the reciprocal of spatial distance and a weight of 0.4 for the difference in local distribution density). This calculation is performed on the plastic products industry and all cluster centers to find the cluster center with the highest density peak reachability. The electricity consumption data of the plastic products industry is then assigned to the electricity density peak cluster of that cluster center. The above operation is performed on the electricity consumption data of all industries in the park. Based on their respective density peak reachability with the cluster center, the industry electricity consumption data is accurately assigned to different electricity density peak clusters, ultimately resulting in electricity consumption data clusters for each industry. This achieves industry electricity consumption data classification based on association cluster analysis, making the characteristics of industry electricity consumption data similar within the same cluster and showing significant differences in characteristics between different clusters.
[0128] Furthermore, step S35 includes the following steps:
[0129] The corresponding industrial electricity consumption cluster centers are determined based on the spatial distance and local distribution density between the characteristic vectors of electricity consumption of each industry.
[0130] In this embodiment of the invention, industrial electricity consumption cluster centers are determined based on the spatial distance and local distribution density between the characteristic vectors of each industry's electricity consumption. When calculating the spatial distance, the Euclidean distance formula is used, based on two industrial electricity consumption characteristic vectors. For example, within the power grid area, the electricity consumption characteristic vector of an electronics manufacturing industry is [0.3, 70kW, 6h, 0.6, 0.1, 0.3, 20kW, 15kW, 22kW·h, 4kW·h, 0.25], and the electricity consumption characteristic vector of a machinery processing industry is [0.25, 80kW, 7h, 0.5, 0.15, 0.35, 25kW, 18kW, 25kW·h, 5kW·h, 0.2]. The corresponding dimensions are substituted into the formula to calculate the spatial distance between the two, such as calculating the squared difference of the first dimension, the squared difference of the second dimension, and so on. The spatial distance between the two industry electricity consumption feature vectors is obtained by summing and square rooting. The spatial distance between each pair of industry electricity consumption feature vectors within the park is then calculated to construct a distance matrix. When calculating the local distribution density, a distance threshold is set, such as 0.2 (this threshold is determined based on statistical analysis of all spatial distances). Taking a food processing industry as an example, the number of other industries with a spatial distance less than 0.2 is counted. Assuming there are 5 industries meeting this condition, the local distribution density of the food processing industry is 5. This calculation is performed on all industries to obtain their local distribution densities. Cluster centers are determined by combining spatial distance and local distribution density. Industry electricity consumption feature vectors with relatively high local distribution densities and relatively large spatial distances from other high-density points are selected as cluster centers. For example, if a textile industry has a local distribution density of 8, and there are no other industries with higher local distribution densities and closer spatial distances within a certain range around it, then the electricity consumption feature vector of the textile industry can be determined as an industry electricity consumption cluster center. This process is repeated to determine all cluster centers.
[0131] Preferably, the electricity consumption data distribution density of each industry is statistically analyzed to obtain the electricity consumption data distribution density of each industry.
[0132] In this embodiment of the invention, when statistically analyzing the electricity consumption data distribution density of various industries, the entire industrial electricity consumption data space is divided into several small regional units based on the industrial electricity consumption feature vector. For example, the frequency range of electricity load distribution is divided into intervals such as 0-0.1 and 0.1-0.2, and the amplitude range of electricity load distribution is divided into intervals such as 0-20kW and 20-40kW, and so on, forming a multi-dimensional regional division. Taking a certain chemical industry as an example, the number of times its electricity consumption feature vector falls into each regional unit is counted. Assuming that the electricity load distribution frequency is 0.1-0.2, the electricity load distribution amplitude is 30-40kW, and the peak-hour electricity consumption ratio is 0.4-0.5... Within a specific regional unit, the electricity consumption characteristic vector of this chemical industry appeared 3 times. Throughout the entire statistical period, the chemical industry recorded a total of 10 electricity consumption characteristic vector data points. Therefore, the electricity consumption data distribution density corresponding to this regional unit is 3 ÷ 10 = 0.3. This calculation is performed on all regional units of the chemical industry to obtain its complete electricity consumption data distribution density. This method is applied to all industries within the park, statistically analyzing the frequency of electricity consumption characteristic vector occurrences within each divided regional unit to calculate the electricity consumption data distribution density for each industry, comprehensively reflecting the density of electricity consumption data distribution across different dimensions for each industry.
[0133] Preferably, based on the distribution density of electricity consumption data corresponding to each industry and combined with the industry electricity consumption cluster center, density peak clustering is performed on the industry electricity consumption data corresponding to each industry. Based on the density peak reachability between each electricity consumption data distribution density and the industry electricity consumption cluster center, the corresponding industry electricity consumption data is divided into different electricity density peak clusters to obtain each industry electricity consumption data cluster.
[0134] In this embodiment of the invention, density peak clustering is performed based on the distribution density of electricity consumption data corresponding to each industry and combined with industry electricity consumption cluster centers. Taking the electricity consumption data of a furniture manufacturing industry as an example, the density peak accessibility between the industry's electricity consumption data and each industry electricity consumption cluster center is first calculated. The calculation of density peak accessibility considers two factors: first, the spatial distance between the industry and the cluster center; the closer the distance, the higher the accessibility; second, the difference in electricity consumption data distribution density between the two. If the distribution density of the industry's electricity consumption data is higher than the density of the area where the cluster center is located, the accessibility is also correspondingly improved. Assuming that the spatial distance between a furniture manufacturing industry and a cluster center (the electricity consumption feature vector of a metal products industry) is 0.15, and the distribution density of the furniture manufacturing industry's electricity consumption data in the relevant area is 0.4, while the density of the area where the cluster center is located is 0.3, through specific calculation rules (such as incorporating distance and density differences into the data distribution density), the accessibility is further improved. (Through weighted combination calculations), the peak density accessibility value of the furniture manufacturing industry to this cluster center is obtained. This calculation is performed on the furniture manufacturing industry and all cluster centers to find the cluster center with the highest peak density accessibility. The electricity consumption data of the furniture manufacturing industry is then assigned to the electricity density peak cluster corresponding to the cluster center with the highest peak density accessibility. For example, if the calculation finds that the furniture manufacturing industry has the highest peak density accessibility to a certain cluster center, it is assigned to the cluster containing that cluster center. This operation is performed on the electricity consumption data of all industries in the park. Based on their respective peak density accessibility with the cluster centers, the industry electricity consumption data is assigned to different electricity density peak clusters, ultimately obtaining the electricity consumption data clusters of each industry. This achieves effective classification of industry electricity consumption data, making the industry electricity consumption data within the same cluster highly similar, and the industry electricity consumption data between different clusters significantly different.
[0135] Furthermore, step S4 includes the following steps:
[0136] Step S41: Discretize the industrial electricity consumption data within each industrial electricity consumption data cluster to obtain each industrial electricity consumption discretization cluster;
[0137] In this embodiment of the invention, when discretizing the industrial electricity consumption data within each industrial electricity consumption data cluster, taking electricity load data as an example, the equal-width binning method is adopted. Assuming a certain industrial electricity consumption data cluster (high-energy-consuming heavy industry category) contains electricity consumption data from 100 industries, with electricity load values ranging from 20kW to 200kW, this range is divided into 5 equal-width intervals, each with a width of (200-20)÷5=36kW. That is, the intervals are 20kW-56kW, 56kW-92kW, 92kW-128kW, 128kW-164kW, and 164kW-200kW. For the electricity load data of each industry within the cluster, its corresponding interval is determined. For example, for a certain steel manufacturing... The industrial electricity load is 135kW, which is categorized into the 128kW-164kW range and marked as a discrete value within this range. Similarly, the power consumption characteristics of equipment start-up and shutdown within this cluster are discretized. Taking the increase in power load when equipment starts up as an example, if the range is 5kW-50kW, it is divided into 4 intervals, each with a width of (50-5)÷4=11.25kW. When a forging equipment starts up, the power load increases by 23kW, so it is categorized into the 16.25kW-27.5kW range. The above discretization operation is performed on all industrial electricity consumption data clusters in the park, transforming continuous electricity consumption data into discrete interval values, and finally obtaining the discrete clusters of electricity consumption for each industry, making the data easier to perform correlation analysis.
[0138] Step S42: Perform association clustering mining on each industrial electricity consumption discretization cluster to obtain an industrial electricity consumption association clustering dataset;
[0139] In this embodiment of the invention, when mining association clusters among various industrial electricity consumption discretization clusters, the basic principle of the Apriori algorithm is adopted. Taking two industrial electricity consumption discretization clusters (a high-energy-consuming heavy industry discretization cluster and a light industry discretization cluster) as an example, a power load discretization range (e.g., 128kW-164kW) is selected from the high-energy-consuming heavy industry discretization cluster and combined with a device start-stop frequency discretization range (e.g., 5-10 times per day) from the light industry discretization cluster to form a candidate set of association clusters: "Power load between 128kW-164kW and device start-stop 5-10 times per day". All industrial electricity consumption data are traversed, and the frequency of this association cluster is counted. Assuming that this association cluster appears 80 times in 1000 industrial electricity consumption data records, other candidate sets of association clusters with different combinations are generated, and the same statistical operation is performed. By continuously generating candidate sets and counting the occurrences, all possible associated clusters are mined out, and finally, an industrial electricity consumption associated cluster dataset is obtained. This dataset contains associated clusters with various combinations of features and their frequency of occurrence.
[0140] Step S43: Obtain the support and confidence of each associated cluster item through the industrial electricity consumption associated cluster item dataset;
[0141] In this embodiment of the invention, the support and confidence of each associated cluster item are obtained through the industrial electricity consumption associated cluster item dataset. The support calculation formula is: support = number of times associated cluster item appears ÷ total number of data records. Taking the associated cluster item "electricity load is between 128kW and 164kW and equipment starts and stops 5-10 times a day" as an example in the previous step, it appears 80 times and the total number of data records is 1000. Then the support of this associated cluster item is 80 ÷ 1000 = 0.08, that is, 8%, which means that the proportion of this associated cluster item in all industrial electricity consumption data is 8%. The confidence calculation is based on conditional probability. Taking the clustered item "If the power load is between 128kW and 164kW, the equipment starts and stops 5-10 times a day" as an example, we first count the number of records with power load between 128kW and 164kW, assuming there are 200 records. Among them, the number of records with equipment starting and stopping 5-10 times a day is 80. The confidence score is calculated as follows: Confidence score = (Number of records with power load between 128kW and 164kW and equipment starting and stopping 5-10 times a day) ÷ (Number of records with power load between 128kW and 164kW), which is 80 ÷ 200 = 0.4, or 40%. This means that when the industrial power load is between 128kW and 164kW, the probability that the equipment starts and stops 5-10 times a day is 40%. We perform this calculation on all clustered items in the industrial power load clustered item dataset to obtain the support and confidence scores for each clustered item.
[0142] Step S44: Based on the support and confidence of each associated cluster item, classify and filter the associated cluster items in the industrial electricity consumption associated cluster item dataset to obtain the industrial electricity consumption associated cluster item class set;
[0143] In this embodiment of the invention, the industrial electricity consumption associated cluster dataset is classified and filtered based on the support and confidence of each associated cluster. The support threshold is set to 0.05 and the confidence threshold is set to 0.3. Taking the associated cluster "electricity load between 128kW and 164kW and equipment starts and stops 5-10 times per day" as an example, its support is 0.08, which is greater than 0.05, and its confidence is 0.4, which is greater than 0.3, meeting the threshold conditions, so it is retained. For the associated cluster "electricity load between 20kW and 56kW and equipment failures 0-2 times per month", if its support is 0.03, which is less than 0.05, even if the confidence reaches 0.35, it is removed from the dataset. By judging whether the support and confidence of each associated cluster meet the threshold requirements one by one, the associated clusters that meet the conditions are filtered out. These associated clusters are organized into an industrial electricity consumption associated cluster category. The associated clusters in this category have a high frequency of occurrence and strong association reliability.
[0144] Step S45: Classify the industrial electricity consumption data corresponding to each industry based on the industrial electricity consumption association cluster category set to obtain the industrial electricity consumption data classification results.
[0145] In this embodiment of the invention, industrial electricity consumption data corresponding to each industry is classified based on an industrial electricity consumption association cluster set. Taking a certain emerging chemical industry as an example, its electricity consumption data is an electricity load of 140kW and equipment start-up and shutdown 8 times per day. The electricity consumption data of this industry is matched with the association clusters in the industrial electricity consumption association cluster set. It is found that it matches the association cluster in the set that is "electricity load between 128kW and 164kW and equipment start-up and shutdown 5-10 times per day". According to the category to which the association cluster belongs (assuming that the association cluster belongs to...), The electricity consumption data of the chemical industry is classified into the category of "high energy consumption and frequent equipment start-up and shutdown". This matching and classification operation is performed on the electricity consumption data of all industries in the park. For example, a food and beverage industry with an electricity load of 45kW and equipment start-up and shutdown 3 times a day matches the associated cluster item "electricity load between 20kW and 56kW and equipment start-up and shutdown 1-5 times a day" in the class set. It is classified into the category of "low energy consumption and fewer equipment start-up and shutdown". In this way, the complete classification results of industrial electricity consumption data are finally obtained, so as to achieve effective classification management of industrial electricity consumption data.
[0146] Furthermore, step S45 includes the following steps:
[0147] Step S451: Based on the industrial electricity consumption association cluster set, perform industrial electricity consumption association classification on the industrial electricity consumption data corresponding to each industry to obtain the initial classification results of industrial electricity consumption;
[0148] In this embodiment of the invention, when classifying industrial electricity consumption data based on the industrial electricity consumption association cluster set, the industrial electricity consumption feature vector is used as the basis. It is assumed that the industrial electricity consumption association cluster set has identified three peak electricity density clusters, namely cluster A (high energy-consuming heavy industry), cluster B (light industry), and cluster C (electronic information industry). Each cluster contains typical feature descriptions of the corresponding industrial electricity consumption feature vector. Taking a newly established precision instrument manufacturing industry as an example, its industrial electricity consumption feature vector is [0.3, 70kW, 6h, 0.6, 0.1, 0.3, 20kW, 15kW, 23kW·h, 3.5kW·h, 0.23]. The vector was compared one by one with the typical features of clusters A, B, and C. The comparison revealed that the vector's power load distribution amplitude and equipment start-up and shutdown power consumption characteristics were more similar to those of high-energy-consuming heavy industries in cluster A, such as a larger power load distribution amplitude and a significant power load impact during equipment startup. Through this feature comparison method, the power consumption data of all industries in the park were analyzed and judged. For a certain food packaging industry, its power consumption feature vector showed a relatively uniform power load distribution frequency and small differences in power consumption during peak and valley periods, which matched the characteristics of light industry in cluster B. Therefore, it was classified into cluster B. After evaluating and classifying the power consumption data of all industries, the initial classification results of industrial power consumption were finally obtained, so that the power consumption data of each industry was assigned to the peak power density cluster that best matched its characteristics, thus completing the initial classification.
[0149] Step S452: Dynamically update and monitor the initial classification results of industrial electricity consumption to collect corresponding new electricity consumption data characteristics in real time;
[0150] In this embodiment of the invention, when dynamically updating and monitoring the initial classification results of industrial electricity consumption, electricity consumption data is continuously collected through smart meters and equipment sensors. Taking a certain mechanical processing industry as an example, the smart meters collect power load data such as active power and reactive power every 15 minutes, and the sensors on the equipment monitor the equipment's start-up and shutdown status, operating speed, and other equipment operation data in real time, and record the corresponding timestamps simultaneously. When the industry's production situation changes, such as when the mechanical processing industry adds an automated production line and the production plan increases from 100 products per day to 150 products, the electricity consumption data will change significantly. These changes are captured in real time. Under the new production conditions, the power load of the industry increases from 20kW to 30kW when the equipment starts up, and the average power load during the production process rises from 50kW to 70kW. These newly generated power consumption data characteristics, such as the magnitude of power load changes and changes in power consumption characteristics during equipment start-up and shutdown, are extracted and organized to form new power consumption data characteristics. This real-time collection and monitoring method is used for all industries in the park. Once a significant change in power consumption data is detected, the corresponding new power consumption data characteristics are obtained in a timely manner to provide a basis for subsequent classification and adjustment.
[0151] Step S453: Calculate the similarity between the new electricity consumption data features and the existing industrial electricity consumption data clusters. If the similarity is lower than a preset threshold, calculate the distance between the new electricity consumption data features and the corresponding industrial electricity consumption cluster centers to obtain the distance value between the new electricity consumption features and the cluster centers.
[0152] In this embodiment of the invention, when calculating the similarity between new electricity consumption data features and existing industrial electricity consumption data clusters, a cosine similarity algorithm is used. Taking the new electricity consumption data feature vector [0.32, 75kW, 6.5h, 0.62, 0.08, 0.32, 22kW, 16kW, 24kW·h, 4kW·h, 0.24] generated by a certain chemical industry as an example, the cosine similarity is calculated between it and the typical feature vectors of three existing industrial electricity consumption data clusters (cluster A, cluster B, and cluster C). The cosine similarity formula is as follows: ,in This is the feature vector of new electricity consumption data. Given a typical feature vector of a certain industry's electricity consumption data cluster, the cosine similarity between this new feature vector and cluster A is calculated to be 0.75, with cluster B 0.4, and with cluster C 0.3. Assuming a preset threshold of 0.6, since the similarity between the new electricity consumption data feature of the chemical industry and clusters B and C is less than 0.6, the distance between it and the corresponding cluster centers of clusters B and C is further calculated. Using the Euclidean distance formula, the distance between the new electricity consumption feature vector of the chemical industry and the cluster center of cluster B is calculated to be 0.15, and the distance between it and the cluster center of cluster C is 0.2. In this way, similarity and distance calculations are performed for all industries that generate new electricity consumption data features, clarifying the relationship between the new electricity consumption data features and existing industry electricity consumption data clusters and cluster centers.
[0153] Step S454: Determine the corresponding temporary electricity consumption cluster based on the distance between the new electricity consumption characteristics and the cluster center, and reclassify and adjust the initial classification results of industrial electricity consumption based on the temporary electricity consumption cluster to obtain the classification results of industrial electricity consumption data.
[0154] In this embodiment of the invention, when determining the temporary electricity consumption cluster based on the distance between the new electricity consumption characteristics and the cluster center, the cluster corresponding to the cluster center with the smallest distance value is selected as the temporary cluster. Taking the chemical industry in the previous step as an example, since the distance value of its new electricity consumption characteristics to the cluster center of cluster B (0.15) is less than the distance value to the cluster center of cluster C (0.2), the new electricity consumption data characteristics of the chemical industry are temporarily assigned to cluster B. Based on the temporary electricity consumption cluster, the initial classification results of the industry's electricity consumption are reclassified and adjusted. In the initial classification results of the industry's electricity consumption, the chemical industry originally belonged to cluster A. Now, according to the temporary assignment, it is removed from cluster A and assigned to cluster B. All industries in the park that generate new electricity consumption data characteristics and whose similarity is lower than the threshold are adjusted in this way. For example, when a new energy battery manufacturing industry generates new electricity consumption data characteristics, its similarity to existing clusters and its distance from the cluster center are calculated, and it is determined that it temporarily belongs to cluster C. The industry is then adjusted from other clusters in the initial classification to cluster C. Through such reclassification and adjustment, the final classification result of industrial electricity consumption data is obtained, ensuring that the classification of industrial electricity consumption data can reflect changes in industrial electricity consumption characteristics in a timely manner, and improving the accuracy and timeliness of classification.
[0155] Furthermore, the present invention also provides an industrial electricity consumption data classification system based on association clustering analysis, used to execute the industrial electricity consumption data classification method based on association clustering analysis as described above. The industrial electricity consumption data classification system based on association clustering analysis includes:
[0156] The industrial electricity consumption filling module is used to obtain the industrial electricity consumption data corresponding to each industry in the industrial park, and to perform missing interpolation filling based on the spatiotemporal distribution of the industrial electricity consumption data corresponding to each industry to obtain the industrial electricity consumption interpolation filling data corresponding to each industry.
[0157] The electricity consumption feature vector analysis module is used to perform electricity consumption feature vector analysis on the interpolated data of electricity consumption for each industry, so as to obtain the industry electricity consumption feature vector corresponding to the electricity load distribution characteristics, peak-valley-normal period electricity consumption ratio characteristics, equipment start-up and shutdown electricity consumption characteristics, and production cycle electricity consumption characteristics of each industry.
[0158] The density peak clustering module is used to obtain the spatial distance and local distribution density between the electricity consumption feature vectors of each industry through the electricity consumption feature vectors of each industry, and to perform density peak clustering on the electricity consumption data of each industry based on the spatial distance and local distribution density between the electricity consumption feature vectors of each industry to obtain the electricity consumption data clusters of each industry.
[0159] The association-aggregated electricity consumption classification module is used to perform association-aggregated mining analysis on electricity consumption data clusters of various industries to obtain industry electricity consumption association-aggregated item categories; based on the industry electricity consumption association-aggregated item categories, the corresponding industry electricity consumption data of each industry is classified to obtain the industry electricity consumption data classification results.
[0160] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.
[0161] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A method for classifying industrial electricity consumption data based on association clustering analysis, characterized in that, Includes the following steps: Step S1: Obtain the electricity consumption data of each industry in the industrial park, and perform missing interpolation based on the spatiotemporal distribution of the electricity consumption data of each industry to obtain the interpolated electricity consumption data of each industry. Step S2: Perform electricity consumption feature vector analysis on the interpolated data of electricity consumption for each industry to obtain the electricity consumption feature vector of each industry; wherein, the electricity consumption feature vector includes the characteristics of electricity load distribution, the proportion of electricity consumption during peak-valley-normal periods, the characteristics of electricity consumption during equipment start-up and shutdown, and the characteristics of electricity consumption during the production cycle. Step S3: Obtain the spatial distance and local distribution density between the electricity consumption feature vectors of each industry through the industry-specific electricity consumption feature vectors, and perform density peak clustering on the electricity consumption data of each industry based on the spatial distance and local distribution density between the industry-specific electricity consumption feature vectors to obtain the electricity consumption data clusters of each industry; wherein, step S3 includes the following steps: Step S31: Obtain the spatial distance between the electricity consumption feature vectors of each industry through the industry electricity consumption feature vectors corresponding to each industry; Step S32: Based on the spatial distance between the corresponding industrial electricity consumption feature vectors, perform spatial distribution analysis on the corresponding industrial electricity consumption feature vectors to obtain the vector spatial distribution between the corresponding industrial electricity consumption feature vectors; Step S33: Calculate the spatial kernel density of the vector space distribution corresponding to the electricity consumption feature vectors of each industry to obtain the magnitude of the spatial kernel density corresponding to the electricity consumption feature vectors of each industry; Step S34: Based on the spatial kernel density between the electricity consumption feature vectors of each industry, perform spatial local distribution estimation between the corresponding industry electricity consumption feature vectors to obtain the local distribution density between the electricity consumption feature vectors of each industry. Step S35: Based on the spatial distance and local distribution density between the electricity consumption feature vectors of each industry, perform density peak clustering on the electricity consumption data of each industry to obtain the electricity consumption data clusters of each industry; Step S4: Perform association and aggregation mining analysis on the electricity consumption data clusters of each industry to obtain the industry electricity consumption association aggregation item category set; classify the industry electricity consumption data corresponding to each industry based on the industry electricity consumption association aggregation item category set to obtain the industry electricity consumption data classification results.
2. The industrial electricity consumption data classification method based on association clustering analysis according to claim 1, characterized in that, Step S1 includes the following steps: Step S11: Obtain the electricity consumption data of each industry in the industrial park, including the electricity load data, electricity time series data and equipment operation data of each industry; Step S12: Perform anomaly identification and removal on the industrial electricity consumption data corresponding to each industry, calculate the moving average and standard deviation of the industrial electricity consumption data, determine the data fluctuation threshold based on the moving average and standard deviation, and remove abnormal data points that exceed the data fluctuation threshold based on the data fluctuation threshold to obtain the industrial electricity consumption anomaly removal data corresponding to each industry. Step S13: Perform spatiotemporal dimension distribution analysis on the electricity consumption anomaly removal data corresponding to each industry to obtain the spatiotemporal dimension distribution of electricity consumption data for each industry; Step S14: Based on the spatiotemporal dimension distribution of electricity consumption data for each industry and combined with the time series autoregressive model and spatial correlation, perform spatiotemporal distribution correlation prediction on the electricity consumption anomaly removal data for each industry to generate missing electricity consumption data for each industry. Step S15: Based on the abnormal electricity consumption data of each industry, perform missing interpolation filling on the corresponding missing electricity consumption data to obtain the interpolated electricity consumption data of each industry.
3. The industrial electricity consumption data classification method based on association clustering analysis according to claim 2, characterized in that, Step S2 includes the following steps: Step S21: Plot the electricity load curves for each industry to generate a sequence of electricity load curves for each industry; perform electricity load distribution statistics on the sequence of electricity load curves for each industry to obtain the electricity load distribution characteristics for each industry, including the electricity load distribution frequency, the electricity load distribution amplitude, and the electricity load distribution phase. Step S22: Obtain the corresponding peak electricity consumption period, valley electricity consumption period, and level electricity consumption period through the electricity consumption time series data corresponding to each industry, and calculate the electricity consumption ratio between the electricity consumption time series data in the corresponding period based on the peak electricity consumption period, valley electricity consumption period, and level electricity consumption period to obtain the peak-valley-normal period electricity consumption ratio characteristics of each industry. Step S23: Based on the equipment operation data corresponding to each industry, perform equipment start-up and shutdown characteristic analysis on the power load data corresponding to each industry to obtain the equipment start-up and shutdown power consumption characteristics corresponding to each industry; Step S24: Evaluate the production cycle fluctuations of the electricity consumption time series data corresponding to each industry to obtain the electricity consumption characteristics of the production cycle for each industry; Step S25: Combine the electricity load distribution characteristics, peak-valley-normal period electricity consumption ratio characteristics, equipment start-up and shutdown electricity consumption characteristics, and production cycle electricity consumption characteristics of each industry to construct the corresponding industry electricity consumption feature vector.
4. The industrial electricity consumption data classification method based on association clustering analysis according to claim 3, characterized in that, Step S23 includes the following steps: Step S231: Obtain the corresponding equipment start-up and shutdown periods through the equipment operation data of each industry; Step S232: Perform equipment operation frequency statistics on the equipment operation data corresponding to each industry to obtain the equipment operation frequency corresponding to each industry; Step S233: Based on the start-up and shutdown periods of equipment in each industry, evaluate the power consumption attenuation of equipment start-up and shutdown for each industry to obtain the power consumption attenuation efficiency of equipment start-up and shutdown for each industry. Step S234: Based on the power consumption attenuation efficiency of equipment start-up and shutdown for each industry, conduct an evaluation and analysis of the impact characteristics of equipment start-up and shutdown on power load for each industry, so as to evaluate and analyze the distribution characteristics of the impact of equipment start-up and shutdown on power load and obtain the power consumption characteristics of equipment start-up and shutdown for each industry.
5. The industrial electricity consumption data classification method based on association clustering analysis according to claim 3, characterized in that, Step S24 includes the following steps: Step S241: Obtain the production stage cycle corresponding to each industry; Step S242: Divide the electricity consumption time series data of each industry into production cycle electricity consumption based on the production stage cycle of each industry, so as to obtain the electricity consumption time series data of each industry under different production cycles; Step S243: Perform a gradient analysis of the electricity consumption distribution between different production cycles for the electricity consumption time series data of each industry under different production cycles, so as to obtain the electricity consumption distribution gradient between different production cycles for each industry. Step S244: Estimate the electricity consumption fluctuation of each industry between different production cycles to obtain the electricity consumption fluctuation coefficient of each industry for the production cycle. Step S245: Based on the electricity consumption fluctuation coefficient of the production cycle corresponding to each industry, perform production cycle fluctuation characteristic analysis on the corresponding electricity consumption time series data to obtain the electricity consumption characteristics of the production cycle corresponding to each industry.
6. The industrial electricity consumption data classification method based on association clustering analysis according to claim 1, characterized in that, Step S35 includes the following steps: The corresponding industrial electricity consumption cluster centers are determined based on the spatial distance and local distribution density between the characteristic vectors of electricity consumption of each industry. Electricity consumption data distribution density is statistically analyzed for each industry to obtain the electricity consumption data distribution density for each industry. Based on the distribution density of electricity consumption data corresponding to each industry and combined with the industry electricity consumption cluster center, the industry electricity consumption data corresponding to each industry is clustered by density peak. Based on the density peak reachability between each electricity consumption data distribution density and the industry electricity consumption cluster center, the corresponding industry electricity consumption data is divided into different electricity density peak clusters to obtain each industry electricity consumption data cluster.
7. The industrial electricity consumption data classification method based on association clustering analysis according to claim 1, characterized in that, Step S4 includes the following steps: Step S41: Discretize the industrial electricity consumption data within each industrial electricity consumption data cluster to obtain each industrial electricity consumption discretization cluster; Step S42: Perform association clustering mining on each industrial electricity consumption discretization cluster to obtain an industrial electricity consumption association clustering dataset; Step S43: Obtain the support and confidence of each associated cluster item through the industrial electricity consumption associated cluster item dataset; Step S44: Based on the support and confidence of each associated cluster item, classify and filter the associated cluster items in the industrial electricity consumption associated cluster item dataset to obtain the industrial electricity consumption associated cluster item class set; Step S45: Classify the industrial electricity consumption data corresponding to each industry based on the industrial electricity consumption association cluster category set to obtain the industrial electricity consumption data classification results.
8. The industrial electricity consumption data classification method based on association clustering analysis according to claim 7, characterized in that, Step S45 includes the following steps: Step S451: Based on the industrial electricity consumption association cluster set, perform industrial electricity consumption association classification on the industrial electricity consumption data corresponding to each industry to obtain the initial classification results of industrial electricity consumption; Step S452: Dynamically update and monitor the initial classification results of industrial electricity consumption to collect corresponding new electricity consumption data characteristics in real time; Step S453: Calculate the similarity between the new electricity consumption data features and the existing industrial electricity consumption data clusters. If the similarity is lower than a preset threshold, calculate the distance between the new electricity consumption data features and the corresponding industrial electricity consumption cluster centers to obtain the distance value between the new electricity consumption features and the cluster centers. Step S454: Determine the corresponding temporary electricity consumption cluster based on the distance between the new electricity consumption characteristics and the cluster center, and reclassify and adjust the initial classification results of industrial electricity consumption based on the temporary electricity consumption cluster to obtain the classification results of industrial electricity consumption data.
9. A classification system for industrial electricity consumption data based on association clustering analysis, characterized in that, For executing the industrial electricity consumption data classification method based on association clustering analysis as described in claim 1, the industrial electricity consumption data classification system based on association clustering analysis includes: The industrial electricity consumption filling module is used to obtain the industrial electricity consumption data corresponding to each industry in the industrial park, and to perform missing interpolation filling based on the spatiotemporal distribution of the industrial electricity consumption data corresponding to each industry to obtain the industrial electricity consumption interpolation filling data corresponding to each industry. The electricity consumption feature vector analysis module is used to perform electricity consumption feature vector analysis on the interpolated data of electricity consumption for each industry to obtain the electricity consumption feature vector for each industry. The electricity consumption feature vector includes the characteristics of electricity load distribution, the proportion of electricity consumption during peak-valley-normal periods, the electricity consumption characteristics of equipment start-up and shutdown, and the electricity consumption characteristics of the production cycle. The density peak clustering module is used to obtain the spatial distance and local distribution density between the electricity consumption feature vectors of each industry through the electricity consumption feature vectors of each industry, and to perform density peak clustering on the electricity consumption data of each industry based on the spatial distance and local distribution density between the electricity consumption feature vectors of each industry to obtain the electricity consumption data clusters of each industry. The association-aggregated electricity consumption classification module is used to perform association-aggregated mining analysis on electricity consumption data clusters of various industries to obtain industry electricity consumption association-aggregated item categories; based on the industry electricity consumption association-aggregated item categories, the corresponding industry electricity consumption data of each industry is classified to obtain the industry electricity consumption data classification results.
Citation Information
Patent Citations
Power customer classification method and device
CN113111924A
Clustering and transfer learning-based proxy electricity purchasing user load prediction method and system
CN118554424A