Energy enterprise dynamic data processing system and method based on multi-dimensional clustering

By constructing a unified data model across all business domains and employing principal component analysis and K-means clustering combined with entropy weighting, the problem of cross-business segment data correlation and dynamic optimization was solved, enabling precise data processing and intelligent decision-making for energy enterprises.

CN121258291APending Publication Date: 2026-01-02BEIJING NARI DIGITAL TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511168028.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing technologies cannot achieve cross-business segment data association and dynamic optimization, resulting in energy companies lacking a global perspective in data processing and decision-making, and making it difficult to meet the data processing requirements of complex scenarios.

Method used

A unified data model is constructed across all business domains. Principal component analysis and K-means multi-dimensional clustering are used, combined with entropy weighting for data benchmarking analysis. The indicator system is continuously and dynamically optimized through data closed-loop model optimization.

Benefits of technology

It improves the accuracy of data models and the precision of analysis, supports cross-sector benchmarking analysis, and enhances enterprise operational efficiency and the level of intelligent decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121258291A_ABST
    Figure CN121258291A_ABST
Patent Text Reader

Abstract

The invention discloses an energy enterprise data processing method and system based on multi-dimensional clustering, and the method comprises the steps: 1, collecting all types of data of an energy enterprise, outputting a standardized data set, and constructing an index dictionary; step 2, based on the index dictionary in the step 1, performing principal component analysis, combining the principal component analysis with a business rule subjected to dynamic feature weighting, and outputting a data matrix subjected to dimension reduction; step 3, based on the data matrix after dimension reduction in the step 2, performing K-means clustering analysis, and outputting clusters after classification; and 4, based on the classified clusters output in the step 3, performing scene mapping, dynamically configuring multiple benchmarking values, performing comprehensive benchmarking analysis, dividing early warning levels, and completing energy enterprise data processing. According to the method, an adaptive weight generation mechanism is introduced, an expert defines an index importance range, the algorithm automatically searches an optimal weight within the range, and the information extraction efficiency is maximized. According to the method, the defects that weight assignment is high in subjectivity and static invariant in a traditional method are effectively overcome, and the model can adapt to data changes in a self-adaptive mode.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of data processing, and particularly relates to an energy enterprise data processing system and method based on multi-dimensional clustering. BACKGROUND

[0002] Large energy enterprises cover wind power, hydropower, new energy power plants, coal mines and other industrial plates, covering production, sales, railway transportation, warehousing and other business links. The industry plate is complex, the business link is complex, and the operation is difficult. Therefore, it involves multiple data types, large data volume, high data dynamic real-time performance, and large data dynamic processing difficulty. In order to protect the optimization of resource allocation of energy enterprises, support benchmarking management business, and realize energy enterprise data processing for complex data, the existing technology mainly researches from three aspects of data modeling, data analysis and data display.

[0003] 1. Data modeling technology

[0004] The traditional scheme defines the index basic attribute through metadata driving (such as IBM InfoSphere), supports index classification expansion, but only stays in the static modeling stage, and does not fuse business scene parameters and cross-plate association.

[0005] However, the data index system of the traditional scheme is fragmented, and the cross-business data association is missing. The existing technology does not establish a unified index language in the whole business domain. The indexes such as "power supply coal consumption" of thermal power and "light energy utilization rate" of new energy belong to different systems, and cannot be analyzed collaboratively across plates, resulting in a lack of global perspective for enterprise-level strategic decision-making.

[0006] 2. Data analysis technology

[0007] In the energy enterprise benchmarking field, the existing data analysis technology GE Digital APM attempts to generate device-level data benchmarking benchmarks combined with working condition parameters, but is limited to a single business plate and lacks a unified benchmark value dynamic adaptation mechanism in the whole business domain. The data model is static, the scene adaptation capability is insufficient, the benchmark is not dynamically bound with the working condition parameters, the data analysis result is low in accuracy under complex working conditions, and the external industry benchmark value is updated with lag, which cannot reflect the technical progress in real time. Moreover, the data analysis technology in the existing technology relies on single-index benchmarking analysis, cannot comprehensively reflect the comprehensive benchmarking level, generally uses chart technology to benchmark target value, trend chart and the like, is excessively single, lacks technology for comprehensive analysis of a scene, affects horizontal analysis and comprehensive evaluation of benchmarking, and is difficult to meet the overall level evaluation of group enterprises.

[0008] 3. Data visualization technology

[0009] The existing method displays the benchmarking result through BI tools, but the early warning rules and scoring formula rely on manual configuration, which is difficult to adapt to changes in energy policy.

[0010] Overall, the existing data processing technology can only display the collected data, cannot correlate between data, is relatively isolated, cannot intelligently feedback dynamic optimization of data model, is difficult to realize complete closed loop of data monitoring, and is difficult to realize data processing requirements of complex scenes of energy enterprises. SUMMARY

[0011] To solve the problems in the prior art, the present application provides an energy enterprise dynamic data processing system and method based on multi-dimensional clustering. The main problems solved by the present application are as follows:

[0012] 1. Constructing a unified data model in the whole business domain

[0013] The present application unifies the units and standardizes the range of indicators of thermal power, hydropower, new energy and mining plate by constructing cross-business data dictionary, eliminates redundant indicators and extracts business core features by principal component analysis method, forms core indicator dimension in combination with K-means multi-dimensional clustering, refines indicators through scene decomposition, and finally establishes a cross-energy plate core indicator system with unified semantics and simplified dimension, eliminates semantic ambiguity caused by indicator fragmentation, and supports cross-plate benchmarking analysis.

[0014] 2. Use entropy weight method and multiple data analysis techniques to improve data analysis accuracy

[0015] Use absolute deviation rate and entropy weight method for multiple mode data benchmarking analysis, support energy plate group, regional company and station three-level vertical index benchmarking, horizontal intra-field and inter-field benchmarking and multiple dimension benchmarking analysis.

[0016] 3. Data model is continuously and dynamically optimized to support intelligent data decision-making

[0017] Build a "data collection-data analysis-data decision optimization" closed loop, bind the data benchmarking results with the management decision depth, realize the continuous optimization of the index system and the benchmarking model, and promote the spiral improvement of the operation level of the energy enterprise.

[0018] The present application adopts the following technical solutions.

[0019] The first aspect of the present application provides an energy enterprise data processing method based on multi-dimensional clustering, comprising:

[0020] Step 1, collect various types of data of energy enterprises, output standardized data set, and construct index dictionary;

[0021] Step 2, based on the index dictionary of step 1, perform principal component analysis, and combine with the dynamic feature weighted business rules to output the reduced dimension data matrix;

[0022] Step 3: Based on the data matrix after dimensionality reduction in Step 2, K-means clustering analysis is performed, and the classified clusters are output.

[0023] Step 4: Based on the classified clusters output in Step 3, scene mapping is performed, multi-benchmark values are dynamically configured, comprehensive benchmarking analysis is performed, early warning levels are divided, and energy enterprise data processing is completed.

[0024] Preferably, it further includes: Step 5, based on the benchmarking analysis results in Step 4, index visualization is performed;

[0025] Step 6: Based on the results of Step 5, the reasons are traced back and improved to continuously optimize.

[0026] Preferably, in Step 1, the standardized data calculation formula in the standardized data set is:

[0027]

[0028] wherein, is the standardized jth index data of the ith energy enterprise, x ij is the jth index data of the ith energy enterprise, μ j is the mean of the jth index, σ j is the standard deviation of the jth index.

[0029] Preferably, the business rules in Step 2 include:

[0030] Rule R1: If "power supply coal consumption > industry benchmark value 10%", trigger "high energy consumption risk", increase the importance evidence weight of coal consumption index in efficiency dimension +0.3;

[0031] Rule R2: If "monthly increase in non-planned equipment downtime >20%", trigger "reliability deterioration", increase the importance evidence weight of downtime coefficient in reliability dimension +0.4;

[0032] Rule R3: If "power generation capacity achievement rate for 3 consecutive months >110%", trigger "overcapacity warning", decrease the importance evidence weight of power generation capacity in capacity dimension -0.2.

[0033] Preferably, when the index value meets the business rules, the corresponding evidence event is automatically triggered to realize dynamic feature weighting:

[0034]

[0035] wherein, is the dynamic weight of the jth index of the energy enterprise in the time window t, b j is the basic importance weight of the jth index of the energy enterprise set by the expert, is the cumulative weight of the jth index of the energy enterprise in the time window t, is the improved dynamic evidence strength coefficient, λ is a compensation coefficient, and C j is the criticality level of the jth index of the energy enterprise.

[0036] Preferably, the improved dynamic evidence strength coefficient is:

[0037]

[0038] wherein β is a basic strength coefficient, γ is a decay factor, and k is a curvature factor.

[0039] Preferably, in step 3, in the K-means clustering analysis, the initial centroid is selected based on the four business dimensions, including production efficiency, energy consumption cost, equipment reliability, and environmental compliance.

[0040] Preferably, in step 4, data drift detection is used when dynamically configuring the multi-benchmark value, and the KL divergence value of the index distribution is calculated. If the KL divergence value exceeds the set threshold, it is determined that there is data drift.

[0041] Preferably, in step 4, the warning level is divided by the absolute deviation rate, and the absolute deviation rate calculation formula is:

[0042]

[0043] The warning levels include very serious, serious, abnormal good, and normal.

[0044] The second aspect of the present application provides an energy enterprise data processing system, which adopts the above method, comprising:

[0045] An energy enterprise data acquisition unit, a principal component analysis unit, a K-means clustering unit, and a comprehensive benchmarking analysis unit;

[0046] The energy enterprise data acquisition unit is used to acquire various types of data of the energy enterprise, output standardized data sets, and construct an index dictionary.

[0047] The principal component analysis unit is used to perform principal component analysis and combine with the dynamically weighted business rules to output a reduced dimension data matrix.

[0048] The K-means clustering unit is used to perform K-means clustering analysis and output classified clusters.

[0049] The comprehensive benchmarking analysis unit is used to perform scene mapping, dynamically configure multi-benchmark values, and perform comprehensive benchmarking analysis to divide warning levels and complete energy enterprise data processing.

[0050] The beneficial effects of the present application are that, compared with the prior art:

[0051] 1. Data model processing precision. PCA-Kmeans model with deep integration of domain knowledge: first business dimension anchoring initialization, for K-means clustering, abandoning the conventional random or Canopy initialization method, innovatively presetting the initial centroid according to the historical optimal performance sample of the core dimension of energy business (production efficiency, energy consumption cost, etc.). This ensures that the clustering result is directly and stably mapped to the key business scene, solves the problem of disconnection between the clustering result and the business meaning, and significantly improves the practicality and interpretability of the result. Dynamic optimization of feature weight, before principal component analysis (PCA), an adaptive weight generation mechanism guided by domain knowledge is introduced. Experts define the importance range of the index, and the algorithm automatically optimizes to generate the optimal weight within the range, maximizing information extraction efficiency. This effectively solves the defects of strong subjectivity and staticity of weight assignment in traditional methods, so that the model can adapt to data changes.

[0052] 2. Improve data analysis accuracy. Define four types of benchmark values and dynamically adapt to business scenarios, calculate index weights combined with entropy weight method, and allocate importance according to data dispersion. In complex working conditions, entropy weight method eliminates the influence of dimension, improves the index discrimination, improves the accuracy of the standard, and helps enterprises accurately identify operational short boards and advantage indicators.

[0053] 3. Data decision intelligence. Entropy weight method dynamically adjusts index weight according to benchmarking results to drive continuous optimization of the model. Through data feedback, real-time calibration of management strategy is realized, and index weight is intelligently updated according to business needs, improving energy enterprise operation efficiency and promoting energy management to fine and intelligent iteration and upgrading. BRIEF DESCRIPTION OF DRAWINGS

[0054] Figure 1 is the architecture diagram provided by the embodiment of the present application. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. The embodiments described in the present application are only a part of the embodiments of the present application, not all embodiments. Based on the spirit of the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the present application.

[0056] In order to solve the problems existing in the prior art, the embodiment of the present application provides a kind of dynamic data processing and continuous optimization system and method of energy enterprise based on multi-dimensional clustering to solve the technical problem of improving energy industry data processing.

[0057] To solve the above technical problems, the application adopts the technical solutions as follows.

[0058] As Figure 1 shown is an architecture diagram provided by an embodiment of the application; the embodiment of the application provides an energy enterprise data processing method based on multi-dimensional clustering, comprising:

[0059] Step 1, collect various types of data of energy enterprises, output a standardized data set, and construct an index dictionary.

[0060] Integrate production data, management data and external industry data of energy enterprises in the power, hydropower, new energy and mine plate blocks, and perform data processing to unify data units (such as ten thousand tons→tons), time granularity (monthly / annual), value range, and output a standardized data set. The production data includes but is not limited to power supply coal consumption, equipment unplanned downtime frequency and power generation, etc.

[0061] Based on the standardized data set, through a graphical interface, new indexes are added, the index name, unit and data caliber are defined, and the business plate is associated to construct an index dictionary. In an embodiment of the application, for example, when a new "light energy utilization rate" index is created, the new energy plate and the power plant level caliber can be selected, and the benchmarking period is set to monthly.

[0062] The process of standardizing various types of data of energy enterprises is as follows:

[0063] The collected various types of data of energy enterprises are standardized to make the mean value of each type of data 0 and the standard deviation 1, which is used to eliminate the influence of different type data characteristic dimension.

[0064] The various types of data matrix of energy enterprises is denoted as X, and the standardized data set is denoted as X std , and the calculation method is as follows:

[0065] The mean value of each index is calculated:

[0066]

[0067] Wherein, μ j is the mean value of the jth index, n is the number of indexes, and x ij is the jth index value of the ith energy enterprise.

[0068] The standard deviation of each index is calculated:

[0069]

[0070] Wherein, σ j is the standard deviation of the jth index.

[0071] The standardized data:

[0072]

[0073] wherein, is the jth index data of the ith energy enterprise after standardization.

[0074] The jth index data of the ith energy enterprise after standardization is calculated to form a standardization data set X std .

[0075] In a preferred but non-limiting embodiment of the present application, the energy enterprise production data is collected through a SCADA (Supervisory Control And Data Acquisition) system, the management data is collected through an ERP (Enterprise Resources Planning) system, the external industry data includes CEA benchmark library, etc., the data units are unified to ten thousand tons or tons, etc., and the time granularity is unified to monthly or annual, etc.

[0076] Step 2, based on the constructed index dictionary, principal component analysis (PCA) is performed, and combined with the dynamic characteristic weighted business rules, a reduced dimension data matrix is output.

[0077] The core goal of principal component analysis is to convert high-dimensional data to low-dimensional space and reduce redundant indexes. Energy business involves numerous operation indexes, and different indexes can reflect various characteristics in various business production processes from different angles. However, the number of these indexes is large and there may be certain correlation between them, which will increase the complexity and calculation amount when directly used for benchmarking analysis.

[0078] Therefore, the embodiment of the present application adopts the method of combining principal component analysis (PCA) and dynamic characteristic weighted business rules to process the index dictionary constructed in step 1, and the processing flow is as follows.

[0079] Step 2.1. Calculate the weighted covariance matrix:

[0080] The feature is weighted on the standardized data set, and the standardized data set is set to p n-dimensional data, denoted as X std_p×n , and the weighted standardized data set χ std_p×n is obtained;

[0081] The covariance matrix S p×n is calculated based on the weighted standardized data set χ std_p×n ;

[0082] Step 2.2. Eigenvalue decomposition:

[0083] Eigenvalue decomposition is performed on the covariance matrix S p×n Eigenvalues are arranged in descending order: Eigenmatrix is obtained: p×n = [phi1, phi2,..., phi p ] and corresponding eigenvectors phi1- phi p .

[0084] Step 2.3. Select principal components:

[0085] According to the size of the eigenvalue , the eigenvectors corresponding to the first m largest eigenvalues are selected as the principal components: p×m = [phi1, phi2,..., phi p ], where m is the dimension after dimension reduction.

[0086] Step 2.4. Data projection:

[0087] The weighted and standardized data set X std is projected onto the selected principal components to obtain the data matrix after dimension reduction:

[0088] Y n×m = chi std_p×n psi p×m

[0089] Where Y n×m is the data matrix after dimension reduction, chi std_p×n is the weighted and standardized data set, psi p×m = [phi1, phi2,..., phi p ], n is the dimension of the weighted and standardized data set, and m is the dimension after dimension reduction.

[0090] Through principal component analysis, various types of data of energy enterprises are processed, the principal components of the data are found, the original high-dimensional index data is converted into low-dimensional comprehensive index, and the simplified index is used as the index data basis of clustering analysis, so that the clustering quality is improved.

[0091] In step 2.1, in order to solve the subjectivity of traditional weight assignment and the conventionality of algorithm combination, the embodiment of the application proposes a dynamic weight generation mechanism based on business rules and data evidence, the core of which is to convert domain knowledge into executable quantitative rules, and automatically adjust the weight through real-time data feedback, and the processing flow includes:

[0092] Step 2.1.1. Business rule base construction: Organize energy experts to define business impact factor rules (non-weight range), and the rule base supports dynamic addition and deletion: ​

[0093] Rule R1: If "power supply coal consumption > industry benchmark value 10%", trigger "high energy consumption risk", increase the importance of the coal consumption index in the efficiency dimension by 0.3;

[0094] Rule R2: If "monthly increase in non-planned shutdown times of equipment > 20%", trigger "reliability deterioration", increase the importance of the shutdown coefficient in the reliability dimension by 0.4;

[0095] Rule R3: If "power generation capacity achievement rate for 3 consecutive months > 110%", trigger "overcapacity warning", decrease the importance of power generation capacity in the capacity dimension by 0.2.

[0096] Step 2.1.2. Real-time evidence collection and weight calculation: The system monitors the index data in real time, and when the index value meets the conditions in the rule library in step 2.1, the corresponding evidence event is automatically triggered.

[0097] The weight of each index in the time window t is determined by its basic importance weight and cumulative evidence weight:

[0098] Basic importance weight: pre-set by experts according to the nature of the business, in one preferred but not limiting embodiment of the invention, such as coal consumption basic value = 1.0, plant power rate = 0.8.

[0099] Cumulative evidence weight: the sum of all triggered evidence values contributed by rules related to the index within the statistical time window (such as the last 3 months).

[0100] Weight synthesis formula:

[0101]

[0102] In the formula, Wj(t) is the weight of the jth index of the energy enterprise in the time window t, α is the evidence intensity coefficient, used to adjust the evidence impact amplitude, b j b is the basic importance weight of the jth index of the energy enterprise set by the expert, Wj(t) is the cumulative evidence weight of the jth index of the energy enterprise in the time window t.

[0103] Step 2.1.3. To improve the flexibility and applicability of the algorithm, improve the evidence intensity coefficient α to be a business rule driven nonlinear dynamic weight generation. Upgrade the fixed coefficient α to a dynamic coefficient that changes nonlinearly with the accumulation of evidence

[0104]

[0105] where, The dynamic evidence intensity coefficient of the jth index of the improved energy enterprise in the time window t is β, the basic intensity coefficient is γ, the attenuation factor is k, and the curvature factor is The cumulative evidence weight of the jth index of the energy enterprise in the time window t.

[0106] In a preferred but non-limiting embodiment of the present application, the energy scene specific parameter is designed as shown below:

[0107] The basic intensity coefficient β is 0.6 by default, 0.7 for thermal power / water power, and 0.5 for new energy;

[0108] The attenuation factor γ is 0.8 by default, and 0.5 for high volatility indicators such as coal consumption;

[0109] The curvature factor k is 1.2 by default, and 1.5 for equipment class indicators.

[0110] At the same time, the dynamic weight synthesis formula is upgraded to:

[0111]

[0112] Among them, The dynamic weight of the jth index of the energy enterprise in the time window t is b j The basic importance weight of the jth index of the energy enterprise set by the expert is The cumulative evidence weight of the jth index of the energy enterprise in the time window t is The improved dynamic evidence intensity coefficient is λ, the compensation coefficient is C j The criticality grade of the jth index of the energy enterprise.

[0113] In a preferred but non-limiting embodiment of the present application, a new energy business compensation item is added:

[0114] The criticality grade C of the jth index of the energy enterprise j Set by the expert: power generation = 3, coal consumption = 3, auxiliary power rate = 2, calorific value difference = 1;

[0115] The compensation coefficient λ is 0.05 by default, which ensures that the basic weight of the core index is not diluted.

[0116] Among them, the formula parameter definition and energy scene value rule:

[0117] b j The basic importance weight of the jth index of the energy enterprise set by the expert is set by the expert according to the nature of the index business, and the power coal consumption (thermal power) = 1.2, the light energy utilization rate (photovoltaic) = 1.0;

[0118] is the cumulative evidence weight of the jth index of the energy enterprise in the time window t, the sum of the trigger rule evidence values in the last 3 months, and the coal consumption exceeds the standard for 3 consecutive months:

[0119] β is a basic strength coefficient, reflecting the sensitivity of the index type, thermal power / water power: 0.7 (high inertia system needs fast response), new energy: 0.5 (allowing fluctuation), thermal power coal consumption: 0.7, and fan vibration: 0.5;

[0120] γ is a decay factor for controlling the weight growth rate, high fluctuation index (coal consumption): 0.5 (fast promotion), stable index (power generation): 0.8 (gentle promotion), coal consumption: γ = 0.5, and power generation: γ = 0.8;

[0121] k is a curvature factor for adjusting the inflection point position of the S-shaped curve: equipment health index: 1.5 (early warning), economic index: 1.0 (balanced response), unplanned shutdown: k = 1.5, and electricity sales revenue: k = 1.0;

[0122] λ is a compensation coefficient, a fixed value of 0.05, and a core index weight;

[0123] C j is the criticality grade of the jth index of the energy enterprise, reflecting the importance of the energy production chain: core production index (power generation / coal consumption), efficiency index (auxiliary power rate), and auxiliary index (heat value difference) | power supply coal consumption C = 3, heat value difference C = 1.

[0124] Step 3, based on the reduced data matrix output by step 2, K-means algorithm clustering analysis is carried out, and the classified clusters are output.

[0125] In view of the problems that the clustering results may be unstable and disconnected with the business meaning caused by the random or Canopy initialization of the K-means algorithm, the core dimension knowledge of the energy business is innovatively used to preset the initial centroid:

[0126] Step 3.1, define the core business dimension: based on the deep understanding of the operation of the energy enterprise, four core analysis dimensions are defined: production efficiency, energy consumption cost, equipment reliability and environmental compliance.

[0127] Step 3.2, business dimension anchor initial point:

[0128] The initial centroid of the production efficiency dimension: calculate the average value of the sample points with the best performance (such as the top 20%) of “power generation” and “utilization hours” in the historical data as the initial centroid of this dimension. This anchors the high-capacity state.

[0129] Energy cost dimension initial centroid: Calculate the average of the sample points with the best performance (i.e., the lowest values in the last 20%) of "power supply coal consumption" and "overall plant power consumption rate" in the historical data as the initial centroid of this dimension. This anchors the low energy consumption state.

[0130] Equipment reliability dimension initial centroid: Calculate the average of the sample points with the best performance (i.e., the lowest values in the first 10%) of the equipment health indicators such as "unscheduled downtime coefficient" in the historical data as the initial centroid of this dimension. This anchors the high reliability state.

[0131] Environmental compliance dimension initial centroid: Calculate the average of the sample points with the best performance (e.g., the first 15%) of the relevant emission indicators in the historical data as the initial centroid of this dimension (the specific indicators are determined according to the business board). This anchors the environmental compliance state.

[0132] Step 3.3, K-means clustering execution: Use the four initial centroid points preset based on the business dimensions above to start the K-means clustering algorithm for subsequent iterative calculation and cluster division. Similar indicators or data points are divided into the same cluster, thereby identifying different business patterns or indicator categories, and outputting the K-means clustering classified clusters.

[0133] According to the characteristics and analysis needs of energy enterprise business, the indicators are divided into four core dimensions: production efficiency, energy cost, equipment reliability, and environmental compliance. Therefore, the number of cluster centers is set to 4. When calculating the distance between samples and cluster centers, the error of production efficiency is amplified by 2 times. And use Canopy algorithm for initialization, by merging close distance data points to determine stable initial center, avoid randomness.

[0134] After the above work is completed, the K-means clustering specific algorithm is as follows:

[0135] 3-1. Set the data set as Y nK×mK , to be divided into K clusters C = {C1, C2,..., CK}, the centroid is μ = {μ1, μ2,..., μK}, then the objective function is: K K

[0136]

[0137] Where ||x jK - μ iK || represents the distance from data point x jK to centroid μ iK , usually using Euclidean distance.

[0138] ​​3-2. Initialization of the centroid: According to the characteristics of the energy enterprise business and the analysis requirements, four data points are selected as the initial centroids μ1, μ2, μ3, μ4 according to the four core dimensions

[0139] 3-3. Assign data points: For each data point x in the data set jK , calculate the distance from the four centroids, and assign it to the cluster where the nearest centroid is located.

[0140] 3-4. Update the centroid: For each cluster C iK , calculate the mean of all data points in the cluster, and take this mean as the new centroid μ iK_NEW

[0141] 3-5. Repeat iteration: Repeat 3-3 and 3-4 until the centroid no longer changes significantly or reaches the preset number of iterations, then complete the K-means clustering.

[0142] Step 4, based on the classified clusters output by step 3, perform scenario mapping, dynamically configure multiple benchmark values, and perform comprehensive benchmark analysis to divide the warning level and complete the energy enterprise data processing.

[0143] Step 4.1, scenario mapping, the clusters classified by clustering analysis in step 3 are mapped to scenarios, scenario mapping is to decompose complex business systems into independent analysis sub-scenarios (such as production efficiency, energy cost, equipment reliability and environmental compliance in step 3), and to clearly define the core indicator set of each sub-scenario, so that the specific scenario can be combined with the scene for benchmarking and associated with the business scene.

[0144] In a preferred but non-limiting embodiment of the present application, the indicator set after scenario mapping is the core indicator of each scenario, and the new scenario mapping includes multiple layers, scenario name, associated indicator code, indicator name, indicator calculation, etc., forming a more business-oriented comprehensive dimension.

[0145] Step 4.2, dynamically configure multiple benchmark values based on the scenario mapping results of step 4.1, in a preferred but non-limiting embodiment of the present application, define four types of benchmark values and establish calculation rules, including target value, internal advanced value, industry average value and same period value:

[0146] Target value: based on equipment design parameters, historical optimal value or industry standard;

[0147] Internal advanced value: same as the top 10% unit indicator value in the same block;

[0148] Industry average value: obtained through API interface with external libraries such as CEDRAB, International Energy Agency, etc.;

[0149] Same period value: automatically match historical same period data.

[0150] In addition, data drift detection technology is introduced when configuring multiple benchmark values, and a benchmark value dynamic updating triggering mechanism is established:

[0151] 1) Data drift detection:

[0152] The KL divergence of the index distribution is calculated monthly: D KL (P t ||P t-1 );

[0153] If D KL (P t ||P t-1 )>θ (θ=0.1), θ is the set KL divergence threshold, it is determined that the data distribution has a significant deviation, that is, data drift occurs;

[0154] 2) Automatic retraining:

[0155] When data drift is detected, the following is automatically performed: update the industry average value (for the ITU API), recalculate the four types of benchmark values for all indicators;

[0156] Step 4.3, based on the multiple benchmark values dynamically configured in step 4.2, the memory comprehensive benchmarking analysis is performed, and the warning level is divided.

[0157] For benchmarking business needs, scene business particularity, the core index data after clustering analysis is subjected to benchmarking analysis, the difference of key indicators is reflected, and whether each benchmarking object is excellent or insufficient is compared. The difference comparison calculation method is as follows:

[0158]

[0159] According to different absolute deviation rates, different warning levels are divided: very serious, serious, abnormal good and normal.

[0160] In a preferred but non-limiting embodiment of the present application, the benchmark is first classified and the weight is initialized; according to the multiple benchmark values (such as device type, region), the data is divided into multiple benchmark groups (such as thermal power plant group, hydropower plant group); the entropy weight method weight of each benchmark group is initialized, and the index entropy value and difference coefficient in each group are calculated. And introduce feedback mechanism, according to the benchmarking result (such as equipment health degree score) dynamically adjust the importance coefficient of benchmark. If a certain benchmark group has a long-term deviation (such as the equipment failure rate of the thermal power plant group is higher than the benchmark), the corresponding coefficient is increased; through the gradient descent algorithm, it is updated to minimize the comprehensive benchmarking error.

[0161] Step 5, data visualization display, based on the benchmarking analysis results in step 4, the index visualization display is performed.

[0162] Index visualization display. The results of the benchmarking analysis are displayed on the control center large screen, and the analysis results are displayed using diversified charts. In a preferred but non-limiting embodiment of the present application, the diversified charts include: longitudinal analysis, transverse benchmarking, and cross-business analysis; longitudinal analysis: display the trend chart of a certain index in the past 12 months, etc. embody single index longitudinal, and mark the target value; transverse benchmarking: single / comprehensive index benchmarking result embodied by inter-field benchmarking ranking of a certain index; cross-business analysis: comprehensive benchmarking result embodied by radar chart, etc.

[0163] Step 6, continuous improvement and optimization, based on the display results in step 5, reason tracing and improvement, continuous optimization.

[0164] In a preferred but non-limiting embodiment of the present application, the continuous optimization operation includes but is not limited to: data analysis benchmarking result feedback data model: if the scene index warning is frequent, the following optimization is automatically triggered: adjust the weight of the index; update the benchmark value. Continuously dynamically optimize and continuously improve the overall operation efficiency of the overall power group enterprise, regional company and power plant, and guide the overhaul management and rectification activities of the power plant, and help the fine and intelligent management of data benchmarking management.

[0165] The beneficial effects of the embodiments of the present application are as follows:

[0166] 1. Multi-dimensional clustering simplifies the complexity of data index model. In the index management of energy enterprises, there are many indexes and they are related to each other, which leads to complex data processing and low benchmarking efficiency. The present application innovatively combines PCA and field knowledge weighting method to realize k-means clustering. First, the PCA algorithm reduces the dimension of the original index data, removes redundant information, reduces the complexity of the data, and at the same time preserves the key information, and relies on the field knowledge weighting method in the process to make the related indexes more in line with the actual needs. The K-means algorithm clusters the reduced data. This way greatly simplifies the index system, makes the originally complex data more orderly, improves the clustering quality, and enables enterprises to more efficiently perform benchmarking analysis and accurately grasp their business situation.

[0167] 2. Dynamic Adjustment of Data Analysis. Traditional benchmarking analysis often uses fixed benchmark values, which are difficult to adapt to the dynamically changing business environment of energy companies. This invention defines four types of benchmark values: target value, internal advanced value, industry average, and value compared to previous periods. By combining API integration with historical data matching, dynamic adjustment of benchmark values ​​is achieved. The latest industry data can be obtained in real time through the API. Combined with historical data analysis, a process of "benchmark classification - weight initialization - cross-benchmark fusion" is used to customize weights for different benchmark groups, improving benchmarking accuracy. A feedback mechanism and gradient descent algorithm are introduced to dynamically adjust the benchmark importance coefficient, ensuring that weights and benchmarking effects are optimized in tandem. Furthermore, by integrating the "benchmark importance coefficient" into business rules, the weighting results are interpretable. This allows benchmarking analysis to more accurately reflect the actual situation of the enterprise, improving benchmarking accuracy, providing a more reliable basis for enterprise decision-making, and helping enterprises maintain competitiveness in a constantly changing market environment.

[0168] 3. The entropy weighting method and domain knowledge are fully integrated, reducing human intervention in quantitative analysis. Energy companies often have complex indicator systems, making it difficult to scientifically measure the importance of each indicator. Often, only a single indicator is considered, or important indicators are judged based on experience. This invention, by organically combining the entropy weighting method and domain knowledge, can objectively and accurately allocate the weights of each indicator, avoiding interference from subjective factors. In comprehensive indicator benchmarking analysis, reasonable indicator weights can more accurately reflect the actual level of each evaluated object, improving the scientific validity and credibility of the benchmarking analysis results and providing stronger support for corporate decision-making.

[0169] 4. Continuous optimization of data indicator models leads to more reliable intelligent decision-making. To ensure the effectiveness and adaptability of the benchmarking management system for energy enterprises, this invention establishes a closed-loop continuous optimization mechanism of "data analysis - result application - model optimization". During the model optimization phase, indicator weights are adjusted and benchmark values ​​are updated based on the benchmarking results, continuously optimizing the system model. This closed-loop mechanism enables the system to continuously adapt to the development and changes of the enterprise, promoting the management upgrade of energy enterprises and achieving the goal of sustainable development.

[0170] 5. Weight adaptation and business anchoring improve model performance

[0171] Addressing the issue of subjective weighting: Precisely responding to persistent anomalies: When indicators continue to deteriorate (such as continuous exceedances of coal consumption standards), the S-curve coefficient is used to... Achieve a gradual increase in weighting, starting with a sharp rise and then stabilizing, to avoid early underestimation or later overreaction. Energy scenario parameter optimization: Formula parameters are configured differently according to the energy sector (thermal power / hydropower / new energy), such as: thermal power k = 1.2 (rapid response to sudden changes in coal consumption); new energy k = 0.8 (smooth adaptation to fluctuations in sunlight); minimum weighting for key indicators: compensation term λ·C jEnsure that the core indicators such as power generation, coal consumption, etc. are not diluted by short-term evidence values, and protect the business logic. Introduce a nonlinear mechanism in the core role of the energy scene: solve the problem of "sudden abnormal underestimation" of thermal power indicators: when the coal quality suddenly changes, causing the coal consumption to exceed the standard for the first time , the decay factor γ = 0.5 makes the intensity coefficient rapidly rise to 0.42 (compared to the linear fixed α = 0.5, only 0.15 increase), triggering an early warning 1-2 days in advance, avoiding a sudden drop in boiler efficiency. Technical effect: After application in a certain power plant, the number of non-stops caused by abnormal coal quality is reduced. Smooth out the "volatility false alarm" problem of new energy indicators: photovoltaic "light energy utilization rate" suddenly drops by 20% in a single day due to weather fluctuations, , β = 0.5, γ = 0.8 makes (only 67% of the same evidence value for thermal power), avoiding excessive weight increase, preventing the fluctuation of automatic recovery after sunny days from being misjudged as equipment failure. Technical effect: false alarm rate decreased by 5, operation and maintenance efficiency improved. Ensure the "continuous deterioration tracking" of equipment health indicators: when the fan vibration value exceeds the standard for 3 consecutive months , the curvature factor k = 1.5 makes the intensity coefficient accelerated growth from 0.3 (1st month) to 0.6 (3rd month), driving weight increase priority, accurately matching the progressive characteristics of gear wear. Technical effect: a certain wind farm early warning of gear box damage, saving maintenance cost.

[0172] 6. Improve the fit of clustering business

[0173] An improved strategy based on business dimension preset initial centroid is adopted, which directly injects the knowledge of energy enterprise core operation dimensions (production efficiency, energy cost, etc.) into the K-means clustering process. Compared with conventional random or Canopy initialization, this method makes the clustering results correspond to key business scenarios organically, greatly enhancing the relevance of model output and business needs.

[0174] 7. Nonlinear weight function customized for energy scene

[0175] Unique parameter response logic: decay factor γ is configured according to indicator volatility: small value (γ = 0.5) for sudden abnormality of thermal power to achieve rapid response, large value (γ = 0.8) for new energy fluctuations to suppress false alarms. Curvature factor is configured according to fault development mode: large value (k = 1.5) for progressive faults (such as equipment wear) to strengthen continuous tracking, small value (k = 1.0) for sudden faults to balance response. Technical breakthrough: the same evidence value produces different weight increase amplitudes for thermal power and new energy scenes (thermal power +54% vs. new energy +32%), overcoming the defect that general algorithms cannot adapt to the characteristics of energy multi-industry.

[0176] Example 2

[0177] Application scenario: Based on the production index data of A, B, C, and D four thermal power plants under a certain energy group in October 2023, focus on the clustering analysis of core indicators such as power generation and power supply coal consumption, and select the core dimensions and index items of the energy group's thermal power board. In terms of fuel efficiency, conduct horizontal benchmarking analysis and continuous improvement optimization on A, B, C, and D four thermal power plants in October 2023.

[0178] Step 1: Data collection, construction of index dictionary.

[0179] From the SCADA system, collect the unit operation data of A, B, C, and D four thermal power plants from January to October 2023: 8 production indicators such as power generation, utilization hours, equivalent availability coefficient, unplanned downtime coefficient, planned downtime coefficient, power reduction downtime coefficient, power coefficient, and standby coefficient;

[0180] From the ERP system, extract historical data from 2018 to 2023: 4 management indicators such as power supply coal consumption, comprehensive plant power consumption rate, electricity sales revenue, and calorific value difference, a total of 12 index data.

[0181] Standardize and normalize the collected data, unify the time granularity to monthly (aggregate real-time minute-level data into monthly average), and standardize the unit (power supply coal consumption unified to g / kWh, power generation unified to 10,000 kWh);

[0182] Define the index name and unit, power generation (unit: kWh), utilization hours (unit: h), equivalent availability coefficient (unit: %), unplanned downtime coefficient (%), power supply coal consumption (g / kWh), comprehensive plant power consumption rate (%), electricity sales revenue (yuan), etc. 12 indicators to build an index dictionary, clearly define the index, unit, and monthly range, and associate the business board with thermal power to form the original index set.

[0183] Step 2: Principal Component Analysis (PCA).

[0184] Perform principal component analysis on the data within the index dictionary range to reduce index redundancy.

[0185] Calculation result: the first three principal components (cumulative variance contribution rate 93%):

[0186] Principal component 1 (load response capability, 48%): high load indicators include power generation (0.95), utilization hours (0.92), and equivalent availability coefficient (0.88), reflecting the response capability and capacity release level of the unit in load dispatching.

[0187] Principal Component 2 (Fuel Efficiency, 32%): High load indicators are power supply coal consumption (-0.90), comprehensive plant power consumption rate (-0.85), and calorific value difference (-0.80). The smaller the value (the higher the negative score), the higher the energy conversion efficiency and the lower the plant energy consumption.

[0188] Principal Component 3 (Equipment Health, 13%): High load indicators are unplanned downtime coefficient (0.85) and reduced output downtime coefficient (0.80), reflecting the frequency of equipment failure or forced downtime.

[0189] Through the above PCA algorithm, we convert the original 12 production indicators into 3 comprehensive indicators (principal components), representing "load response capability", "fuel efficiency" and "equipment health", providing more concise and effective data for subsequent cluster analysis and benchmarking analysis.

[0190] Step 3: K-means algorithm cluster analysis.

[0191] After PCA analysis, the indicators are clustered to analyze the core dimensions of the energy group that affect the overall operating efficiency.

[0192] After 10 iterations, the centroid is stable, and finally divided into 3 core dimensions, each dimension contains indicators and centroid physical meaning as follows:

[0193] Core Dimension 1: Production Efficiency; Contains indicators: power generation, utilization hours, equivalent availability coefficient.

[0194] Core Dimension 2: Energy Consumption Cost; Contains indicators: power supply coal consumption, comprehensive plant power consumption rate, calorific value difference

[0195] Core Dimension 3: Reliability; Contains unplanned downtime coefficient, reduced output downtime coefficient, and planned downtime coefficient Step 4: Scene mapping, dynamic adaptation of multiple benchmark values, and comprehensive benchmarking analysis using entropy weight method.

[0196] The cluster analysis indicators are mapped to the scene, and the scene mapping results are as follows:

[0197] Scene 1: Production; Core Dimension: Production Efficiency; Contains indicators: power generation, utilization hours, equivalent availability coefficient

[0198] Scene 2: Energy Consumption; Core Dimension: Energy Consumption Cost; Contains indicators: power supply coal consumption, comprehensive plant power consumption rate, calorific value difference

[0199] Scene 3: Equipment; Core Dimension: Reliability; Unplanned downtime coefficient, reduced output downtime coefficient, and planned downtime coefficient

[0200] Take power supply coal consumption as an example, the four types of benchmark values are defined as follows:

[0201] Target value

[0202] Calculation rule: 300MW unit design value (70% load rate)

[0203] A / D power plant target value: 300g / kWh

[0204] B / C power plant target value: 300g / kWh

[0205] Industry average

[0206] China Electricity Council 2023 average of the same type of unit (API real-time access)

[0207] Industry average: 305g / kWh

[0208] Same as previous value

[0209] October 2022 data from the plant

[0210] A / D power plant same period value: 310g / kWh

[0211] B / C power plant same period value: 295g / kWh

[0212] Internal advanced value

[0213] Group top 10% power plant average in October 2023 (B power plant)

[0214] Internal advanced value: 290g / kWh

[0215] The fuel efficiency of A, B, C and D power plants owned by the enterprise is compared using the entropy weight method. The power plant with higher fuel efficiency has higher score.

[0216] After calculation, the benchmarking results are obtained:

[0217] In the energy consumption cost scenario, the index weights of the three indicators of power supply coal consumption, comprehensive plant power consumption rate and calorific value difference are: 38.1%, 33.3% and 28.6% respectively

[0218] Weighted score of each power plant

[0219] A power plant weighted score: 0.5715

[0220] B power plant weighted score: 0.7068

[0221] C power plant weighted score: 0.6296

[0222] D power plant weighted score: 0.5788

[0223] It is concluded that:

[0224] The power plant with the highest fuel efficiency is B power plant, with a score of 0.7068

[0225] The power plant with the lowest fuel efficiency is: A Power Plant, score: 0.5715

[0226] The deviation rate of each power plant index relative to the industry benchmark:

[0227] A Power Plant: deviation rate of supply coal consumption 5.00%, deviation rate of comprehensive plant power consumption 11.76%, deviation rate of heat value difference -25.00%

[0228] B Power Plant: deviation rate of supply coal consumption -3.33%, deviation rate of comprehensive plant power consumption -5.88%, deviation rate of heat value difference 10.00%

[0229] C Power Plant: deviation rate of supply coal consumption 1.67%, deviation rate of comprehensive plant power consumption 3.53%, deviation rate of heat value difference -10.00%

[0230] D Power Plant: deviation rate of supply coal consumption 6.67%, deviation rate of comprehensive plant power consumption 8.24%, deviation rate of heat value difference -30.00%

[0231] Step 5: Visualization of benchmarking results

[0232] Visualization of benchmarking results:

[0233] 1. Weighted score column chart of each power plant

[0234] Draw a column chart to visually display the weighted score of fuel efficiency of each power plant. It can be seen directly: B > C > D > A, the power plant with the highest fuel efficiency is: B Power Plant, score: 0.7068, the power plant with the lowest fuel efficiency is: A Power Plant, score: 0.5715. A / D power plants as low-score power plants need to be focused on, improve fuel efficiency.

[0235] 2. Pie chart of index weight

[0236] Under the weighted score column chart of the power plant, display the index weight pie chart of this benchmarking. The index weights of the three indexes supply coal consumption, comprehensive plant power consumption, and heat value difference are: 38.1%, 33.3%, and 28.6% respectively.

[0237] It can be seen that:

[0238] Supply coal consumption (38.1%): highest weight, reflects the core influence of fuel conversion efficiency, is the key indicator to reduce power generation cost.

[0239] Comprehensive plant power consumption (33.3%): second weight, reflects the energy utilization efficiency within the plant, and high plant power consumption will directly reduce the amount of external power sales.

[0240] Heat value difference (28.6%): slightly low weight, but still affects the evaluation of fuel quality, and insufficient heat value will lead to increased coal consumption.

[0241] 3. Each power plant index value and target value column chart

[0242] Display the power supply coal consumption, comprehensive plant power consumption rate, and heat difference three index column charts, respectively display the values and target values of A, B, C, and D four power plants. Intuitively show the horizontal comparison difference of each index and the difference with the target value.

[0243] 4. Each index deviation rate radar chart

[0244] The deviation rate of each power plant reflects their gap with the industry benchmark in each index. For negative indexes (power supply coal consumption, comprehensive plant power consumption rate, heat value difference, all the lower the better): negative deviation rate (value <0) indicates better than the industry benchmark;

[0245] positive deviation rate (value >0) indicates inferior to the industry benchmark.

[0246] Draw a radar chart to show the deviation rate of each power plant index relative to the industry benchmark.

[0247] You can see:

[0248] A power plant

[0249] Advantages: heat value difference is significantly better than the benchmark (-25%), indicating higher fuel energy utilization efficiency.

[0250] Disadvantages: power supply coal consumption and plant power consumption are higher than the benchmark (+5%, +11.76%), reflecting higher fuel consumption and plant power consumption costs, which need to be optimized.

[0251] B power plant

[0252] Advantages: power supply coal consumption (-3.33%) and plant power consumption (-5.88%) are better than the benchmark, reflecting lower fuel consumption and power efficiency, which is the core reason for fuel efficiency.

[0253] Disadvantages: heat value difference is higher than the benchmark (+10%), which may indicate that fuel energy is not fully utilized, and attention should be paid to fuel quality or combustion efficiency.

[0254] C power plant

[0255] Advantages: heat value difference is better than the benchmark (-10%), fuel energy utilization efficiency is medium.

[0256] Disadvantages: power supply coal consumption (+1.67%) and plant power consumption (+3.53%) are slightly higher than the benchmark, although the gap is small, there is still room for improvement, and fine management is needed to reduce energy consumption.

[0257] D power plant

[0258] Advantages: The heat value difference is significantly better than the benchmark (-30%), which is the best among the four power plants, and the fuel energy utilization efficiency is the best.

[0259] Disadvantages: The power supply coal consumption (+6.67%) and the plant power rate (+8.24%) are the highest among the four plants, far exceeding the benchmark, which is the main reason for the lowest fuel efficiency, and the energy consumption management needs to be systematically improved.

[0260] Step 6: Continuous improvement and optimization

[0261] Improvement measures

[0262] According to the results of the benchmark data analysis, management makes different improvement suggestions for high-score power plants, midstream power plants, and low-score power plants:

[0263] A / D power plant (low-score power plant):

[0264] Prioritize optimizing power supply coal consumption (such as boiler combustion efficiency optimization) and plant power rate (such as eliminating high-energy-consuming equipment), as the improvement of these two indicators has the most significant impact on the total score.

[0265] Maintain the advantage of heat value difference (D power plant) or further improve (A power plant), but pay attention to the low weight, and the improvement priority is lower than the first two.

[0266] B power plant (high-score power plant):

[0267] Pay attention to the fluctuation of heat value difference to avoid the rebound of coal consumption due to the decline of coal quality, and establish a real-time monitoring mechanism for coal heat value.

[0268] C power plant (midstream power plant):

[0269] For the slightly over-standard plant power rate (+3.53%), carry out equipment energy efficiency diagnosis to achieve the benchmark level to improve the total score.

[0270] Model optimization

[0271] Due to the method of reducing human intervention, the system automatically updates the indicator weight and benchmark data analysis results based on the latest indicators and data, and continuously optimizes the model.

[0272] This embodiment realizes intelligent horizontal benchmark data analysis of thermal power plant production indicators through the patent technical scheme of "data collection - dimensionality reduction clustering - data analysis - continuous optimization". Experiments prove that this method can accurately locate key problems such as high energy consumption and low reliability, provide data-driven decision support for group-level production management, and significantly improve energy utilization efficiency and equipment management level.

[0273] Example 3

[0274] The embodiment of the present application also provides a data processing system using the energy enterprise data processing method based on multi-dimensional clustering, comprising:

[0275] The energy enterprise data acquisition unit, the principal component analysis unit, the K-means clustering unit and the comprehensive benchmarking analysis unit;

[0276] The energy enterprise data acquisition unit is used for acquiring various types of data of the energy enterprise, outputting a standardized data set, and constructing an index dictionary;

[0277] The principal component analysis unit is used for performing principal component analysis and combining with the business rules after dynamic characteristic weighting, and outputting a reduced dimension data matrix;

[0278] The K-means clustering unit is used for performing K-means clustering analysis and outputting classified clusters;

[0279] The comprehensive benchmarking analysis unit is used for performing scene mapping, dynamically configuring multi-benchmark values, performing comprehensive benchmarking analysis, dividing early warning levels, and completing energy enterprise data processing.

[0280] The present disclosure can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.

[0281] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a magneto-optical storage device, or any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0282] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0283] Computer readable program instructions for carrying out operations of the present disclosure can be assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computing device, partly on the user's computing device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device, for example, through the Internet using an Internet Service Provider. In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0284] Finally, it should be noted that the above-mentioned embodiments are merely intended for describing and illustrating, not limiting the technical solutions of the present application. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the specific embodiments of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the present application, and any modification or equivalent replacement without departing from the spirit and scope of the present application should be covered within the protection scope of the claims of the present application.

Claims

1. A method for processing data of an energy enterprise based on multi-dimensional clustering, characterized in that, Comprising: Step 1, collecting various types of data of energy enterprises, outputting standardized data sets, and constructing an index dictionary; Step 2, based on the index dictionary of step 1, performing principal component analysis, and combining with the business rules after dynamic characteristic weighting, outputting the reduced data matrix; Step 3, based on the reduced data matrix of step 2, performing K-means clustering analysis, and outputting the classified clusters; Step 4, based on the classified clusters output by step 3, performing scene mapping, dynamically configuring multi-benchmark values, and performing comprehensive benchmarking analysis to divide the warning level and complete the energy enterprise data processing.

2. The multi-dimensional clustering based energy enterprise data processing method according to claim 1, wherein, Also comprising: Step 5, based on the benchmarking analysis results in step 4, performing index visualization display; Step 6, based on the display results in step 5, performing reason tracing and improvement, and continuous optimization.

3. The multi-dimensional clustering-based energy enterprise data processing method according to claim 1, characterized in that: In step 1, the standardized data calculation formula in the standardized data set is: wherein, is the standardized jthindex data of the ithenergy enterprise, x ij is the jthindex data of the ithenergy enterprise, μ j is the mean of the jthindex, σ j is the standard deviation of the jthindex.

4. The multi-dimensional clustering-based energy enterprise data processing method according to claim 1, characterized in that: The business rules in step 2 include: Rule R1: If "power supply coal consumption > industry benchmark value 10%", trigger "high energy consumption risk", increase the importance evidence weight of coal consumption index in efficiency dimension +0.3; Rule R2: If "equipment unplanned downtime frequency monthly increase >20%", trigger "reliability deterioration", increase the importance evidence weight of downtime coefficient in reliability dimension +0.4; Rule R3: If "power generation capacity continuous 3 months achievement rate >110%", trigger "overcapacity warning", decrease the importance evidence weight of power generation capacity in capacity dimension -0.

2.

5. The multi-dimensional clustering-based energy enterprise data processing method according to claim 4, characterized in that: When the index value meets the business rules, the corresponding evidence event is automatically triggered to realize dynamic characteristic weighting: wherein, is the dynamic weight of the jth indicator of the energy enterprise in the time window t, b j is the basic importance weight of the jth indicator of the energy enterprise set by the expert, is the cumulative evidence weight of the jth indicator of the energy enterprise in the time window t, is the improved dynamic evidence strength coefficient, λ is a compensation coefficient, C j is the criticality grade of the jth indicator of the energy enterprise.

6. The multi-dimensional clustering-based energy enterprise data processing method according to claim 5, characterized in that: The improved dynamic evidence intensity coefficient is: Wherein, β is the basic intensity coefficient, γ is the attenuation factor, and k is the curvature factor.

7. The multi-dimensional clustering-based energy enterprise data processing method according to claim 1, characterized in that: In step 3, in the K-means clustering analysis, the initial centroid is selected based on four business dimensions, including production efficiency, energy consumption cost, equipment reliability, and environmental compliance.

8. The multi-dimensional clustering-based energy enterprise data processing method according to claim 1, characterized in that: In step 4, when dynamically configuring multi-benchmark values, data drift detection is adopted to calculate the KL divergence value of index distribution. If the KL divergence value exceeds the set threshold, it is determined that there is data drift.

9. The multi-dimensional clustering-based energy enterprise data processing method according to claim 1, characterized in that: In step 4, the warning level is divided by the absolute deviation rate, and the absolute deviation rate calculation formula is: The warning levels include very serious, serious, abnormal good, and normal.

10. An energy enterprise data processing system employing the multi-dimensional clustering based energy enterprise data processing method according to any one of claims 1-9, characterized in that, Comprising: The energy enterprise data acquisition unit, the principal component analysis unit, the K-means clustering unit and the comprehensive benchmarking analysis unit; The energy enterprise data acquisition unit is configured to acquire various types of data of an energy enterprise, output a standardized data set, and construct an index dictionary; The principal component analysis unit is configured to perform principal component analysis and combine the business rules weighted by the dynamic characteristics to output a data matrix after dimension reduction; The K-means clustering unit is configured to perform K-means clustering analysis and output clusters after classification; The comprehensive benchmarking analysis unit is configured to perform scene mapping, dynamically configure multiple benchmark values, perform comprehensive benchmarking analysis, divide early warning levels, and complete energy enterprise data processing.