Hydroelectric energy consumption data monitoring method and system based on edge calculation
By clustering water meters and electricity meters on edge nodes and training an isolated forest model, the problems of edge node computing power limitations and device mode differences are solved, achieving efficient and accurate detection of water and electricity energy consumption anomalies and optimizing the real-time performance and reliability of energy consumption management.
Patent Information
- Application Number
- CN202511484391.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-10-17
AI Technical Summary
Existing edge nodes have limited computing power, making it difficult to deploy complex models. Furthermore, the isolated forest algorithm does not fully consider the differences in device modes in multi-device scenarios, resulting in low efficiency and accuracy in detecting anomalies in water and electricity meters.
By clustering historical data of water and electricity meters within the industrial area, forming clusters, and training an isolated forest model on the edge nodes of each cluster, the model weights are updated to adapt to the latest dataset. Hierarchical clustering and average connectivity are used to group the data, thereby improving the model's adaptability and accuracy.
It improves the efficiency and accuracy of detecting abnormal hydropower energy consumption, enhances the model's generalization ability in multi-device scenarios, and ensures the real-time performance and reliability of energy consumption management in industrial areas.
Smart Images

Figure CN120951019A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial energy consumption data monitoring. In particular, it relates to a method and system for monitoring hydropower energy consumption data based on edge computing. Background Technology
[0002] With the application of smart meters, monitoring hydropower energy consumption data is crucial for energy management and fault early warning. Smart meters can collect and transmit hydropower energy consumption data in real time, providing rich data support for energy management and helping to achieve energy conservation, emission reduction, and optimal resource allocation.
[0003] In the monitoring of water and electricity meter data, edge computing technology reduces the latency of data transmission to the cloud or data center by offloading data processing to edge nodes, thereby improving the system's real-time performance and response speed. Edge computing nodes can perform preliminary processing and analysis of water and electricity consumption data locally, quickly identify anomalies and issue early warnings, thus improving the efficiency and reliability of energy management.
[0004] Although edge computing technology can effectively reduce data transmission latency, the limited computing power of existing edge nodes makes it difficult to deploy complex models on edge nodes. This makes it difficult to balance the versatility and accuracy of models in practical applications. In addition, although the isolated forest algorithm is suitable for anomaly detection, it is prone to misjudgment in multi-device scenarios because it does not fully consider the differences in device modes, thus limiting the efficiency and accuracy of anomaly detection for water and electricity meters. Summary of the Invention
[0005] To address the problem that the limited computing power of edge nodes hinders the deployment of complex models, and that the isolated forest algorithm is prone to misjudgment in multi-device scenarios due to insufficient consideration of the differences in modes between devices, thus affecting the efficiency and accuracy of anomaly detection in water and electricity meters, this invention provides solutions in the following aspects.
[0006] In the first aspect, the edge computing-based method for monitoring water and electricity energy consumption data includes: acquiring historical data from multiple water meters and electricity meters to be monitored within an industrial area and preprocessing them to obtain vectors for the water meters and electricity meters to be monitored; clustering all vectors for the water meters and electricity meters to be monitored within the area to obtain clustering results, wherein the clustering results include water meter clusters and electricity meter clusters, with each cluster corresponding to an edge node; using the historical water energy consumption data of all water meters or the historical electricity energy consumption data of electricity meters within the cluster to form training sets for the corresponding clusters, training an isolated forest model, and deploying the trained isolated forest model to the edge nodes of the corresponding clusters; obtaining the similarity between the energy consumption dataset of a preset time period before the current moment and the subset of samples corresponding to all trees in the isolated forest, and updating the weights of the corresponding isolated trees; inputting the real-time collected energy consumption data of the water meters and electricity meters to be monitored into the updated isolated forest model of the corresponding edge nodes, calculating the overall anomaly score of the updated isolated forest model, and evaluating whether the energy consumption data of the water meters and electricity meters to be monitored are abnormal.
[0007] By clustering historical data from multiple water and electricity meters within an industrial area and deploying a trained Isolation Forest model on the edge nodes of each cluster, efficient monitoring of water and electricity consumption data was achieved. Updating the model weights to reflect the similarity between the latest energy consumption dataset and the subset used during model training not only improved the model's adaptability to new data and the accuracy of anomaly detection but also enhanced its generalization ability in multi-device scenarios. This effectively improved the efficiency and accuracy of water and electricity consumption anomaly detection, ensuring the real-time performance and reliability of energy consumption management in industrial areas.
[0008] Preferably, obtaining the clustering results includes: Hierarchical clustering is used to treat each water meter as an independent cluster. The distance between each cluster is calculated pairwise. The average connection method is used to take the average of the Euclidean distances between all samples in one cluster and all samples in another cluster as the distance between the two clusters. The two clusters with the smallest distance are found and merged into a new cluster. This process is repeated until the minimum distance between all clusters is greater than a preset distance threshold, and then the clustering result corresponding to the water meter is obtained. Each electricity meter is treated as an independent cluster. The distance between each cluster is calculated pairwise. The average connection method is used, and the average of the Euclidean distances between all samples in one cluster and all samples in another cluster is taken as the distance between the two clusters. The two clusters with the smallest distance are found and merged into a new cluster. This process is repeated until the minimum distance between all clusters is greater than a preset distance threshold, and then the clustering result corresponding to the electricity meter is obtained.
[0009] By clustering water and electricity meters using hierarchical clustering and average connectivity, similar water and electricity meters can be effectively grouped based on the characteristics of their energy consumption data, forming clusters with similar energy consumption patterns. The reasonableness and accuracy of the clustering results are ensured by calculating the distance between clusters and merging the closest clusters until the minimum distance between clusters exceeds a preset threshold. The final clustering results contribute to more refined energy consumption management and anomaly detection.
[0010] Preferably, the process of training the isolated forest model includes: Based on the clustering results, all historical water energy consumption time series data and all historical electricity energy consumption time series data within a single water meter cluster are extracted and used to form corresponding training sets. Each sample represents multiple energy consumption data points of each water meter or electricity meter at different times, and each sample is a single feature vector. Calculate the local density of each sample in the training set, and randomly sample a predetermined number of samples from the training set without replacement to generate a subset of samples. Construct an isolation tree from the subset of samples. The isolated trees are combined into an isolated forest model. The discrimination of each isolated tree is calculated to evaluate its ability to separate sample patterns. The discrimination of each isolated tree is normalized and used as the initial weight of each isolated tree to complete the training of the isolated forest model.
[0011] By extracting historical data from individual clusters and assembling them into a training set, the training data for the model is ensured to have high relevance and consistency. Calculating local density and constructing isolated trees helps to assess and quantify anomalous behavior in the data, while normalizing the discriminative power as the initial weights of the trees enhances the model's ability to identify different cluster characteristics. This effectively distinguishes between normal and anomalous energy consumption patterns, thereby improving the accuracy and efficiency of hydropower energy consumption anomaly detection.
[0012] Preferably, the method for calculating the local density includes: Take any sample as the target sample, select a preset number of nearest neighbor samples, calculate the average Euclidean distance from the target sample to each nearest neighbor sample, and if the average Euclidean distance is greater than or equal to a preset minimum threshold, then the reciprocal of the average Euclidean distance is used as the local density of the target sample. If the average Euclidean distance is less than the preset minimum threshold, then the preset minimum threshold is used as the average Euclidean distance, and the reciprocal is used as the local density of the target sample.
[0013] Preferably, constructing an isolation tree for the subsample sets includes: The local density and mean local density of all samples in the subset are calculated, and the sample set is divided. The splitting feature value is selected from the feature value range of the split sample set according to a preset probability ratio. The sample is divided into smaller subsets according to recursive splitting until all leaf nodes have only one sample or the preset maximum depth is reached, then the splitting stops and the construction of the isolated tree is completed. The steps for partitioning the sample set include: if the local density is less than or equal to the local density mean, it is partitioned into a low-density sample set; otherwise, if it is greater than the local density mean, it is partitioned into a high-density sample set.
[0014] Preferably, the distinguishability of the isolated tree includes: Using any isolated tree as the target isolated tree, calculate the ratio of the total water energy consumption of all samples in the low-density sample set to the number of samples based on the target isolated tree. The feature mean of the low-density sample set and the feature mean of the high-density sample set are obtained. The absolute value of the difference between the two feature means is used as the discriminant of the target isolated tree.
[0015] Preferably, updating the weights of the corresponding isolated tree includes: Taking any edge node as the target node, obtain the energy consumption dataset of all water meters or electricity meters within the target node within a preset time period, calculate the average Euclidean distance between the energy consumption dataset and each subset of samples of any tree in the corresponding isolated forest, and perform normalization processing. Then, normalize the product between the normalized result and the initial weights again to use as the weights of the updated isolated tree.
[0016] By dynamically adjusting the weights of each isolated tree in the isolated forest model, the similarity between the latest energy consumption dataset of water or electricity meters within the edge nodes and the subset used during model training is reflected. A similarity metric is obtained by calculating and normalizing the average Euclidean distance between the energy consumption dataset and the subset, and the weights of the isolated trees are then updated accordingly. This approach adapts to changes in data distribution, improving the model's adaptability to current data patterns and the accuracy of anomaly detection. This enhances the model's flexibility and robustness in practical applications, ensuring more accurate and reliable anomaly detection results.
[0017] Preferably, the overall anomaly score of the updated isolated forest model includes: Traverse each isolated tree in the isolated forest, obtain the path length of the real-time energy consumption data sample in the corresponding isolated tree and the average path length of all samples in each isolated tree, calculate the anomaly score of the real-time energy consumption data sample in each isolated tree, and sum the products of the anomaly scores of all isolated trees and the corresponding updated weights to obtain the overall anomaly score of the real-time energy consumption data.
[0018] Preferably, the methods for obtaining the vectors of the water meter to be monitored and the vectors of the electricity meter to be monitored include: The historical data of the water meters and electricity meters to be monitored were cleaned separately to remove duplicate and invalid values, fill in missing data, and align them according to the time axis. Feature vectors are extracted from the processed historical data of the water meters to be monitored: [average daily water consumption, standard deviation of daily water consumption, maximum hourly water consumption]; feature vectors are extracted from the processed historical data of the electricity meters to be monitored: [average daily electricity consumption, standard deviation of daily electricity consumption, maximum hourly electricity consumption]; and each feature vector is standardized; one water meter vector corresponds to one water meter, and one electricity meter vector corresponds to one electricity meter.
[0019] Secondly, a hydropower energy consumption data monitoring system based on edge computing includes a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned hydropower energy consumption data monitoring method based on edge computing is implemented.
[0020] The present invention has the following effects: 1. This invention trains isolated forest models in each water meter cluster and electricity meter cluster separately, and deploys the models to the corresponding edge nodes according to the clustering results. This enables the models to better adapt to the energy consumption patterns of specific equipment clusters. The weights of the isolated trees are updated by calculating the similarity between the new dataset and the subsample set, which further improves the adaptability of the models to the current data patterns and thus improves the accuracy of anomaly detection in water and electricity energy consumption data.
[0021] 2. This invention leverages the advantages of edge computing by deploying and running an isolated forest model on edge nodes, reducing data transmission latency and alleviating the computational burden on the central server. Simultaneously, through hierarchical clustering and average connectivity methods, it effectively processes historical data from multiple water and electricity meters to be monitored, making model training and deployment more efficient and optimizing the utilization of edge computing resources. Attached Figure Description
[0022] Figure 1 This is a flowchart of steps S1-S5 in the edge computing-based hydropower energy consumption data monitoring method of this invention.
[0023] Figure 2 This is a structural block diagram of the hydropower energy consumption data monitoring system based on edge computing, according to an embodiment of the present invention. Detailed Implementation
[0024] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0025] Reference Figure 1 The edge computing-based hydropower energy consumption data monitoring method includes steps S1-S5, as follows: S1: Obtain historical data of multiple water meters and electricity meters to be monitored within the industrial area and preprocess them to obtain the vectors of the water meters and electricity meters to be monitored.
[0026] For example, the data collection frequency for the water meters and electricity meters to be monitored is 5 minutes / time, the water energy consumption unit is cubic meters, and the electricity energy consumption unit is kilowatt-hours. The historical data of the industrial area recorded in the cloud includes data from the most recent 30 days. The historical data of the water meters and electricity meters to be monitored are cleaned to remove duplicate and invalid values. A three-dimensional feature vector [average daily water consumption, standard deviation of daily water consumption, maximum hourly water consumption] is extracted from the cleaned historical data of the water meters to be monitored. The average daily water consumption is the average of the total daily water consumption over 30 days, the standard deviation of daily water consumption is the standard deviation of daily water consumption over 30 days, and the maximum hourly water consumption is the maximum water consumption in all 1-hour periods over 30 days. A three-dimensional feature vector [average daily electricity consumption, standard deviation of daily electricity consumption, maximum hourly electricity consumption] is extracted from the cleaned electricity meter data. The average daily electricity consumption is the average of the total daily electricity consumption over 30 days, the standard deviation of daily electricity consumption is the standard deviation of daily electricity consumption over 30 days, and the maximum hourly electricity consumption is the maximum electricity consumption in all 1-hour periods over 30 days.
[0027] The three-dimensional feature vectors are standardized using the Z-score (Z-Standard Score) standardization method to obtain standardized feature vectors, where one three-dimensional feature vector corresponds to a water meter or electricity meter.
[0028] S2: Cluster all water meter vectors and electricity meter vectors to be monitored within the region to obtain clustering results. The clustering results include water meter clusters and electricity meter clusters, with each cluster corresponding to an edge node.
[0029] Hierarchical clustering is used to treat each water meter as an independent cluster. The distance between each cluster is calculated pairwise. The average connection method is used to take the average of the Euclidean distances between all samples in one cluster and all samples in another cluster as the distance between the two clusters. The two clusters with the smallest distance are found and merged into a new cluster. This process is repeated until the minimum distance between all clusters is greater than a preset distance threshold, and then the clustering result corresponding to the water meter is obtained. Each electricity meter is treated as an independent cluster. The distance between each cluster is calculated pairwise. The average connection method is used, and the average of the Euclidean distances between all samples in one cluster and all samples in another cluster is taken as the distance between the two clusters. The two clusters with the smallest distance are found and merged into a new cluster. This process is repeated until the minimum distance between all clusters is greater than a preset distance threshold, and then the clustering result corresponding to the electricity meter is obtained.
[0030] For example, the preset distance threshold is 0.3, which can be adjusted according to specific circumstances.
[0031] If the average Euclidean distance between two calculated water meter clusters drops to 0.25 or below, the two clusters will be merged into one water meter cluster. The purpose of this process is to ensure that the members within each cluster in the clustering results are highly similar, so that the subsequent training of the isolated forest model can more accurately reflect the energy consumption pattern of a specific water meter cluster.
[0032] Specifically, the mean Euclidean distance is a key metric for measuring the similarity between two clusters. It is obtained by calculating the average Euclidean distance between all samples in one cluster and all samples in another cluster. Euclidean distance is a method for measuring the actual distance between two points in multidimensional space, and its calculation formula is the square root of the sum of the squares of the differences in the coordinates of the two points in each dimension. In this invention, the preset distance threshold of 0.25 is an example value and can be adjusted according to the actual application scenario and data characteristics.
[0033] When the average Euclidean distance between two water meter clusters is less than or equal to 0.25, it indicates that the two clusters are relatively close in the feature space and have similar water use patterns. In this case, merging the two clusters into one cluster allows the isolated forest model to focus more on learning the energy consumption characteristics of this specific water use pattern during training, thereby improving the model's accuracy in detecting abnormal water use behavior.
[0034] Furthermore, merging similar water meter clusters can reduce model complexity, improve generalization ability, and avoid overfitting. This is particularly important for models deployed on edge nodes, where computational resources are relatively limited, requiring a simplified model structure while maintaining performance.
[0035] In summary, by merging water meter clusters with an average Euclidean distance reduced to 0.25 or below during hierarchical clustering, the method described in this invention can generate more reasonable and accurate clustering results, providing a solid foundation for subsequent isolated forest model training and energy consumption data anomaly detection. This not only improves the accuracy and reliability of the model but also optimizes the utilization of edge computing resources.
[0036] S3: Take the historical water energy consumption data of all water meters or the historical electricity energy consumption data of electricity meters in the corresponding cluster and form the training set of the corresponding cluster. Train the isolated forest model and deploy the trained isolated forest model to the edge node of the corresponding cluster.
[0037] The process of training an isolation forest model includes: Based on the clustering results, all historical water energy consumption time series data and all historical electricity energy consumption time series data within a single water meter cluster are extracted and used to form corresponding training sets. Each sample represents multiple energy consumption data points of each water meter or electricity meter at different times, and each sample is a single feature vector. Calculate the local density of each sample in the training set, and randomly sample a predetermined number of samples from the training set without replacement to generate a subset of samples. Construct an isolation tree from the subset of samples. Methods for calculating local density include: Take any sample as the target sample, select a preset number of nearest neighbor samples, calculate the average Euclidean distance from the target sample to each nearest neighbor sample, and if the average Euclidean distance is greater than or equal to a preset minimum threshold, then the reciprocal of the average Euclidean distance is used as the local density of the target sample. If the average Euclidean distance is less than the preset minimum threshold, then the preset minimum threshold is used as the average Euclidean distance, and the reciprocal is used as the local density of the target sample.
[0038] Calculating local density is beneficial for identifying outliers in a dataset. In isolated forests, outliers are typically those points that exhibit significantly different characteristics compared to most other data points. By calculating local density, these outliers can be identified because they tend to have low local density values. This enhances the model's adaptability and robustness to different data distributions, optimizes the utilization of limited computing resources in edge computing environments, and enables the model to process dynamically changing data more efficiently.
[0039] For example, eight neighboring samples are selected for analysis, but this can be adjusted according to specific circumstances.
[0040] Specifically, the local density satisfies the following relationship: ; in, It refers to the first Local density of a sample It refers to the first One sample to In the nearest neighbor sample, the first Euclidean distance of each sample This indicates the total number of nearest neighbor samples. In this embodiment, the minimum threshold is represented. This ensures that when calculating the local density of samples, the characteristic of high clustering of samples with the same value can be preserved, while effectively avoiding numerical calculation anomalies.
[0041] The local density and mean local density of all samples in the subset are calculated, and the sample set is divided. The splitting feature value is selected from the feature value range of the split sample set according to a preset probability ratio. The sample is divided into smaller subsets according to recursive splitting until all leaf nodes have only one sample or the preset maximum depth is reached, then the splitting stops and the construction of the isolated tree is completed. The steps for partitioning the sample set include: if the local density is less than or equal to the local density mean, it is partitioned into a low-density sample set; otherwise, if it is greater than the local density mean, it is partitioned into a high-density sample set.
[0042] In other words, based on the divided low-density sample set and high-density sample set, when determining the segmentation feature value, it can be selected according to a preset probability distribution. For example, a value can be randomly selected from the feature value range of the low-density sample set with a 70% probability, and a value can be randomly selected from the feature value range of the high-density sample set with a 30% probability.
[0043] By distinguishing between low-density and high-density sample sets, the distribution characteristics of the data can be reflected more accurately. Low-density regions typically contain outliers or sparse data areas, while high-density regions contain more normal data points. Choosing the splitting feature value is a crucial step when constructing an isolation tree. By selecting the splitting feature value based on the density distribution of the sample set, outliers can be isolated more effectively, improving the model's anomaly detection capability.
[0044] By prioritizing feature values from low-density regions as segmentation points, outliers can be isolated from normal data more quickly, thereby improving the accuracy of anomaly detection. In cases of uneven data distribution, this method can reduce invalid segmentation during model training, thereby improving training efficiency. By selecting segmentation feature values more precisely, it can reduce the occurrence of misclassifying normal data as anomalies and the failure to detect actual anomalies or missed detections.
[0045] The isolated trees are combined into an isolated forest model. The discrimination of each isolated tree is calculated to evaluate its ability to separate sample patterns. The discrimination of each isolated tree is normalized and used as the initial weight of each isolated tree to complete the training of the isolated forest model.
[0046] Using any isolated tree as the target isolated tree, calculate the ratio of the total water energy consumption of all samples in the low-density sample set to the number of samples based on the target isolated tree. The feature mean of the low-density sample set and the feature mean of the high-density sample set are obtained. The absolute value of the difference between the two feature means is used as the discriminant of the target isolated tree.
[0047] Specifically, the discrimination index satisfies the following relationship: ; In the formula, The mean difference, or discriminant, of an isolated tree reflects its ability to separate high-density and low-density sample sets. This indicates the number of samples in a low-density sample set. Indicates the first in the low-density sample set Water energy consumption value for each sample This indicates the number of samples in a high-density sample set. Indicates the first high-density sample set Water energy consumption value for each sample and These refer to the feature means of the low-density sample set and the high-density sample set, respectively. Mean difference. The larger the value, the more significant the difference in features between the normal pattern (high density) and the potentially abnormal pattern (low density) in the tree, and the higher the tree's discriminative power.
[0048] S4: Obtain the similarity between the energy consumption dataset for the preset time period before the current moment and the corresponding subset of all trees in the isolated forest, and update the weight of the corresponding isolated tree.
[0049] Taking any edge node as the target node, obtain the energy consumption dataset of all water meters or electricity meters within the target node within a preset time period, calculate the average Euclidean distance between the energy consumption dataset and each subset of samples of any tree in the corresponding isolated forest, and perform normalization processing. Then, normalize the product between the normalized result and the initial weights again to use as the weights of the updated isolated tree.
[0050] For example, the preset time period is the three days prior to the current time, which can be adjusted according to specific circumstances.
[0051] Specifically, the weights of the updated isolated tree satisfy the following relationship: ; In the formula, Indicates the first The updated weights of the isolated trees Indicates the first The initial weights of the isolated trees, This represents the energy consumption dataset within a preset time period prior to the current moment. In the isolated forest Subset of trees The average Euclidean distance, This represents the energy consumption dataset within a preset time period prior to the current moment. With the first in the isolated forest The normalized similarity of subsets of trees. A smaller average Euclidean distance in the energy consumption dataset indicates a higher similarity, corresponding to higher weights for trained isolated trees, thus enhancing their influence on the current anomaly detection results. This represents the total number of isolated trees in an isolated forest. This represents the normalization function.
[0052] S5: Input the real-time energy consumption data of the water meters and electricity meters to be monitored into the updated isolated forest model of the corresponding edge nodes, calculate the overall anomaly score of the updated isolated forest model, and evaluate whether the energy consumption data of the water meters and electricity meters to be monitored are abnormal.
[0053] Traverse each isolated tree in the isolated forest, obtain the path length of the real-time energy consumption data sample in the corresponding isolated tree and the average path length of all samples in each isolated tree, calculate the anomaly score of the real-time energy consumption data sample in each isolated tree, and sum the products of the anomaly scores of all isolated trees and the corresponding updated weights to obtain the overall anomaly score of the real-time energy consumption data.
[0054] If the overall anomaly score is less than or equal to the preset anomaly threshold, the real-time energy consumption data is considered normal; otherwise, if it is greater than the preset anomaly threshold, the real-time energy consumption data is considered abnormal.
[0055] For example, the preset abnormal threshold is 0.6, which can be adjusted according to specific circumstances.
[0056] Specifically, the overall anomaly score satisfies the following relationship: ; In the formula, This represents the overall anomaly score of real-time energy consumption data. This represents the total number of isolated trees in an isolated forest. Indicates the first The updated weights of the isolated trees Indicates the sample in the th Path length in an isolated tree Indicates the first The average path length of all samples in an isolated tree.
[0057] Real-time energy consumption data from water and electricity meters is input into the updated Isolation Forest model on the corresponding edge nodes, enabling the model to adapt to the specific types and patterns of data at the edge nodes. By running the updated model on the edge nodes, the model's decision-making process is ensured to closely correspond to the data characteristics of the edge nodes, thereby improving the accuracy and efficiency of anomaly detection. Leveraging the advantages of edge computing reduces data transmission latency and allows the model to quickly respond to new data inputs, enabling real-time monitoring and analysis of energy consumption data, thereby optimizing energy management and improving system reliability.
[0058] This invention also provides a hydropower energy consumption data monitoring system based on edge computing. For example... Figure 2 As shown, the system includes a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement the edge computing-based hydropower energy consumption data monitoring method according to the first aspect of the present invention. The system also includes other components well known to those skilled in the art, such as a communication bus and a communication interface, the setup and functions of which are known in the art and will not be described further here.
[0059] It should be noted that those skilled in the art can make various modifications and improvements without departing from the inventive concept, and these all fall within the scope of protection of this invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A method for monitoring hydropower energy consumption data based on edge computing, characterized in that, include: Historical data of multiple water meters and electricity meters to be monitored within the industrial area were acquired and preprocessed to obtain the vectors of the water meters and electricity meters to be monitored. Cluster all water meter vectors and electricity meter vectors to be monitored within the region to obtain clustering results. The clustering results include water meter clusters and electricity meter clusters, with each cluster corresponding to an edge node. The historical water energy consumption data of all water meters or the historical electricity energy consumption data of electricity meters in the corresponding cluster are used to form the training set of the corresponding cluster, and the isolated forest model is trained. The trained isolated forest model is then deployed to the edge node of the corresponding cluster. Obtain the similarity between the energy consumption dataset for a preset time period before the current moment and the corresponding subset of all trees in the isolated forest, and update the weight of the corresponding isolated tree. The energy consumption data of the water meters and electricity meters to be monitored are collected in real time and input into the updated isolated forest model of the corresponding edge nodes. The overall anomaly score of the updated isolated forest model is calculated to evaluate whether the energy consumption data of the water meters and electricity meters to be monitored are abnormal.
2. The method for monitoring hydropower energy consumption data based on edge computing according to claim 1, characterized in that, Obtaining the clustering results includes: Hierarchical clustering is used to treat each water meter as an independent cluster. The distance between each cluster is calculated pairwise. The average connection method is used to take the average of the Euclidean distances between all samples in one cluster and all samples in another cluster as the distance between the two clusters. The two clusters with the smallest distance are found and merged into a new cluster. This process is repeated until the minimum distance between all clusters is greater than a preset distance threshold, and then the clustering result corresponding to the water meter is obtained. Each electricity meter is treated as an independent cluster. The distance between each cluster is calculated pairwise. The average connection method is used, and the average of the Euclidean distances between all samples in one cluster and all samples in another cluster is taken as the distance between the two clusters. The two clusters with the smallest distance are found and merged into a new cluster. This process is repeated until the minimum distance between all clusters is greater than a preset distance threshold, and then the clustering result corresponding to the electricity meter is obtained.
3. The method for monitoring hydropower energy consumption data based on edge computing according to claim 1, characterized in that, The process of training the isolated forest model includes: Based on the clustering results, all historical water energy consumption time series data and all historical electricity energy consumption time series data within a single water meter cluster are extracted and used to form corresponding training sets. Each sample represents multiple energy consumption data points of each water meter or electricity meter at different times, and each sample is a single feature vector. Calculate the local density of each sample in the training set, and randomly sample a predetermined number of samples from the training set without replacement to generate a subset of samples. Construct an isolation tree from the subset of samples. The isolated trees are combined into an isolated forest model. The discrimination of each isolated tree is calculated to evaluate its ability to separate sample patterns. The discrimination of each isolated tree is normalized and used as the initial weight of each isolated tree to complete the training of the isolated forest model.
4. The method for monitoring hydropower energy consumption data based on edge computing according to claim 3, characterized in that, The calculation methods for the local density include: Take any sample as the target sample, select a preset number of nearest neighbor samples, calculate the average Euclidean distance from the target sample to each nearest neighbor sample, and if the average Euclidean distance is greater than or equal to a preset minimum threshold, then the reciprocal of the average Euclidean distance is used as the local density of the target sample. If the average Euclidean distance is less than the preset minimum threshold, then the preset minimum threshold is used as the average Euclidean distance, and the reciprocal is used as the local density of the target sample.
5. The method for monitoring hydropower energy consumption data based on edge computing according to claim 3, characterized in that, The construction of the isolated tree for the paired sample sets includes: The local density and mean local density of all samples in the subset are calculated, and the sample set is divided. The splitting feature value is selected from the feature value range of the split sample set according to a preset probability ratio. The sample is divided into smaller subsets according to recursive splitting until all leaf nodes have only one sample or the preset maximum depth is reached, then the splitting stops and the construction of the isolated tree is completed. The steps for partitioning the sample set include: if the local density is less than or equal to the local density mean, it is partitioned into a low-density sample set; otherwise, if it is greater than the local density mean, it is partitioned into a high-density sample set.
6. The method for monitoring hydropower energy consumption data based on edge computing according to claim 3, characterized in that, The discriminative power of the isolated tree includes: Using any isolated tree as the target isolated tree, calculate the ratio of the total water energy consumption of all samples in the low-density sample set to the number of samples based on the target isolated tree. The feature mean of the low-density sample set and the feature mean of the high-density sample set are obtained. The absolute value of the difference between the two feature means is used as the discriminant of the target isolated tree.
7. The method for monitoring hydropower energy consumption data based on edge computing according to claim 1, characterized in that, The update of the weights corresponding to the isolated tree includes: Taking any edge node as the target node, obtain the energy consumption dataset of all water meters or electricity meters within the target node within a preset time period, calculate the average Euclidean distance between the energy consumption dataset and each subset of samples of any tree in the corresponding isolated forest, and perform normalization processing. Then, normalize the product between the normalized result and the initial weights again to use as the weights of the updated isolated tree.
8. The method for monitoring hydropower energy consumption data based on edge computing according to claim 1, characterized in that, The overall anomaly score of the updated isolated forest model includes: Traverse each isolated tree in the isolated forest, obtain the path length of the real-time energy consumption data sample in the corresponding isolated tree and the average path length of all samples in each isolated tree, calculate the anomaly score of the real-time energy consumption data sample in each isolated tree, and sum the products of the anomaly scores of all isolated trees and the corresponding updated weights to obtain the overall anomaly score of the real-time energy consumption data.
9. The method for monitoring hydropower energy consumption data based on edge computing according to claim 1, characterized in that, The methods for obtaining the vectors of the water meter and the electricity meter to be monitored include: The historical data of the water meters and electricity meters to be monitored were cleaned separately to remove duplicate and invalid values, fill in missing data, and align them according to the time axis. Feature vectors are extracted from the processed historical data of the water meters to be monitored: [average daily water consumption, standard deviation of daily water consumption, maximum hourly water consumption]; feature vectors are extracted from the processed historical data of the electricity meters to be monitored: [average daily electricity consumption, standard deviation of daily electricity consumption, maximum hourly electricity consumption]; and each feature vector is standardized; one water meter vector corresponds to one water meter, and one electricity meter vector corresponds to one electricity meter.
10. A hydropower energy consumption data monitoring system based on edge computing, characterized in that, include: A processor and a memory, wherein the memory stores computer program instructions that, when executed by the processor, implement the edge computing-based hydropower energy consumption data monitoring method according to any one of claims 1-9.
Citation Information
Patent Citations
A method and a device for generating a malicious software benchmark test set
CN109241740A
Unsupervised anomaly detection method based on traffic log feature extraction
CN114528909A
Clustering federated learning method oriented to heterogeneous statistics
CN115952860A
Power grid data acquisition method, device and equipment based on digital twinning and medium
CN116090614A
Image clustering method and system based on improved density peak
CN117115492A
Cited By
Electric meter terminal group fault detection method and system based on edge calculation
CN121348213A