Data analysis method and system based on cloud computing distributed data platform
By adopting a data analysis method of multi-node collaborative work on the cloud computing distributed data platform, the problem of inefficiency in traditional methods when dealing with large-scale power consumption data is solved, and high-precision and high-efficiency power consumption anomaly analysis is achieved.
Patent Information
- Application Number
- CN202510145669.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional data analysis methods for abnormal electricity consumption are difficult to quickly process large-scale data, resulting in low analysis efficiency and inability to detect potential electricity consumption abnormalities in time, and there are defects in the accuracy and comprehensiveness of data processing.
The data analysis method based on the cloud computing distributed data platform is adopted to store, filter, transform, feature analysis and abnormal identification through the collaborative work of multiple nodes, and the power consumption anomaly analysis is performed using clustering analysis, European-style distance formulas and fuzzy neural networks.
It improves the accuracy and efficiency of power consumption abnormality analysis, reduces analysis errors and costs, can promptly detect and deal with power consumption abnormalities, and supports real-time monitoring and decision-making of the power system.
Smart Images

Figure CN120067940A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and particularly to a data analysis method and system based on a cloud computing distributed data platform. Background Art
[0002] In today's digital age, power data has grown explosively. Analyzing abnormal electricity consumption behavior data is of crucial significance for the stable operation of the power system, the rational allocation of power resources, and the guarantee of user electricity consumption safety.
[0003] Traditional methods for analyzing abnormal electricity consumption behavior expose many deficiencies when faced with massive, high-dimensional, and complex and changeable power data. These methods often have difficulty quickly processing large-scale data, resulting in low analysis efficiency and being unable to detect potential electricity consumption anomalies in a timely manner, thus affecting the real-time monitoring and decision-making of the power system. In addition, traditional methods also have defects in the accuracy and comprehensiveness of data processing, and are prone to missing some key information, making the analysis of electricity consumption behavior inaccurate. For example, when dealing with data at the TB level or even PB level, the efficiency of traditional methods is very low. Summary of the Invention
[0004] The present invention provides a data analysis method and a computer-readable storage medium based on a cloud computing distributed data platform, and its main purpose is to reduce the error and cost of abnormal electricity consumption analysis and improve the accuracy of abnormal electricity consumption analysis.
[0005] To achieve the above object, a data analysis method based on a cloud computing distributed data platform provided by the present invention includes:
[0006] Collecting users' electricity consumption behavior data based on a preset data collection instruction, and transmitting the electricity consumption behavior data to a preset distributed data platform for storage using the distributed data platform to obtain an electricity consumption behavior data set, where the distributed data platform includes multiple nodes, and the multiple nodes include a first part of nodes, a second part of nodes, and a third part of nodes;
[0007] Performing a data filtering operation on the electricity consumption behavior data set to obtain a denoised electricity consumption behavior data set, and converting the denoised electricity consumption behavior data in the denoised electricity consumption behavior data set into a preset standard format to obtain a standard electricity consumption behavior data set, where the data filtering operation includes removing duplicate data and removing noise data;
[0008] According to the distributed data platform, constructing an abnormal electricity consumption feature analysis model using the standard electricity consumption behavior data set, and generating a typical load characteristic curve using the standard electricity consumption behavior data in the standard electricity consumption behavior data set;
[0009] Obtain the electricity consumption behavior data to be analyzed, and perform clustering analysis on the electricity consumption behavior data to be analyzed by using a preset clustering method according to the first part of nodes in the distributed data platform to obtain a user electricity consumption characteristic curve;
[0010] Calculate the similarity between the user electricity consumption characteristic curve and the typical load characteristic curve by using a preset Euclidean distance formula according to the second part of nodes in the distributed data platform to obtain an electricity consumption anomaly similarity value;
[0011] Identify the electricity consumption behavior anomaly characteristics in the electricity consumption behavior data to be analyzed by using the electricity consumption anomaly characteristic analysis model according to the third part of nodes in the distributed data platform, and perform electricity consumption anomaly analysis based on the electricity consumption behavior anomaly characteristics and the electricity consumption anomaly similarity value to obtain an electricity consumption anomaly analysis result.
[0012] Optionally, the performing clustering analysis on the electricity consumption behavior data to be analyzed by using a preset clustering method to obtain a user electricity consumption characteristic curve includes:
[0013] Obtain a preset self-organizing mapping network, and perform an initialization operation on the self-organizing mapping network to obtain an initialized self-organizing mapping network;
[0014] Perform initial clustering on the electricity consumption behavior data to be analyzed by using the initialized self-organizing mapping network to obtain self-organizing clustering centers;
[0015] Perform K-means clustering based on the self-organizing clustering centers and the electricity consumption behavior data to be analyzed to obtain a clustering result, and generate a curve according to the clustering result to obtain a user electricity consumption characteristic curve, where the electricity consumption behavior data to be analyzed includes a plurality of data points to be analyzed.
[0016] Optionally, the performing initial clustering on the electricity consumption behavior data to be analyzed by using the initialized self-organizing mapping network to obtain self-organizing clustering centers includes:
[0017] Obtain the clustering objects in the electricity consumption behavior data to be analyzed to obtain a plurality of characteristic curve clustering objects;
[0018] Obtain the weight vectors in the initialized self-organizing mapping network to obtain a plurality of network weight vectors;
[0019] Sequentially extract characteristic curve clustering objects from the plurality of characteristic curve clustering objects, and calculate the distances between the extracted characteristic curve clustering objects and the plurality of network weight vectors respectively to obtain a set of clustering distances;
[0020] Summarize the set of clustering distances to obtain a set of sets of clustering distances corresponding to the plurality of characteristic curve clustering objects;
[0021] Iterate the initialized self-organizing mapping network using the set of clustering distances to obtain a sub-optimal self-organizing mapping network, calculate the iterative clustering centers of the sub-optimal self-organizing mapping network, and obtain the self-organizing clustering centers.
[0022] Optionally, the calculating the similarity between the user's electricity consumption characteristic curve and the typical load characteristic curve using a preset Euclidean distance formula to obtain an electricity consumption anomaly similarity value includes:
[0023] Calculate the correlation coefficient between the user's electricity consumption characteristic curve and the typical load characteristic curve to obtain a target correlation coefficient;
[0024] Calculate the Euclidean distance between the user's characteristic curve and the typical load characteristic curve to obtain a curve Euclidean distance;
[0025] Based on the target correlation coefficient, the curve Euclidean distance, and a preset weight coefficient, calculate the similarity between the user's electricity consumption characteristic curve and the typical load characteristic curve to obtain an electricity consumption anomaly similarity value.
[0026] Optionally, the calculation formula of the target correlation coefficient is as follows:
[0027]
[0028] where p represents the target correlation coefficient, β i represents the i-th feature point in the user's electricity consumption characteristic curve, τ i represents the i-th feature point in the typical load characteristic curve, and respectively represent the average values of the feature points in the user's electricity consumption characteristic curve and the typical load characteristic curve, and n represents the number of feature points.
[0029] Optionally, the constructing an electricity consumption anomaly feature analysis model using the standard electricity consumption behavior dataset includes:
[0030] Reconstruct the standard electricity consumption behavior dataset to obtain a reconstructed electricity consumption behavior dataset, and obtain the abnormal features in the reconstructed electricity consumption behavior dataset to obtain electricity consumption abnormal features;
[0031] Obtain the standard electricity consumption behavior data corresponding to the electricity consumption abnormal features in the standard electricity consumption behavior dataset to obtain abnormal feature data;
[0032] Divide the electricity consumption abnormal features and the abnormal feature data into a training set and a test set to obtain a feature training set and a feature test set, train a preset fuzzy neural network using the feature training set, and test the trained fuzzy neural network using the feature test set to obtain an electricity consumption anomaly feature analysis model.
[0033] Optionally, perform K-means clustering based on the self-organizing clustering centers and the power consumption behavior data to be analyzed, obtain a clustering result, and generate a curve according to the clustering result to obtain a user power consumption characteristic curve, where the power consumption behavior data to be analyzed includes a plurality of data points to be analyzed, including:
[0034] Use the self-organizing clustering centers as the initial clustering centers of a preset K-means clustering algorithm to obtain K-means initial clustering centers, where the K-means initial clustering centers include a plurality of initial clustering centers;
[0035] Sequentially extract initial clustering centers from the plurality of initial clustering centers, and calculate the distances from the plurality of data points to be analyzed to the extracted initial clustering centers to obtain a K-means clustering distance group, summarize the K-means clustering distance groups to obtain a K-means clustering distance group set, where the K-means clustering distance group includes a plurality of clustering distances, and the clustering distance is the distance from the data point to be analyzed to the extracted initial clustering center, and the K-means clustering distance group set includes a plurality of K-means clustering distance groups;
[0036] According to the K-means clustering distance group set, divide each data point to be analyzed among the plurality of data points to be analyzed into the initial clustering center with the closest distance to obtain a plurality of K-means clustering clusters;
[0037] Sequentially extract K-means clustering clusters from the plurality of K-means clustering clusters, obtain the data mean points in the K-means clustering clusters, and use the data mean points to update the initial clustering centers of the extracted K-means clustering clusters to obtain updated clustering centers, calculate the distances between the updated clustering centers and the initial clustering centers corresponding to the K-means clustering clusters to obtain a clustering center change value, where the value of the data mean point is the average value of the numerical values corresponding to all data points to be analyzed;
[0038] Summarize the updated clustering centers to obtain an updated clustering center group, use the updated clustering center group as the plurality of initial clustering centers, and return to the step of sequentially extracting initial clustering centers from the plurality of initial clustering centers until the clustering center change value is less than a preset threshold, obtain the number of times of the step of returning to the step of sequentially extracting initial clustering centers from the plurality of initial clustering centers to obtain the iteration number, if it is confirmed that the iteration number is greater than a preset number, obtain the final clustering clusters;
[0039] Extract the representative features in the final clustering clusters to obtain cluster representative features, use time as the horizontal axis and the cluster representative features as the vertical axis to construct a user power consumption characteristic curve, where the cluster representative features are the features represented by the clustering centers of the final clustering clusters.
[0040] Optionally, generating a typical load characteristic curve by using the standard electricity consumption behavior data in the standard electricity consumption behavior dataset includes:
[0041] Obtaining service requirements based on the electricity consumption behavior data to obtain data analysis requirements;
[0042] Determining the data aggregation granularity by using the data analysis requirements to obtain the electricity consumption data aggregation granularity;
[0043] Calculating the standard electricity consumption behavior data according to the electricity consumption data aggregation granularity to obtain granularity adjustment data, and generating an image according to the granularity adjustment data to obtain a typical load characteristic curve.
[0044] Optionally, the calculation formulas for calculating the Euclidean distance between the user characteristic curve and the typical load characteristic curve and calculating the similarity between the user electricity consumption characteristic curve and the typical load characteristic curve are as follows:
[0045]
[0046] Among them, δ represents similarity, D * represents the Euclidean distance between the user characteristic curve and the typical load characteristic curve, W 1 and W 2 are weight coefficients.
[0047] To achieve the above object, the present invention also provides a data analysis system based on a cloud computing distributed data platform, including:
[0048] A data collection module, configured to collect the electricity consumption behavior data of users based on a preset data collection instruction, and transmit the electricity consumption behavior data to a preset distributed data platform for storage by using the distributed data platform to obtain an electricity consumption behavior dataset, where the distributed data platform includes multiple nodes, and the multiple nodes include a first part of nodes, a second part of nodes, and a third part of nodes;
[0049] A model construction module, configured to perform a data filtering operation on the electricity consumption behavior dataset to obtain a denoised electricity consumption behavior dataset, and convert the denoised electricity consumption behavior data in the denoised electricity consumption behavior dataset into a preset standard format to obtain a standard electricity consumption behavior dataset, where the data filtering operation includes removing duplicate data and removing noise data;
[0050] Constructing an electricity consumption anomaly feature analysis model by using the standard electricity consumption behavior dataset according to the distributed data platform, and generating a typical load characteristic curve by using the standard electricity consumption behavior data in the standard electricity consumption behavior dataset;
[0051] A similarity calculation module, configured to obtain power consumption behavior data to be analyzed, perform clustering analysis on the power consumption behavior data to be analyzed by using a preset clustering method according to the first part of nodes in the distributed data platform, and obtain a user power consumption feature curve;
[0052] According to the second part of nodes in the distributed data platform, calculate the similarity between the user power consumption feature curve and the typical load feature curve by using a preset Euclidean distance formula, and obtain a power consumption anomaly similarity value;
[0053] An anomaly analysis module, configured to identify power consumption behavior anomaly features in the power consumption behavior data to be analyzed by using the power consumption anomaly feature analysis model according to the third part of nodes in the distributed data platform, and perform power consumption anomaly analysis based on the power consumption behavior anomaly features and the power consumption anomaly similarity value, so as to obtain a power consumption anomaly analysis result.
[0054] To solve the above problems, the present invention further provides an electronic device, where the electronic device includes:
[0055] A memory, storing at least one instruction; and a processor, executing the instruction stored in the memory to implement the above-mentioned data analysis method based on a cloud computing distributed data platform.
[0056] To solve the above problems, the present invention further provides a computer-readable storage medium, where at least one instruction is stored in the computer-readable storage medium, and the at least one instruction is executed by a processor in an electronic device to implement the above-mentioned data analysis method based on a cloud computing distributed data platform.
[0057] To solve the problems described in the background art, the present invention collects the user's electricity consumption behavior data based on a preset data collection instruction, and transmits the electricity consumption behavior data to a preset distributed data platform for storage using multiple nodes of the distributed data platform to obtain an electricity consumption behavior data set. Among them, the distributed data platform includes multiple nodes, and the multiple nodes include a first part of nodes, a second part of nodes, and a third part of nodes. It can be seen that the present invention adopts a distributed storage method to ensure data security; perform a data filtering operation on the electricity consumption behavior data set to obtain a denoised electricity consumption behavior data set, and convert the data in the denoised electricity consumption behavior data set into a preset standard format to obtain a standard electricity consumption behavior data set. Among them, the data filtering operation includes removing duplicate data and removing noise data; according to the distributed data platform, use the standard electricity consumption behavior data set to construct an abnormal feature analysis model to obtain an electricity consumption abnormal feature analysis model, and use the standard electricity consumption behavior data in the standard electricity consumption behavior data set to generate image data to obtain a typical load characteristic curve; obtain the electricity consumption behavior data to be analyzed, and according to the first part of nodes in the distributed data platform, use a preset clustering method to perform clustering analysis on the electricity consumption behavior data to be analyzed to obtain a user electricity consumption characteristic curve. It can be seen that the present invention uses node division in the distributed data platform to work, which can effectively improve the processing speed; according to the second part of nodes in the distributed data platform, use a preset Euclidean distance formula to calculate the similarity between the user electricity consumption characteristic curve and the typical load characteristic curve to obtain an electricity consumption abnormal similarity value; according to the third part of nodes in the distributed data platform, use the electricity consumption abnormal feature analysis model to identify the abnormal features in the electricity consumption behavior data to be analyzed to obtain electricity consumption behavior abnormal features, and perform electricity consumption abnormal analysis based on the electricity consumption behavior abnormal features and the electricity consumption abnormal similarity value to obtain the analysis result of the electricity consumption behavior data to be analyzed, and obtain an electricity consumption abnormal analysis result. It can be seen that the present invention combines features and similarity values for abnormal analysis to improve the accuracy of abnormal analysis. Therefore, the present invention can reduce the error and cost of electricity consumption abnormal analysis and improve the accuracy of electricity consumption abnormal analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 FIG. is a schematic flow chart of a data analysis method based on a cloud computing distributed data platform provided by an embodiment of the present invention;
[0059] Figure 2 FIG. is a functional module diagram of a data analysis system based on a cloud computing distributed data platform provided by an embodiment of the present invention;
[0060] Figure 3 FIG. is a schematic structural diagram of an electronic device for implementing the data analysis method based on the cloud computing distributed data platform provided by an embodiment of the present invention.
[0061] Description of reference numerals:
[0062] 1. Electronic device; 10. Processor; 11. Memory; 12. Bus.
[0063] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0064] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0065] The embodiment of the present application provides a data analysis method based on a cloud computing distributed data platform. The execution subject of the data analysis method based on the cloud computing distributed data platform includes but is not limited to at least one of the electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the data analysis method based on the cloud computing distributed data platform can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc.
[0066] Reference Figure 1 FIG. 1 is a flow chart of a data analysis method based on a cloud computing distributed data platform provided by an embodiment of the present invention. In this embodiment, the data analysis method based on a cloud computing distributed data platform includes:
[0067] S1. Collect the user's electricity usage behavior data based on preset data collection instructions, and transmit the electricity usage behavior data to a preset distributed data platform, and use the distributed data platform to store it to obtain an electricity usage behavior data set, wherein the distributed data platform includes multiple nodes, and the multiple nodes include a first part of nodes, a second part of nodes, and a third part of nodes.
[0068] It is understandable that the data collection instruction is issued by a person engaged in the work related to the analysis of abnormal power consumption characteristics, and the person confirms the necessary environment for the analysis of abnormal power consumption characteristics, where the necessary environment is a big data collection environment that can obtain large-scale load characteristic curve data, reduce data missing, outliers and noise, and ensure data quality. The data collected in this environment will be used for the subsequent processing and analysis of power consumption behavior data to achieve the goals of screening abnormal power users and automatic analysis of abnormal power consumption characteristics, providing support for the safe and stable operation of the power system and the intelligent power management.
[0069] It should be understood that electricity consumption behavior data refers to the relevant data reflecting various behaviors and states of users during the electricity consumption process, which is of great value for the management, analysis of the power system and the optimization of user electricity consumption. Electricity consumption behavior data includes basic information data, electricity quantity and power data, electricity consumption time data, electricity quality data, cost-related data, etc.
[0070] It should be explained that by using multiple nodes of the distributed data platform for storage to obtain the electricity consumption behavior dataset, the reliability and security of the data can be improved by virtue of multi-node storage, preventing data loss caused by single-point failures; by leveraging the scalability of the distributed architecture, it can flexibly cope with the continuous growth of electricity consumption data volume; with the help of the parallel processing capabilities of multiple nodes, the read / write speed and processing efficiency of the data can be accelerated, providing a solid data foundation for subsequent electricity consumption behavior analysis, anomaly detection, pattern mining, etc., and strongly supporting the intelligent management and decision-making of the power system.
[0071] Furthermore, the distributed data platform is a system architecture that realizes the storage, processing and analysis of electricity consumption behavior data through the collaborative work of multiple nodes. The distributed data platform collects electricity consumption behavior data according to preset instructions and dispersedly stores it in multiple nodes, providing an operating environment and resource support for removing duplicate and noise data, constructing an anomaly feature analysis model, performing clustering analysis, calculating curve similarity, etc., and completing complex electricity consumption anomaly analysis tasks through the efficient cooperation between nodes, improving the reliability, efficiency and intelligent level of power system data processing.
[0072] It should be explained that by using multiple nodes of the distributed data platform for storage to obtain the electricity consumption behavior dataset, the reliability of the data can be improved by virtue of the redundant storage of multiple nodes, avoiding data loss caused by single-point failures.
[0073] It should be understood that electricity consumption behavior data can be collected through devices such as smart meters and sensors, or can be input by users independently. For example, users manually record and upload information such as the usage duration, start and stop times of specific electrical equipment on relevant power service APPs; or in the home energy management system, users actively fill in data such as the power and usage frequency of newly added electrical appliances. These data input by users independently can make up for the deficiencies of device-collected data, provide richer information for a more comprehensive and accurate analysis of user electricity consumption behavior, and help power companies formulate energy management plans and service strategies that better meet user needs.
[0074] S2. Perform a data filtering operation on the electricity consumption behavior dataset to obtain a denoised electricity consumption behavior dataset, and convert the denoised electricity consumption behavior data in the denoised electricity consumption behavior dataset into a preset standard format to obtain a standard electricity consumption behavior dataset, where the data filtering operation includes removing duplicate data and removing noise data.
[0075] It is understandable that removing duplicate data and noise data from the electricity consumption behavior dataset and converting them into a preset standard format can greatly improve the data quality. Removing duplicate data can avoid the waste of storage space caused by data redundancy and the interference with analysis results, ensuring the simplicity of data; clearing noise data can reduce the impact of outliers on the accuracy of data analysis, making the data more truly reflect the electricity consumption behavior pattern. Converting to the standard format unifies the data expression form, facilitating the integration and interaction of data from different sources and of different types, laying a foundation for the subsequent construction of an accurate abnormal feature analysis model, and effectively improving the reliability of analysis results and the scientific nature of decision-making.
[0076] Among them, duplicate data refers to data records in the electricity consumption behavior dataset that are completely the same or have some key information repeated. These duplicate data may be caused by data acquisition device failures, network transmission errors, or data storage problems. For example, when collecting the electricity consumption data of a user at a certain moment, due to sensor failures, the same electricity consumption value was recorded continuously for multiple times; noise data refers to data that does not conform to the actual electricity consumption behavior and significantly deviates from the normal range, which may be caused by interference during data acquisition (such as electromagnetic interference, environmental noise, etc.), data transmission errors, or abnormal electricity consumption events (such as abnormal power fluctuations caused by momentary electrical failures). For example, under normal electricity consumption conditions, the electricity consumption of a user usually fluctuates within a certain reasonable range, but if affected by strong electromagnetic interference, the electricity consumption data collected at a certain moment may show extremely large or extremely small abnormal values, which are noise data.
[0077] It should be explained that the preset standard format is a unified data representation form stipulated in advance before processing and analyzing the electricity consumption behavior data, covering data type specifications, such as the electricity consumption time being of the date and time type, and the power being of the numerical type, making the data more orderly during storage and operation; stipulating the data length, such as setting the user number to a fixed number of digits to avoid uneven data; unifying the data encoding to ensure that there is no garbled code during data transmission and processing between systems; determining the data structure, whether it is the table structure of a relational database or a non-relational structure such as JSON, can make data storage and call more efficient.
[0078] S3. According to the distributed data platform, use the standard electricity consumption behavior dataset to construct an electricity consumption abnormal feature analysis model, and use the standard electricity consumption behavior data in the standard electricity consumption behavior dataset to generate a typical load characteristic curve.
[0079] It is understandable that the abnormal electricity consumption feature analysis model aims to accurately identify and analyze abnormal electricity consumption patterns and features from electricity consumption behavior data. Based on the standard electricity consumption behavior dataset, with the powerful computing power of the distributed data platform, through learning and mining a large amount of historical electricity consumption data, the rules and patterns of normal electricity consumption behavior are discovered. In actual application, the abnormal electricity consumption feature analysis model can quickly compare the real-time collected electricity consumption data with the mastered normal patterns, and use specific algorithms to detect data points or trend changes that do not conform to the conventional patterns, so as to determine the abnormal electricity consumption features.
[0080] It should be explained that the typical load characteristic curve is a characteristic curve constructed based on a large amount of user electricity consumption behavior data. After these data are processed by deduplication, noise reduction and format standardization, clustering algorithms are used to extract typical features, which show the variation law of power load over time under a specific electricity consumption pattern. In the power system, it can be used to detect abnormal user electricity consumption situations, assist power enterprises in resource allocation and planning, and is of great significance for ensuring the stable operation and efficient management of the power system.
[0081] It should be understood that constructing the abnormal electricity consumption feature analysis model using the standard electricity consumption behavior dataset includes:
[0082] Reconstruct the standard electricity consumption behavior dataset to obtain a reconstructed electricity consumption behavior dataset, and obtain the abnormal features in the reconstructed electricity consumption behavior dataset to get the abnormal electricity consumption features;
[0083] Obtain the standard electricity consumption behavior data corresponding to the abnormal electricity consumption features in the standard electricity consumption behavior dataset to get the abnormal feature data;
[0084] Divide the abnormal electricity consumption features and the abnormal feature data into a training set and a test set to obtain a feature training set and a feature test set. Use the feature training set to train a preset fuzzy neural network, and use the feature test set to test the trained fuzzy neural network to obtain the abnormal electricity consumption feature analysis model.
[0085] It should be explained that reconstructing the standard electricity consumption behavior data can change the arrangement and combination methods of the original data to better mine the hidden information therein. This may include operations such as grouping, aggregating, and transforming the data, which can make the data structure more suitable for subsequent abnormal feature extraction work.
[0086] It is understandable that the abnormal electricity consumption characteristics refer to the data characteristics identified through data reconstruction and specific abnormal feature extraction methods during the in-depth analysis and processing of the standard electricity consumption behavior dataset, which are different from the normal electricity consumption behavior patterns. These characteristics are manifested as deviations from the normal electricity consumption rules in aspects such as time, electricity quantity, power, power factor, current, and the three-phase balance of voltage in electricity consumption data. It may be a single abnormal situation such as a sudden change in power at a certain moment, electricity consumption exceeding a reasonable range, or abnormal three-phase unbalance rate, or a combination of multiple abnormal situations.
[0087] It should be explained that the fuzzy neural network is an intelligent computing model that combines the advantages of fuzzy logic and neural networks. It combines the ability of fuzzy logic to process uncertain and fuzzy information and the self-learning and adaptive characteristics of neural networks. In the analysis of abnormal electricity consumption characteristics, the fuzzy neural network can handle the fuzziness and uncertainty of power data, realize the automatic analysis of abnormal electricity consumption, and provide support for the stable operation and management of the power system.
[0088] It should be understood that generating the typical load characteristic curve using the standard electricity consumption behavior data in the standard electricity consumption behavior dataset includes:
[0089] Obtaining the business requirements based on the electricity consumption behavior data to get the data analysis requirements;
[0090] Using the data analysis requirements to determine the data aggregation granularity to get the electricity consumption data aggregation granularity;
[0091] Calculating the standard electricity consumption behavior data according to the electricity consumption data aggregation granularity to get the granularity adjustment data, and generating an image based on the granularity adjustment data to obtain the typical load characteristic curve.
[0092] It should be explained that the data aggregation granularity is the degree of subdivision of dimensions such as time or space based on which data aggregation operations are performed during data processing and analysis. When processing electricity consumption behavior data, if considering the time dimension, aggregating by hour means combining the electricity consumption data within each hour for calculation, such as calculating the total electricity consumption and average power per hour. At this time, "hour" is the aggregation granularity; if aggregating by day, that is, integrating and processing the electricity consumption data within a day, "day" becomes the aggregation granularity. From the spatial dimension, for the electricity consumption data in different regions, aggregating by community unit, then "community" is the aggregation granularity. A smaller aggregation granularity can retain more data details and is suitable for analyzing short-term and fine electricity consumption changes; a larger aggregation granularity can present macroscopic trends and is conducive to grasping long-term and overall electricity consumption situations.
[0093] S4. Obtain the electricity consumption behavior data to be analyzed, and perform clustering analysis on the electricity consumption behavior data to be analyzed by using a preset clustering method according to the first part of nodes in the distributed data platform, so as to obtain the user electricity consumption characteristic curve.
[0094] It can be understood that after obtaining the electricity consumption behavior data to be analyzed, by means of the first part of nodes in the distributed data platform and the preset clustering method for clustering analysis and obtaining the user electricity consumption characteristic curve, representative features and rules can be extracted from the complex and diverse electricity consumption behavior data. Through clustering, users with similar electricity consumption behaviors are grouped into one category, intuitively showing the electricity consumption patterns of different user groups, such as peak-valley electricity consumption periods, electricity consumption fluctuation rules, etc. This not only helps power enterprises accurately grasp the electricity consumption characteristics of users, achieve refined management and personalized services, but also provides a comparison benchmark for subsequent electricity consumption anomaly judgment, improves the accuracy and efficiency of anomaly detection, and helps the rational allocation of power resources and the stable operation of the power system.
[0095] It should be explained that in the processing and anomaly analysis of electricity consumption behavior data, different nodes of the distributed data platform handle different tasks, which can significantly improve the processing efficiency, enable each task to be executed in parallel to quickly obtain results; enhance the system reliability, a single node failure does not affect the overall process; optimize resource utilization, and reasonably allocate tasks according to the node performance; facilitate system expansion and maintenance, nodes can be added as needed and maintained and upgraded separately; and can also improve the analysis accuracy, each node adopts an adapted algorithm for specific tasks to ensure the accuracy and timeliness of the electricity consumption anomaly analysis results.
[0096] It should be understood that the performing clustering analysis on the electricity consumption behavior data to be analyzed by using a preset clustering method and obtaining the user electricity consumption characteristic curve includes:
[0097] Obtain a preset self-organizing mapping network, and perform an initialization operation on the self-organizing mapping network to obtain an initialized self-organizing mapping network;
[0098] Use the initialized self-organizing mapping network to perform initial clustering on the electricity consumption behavior data to be analyzed to obtain self-organizing clustering centers;
[0099] Perform K-means clustering based on the self-organizing clustering centers and the electricity consumption behavior data to be analyzed to obtain a clustering result, and generate a curve according to the clustering result to obtain the user electricity consumption characteristic curve, where the electricity consumption behavior data to be analyzed includes multiple data points to be analyzed.
[0100] It should be noted that the Self-Organizing Map (SOM) is an unsupervised system artificial neural network that simulates the characteristics of human brain signal processing. It consists of an input layer and an output layer of neurons to form a data processing array, and the output nodes are connected to the input nodes through strength values. During operation, its weight vectors will be continuously adjusted according to specific rules during the training process. As the training progresses, the network gradually reaches a stable state. At this time, the nodes in each area will produce similar outputs for specific inputs, thereby realizing the clustering of data, dividing similar data into adjacent areas, and being widely used in fields such as data mining and pattern recognition. For example, in the analysis of electricity consumption behavior data, it is used to optimize the selection of the initial clustering center of the K-means algorithm to improve the clustering accuracy.
[0101] Furthermore, the self-organizing clustering center is the representative data point of each clustering cluster after the self-organizing map network performs initial clustering on the electricity consumption behavior data. The self-organizing map network simulates the characteristics of human brain signal processing and classifies multiple data points to be analyzed in the electricity consumption behavior data according to similarity. During the initial clustering process, the network adjusts the strength values of the connections between the input layer and the output layer neurons, so that similar data points gather together, and the data points at the core position or the average position of these clustering clusters form the self-organizing clustering center.
[0102] It should be noted that K-means clustering is performed based on the self-organizing clustering center and the electricity consumption behavior data to be analyzed to obtain a clustering result, and a curve is generated according to the clustering result to obtain a user electricity consumption feature curve. Among them, the electricity consumption behavior data to be analyzed includes multiple data points to be analyzed, including:
[0103] The self-organizing clustering center is used as the initial clustering center of the preset K-means clustering algorithm to obtain the K-means initial clustering center, where the K-means initial clustering center includes multiple initial clustering centers;
[0104] The initial clustering centers are sequentially extracted from the multiple initial clustering centers, and the distances from the multiple data points to be analyzed to the extracted initial clustering centers are calculated to obtain a K-means clustering distance group. The K-means clustering distance groups are summarized to obtain a K-means clustering distance group set. Among them, the K-means clustering distance group includes multiple clustering distances, and the clustering distance is the distance from the data point to be analyzed to the extracted initial clustering center. The K-means clustering distance group set includes multiple K-means clustering distance groups;
[0105] Each data point to be analyzed among the multiple data points to be analyzed is divided into the initial clustering center with the closest distance according to the K-means clustering distance group set to obtain multiple K-means clustering clusters;
[0106] Extract K-means clustering clusters sequentially from multiple K-means clustering clusters, obtain the data mean points in the K-means clustering clusters, and use the data mean points to update the initial clustering centers of the extracted K-means clustering clusters to obtain updated clustering centers. Calculate the distances between the updated clustering centers and the initial clustering centers corresponding to the K-means clustering clusters to obtain clustering center change values. Here, the value of the data mean point is the average of the numerical values corresponding to all data points to be analyzed;
[0107] Summarize the updated clustering centers to obtain an updated clustering center group. Use the updated clustering center group as the multiple initial clustering centers, and return the step of sequentially extracting the initial clustering centers from the multiple initial clustering centers until the clustering center change value is less than a preset threshold. Obtain the number of times of the step of returning the step of sequentially extracting the initial clustering centers from the multiple initial clustering centers to obtain the iteration times. If it is confirmed that the iteration times are greater than the preset times, obtain the final clustering clusters;
[0108] Extract the representative features in the final clustering clusters to obtain cluster representative features, and use time as the horizontal axis and the cluster representative features as the vertical axis to construct a user electricity consumption feature curve. Here, the cluster representative feature is the feature represented by the clustering center of the final clustering cluster.
[0109] Furthermore, the average load value, median load value of the final clustering clusters or other statistics that can represent the characteristics of the cluster can also be used as the cluster representative features, and these statistics are corresponding to the corresponding time points. Connecting these points in sequence forms a user electricity consumption feature curve. This curve can intuitively present the changing trend of the user's electricity consumption behavior over time, helping analysts understand the user's electricity consumption pattern and judge whether there is abnormal electricity consumption behavior.
[0110] Furthermore, the initial clustering of the electricity consumption behavior data to be analyzed by using the initialized self-organizing mapping network to obtain self-organizing clustering centers includes:
[0111] Obtain the clustering objects in the electricity consumption behavior data to be analyzed to obtain multiple feature curve clustering objects;
[0112] Obtain the weight vectors in the initialized self-organizing mapping network to obtain multiple network weight vectors;
[0113] Extract the feature curve clustering objects sequentially from the multiple feature curve clustering objects, and calculate the distances between the extracted feature curve clustering objects and the multiple network weight vectors respectively to obtain a clustering distance group;
[0114] Summarize the clustering distance group to obtain a set of clustering distance groups corresponding to multiple feature curve clustering objects;
[0115] Iterate the initialized self-organizing mapping network using the set of clustering distances to obtain a sub-optimal self-organizing mapping network, and calculate the iterative clustering centers of the sub-optimal self-organizing mapping network to obtain self-organizing clustering centers.
[0116] Among them, the clustering object refers to an individual or data unit with specific characteristics in the electricity consumption behavior data to be analyzed that needs to be clustered, such as the set of electricity consumption data of each user at different times within a period, or the electricity consumption data of different regions at the same time period, etc. The sets in the clustering objects constitute the basic objects for clustering. These basic objects have a series of characteristics that can be used for analysis and comparison. By analyzing these clustering objects, patterns and rules in the data are discovered, and then clustering is achieved.
[0117] It should be explained that the weight vector is a set of numerical values used to represent the connection strength between neurons and input data in the initialized self-organizing mapping network. In the self-organizing mapping network, each neuron has a corresponding weight vector, and its dimension is the same as the feature dimension of the input data. The weight vector will be continuously adjusted during the network training process, which can determine the response degree of the neuron to different input data, reflect the mapping relationship from the input data space to the output neuron space, and determine the clustering result and the iterative direction of the network by calculating operations such as the distance between the weight vector and the clustering object.
[0118] S5. According to the second part of nodes in the distributed data platform, use the preset Euclidean distance formula to calculate the similarity between the user's electricity consumption feature curve and the typical load feature curve to obtain the electricity consumption anomaly similarity value.
[0119] It can be understood that calculating the similarity between the user's electricity consumption feature curve and the typical load feature curve using the Euclidean distance formula through the second part of nodes in the distributed data platform and obtaining the electricity consumption anomaly similarity value has many significant effects. First, it can intuitively reflect the deviation degree of the user's electricity consumption behavior from the typical pattern with a quantitative value, which is convenient for power personnel to quickly judge the possibility of anomalies. The more the value deviates from the normal range, the higher the possibility of anomalies. Second, it provides a feasible method for the rapid analysis of large-scale electricity consumption data. The parallel computing ability of the distributed platform nodes greatly improves the processing efficiency, can detect anomalies in a timely manner, and helps the real-time monitoring and early warning of the power system. Third, it helps power enterprises deeply understand the user's electricity consumption behavior, discover abnormal behavior patterns, provide a basis for personalized services, energy-saving guidance, and optimizing power grid scheduling, and then improve the quality of power services and ensure the stable and efficient operation of the power system.
[0120] It should be understood that the use of the preset Euclidean distance formula to calculate the similarity between the user's electricity consumption feature curve and the typical load feature curve to obtain the electricity consumption anomaly similarity value includes:
[0121] Calculate the correlation coefficient between the user's electricity consumption characteristic curve and the typical load characteristic curve to obtain the target correlation coefficient;
[0122] Calculate the Euclidean distance between the user characteristic curve and the typical load characteristic curve to obtain the curve Euclidean distance;
[0123] Based on the target correlation coefficient, the curve Euclidean distance, and a preset weight coefficient, calculate the similarity between the user's electricity consumption characteristic curve and the typical load characteristic curve to obtain the electricity consumption anomaly similarity value.
[0124] It should be noted that the electricity consumption anomaly similarity value is a quantitative index for comprehensively evaluating the similarity degree between the user's electricity consumption characteristic curve and the typical load characteristic curve, and is used to judge whether the user's electricity consumption behavior is abnormal.
[0125] Furthermore, the calculation formula for the target correlation coefficient is as follows:
[0126]
[0127] where p represents the target correlation coefficient, β i represents the i-th feature point in the user's electricity consumption characteristic curve, τ i represents the i-th feature point in the typical load characteristic curve, and respectively represent the average values of the feature points in the user's electricity consumption characteristic curve and the typical load characteristic curve, and n represents the number of feature points.
[0128] Furthermore, the feature point represents the power load data point at a specific time point, which is used to measure the user's electricity consumption situation or the power load under the typical electricity consumption mode. Among them, the power load data point for measuring the user's electricity consumption situation refers to the k-th user load point on the inspection day, representing the actual power load value of the user at the k-th moment recorded in chronological order on the inspection day. These values form the user load curve, reflecting the user's electricity consumption changes at different times of the day. For example, when the user turns on a high-power electrical appliance at a certain moment, the value of the user load point at that moment will increase, which can intuitively show the user's real-time electricity consumption status; the power load data point under the typical electricity consumption mode represents the load value at that moment under the normal electricity consumption mode. It is a representative value obtained through the analysis and processing of a large amount of normal electricity consumption data, reflecting the electricity consumption level of the user at each time point under normal circumstances.
[0129] It should be noted that the calculation formulas for calculating the Euclidean distance between the user characteristic curve and the typical load characteristic curve and calculating the similarity between the user's electricity consumption characteristic curve and the typical load characteristic curve are as follows:
[0130]
[0131] where δ represents the similarity, and D * represents the Euclidean distance between the user feature curve and the typical load feature curve, and W 1 and W 2 are weight coefficients.
[0132] It should be explained that the electricity consumption anomaly similarity value is a quantitative indicator used to measure the similarity between the user's electricity consumption feature curve and the typical load feature curve. It is calculated through specific algorithms, such as the Euclidean distance formula. This value reflects the deviation between the user's actual electricity consumption behavior and the normal or expected electricity consumption pattern. The lower the value, the more similar the user's electricity consumption characteristics are to the typical load characteristics, and the more normal the electricity consumption behavior; the higher the value, the greater the difference between the two, and the higher the possibility of abnormal electricity consumption behavior of the user, which helps the operation personnel of the power system to timely discover and judge whether there is an electricity consumption anomaly for the user.
[0133] S6. According to the third - part nodes in the distributed data platform, use the electricity consumption anomaly feature analysis model to identify the electricity consumption anomaly features in the to - be - analyzed electricity consumption behavior data, and conduct electricity consumption anomaly analysis based on the electricity consumption anomaly features and the electricity consumption anomaly similarity value to obtain the electricity consumption anomaly analysis result.
[0134] It can be understood that using the electricity consumption anomaly feature analysis model to identify the anomaly features in the to - be - analyzed electricity consumption behavior data, and conducting electricity consumption anomaly analysis in combination with the electricity consumption anomaly similarity value can accurately locate the electricity consumption anomaly, comprehensively capture the anomaly features and quantify the degree of anomaly; it can timely give early warnings and preventions, achieve real - time monitoring and rapid response, and prevent potential risks; it helps to optimize power management, formulate personalized strategies and reasonably allocate resources; it can also provide a basis for policy - making and enterprise strategic planning, support scientific decision - making, ensure the stable operation of the power system, and improve the quality of power services and industry benefits.
[0135] It should be understood that the electricity consumption anomaly analysis result is a comprehensive conclusion obtained by using the electricity consumption anomaly feature analysis model to identify the anomaly features in the to - be - analyzed electricity consumption behavior data through the distributed data platform and in combination with the electricity consumption anomaly similarity value. It not only accurately presents specific anomaly features such as power mutation and abnormal electricity consumption in the electricity consumption behavior, but also intuitively reflects the degree and trend of the anomaly with the quantified similarity value.
[0136] It is understandable that this data analysis method is not only applicable to abnormal power consumption analysis, but also has wide applicability and strong expandability. In the field of energy management, it can be used for abnormal monitoring and refined management of natural gas, water resources, etc.; in industrial production, it can monitor equipment status and optimize production processes; in commercial operations, it can analyze customer consumption behaviors to formulate precise marketing strategies; in transportation, it can evaluate driving behaviors and assist enterprise management decisions. In addition, it can also deeply analyze data, mine value, and discover anomalies in multiple fields such as medical care, education, and finance, providing strong support for the development and management of various industries.
[0137] To solve the problems described in the background art, the present invention collects the power consumption behavior data of users based on a preset data collection instruction, and transmits the power consumption behavior data to a preset distributed data platform for storage using multiple nodes of the distributed data platform to obtain a power consumption behavior data set. Among them, the distributed data platform includes multiple nodes, and the multiple nodes include a first part of nodes, a second part of nodes, and a third part of nodes. It can be seen that the present invention adopts a distributed storage method to ensure data security; perform a data filtering operation on the power consumption behavior data set to obtain a denoised power consumption behavior data set, and convert the data in the denoised power consumption behavior data set into a preset standard format to obtain a standard power consumption behavior data set. Among them, the data filtering operation includes removing duplicate data and removing noise data; according to the distributed data platform, use the standard power consumption behavior data set to construct an abnormal feature analysis model to obtain a power consumption abnormal feature analysis model, and use the standard power consumption behavior data in the standard power consumption behavior data set to generate image data to obtain a typical load characteristic curve; obtain the power consumption behavior data to be analyzed, and according to the first part of nodes in the distributed data platform, use a preset clustering method to perform clustering analysis on the power consumption behavior data to be analyzed to obtain a user power consumption characteristic curve. It can be seen that the present invention uses node division in the distributed data platform to work, which can effectively improve the processing speed; according to the second part of nodes in the distributed data platform, use a preset Euclidean distance formula to calculate the similarity between the user power consumption characteristic curve and the typical load characteristic curve to obtain a power consumption abnormal similarity value; according to the third part of nodes in the distributed data platform, use the power consumption abnormal feature analysis model to identify the abnormal features in the power consumption behavior data to be analyzed to obtain power consumption behavior abnormal features, and perform power consumption abnormal analysis based on the power consumption behavior abnormal features and the power consumption abnormal similarity value to obtain the analysis result of the power consumption behavior data to be analyzed, obtaining a power consumption abnormal analysis result. It can be seen that the present invention combines features and similarity values for abnormal analysis to improve the accuracy of abnormal analysis. Therefore, the present invention can reduce the error and cost of power consumption abnormal analysis and improve the accuracy of power consumption abnormal analysis.
[0138] Such as Figure 2As shown, it is a functional module diagram of a data analysis system based on a cloud computing distributed data platform provided by an embodiment of the present invention.
[0139] The data analysis system 100 based on the cloud computing distributed data platform of the present invention can be installed in an electronic device. According to the implemented functions, the data analysis system 100 based on the cloud computing distributed data platform can include a data acquisition module 101, a model construction module 102, a similarity calculation module 103, and an anomaly analysis module 104. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0140] The data acquisition module 101 is used to collect the electricity consumption behavior data of users based on a preset data acquisition instruction, and transmit the electricity consumption behavior data to a preset distributed data platform, and store it using multiple nodes of the distributed data platform to obtain an electricity consumption behavior data set;
[0141] The data acquisition module 101 is used to collect the electricity consumption behavior data of users based on a preset data acquisition instruction, and transmit the electricity consumption behavior data to a preset distributed data platform for storage to obtain an electricity consumption behavior data set, where the distributed data platform includes multiple nodes, and the multiple nodes include a first part of nodes, a second part of nodes, and a third part of nodes;
[0142] The model construction module 102 is used to perform a data filtering operation on the electricity consumption behavior data set to obtain a denoised electricity consumption behavior data set, and convert the denoised electricity consumption behavior data in the denoised electricity consumption behavior data set into a preset standard format to obtain a standard electricity consumption behavior data set, where the data filtering operation includes removing duplicate data and removing noise data;
[0143] According to the distributed data platform, an electricity consumption anomaly feature analysis model is constructed using the standard electricity consumption behavior data set, and a typical load feature curve is generated using the standard electricity consumption behavior data in the standard electricity consumption behavior data set;
[0144] The similarity calculation module 103 is used to obtain the electricity consumption behavior data to be analyzed, and perform clustering analysis on the electricity consumption behavior data to be analyzed using a preset clustering method according to the first part of nodes in the distributed data platform to obtain a user electricity consumption feature curve;
[0145] According to the second part of nodes in the distributed data platform, the similarity between the user electricity consumption feature curve and the typical load feature curve is calculated using a preset Euclidean distance formula to obtain an electricity consumption anomaly similarity value;
[0146] The abnormal analysis module 104 is configured to identify abnormal characteristics of the electricity consumption behavior in the to-be-analyzed electricity consumption behavior data by using the third part of nodes in the distributed data platform, and perform abnormal analysis of electricity consumption based on the abnormal characteristics of the electricity consumption behavior and the abnormal similarity value of the electricity consumption, so as to obtain an abnormal analysis result of the electricity consumption.
[0147] Specifically, when the modules in the data analysis system 100 based on the cloud computing distributed data platform in the embodiments of the present invention are used, they adopt the same technical means as those in the Figure 1 data analysis method based on the cloud computing distributed data platform described above, and can produce the same technical effects, which will not be elaborated here.
[0148] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing the data analysis method based on the cloud computing distributed data platform provided by an embodiment of the present invention.
[0149] The electronic device 1 may include a processor 10, a memory 11, and a bus 12, and may further include a computer program stored in the memory 11 and executable on the processor 10, such as a data analysis method program based on the cloud computing distributed data platform.
[0150] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In other embodiments, the memory 11 may also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 further includes an internal storage unit of the electronic device 1 and also includes an external storage device. The memory 11 can not only be used to store application software installed in the electronic device 1 and various types of data, such as the code of the data analysis method program based on the cloud computing distributed data platform, but also be used to temporarily store data that has been output or will be output.
[0151] In some embodiments, the processor 10 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as data analysis method programs based on cloud computing distributed data platforms, etc.), and calling the data stored in the memory 11, to execute various functions of the electronic device 1 and process data.
[0152] The bus 12 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus 12 can be divided into an address bus, a data bus, a control bus, etc. The bus 12 is set to enable connection communication between the memory 11 and at least one processor 10, etc.
[0153] Figure 3 Only the electronic device with components is shown. Those skilled in the art can understand that, Figure 3 the shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0154] For example, although not shown, the electronic device 1 may further include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device 1 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0155] Furthermore, the electronic device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the electronic device 1 and other electronic devices.
[0156] Optionally, the electronic device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.
[0157] The data analysis method program based on the cloud computing distributed data platform stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can implement:
[0158] Collect the user's electricity consumption behavior data based on a preset data collection instruction, and transmit the electricity consumption behavior data to a preset distributed data platform for storage using the distributed data platform to obtain an electricity consumption behavior data set. Among them, the distributed data platform includes multiple nodes, and the multiple nodes include a first part of nodes, a second part of nodes, and a third part of nodes;
[0159] Perform a data filtering operation on the electricity consumption behavior data set to obtain a denoised electricity consumption behavior data set, and convert the denoised electricity consumption behavior data in the denoised electricity consumption behavior data set into a preset standard format to obtain a standard electricity consumption behavior data set. Among them, the data filtering operation includes removing duplicate data and removing noise data;
[0160] According to the distributed data platform, use the standard electricity consumption behavior data set to construct an electricity consumption anomaly feature analysis model, and use the standard electricity consumption behavior data in the standard electricity consumption behavior data set to generate a typical load feature curve;
[0161] Obtain the electricity consumption behavior data to be analyzed, and perform clustering analysis on the electricity consumption behavior data to be analyzed using a preset clustering method according to the first part of nodes in the distributed data platform to obtain a user electricity consumption feature curve;
[0162] According to the second part of nodes in the distributed data platform, calculate the similarity between the user electricity consumption feature curve and the typical load feature curve using a preset Euclidean distance formula to obtain an electricity consumption anomaly similarity value;
[0163] Based on the third - part nodes in the distributed data platform, use the electricity - consumption abnormal - feature analysis model to identify the electricity - consumption behavior abnormal features in the to - be - analyzed electricity - consumption behavior data, and perform electricity - consumption abnormal analysis based on the electricity - consumption behavior abnormal features and the electricity - consumption abnormal similarity value to obtain the electricity - consumption abnormal analysis result.
[0164] Specifically, the specific implementation method of the above instructions by the processor 10 can refer to Figures 1 to 3 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.
[0165] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer - readable storage medium. The computer - readable storage medium can be volatile or non - volatile. For example, the computer - readable medium can include: any entity or system that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read - only memory (ROM, Read - Only Memory).
[0166] The present invention also provides a computer - readable storage medium. The readable storage medium stores a computer program, and when the computer program is executed by the processor of the electronic device, it can implement:
[0167] Collect the electricity - consumption behavior data of users based on a preset data - collection instruction, and transmit the electricity - consumption behavior data to a preset distributed data platform for storage by the distributed data platform to obtain an electricity - consumption behavior data set. Among them, the distributed data platform includes multiple nodes, and the multiple nodes include the first - part nodes, the second - part nodes, and the third - part nodes;
[0168] Perform a data - filtering operation on the electricity - consumption behavior data set to obtain a denoised electricity - consumption behavior data set, and convert the denoised electricity - consumption behavior data in the denoised electricity - consumption behavior data set into a preset standard format to obtain a standard electricity - consumption behavior data set. Among them, the data - filtering operation includes removing duplicate data and removing noise data;
[0169] Based on the distributed data platform, use the standard electricity - consumption behavior data set to construct an electricity - consumption abnormal - feature analysis model, and use the standard electricity - consumption behavior data in the standard electricity - consumption behavior data set to generate a typical load - feature curve;
[0170] Obtain the to - be - analyzed electricity - consumption behavior data, and based on the first - part nodes in the distributed data platform, use a preset clustering method to perform clustering analysis on the to - be - analyzed electricity - consumption behavior data to obtain a user electricity - consumption feature curve;
[0171] According to the second part of nodes in the distributed data platform, use the preset Euclidean distance formula to calculate the similarity between the user's electricity consumption characteristic curve and the typical load characteristic curve, and obtain the electricity consumption anomaly similarity value;
[0172] According to the third part of nodes in the distributed data platform, use the electricity consumption anomaly feature analysis model to identify the electricity consumption behavior anomaly features in the to-be-analyzed electricity consumption behavior data, and perform electricity consumption anomaly analysis based on the electricity consumption behavior anomaly features and the electricity consumption anomaly similarity value, so as to obtain the electricity consumption anomaly analysis result.
[0173] In several embodiments provided by the present invention, it should be understood that the disclosed devices, systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative, and there can be other division methods in actual implementation.
[0174] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0175] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.
[0176] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-mentioned exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.
[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A data analysis method based on a cloud computing distributed data platform, characterized in that: The method comprises: Collecting the user's electricity usage behavior data based on a preset data collection instruction, and transmitting the electricity usage behavior data to a preset distributed data platform, storing the data using the distributed data platform, and obtaining an electricity usage behavior data set, wherein the distributed data platform includes a plurality of nodes, and the plurality of nodes include a first part of nodes, a second part of nodes, and a third part of nodes; Performing a data filtering operation on the electricity usage behavior data set to obtain a denoised electricity usage behavior data set, and converting the denoised electricity usage behavior data in the denoised electricity usage behavior data set into a preset standard format to obtain a standard electricity usage behavior data set, wherein the data filtering operation includes removing duplicate data and removing noise data; According to the distributed data platform, a power consumption abnormality feature analysis model is constructed using the standard power consumption behavior data set, and a typical load characteristic curve is generated using the standard power consumption behavior data in the standard power consumption behavior data set; Acquire the electricity consumption behavior data to be analyzed, and perform cluster analysis on the electricity consumption behavior data to be analyzed by using a preset clustering method according to the first part of nodes in the distributed data platform to obtain a user electricity consumption characteristic curve; According to the second part of nodes in the distributed data platform, a preset Euclidean distance formula is used to calculate the similarity between the user power consumption characteristic curve and the typical load characteristic curve to obtain a power consumption anomaly similarity value; According to the third part nodes in the distributed data platform, the abnormal power consumption feature analysis model is used to identify the abnormal power consumption behavior features in the power consumption behavior data to be analyzed, and the power consumption anomaly analysis is performed based on the abnormal power consumption behavior features and the power consumption anomaly similarity values to obtain the abnormal power consumption analysis results.
2. The data analysis method based on cloud computing distributed data platform according to claim 1, characterized in that: The method of performing cluster analysis on the power consumption behavior data to be analyzed by using a preset clustering method to obtain a user power consumption characteristic curve includes: Acquire a preset self-organizing map network, and perform an initialization operation on the self-organizing map network to obtain an initialized self-organizing map network; Using the initialized self-organizing map network to perform initial clustering on the power consumption behavior data to be analyzed, and obtaining a self-organizing clustering center; K-means clustering is performed based on the self-organizing cluster center and the electricity consumption behavior data to be analyzed to obtain a clustering result, and a curve is generated according to the clustering result to obtain a user electricity consumption characteristic curve, wherein the electricity consumption behavior data to be analyzed includes multiple data points to be analyzed.
3. The data analysis method based on cloud computing distributed data platform according to claim 2, characterized in that: The using the initialized self-organizing map network to initially cluster the power consumption behavior data to be analyzed to obtain a self-organizing clustering center includes: Acquire cluster objects in the electricity consumption behavior data to be analyzed to obtain a plurality of characteristic curve cluster objects; Obtaining a weight vector in the initialized self-organizing map network to obtain a plurality of network weight vectors; Extracting characteristic curve clustering objects from the plurality of characteristic curve clustering objects in sequence, and calculating distances between the extracted characteristic curve clustering objects and the plurality of network weight vectors respectively based on the extracted characteristic curve clustering objects to obtain a clustering distance group; Summarizing the cluster distance groups to obtain cluster distance group sets corresponding to multiple characteristic curve cluster objects; The initialized self-organizing map network is iterated using the cluster distance group set to obtain a suboptimal self-organizing map network, and the iterative clustering center of the suboptimal self-organizing map network is calculated to obtain a self-organizing clustering center.
4. The data analysis method based on cloud computing distributed data platform according to claim 1, characterized in that: The method of calculating the similarity between the user power consumption characteristic curve and the typical load characteristic curve by using a preset Euclidean distance formula to obtain a power consumption anomaly similarity value includes: Calculating the correlation coefficient between the user power consumption characteristic curve and the typical load characteristic curve to obtain a target correlation coefficient; Calculating the Euclidean distance between the user characteristic curve and the typical load characteristic curve to obtain the curve Euclidean distance; Based on the target correlation coefficient, the curve Euclidean distance and a preset weight coefficient, the similarity between the user power consumption characteristic curve and the typical load characteristic curve is calculated to obtain a power consumption anomaly similarity value.
5. The data analysis method based on cloud computing distributed data platform according to claim 4, characterized in that: The calculation formula of the target correlation coefficient is as follows: Among them, p represents the target correlation coefficient, β i represents the i-th characteristic point in the user's electricity consumption characteristic curve, τ i represents the i-th characteristic point in the typical load characteristic curve, and They represent the average values of characteristic points in the user's power consumption characteristic curve and the typical load characteristic curve respectively, and n represents the number of characteristic points.
6. The data analysis method based on cloud computing distributed data platform according to claim 1, characterized in that: The method of using the standard electricity consumption behavior data set to construct an abnormal electricity consumption feature analysis model includes: Reconstructing the standard electricity usage behavior data set to obtain a reconstructed electricity usage behavior data set, and obtaining abnormal features in the reconstructed electricity usage behavior data set to obtain abnormal electricity usage features; Obtaining standard electricity usage behavior data corresponding to the abnormal electricity usage characteristics in the standard electricity usage behavior data set to obtain abnormal characteristic data; The abnormal power consumption characteristics and the abnormal characteristic data are divided into a training set and a test set to obtain a characteristic training set and a characteristic test set. The characteristic training set is used to train a preset fuzzy neural network, and the characteristic test set is used to test the trained fuzzy neural network to obtain an abnormal power consumption characteristic analysis model.
7. The data analysis method based on cloud computing distributed data platform according to claim 2, characterized in that: The K-means clustering is performed based on the self-organizing cluster center and the power consumption behavior data to be analyzed to obtain a clustering result, and a curve is generated according to the clustering result to obtain a user power consumption characteristic curve, wherein the power consumption behavior data to be analyzed includes multiple data points to be analyzed, including: The self-organizing cluster center is used as the initial cluster center of the preset K-means clustering algorithm to obtain the K-means initial cluster center, wherein the K-means initial cluster center includes multiple initial cluster centers; Extracting initial cluster centers from multiple initial cluster centers in sequence, and calculating the distances from the multiple data points to be analyzed to the extracted initial cluster centers to obtain a K-means cluster distance group, summarizing the K-means cluster distance groups to obtain a K-means cluster distance group set, wherein the K-means cluster distance group includes multiple cluster distances, and the cluster distance is the distance from the data points to be analyzed to the extracted initial cluster center, and the K-means cluster distance group set includes multiple K-means cluster distance groups; According to the K-means clustering distance group set, each of the multiple data points to be analyzed is divided into an initial cluster center with the closest distance, so as to obtain multiple K-means clustering clusters; Extract K-means clusters from multiple K-means clusters in sequence, obtain data mean points in the K-means clusters, and use the data mean points to update the initial cluster centers of the extracted K-means clusters to obtain updated cluster centers, calculate the distance between the updated cluster centers and the initial cluster centers corresponding to the K-means clusters, and obtain cluster center change values, wherein the value of the data mean point is the average value of the values corresponding to all the data points to be analyzed; Summarize and update the cluster centers to obtain an updated cluster center group, update the cluster center group to the multiple initial cluster centers, return to the step of sequentially extracting the initial cluster centers from the multiple initial cluster centers until the cluster center change value is less than a preset threshold, obtain the number of times the step of sequentially extracting the initial cluster centers from the multiple initial cluster centers is returned, and obtain the number of iterations. If it is confirmed that the number of iterations is greater than the preset number, obtain the final cluster cluster; The representative features in the final cluster are extracted to obtain cluster representative features, and the user power consumption characteristic curve is constructed with time as the horizontal axis and the cluster representative features as the vertical axis, wherein the cluster representative features are features represented by the cluster center of the final cluster.
8. The data analysis method based on cloud computing distributed data platform according to claim 1, characterized in that: The generating a typical load characteristic curve by using the standard electricity consumption behavior data in the standard electricity consumption behavior data set includes: Acquire business requirements based on the electricity consumption behavior data to obtain data analysis requirements; Determine the data aggregation granularity by using the data analysis requirements to obtain the electricity consumption data aggregation granularity; The standard electricity consumption behavior data is calculated according to the electricity consumption data aggregation granularity to obtain granularity adjustment data, and an image is generated according to the granularity adjustment data to obtain a typical load characteristic curve.
9. The data analysis method based on cloud computing distributed data platform according to claim 4, characterized in that: The calculation formula for calculating the Euclidean distance between the user characteristic curve and the typical load characteristic curve and the similarity between the user power consumption characteristic curve and the typical load characteristic curve is as follows: Among them, δ represents the similarity, D * It represents the Euclidean distance between the user characteristic curve and the typical load characteristic curve, and W1 and W2 are weight coefficients.
10. A data analysis system based on a cloud computing distributed data platform, characterized in that: The system comprises: A data collection module, used to collect the user's electricity usage behavior data based on a preset data collection instruction, and transmit the electricity usage behavior data to a preset distributed data platform, and use the distributed data platform to store it to obtain an electricity usage behavior data set, wherein the distributed data platform includes a plurality of nodes, and the plurality of nodes include a first part of nodes, a second part of nodes, and a third part of nodes; A model building module, configured to perform a data filtering operation on the electricity usage behavior data set to obtain a denoised electricity usage behavior data set, and convert the denoised electricity usage behavior data in the denoised electricity usage behavior data set into a preset standard format to obtain a standard electricity usage behavior data set, wherein the data filtering operation includes removing duplicate data and removing noise data; According to the distributed data platform, a power consumption abnormality feature analysis model is constructed using the standard power consumption behavior data set, and a typical load characteristic curve is generated using the standard power consumption behavior data in the standard power consumption behavior data set; A similarity calculation module is used to obtain the power consumption behavior data to be analyzed, and perform cluster analysis on the power consumption behavior data to be analyzed using a preset clustering method according to the first part of nodes in the distributed data platform to obtain a user power consumption characteristic curve; According to the second part of nodes in the distributed data platform, a preset Euclidean distance formula is used to calculate the similarity between the user power consumption characteristic curve and the typical load characteristic curve to obtain a power consumption anomaly similarity value; The anomaly analysis module is used to identify the abnormal power usage characteristics in the power usage behavior data to be analyzed based on the third part nodes in the distributed data platform using the power usage abnormality feature analysis model, and perform power usage abnormality analysis based on the power usage abnormality characteristics and the power usage abnormality similarity value to obtain the power usage abnormality analysis result.
Citation Information
Cited By
Fuel gas state early warning and monitoring method and system based on cloud-side cooperation
CN121884535A