Virtual power plant data processing method and system

Through edge data node preprocessing and cloud data integration combined with C4.5 decision tree algorithm, the calculation pressure and abnormal calibration problems in virtual power plant data processing are solved, and efficient data processing and accurate use are achieved.

CN120354059APending Publication Date: 2025-07-22STATE GRID GANSU ELECTRIC POWER RESEARCH INSTITUTE

Patent Information

Application Number
CN202510850841.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing virtual power plants have problems such as high computational pressure and poor data mining and abnormal calibration in data processing, which affects the accurate use of data.

Method used

Data preprocessing is performed through multiple edge data nodes, effectively collect data, and transmit it to the cloud for data integration and standardization. C4.5 decision tree algorithm is used for classification mining and outlier detection, and local outlier factors are calculated for abnormal verification.

Benefits of technology

It reduces the computing pressure in the cloud, improves data processing efficiency, and ensures data accuracy and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354059A_ABST
    Figure CN120354059A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of virtual power plants, and provides a virtual power plant data processing method and system. According to the method, data acquisition is carried out on a plurality of application ends, and data preprocessing is carried out through a plurality of edge data nodes, so that a plurality of pieces of effective acquisition data are obtained; transmitting the plurality of effective acquisition data to a cloud end, and performing data integration and standardization through the cloud end to generate standard integrated data; neglecting defect attribute data, and adopting a C4.5 decision tree algorithm to perform classification mining on defect-free attribute data to obtain classification mining data; and performing outlier detection of data points on the classified mining data, calculating a plurality of local outlier factors, and performing abnormity checking and processing. According to the method, cloud edge collaboration can be carried out, data preprocessing is carried out through edge data nodes, the computing pressure of a cloud end is effectively reduced, the data processing efficiency is improved, a C4.5 decision tree algorithm is adopted, data classification mining is carried out, outlier detection and anomaly checking and processing are carried out, and accurate use of data is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of virtual power plants, and in particular relates to a virtual power plant data processing method and system. Background Art

[0002] Traditional power grids control distributed photovoltaics individually, which requires high communication requirements, has relatively poor economy, and the control effect is average; after the large-scale access of small-capacity and decentralized distributed photovoltaics, the economic benefits of the distribution network are more affected, and it is more difficult for traditional distribution network control to maximize the utilization of resources.

[0003] As an effective technical means, a virtual power plant can solve the problems of power grid peak shaving and consumption brought about by the large-scale access of renewable energy. It can not only make full use of distributed resources without changing the existing power grid, but also achieve a situation of multi-energy complementarity on the power supply side and flexible interaction on the load side.

[0004] In the prior art, there are still some defects in the data processing of virtual power plants: (1) mainly through the central cloud server for unified and large-scale data processing, resulting in huge computing pressure and affecting data processing efficiency; (2) the data mining and anomaly checking effects are not good, affecting the accurate use of data. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide a virtual power plant data processing method and system, aiming to solve the technical problems existing in the prior art mentioned in the background art.

[0006] The embodiments of the present invention are implemented as follows: A virtual power plant data processing method, the method includes the following steps: Collect data from multiple application terminals according to multiple preset data types, and perform dynamic data sharing and preprocessing through multiple edge data nodes to obtain multiple effective collected data; Transmit the multiple effective collected data to the cloud, and perform data integration and standardization through the cloud to generate standard integrated data; Perform attribute recognition on the standard integrated data, ignore the missing attribute data, and use the C4.5 decision tree algorithm to classify and mine the non-missing attribute data to obtain classified and mined data; Perform outlier detection on the classified and mined data, calculate multiple local outlier factors, and perform anomaly checking and processing.

[0007] As a further limitation of the technical solution of the embodiments of the present invention, the step of collecting data from multiple application terminals according to multiple preset data types, and performing dynamic data sharing and preprocessing through multiple edge data nodes to obtain multiple effective collected data specifically includes the following steps: Collect data from multiple application terminals of the virtual power plant according to multiple preset data types to obtain multiple relevant collected data; Determine multiple online edge data nodes; Monitor the status of multiple said edge data nodes to obtain multiple node status data; Analyze multiple said node status data and match multiple corresponding sharing ratios; According to multiple said sharing ratios, share and transmit multiple said relevant collected data among multiple said edge data nodes; Through multiple said edge data nodes, perform data preprocessing on multiple corresponding relevant collected data, including missing value filling, outlier removal, and duplicate deletion, to obtain multiple valid collected data.

[0008] As a further limitation of the technical solution of the embodiment of the present invention, the step of transmitting multiple said valid collected data to the cloud and performing data integration and standardization through the cloud to generate standard integrated data specifically includes the following steps: Determine the central cloud node corresponding to the cloud; Upload multiple said valid collected data to the central cloud node; Through the central cloud node, use a data warehouse or a data lake to perform data integration on multiple said valid collected data to generate integrated collected data; Perform standardization processing on the integrated collected data to generate standard integrated data.

[0009] As a further limitation of the technical solution of the embodiment of the present invention, the step of performing attribute recognition on the standard integrated data, ignoring defective attribute data, and using the C4.5 decision tree algorithm to perform classification mining on non-defective attribute data to obtain classification mining data specifically includes the following steps: Perform attribute recognition on the standard integrated data and record the attribute recognition results; According to the attribute recognition results, ignore defective attribute data and retain non-defective attribute data from the standard integrated data; Use the C4.5 decision tree algorithm to perform classification mining on the non-defective attribute data to obtain classification mining data.

[0010] As a further limitation of the technical solution of the embodiment of the present invention, the step of performing classification mining on the non-defective attribute data to obtain classification mining data specifically includes the following steps: Select multiple information features from the non-defective attribute data; Calculate the information gain of multiple said information features; Supplement multiple said information gains and calculate multiple information gain ratios; Compare the multiple information gain rates, select split nodes from multiple corresponding information features, construct a decision tree, and obtain classified mining data.

[0011] As a further limitation of the technical solution of the embodiment of the present invention, the outlier detection of data points for the classified mining data, calculating multiple local outlier factors, and performing anomaly verification and processing specifically include the following steps: Input the classified mining data and specify the data points to be evaluated that have not been traversed ; Calculate the reachability distance between the data point to be evaluated and its neighbor data points , where is the number of neighbors of the data point to be evaluated ; Calculate the local reachability density of the data point to be evaluated based on the reachability distance; Calculate the local outlier factor of the data point to be evaluated based on the local reachability density; Arrange the multiple local outlier factors in descending order, determine multiple outlier data points, and perform data elimination processing.

[0012] As a further limitation of the technical solution of the embodiment of the present invention, the formula for calculating the reachability distance between the data point to be evaluated and its neighbor data points is: ; where is the reachability distance between the data point to be evaluated and its neighbor data points , is the distance of the nearest neighbor of the data point to be evaluated, is the distance between the data point to be evaluated and its neighbor data point ; The formula for calculating the local reachability density of the data point to be evaluated is: ; where is the distance neighborhood of the nearest neighbor of the data point to be evaluated; The The calculation formula for the local outlier factor is as follows: ; where is the local reachability density of the neighbor data point .

[0013] A virtual power plant data processing system for performing any of the virtual power plant data processing methods described above. The system includes an edge acquisition and processing module, a cloud upload and processing module, a data classification and mining module, and an anomaly verification and processing module, where: The edge acquisition and processing module is used to collect data from multiple application terminals according to multiple preset data types, and perform dynamic data sharing and preprocessing through multiple edge data nodes to obtain multiple effective acquisition data; The cloud upload and processing module is used to transmit multiple pieces of the effective acquisition data to the cloud, and perform data integration and standardization through the cloud to generate standard integrated data; The data classification and mining module is used to perform attribute recognition on the standard integrated data, ignore the defective attribute data, and use the C4.5 decision tree algorithm to perform classification and mining on the non-defective attribute data to obtain classification and mining data; The anomaly verification and processing module is used to perform outlier detection on the data points of the classification and mining data, calculate multiple local outlier factors, and perform anomaly verification and processing.

[0014] As a further limitation of the technical solution of the embodiment of the present invention, the cloud upload and processing module specifically includes: A node determination unit for determining the central cloud node corresponding to the cloud; A cloud upload unit for uploading multiple pieces of the effective acquisition data to the central cloud node; A data integration unit for performing data integration on multiple pieces of the effective acquisition data through the central cloud node using a data warehouse or a data lake to generate integrated acquisition data; A standardization processing unit for performing standardization processing on the integrated acquisition data to generate standard integrated data.

[0015] As a further limitation of the technical solution of the embodiment of the present invention, the data classification and mining module specifically includes: An attribute recognition unit for performing attribute recognition on the standard integrated data and recording the attribute recognition result; A data selection unit for ignoring the defective attribute data and retaining the non-defective attribute data from the standard integrated data according to the attribute recognition result; A classification and mining unit for performing classification and mining on the non-defective attribute data using the C4.5 decision tree algorithm to obtain classification and mining data.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: In the embodiment of the present invention, data is collected from multiple application terminals, and data preprocessing is performed through multiple edge data nodes to obtain multiple effective collected data; the multiple effective collected data is transmitted to the cloud, and data integration and standardization are performed through the cloud to generate standard integrated data; the defective attribute data is ignored, and the C4.5 decision tree algorithm is used to classify and mine the non-defective attribute data to obtain classified and mined data; outlier detection of data points is performed on the classified and mined data, multiple local outlier factors are calculated, and anomaly verification and processing are performed. It can perform cloud-edge collaboration, perform data preprocessing through edge data nodes, effectively reduce the computing pressure on the cloud, improve data processing efficiency, and use the C4.5 decision tree algorithm for data classification and mining, and perform outlier detection, anomaly verification and processing to ensure the accurate use of data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 The flowchart of the virtual power plant data processing method provided by the embodiment of the present invention is shown; Figure 2 The application architecture diagram of the virtual power plant data processing system provided by the embodiment of the present invention is shown; Figure 3 The schematic diagram of cloud-edge collaboration provided by the embodiment of the present invention is shown; Figure 4 The schematic diagram of the C4.5 decision tree provided by the embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0019] It can be understood that in the prior art, as an effective technical means, the virtual power plant can solve the problems of power grid peak regulation and consumption brought about by the large-scale access of renewable energy. It can not only make full use of distributed resources without changing the existing power grid, but also achieve the situation of multi-energy complementarity on the power supply side and flexible interaction on the load side. However, there are still some defects in the virtual power plant in terms of data processing: (1) mainly through the central cloud server for unified and large-scale data processing, resulting in huge computing pressure and affecting data processing efficiency; (2) the data mining and anomaly verification effects are not good, affecting the accurate use of data.

[0020] To solve the above problems, a virtual power plant data processing method and system disclosed in an embodiment of the present invention collect data from multiple application terminals according to multiple preset data types, and perform dynamic data sharing and preprocessing through multiple edge data nodes to obtain multiple effective collected data; transmit the multiple effective collected data to the cloud, and perform data integration and standardization through the cloud to generate standard integrated data; perform attribute recognition on the standard integrated data, ignore the defective attribute data, and use the C4.5 decision tree algorithm to classify and mine the non-defective attribute data to obtain classified and mined data; perform outlier detection on the data points of the classified and mined data, calculate multiple local outlier factors, and perform anomaly verification and processing. It can perform cloud-edge collaboration, perform data preprocessing through edge data nodes, effectively reduce the computing pressure on the cloud, improve data processing efficiency, and use the C4.5 decision tree algorithm to perform data classification and mining, and perform outlier detection, anomaly verification and processing to ensure the accurate use of data.

[0021] Specifically, Figure 1 FIG. shows a flowchart of the virtual power plant data processing method provided by an embodiment of the present invention.

[0022] In a preferred embodiment provided by the present invention, a virtual power plant data processing method includes the following steps: Step S101: Collect data from multiple application terminals according to multiple preset data types, and perform dynamic data sharing and preprocessing through multiple edge data nodes to obtain multiple effective collected data.

[0023] In an embodiment of the present invention, data is collected from multiple application terminals of the virtual power plant according to multiple data types such as preset energy storage status, power generation, equipment status, environmental parameters, and compliance requirements to obtain multiple relevant collected data, and multiple edge data nodes in the online state are determined (the multiple edge data nodes include both the edge data nodes corresponding to the multiple application terminals and other edge data nodes). By monitoring the status of the multiple edge data nodes, multiple node status data are obtained, and then the multiple node status data are analyzed. According to the corresponding computing power, memory, and network conditions, the sharing ratios corresponding to the multiple edge data nodes are matched. Then, according to the multiple sharing ratios, the multiple relevant collected data are divided and transmitted to the multiple corresponding edge data nodes. Furthermore, through the multiple edge data nodes, missing value filling (for numerical data, column means are used for missing value filling; for skewed distribution data, column medians are used for missing value filling; for categorical data, the most frequently occurring value in the column is used for missing value filling), outlier removal (using the Z-score method or the IQR method for outlier identification and removal processing), and duplicate deletion and other data preprocessing are performed on the multiple corresponding relevant collected data to obtain multiple effective collected data.

[0024] It can be understood that in the sharing and partitioning, data sharing and preprocessing are preferentially performed through the edge data nodes of the application side corresponding to the node status data itself, and the redundant data after sharing is shared and transmitted to other edge data nodes, and the edge data nodes with the shortest communication transmission path are preferentially allocated.

[0025] Specifically, in another preferred embodiment provided by the present invention, the steps of collecting data from multiple application sides according to multiple preset data types, and performing dynamic data sharing and preprocessing through multiple edge data nodes to obtain multiple effective collected data specifically include the following steps: Collect data from multiple application sides of the virtual power plant according to multiple preset data types to obtain multiple relevant collected data; Determine multiple online edge data nodes; Monitor the status of multiple said edge data nodes to obtain multiple node status data; Analyze multiple said node status data to match multiple corresponding sharing ratios; According to multiple said sharing ratios, perform sharing and transmission of multiple said relevant collected data on multiple said edge data nodes; Perform data preprocessing of missing value filling, abnormal value elimination, and duplicate deletion on multiple corresponding relevant collected data through multiple said edge data nodes to obtain multiple effective collected data.

[0026] Furthermore, the virtual power plant data processing method further includes the following steps: Step S102: Transmit multiple said effective collected data to the cloud, and perform data integration and standardization through the cloud to generate standard integrated data.

[0027] In the embodiment of the present invention, determine the central cloud node corresponding to the cloud, upload multiple effective collected data to the central cloud node, and then through the central cloud node, adopt the method of data warehouse or data lake to perform data integration on multiple effective collected data from different sources and in different formats, form a unified data view to obtain integrated collected data, and then perform standardization processing on the integrated collected data to generate standard integrated data.

[0028] It can be understood that installing numerous application sides on-site will generate a large amount of data, and uploading all of them will cause a huge pressure on the cloud server. To share the pressure on the central cloud node, the edge data nodes need to be responsible for data calculation work within their own scope, and then converge the data to the cloud, and perform data integration and data mining processing through the central cloud node, such as Figure 3The figure shows a schematic diagram of cloud-edge collaboration provided by an embodiment of the present invention. The central cloud node combined with multiple edge data nodes can form a data processing form of cloud-edge collaboration, which can not only disperse the computing pressure of the central cloud node and improve the computing speed, but also ensure that the data in the cloud will not be lost if an edge data node fails and causes an exception or goes offline.

[0029] Specifically, in another preferred embodiment provided by the present invention, the steps of transmitting multiple pieces of the valid collected data to the cloud and performing data integration and standardization through the cloud to generate standard integrated data specifically include the following steps: Determine the central cloud node corresponding to the cloud; Upload multiple pieces of the valid collected data to the central cloud node; Through the central cloud node, use a data warehouse or a data lake to perform data integration on multiple pieces of the valid collected data to generate integrated collected data; Perform standardization processing on the integrated collected data to generate standard integrated data.

[0030] Furthermore, the virtual power plant data processing method further includes the following steps: Step S103: Perform attribute recognition on the standard integrated data, ignore the defective attribute data, and use the C4.5 decision tree algorithm to perform classification mining on the non-defective attribute data to obtain classification mining data.

[0031] In an embodiment of the present invention, perform attribute recognition on the standard integrated data, record the attribute recognition results, and then, according to the attribute recognition results, determine the defective attribute data and the non-defective attribute data from the standard integrated data, then ignore the defective attribute data and retain the non-defective attribute data. After that, use the C4.5 decision tree algorithm to perform classification mining on the non-defective attribute data to obtain classification mining data. For example, Figure 4 The figure shows a schematic diagram of the C4.5 decision tree provided by an embodiment of the present invention. Specifically, select multiple information features from the non-defective attribute data, calculate the information gain of the multiple information features, then supplement the multiple information gains, calculate the multiple information gain rates, and by comparing the multiple information gain rates, select split nodes from the multiple corresponding information features to construct a decision tree, and perform classification and mining on the data to obtain classification mining data.

[0032] It can be understood that the C4.5 decision tree algorithm is used as the core of data mining to classify a large amount of data purposefully, search for effective information beneficial to decision-making, and thus be able to obtain classified and mined data; the C4.5 decision tree algorithm first discretizes the data to be classified with continuous attributes, sorts them before and after according to the data segments, and takes the center point of the data segment as the data classification point. By processing different data in batches, the size of the information gain rate can be calculated to obtain the optimal classification calculation metric.

[0033] Specifically, in another preferred embodiment provided by the present invention, the attribute recognition of the standard integrated data, ignoring the defective attribute data, using the C4.5 decision tree algorithm to classify and mine the non-defective attribute data, and obtaining the classified and mined data specifically includes the following steps: Perform attribute recognition on the standard integrated data and record the attribute recognition results; According to the attribute recognition results, from the standard integrated data, ignore the defective attribute data and retain the non-defective attribute data; Use the C4.5 decision tree algorithm to classify and mine the non-defective attribute data to obtain the classified and mined data.

[0034] Specifically, in another preferred embodiment provided by the present invention, the classification and mining of the non-defective attribute data to obtain the classified and mined data specifically includes the following steps: Select multiple information features from the non-defective attribute data; Calculate the information gain of multiple information features; Supplement multiple information gains and calculate multiple information gain rates; Compare multiple information gain rates, select split nodes from multiple corresponding information features, construct a decision tree, and obtain the classified and mined data.

[0035] Furthermore, the virtual power plant data processing method further includes the following steps: Step S104: Perform outlier detection on the data points of the classified and mined data, calculate multiple local outlier factors, and perform anomaly verification and processing.

[0036] In the embodiment of the present invention, input the classified and mined data and specify the data points to be evaluated that have not been traversed , and then, calculate the th reachable distance between the data point to be evaluated and the neighbor data points , where is the number of neighbors of the data point to be evaluated . Then, according to multiple reachable distances, calculate the data point to be evaluated Local reachability density. Based on the local reachability density, calculate the data points to be evaluated of the local outlier factor. By sorting the local outlier factors corresponding to multiple data points to be evaluated in descending order, determine multiple outlier data points, and perform data elimination processing on the multiple outlier data points. Specifically, the data point to be evaluated and the neighbor data point The calculation formula for the reachability distance of the th is as follows: ; Among them, is the reachability distance of the th neighbor data point of the data point to be evaluated th ; is the th distance of the data point to be evaluated is the data point to be evaluated to the neighbor data point the distance between; The calculation formula for the local reachability density of the data point to be evaluated is as follows: ; Among them, is the th distance neighborhood of the data point to be evaluated; The calculation formula for the local outlier factor of the data point to be evaluated is as follows: ; Among them, is the local reachability density of the neighbor data point .

[0037] It can be understood that by sorting the local outlier factors corresponding to the calculated multiple data points to be evaluated in descending order, if the first data points are separated from the population data, that is, there are anomalies, the purpose of verifying the mined data is achieved.

[0038] Specifically, in another preferred embodiment provided by the present invention, the outlier detection of data points for the classified and mined data, calculating multiple local outlier factors, and performing anomaly verification and processing specifically include the following steps: Input the classified and mined data and specify value and the data points not yet traversed ; Calculate the data point​ With data points The Reachable distance; According to multiple said Reachable distances, calculate the data points Local reachability density; Based on the local reachability density, calculate the data points Local outlier factor; Arrange multiple said local outlier factors in descending order, determine multiple outlier data points and perform data elimination processing.

[0039] Furthermore, Figure 2 Shows the application architecture diagram of the virtual power plant data processing system provided by the embodiments of the present invention.

[0040] Specifically, in another preferred embodiment provided by the present invention, a virtual power plant data processing system includes: Edge acquisition and processing module 101, configured to collect data from multiple application terminals according to multiple preset data types, and perform dynamic data sharing and preprocessing through multiple edge data nodes to obtain multiple effective acquisition data.

[0041] In the embodiments of the present invention, the edge acquisition and processing module 101 collects data from multiple application terminals of the virtual power plant according to multiple preset data types such as energy storage state, power generation, equipment state, environmental parameters, and compliance requirements, obtains multiple relevant acquisition data, and determines multiple edge data nodes in the online state (multiple edge data nodes include both edge data nodes corresponding to multiple application terminals and other edge data nodes). By monitoring the states of multiple edge data nodes, multiple node state data are obtained, and then the multiple node state data are analyzed. According to the corresponding computing power, memory, and network conditions, the sharing ratios corresponding to multiple edge data nodes are matched. Then, according to multiple sharing ratios, multiple relevant acquisition data are divided and shared and then transmitted to multiple corresponding edge data nodes. Furthermore, through multiple edge data nodes, missing value filling (for numerical data, column mean is used for missing value filling; for skewed distribution data, column median is used for missing value filling; for categorical data, the value with the highest frequency in the column is used for missing value filling), outlier elimination (using the Z-score method or the IQR method for outlier identification and elimination processing), and duplicate deletion and other data preprocessing are performed on multiple corresponding relevant acquisition data to obtain multiple effective acquisition data.

[0042] Cloud upload and processing module 102, configured to transmit multiple said effective acquisition data to the cloud and perform data integration and standardization through the cloud to generate standard integrated data.

[0043] In an embodiment of the present invention, the cloud upload processing module 102 determines the central cloud node corresponding to the cloud, uploads multiple valid collected data to the central cloud node, and then, through the central cloud node, integrates the multiple valid collected data from different sources and in different formats in the form of a data warehouse or a data lake to form a unified data view, obtaining integrated collected data, and then performs standardization processing on the integrated collected data to generate standard integrated data.

[0044] Specifically, in another preferred embodiment provided by the present invention, the cloud upload processing module 102 specifically includes: A node determination unit, configured to determine the central cloud node corresponding to the cloud; A cloud upload unit, configured to upload the multiple valid collected data to the central cloud node; A data integration unit, configured to, through the central cloud node, use a data warehouse or a data lake to perform data integration on the multiple valid collected data to generate integrated collected data; A standardization processing unit, configured to perform standardization processing on the integrated collected data to generate standard integrated data.

[0045] Furthermore, the virtual power plant data processing system further includes: A data classification and mining module 103, configured to perform attribute recognition on the standard integrated data, ignore the defective attribute data, and use the C4.5 decision tree algorithm to perform classification and mining on the non-defective attribute data to obtain classification and mining data.

[0046] In an embodiment of the present invention, the data classification and mining module 103 performs attribute recognition on the standard integrated data, records the attribute recognition result, and then, according to the attribute recognition result, determines the defective attribute data and the non-defective attribute data from the standard integrated data, then ignores the defective attribute data and retains the non-defective attribute data. After that, the C4.5 decision tree algorithm is used to perform classification and mining on the non-defective attribute data to obtain classification and mining data, as Figure 4 shows a schematic diagram of the C4.5 decision tree provided by an embodiment of the present invention. Specifically, multiple information features are selected from the non-defective attribute data, the information gain of the multiple information features is calculated, then the multiple information gains are supplemented, the multiple information gain rates are calculated, and by comparing the multiple information gain rates, a split node is selected from the multiple corresponding information features to construct a decision tree to classify and mine the data to obtain classification and mining data.

[0047] Specifically, in another preferred embodiment provided by the present invention, the data classification and mining module 103 specifically includes: An attribute recognition unit, configured to perform attribute recognition on the standard integrated data and record the attribute recognition result; A data selection unit, configured to, according to the attribute recognition result, ignore the defective attribute data in the standard integrated data and retain the non-defective attribute data; A classification and mining unit, configured to perform classification and mining on the non-defective attribute data by using the C4.5 decision tree algorithm to obtain classification and mining data.

[0048] Furthermore, the virtual power plant data processing system further includes: An abnormal checking and processing module 104, configured to perform outlier detection on data points of the classification and mining data, calculate multiple local outlier factors, and perform abnormal checking and processing.

[0049] In an embodiment of the present invention, when inputting the classification and mining data, the abnormal checking and processing module 104 designates the data points to be evaluated that have not been traversed , and then calculates the th reachable distance between the data point to be evaluated and its neighbor data points . Among them, is the number of neighbors of the data point to be evaluated . Then, according to multiple th reachable distances, calculate the local reachability density of the data point to be evaluated . On the basis of the local reachability density, calculate the local outlier factor of the data point to be evaluated . By sorting the local outlier factors corresponding to multiple data points to be evaluated in descending order, determine multiple abnormal data points, and perform data elimination processing on the multiple abnormal data points. Specifically, the th reachable distance between the data point to be evaluated and its neighbor data points is calculated by the following formula: ; Among them, is the th reachable distance between the data point to be evaluated and its neighbor data points , is the th distance of the data point to be evaluated , is the distance from the data point to be evaluated to its neighbor data points ; The formula for calculating the local reachability density of the data point to be evaluated is: ; Among them, is the th Distance neighborhood; Data point to be evaluated The calculation formula of the local outlier factor is as follows: ; Wherein, is the neighbor data point The local reachability density of.

[0050] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0051] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0052] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.

Claims

1. A method for processing virtual power plant data, characterized in that, The method includes the following steps: Collect data from multiple application terminals according to multiple preset data types, and perform dynamic data sharing and preprocessing through multiple edge data nodes to obtain multiple effective collected data; Transmit the multiple effective collected data to the cloud, and perform data integration and standardization through the cloud to generate standard integrated data; Perform attribute recognition on the standard integrated data, ignore the defective attribute data, and use the C4.5 decision tree algorithm to classify and mine the non-defective attribute data to obtain classified and mined data; Perform outlier detection on the data points of the classified and mined data, calculate multiple local outlier factors, and perform anomaly verification and processing.

2. The virtual power plant data processing method according to claim 1, wherein The step of collecting data from multiple application terminals according to multiple preset data types, and performing dynamic data sharing and preprocessing through multiple edge data nodes to obtain multiple effective collected data specifically includes the following steps: Collect data from multiple application terminals of the virtual power plant according to multiple preset data types to obtain multiple relevant collected data; Determine multiple online edge data nodes; Monitor the status of the multiple edge data nodes to obtain multiple node status data; Analyze the multiple node status data and match multiple corresponding sharing ratios; According to the multiple sharing ratios, perform sharing transmission of the multiple relevant collected data on the multiple edge data nodes; Through the multiple edge data nodes, perform data preprocessing of missing value filling, anomaly elimination, and duplicate deletion on the multiple corresponding relevant collected data to obtain multiple effective collected data.

3. The virtual power plant data processing method according to claim 1, wherein The step of transmitting the multiple effective collected data to the cloud, and performing data integration and standardization through the cloud to generate standard integrated data specifically includes the following steps: Determine the central cloud node corresponding to the cloud; Upload the multiple effective collected data to the central cloud node; Through the central cloud node, use a data warehouse or a data lake to perform data integration on the multiple effective collected data to generate integrated collected data; Perform standardization processing on the integrated collected data to generate standard integrated data.

4. The virtual power plant data processing method according to claim 1, wherein The step of performing attribute recognition on the standard integrated data, ignoring the defective attribute data, and using the C4.5 decision tree algorithm to classify and mine the non-defective attribute data to obtain classified and mined data specifically includes the following steps: Perform attribute recognition on the standard integrated data and record the attribute recognition results; According to the attribute recognition results, ignore the defective attribute data in the standard integrated data and retain the non-defective attribute data; Use the C4.5 decision tree algorithm to classify and mine the non-defective attribute data to obtain classified and mined data.

5. The virtual power plant data processing method according to claim 4, characterized in that The step of classifying and mining the non-defective attribute data to obtain classified and mined data specifically includes the following steps: Select multiple information features from the non-defective attribute data; Calculate the information gain of the multiple information features; Supplement the multiple information gains and calculate multiple information gain ratios; Compare the multiple information gain ratios, select split nodes from the multiple corresponding information features, construct a decision tree, and obtain classified and mined data.

6. The virtual power plant data processing method according to claim 1, characterized in that, Perform outlier detection on the data points of the classified and mined data, calculate multiple local outlier factors, and perform anomaly verification and processing, which specifically includes the following steps: Input the classified mining data and specify the data points to be evaluated that have not been traversed ; Calculate the data point to be evaluated with the neighboring data points for the reachability distance, where is the number of neighbors of the data point to be evaluated ; According to the said reachable distance, calculate the data point to be evaluated local reachability density; Calculate the local outlier factor of the data point to be evaluated based on the local reachability density ; Arrange the multiple local outlier factors in descending order, determine multiple outlier data points, and perform data elimination processing.

7. The virtual power plant data processing method according to claim 6, characterized in that The data point to be evaluated and the neighbor data points The calculation formula for the reachable distance is as follows: ; Among them, is the data point to be evaluated and the k-th reachable distance to the neighbor data point The k-th reachable distance is the k-th distance of the data point to be evaluated and the distance between the data point to be evaluated and the neighbor data point is the distance between them; The data point to be evaluated The calculation formula for the local reachability density is as follows: ; Among them, is the th distance neighborhood; The data point to be evaluated The calculation formula for the local outlier factor is as follows: ; Among them, is the local reachability density of the neighbor data points .

8. A virtual power plant data processing system for performing the virtual power plant data processing method as described in any one of claims 1 to 7, characterized in that, The system includes an edge acquisition and processing module, a cloud upload and processing module, a data classification and mining module, and an anomaly verification and processing module, where: The edge acquisition and processing module is used to collect data from multiple application terminals according to multiple preset data types, and perform dynamic data sharing and preprocessing through multiple edge data nodes to obtain multiple effective acquisition data; The cloud upload and processing module is used to transmit the multiple effective acquisition data to the cloud, and perform data integration and standardization through the cloud to generate standard integrated data; The data classification and mining module is used to perform attribute recognition on the standard integrated data, ignore the defective attribute data, and use the C4.5 decision tree algorithm to perform classification and mining on the non-defective attribute data to obtain classified and mined data; The anomaly verification and processing module is used to perform outlier detection on the data points of the classified and mined data, calculate multiple local outlier factors, and perform anomaly verification and processing.

9. The virtual power plant data processing system according to claim 8, wherein The cloud upload and processing module specifically includes: A node determination unit for determining the central cloud node corresponding to the cloud; A cloud upload unit for uploading the multiple effective acquisition data to the central cloud node; A data integration unit for integrating the multiple effective acquisition data through the central cloud node using a data warehouse or a data lake to generate integrated acquisition data; A standardization processing unit for performing standardization processing on the integrated acquisition data to generate standard integrated data.

10. The virtual power plant data processing system according to claim 8, wherein The data classification and mining module specifically includes: An attribute recognition unit for performing attribute recognition on the standard integrated data and recording the attribute recognition results; A data selection unit for ignoring the defective attribute data and retaining the non-defective attribute data from the standard integrated data according to the attribute recognition results; A classification and mining unit for performing classification and mining on the non-defective attribute data using the C4.5 decision tree algorithm to obtain classified and mined data.

Citation Information

Patent Citations

  • A real-time energy consumption abnormity detection method with college building structure characteristics being combined

    CN106250905A

  • Building energy consumption integrated analysis method based on data mining

    CN113408659A

  • Power distribution network data management method under hierarchical collaborative architecture

    CN115566668A

  • Edge computing system based on industrial Internet of Things platform

    CN117097727A

  • Data processing method and device, electronic equipment, system and storage medium

    CN119360605A

Cited By

  • Virtual power plant evaluation index reduction method and system based on hybrid strategy decision tree

    CN122196756A