A data processing method and a data processing system in an energy internet
By conducting multi-level evaluation and correction of data from energy hub stations, the problem of inconsistent data quality has been solved, the data quality and system stability of the energy internet have been improved, and efficient data processing and storage have been achieved.
Patent Information
- Application Number
- CN202011576479.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-28
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2040-12-28
AI Technical Summary
The lack of effective methods for processing data from energy hubs in existing technologies leads to inconsistent data quality, affecting the stable operation of the energy internet system and the fairness of transactions.
A data processing method and system are provided, which receives and classifies static and dynamic data, determines relevant parameters such as accuracy, completeness, consistency and timeliness, performs multi-level evaluation and correction, and realizes quantitative quality assessment and classified storage of data.
This has improved the quality and credibility of data in the energy internet, ensuring the accuracy, integrity, and timeliness of data, and enhancing the stability and fairness of the system and transactions.
Smart Images

Figure CN114691654B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the energy internet, and more particularly, to a method for processing multi-station fusion data in the energy internet. BACKGROUND
[0002] The operation and management of smart cities need energy and information support under the integration of energy information infrastructure, and multi-station fusion is an important way to achieve the integration of energy information infrastructure. It is characterized by unified planning, unified construction and unified operation of energy stations, energy storage stations and data stations for different business applications, and is the main content of smart city construction. Multi-station fusion can jointly manage and optimize the efficiency of energy systems and the operation efficiency of data centers by organically combining distributed data centers with various energy facilities, perform real-time processing and control of data, quickly form system optimization control strategies, provide data center services, and create a coordinated control center of regional "energy flow + information flow + value flow", and finally improve the intelligent operation level of energy information stations through big data analysis and new generation artificial intelligence technology.
[0003] Renewable new energy such as wind, light and water is gradually reducing the degree of dependence of human beings on fossil energy. In the world, countries such as the European Union, the United States and China have successively proposed aggressive targets of achieving 100%, 80% and 50%-70% of renewable new energy in the energy supply structure by 2050, which has promoted the large-scale popularization of distributed new energy stations at the user side. Internet relies on multiple communication connection modes, adopts a hierarchical design structure to shield the complex networking protocol at the bottom, provides great convenience for people to obtain and utilize data information, changes the previous means of communication and information exchange, and also reshapes the production and operation mode of many traditional industries. These technical achievements all contribute to the realization of the structural adjustment goal of the energy industry in the middle of this century. The combination of Internet information technology and renewable new energy generation technology, as an important part of the energy internet, on the one hand, can achieve the goal of clean and low-carbon energy, and on the other hand, can improve the efficiency of energy utilization and fair trade, and become an inevitable way to build the energy system of future smart cities.
[0004] Multi-station fusion is to extend, expand and integrate three stations (energy station, energy storage station and data station) into a regional "energy flow + information flow + value flow" hub station, also known as an energy hub station. In the operation process of the energy hub station, a large amount of data will be collected, and the quality of these data is uneven, which will have an adverse effect on the stable operation of the energy internet system. Especially for the energy internet system based on information-energy coupling, its dependence on data leads to its particular sensitivity to fluctuations in data quality. At present, there is a lack of effective processing method for data from the energy hub station, it is difficult to monitor and ensure the quality of the fusion data from the energy hub station, once the data from the energy hub station has a problem, it will affect the safe operation of supply and demand and the fairness of transaction in the entire energy internet. SUMMARY
[0005] The present disclosure is provided to solve the above problems existing in the prior art.
[0006] There is a need for a data processing method and data processing system in an energy internet, which can perform targeted and efficient quantitative quality assessment on data from an energy hub station for various aspects of static and dynamic main application scenarios, and accordingly correct and classify the storage of data, thereby improving the overall quality and credibility of data in the energy internet.
[0007] According to a first aspect of the present disclosure, a data processing method in an energy internet is provided. The data processing method can include receiving data from a plurality of energy hub stations. The method can also include dividing, with at least one processor, the received data into static data and dynamic data, further dividing the static data into various data including facility-related data, ownership-related data and transaction configuration-related data of the energy hub station, and further dividing the dynamic data into various data including external environment-related data, internal operation-related data and transaction dynamic data of the energy hub station. The accuracy-related parameters, the integrity-related parameters, the consistency-related parameters and the timeliness-related parameters of the various data can be determined respectively with the at least one processor, and the quality evaluation parameters of the various data can be determined by integrating the accuracy-related parameters, the integrity-related parameters, the consistency-related parameters and the timeliness-related parameters.
[0008] According to a second aspect of this disclosure, a data processing system for an energy internet is provided. The data processing system may include an interface and at least one processor. The interface is configured to receive data from multiple energy hub stations. The at least one processor is configured to perform the following steps: The received data can be divided into static data and dynamic data; the static data can be further divided into various types of data including facility-related data, attribution-related data, and transaction configuration-related data of the energy hub stations; the dynamic data can be further divided into various types of data including external environment-related data, internal operation-related data, and transaction dynamic data of the energy hub stations. For each type of data, accuracy-related parameters, integrity-related parameters, consistency-related parameters, and timeliness-related parameters can be determined respectively, and quality assessment parameters for each type of data can be determined by comprehensively considering the accuracy-related parameters, integrity-related parameters, consistency-related parameters, and timeliness-related parameters.
[0009] By utilizing the data processing methods and systems in the energy internet according to various embodiments of this disclosure, targeted and efficient quantitative quality assessments can be performed on data from energy hub stations for key application scenarios in both static and dynamic aspects, and the data can be corrected and classified for storage accordingly, thereby improving the overall quality and reliability of data in the energy internet. Attached Figure Description
[0010] In drawings that are not necessarily drawn to scale, the same reference numerals may describe similar parts in different views. The same reference numerals with or without letter suffixes may indicate different instances of similar parts. The drawings illustrate various embodiments generally by way of example rather than limitation, and are used, together with the description and claims, to illustrate the disclosed embodiments. Such embodiments are illustrative and not intended to be exhaustive or exclusive embodiments of the apparatus or method.
[0011] Figure 1 A schematic diagram of a data processing method and a data processing system in an energy internet according to embodiments of the present disclosure is shown.
[0012] Figure 2 A flowchart illustrating a data processing method in an energy internet according to an embodiment of the present disclosure is shown.
[0013] Figure 3 An exemplary schematic diagram illustrates a sub-process of pre-classifying data from an energy hub station in a data processing method for an energy internet according to embodiments of the present disclosure; and
[0014] Figure 4 A flowchart illustrating a data processing method in an energy internet according to an embodiment of this disclosure is shown. Detailed Implementation
[0015] For those skilled in the art to better understand the technical solutions of the present disclosure, the present disclosure will be described in detail below in combination with the drawings and specific embodiments. The embodiments of the present disclosure will be described in further detail below in combination with the drawings and specific embodiments, but not as a limitation on the present disclosure. The wording of "first", "second" and "third" used in the present disclosure is only intended to distinguish the corresponding features, and does not represent the need for such an order, nor does it necessarily represent only the singular form. The execution order of the processing steps in the text is only as an example, as long as the logical relationship of each step is not affected, the execution order of each step can be appropriately adjusted, and the separate steps can also be integrated and executed.
[0016] Figure 1 The schematic diagram of the data processing method and the data processing system in the energy internet according to the embodiments of the present disclosure is shown. As shown in Figure 1 The energy internet includes energy hub station 1 100, energy hub station 2 100, …, energy hub station n 100 (n is a natural number), each energy hub station 100 generates and collects a large amount of data during operation, which integrates information, energy and value aspects, also known as fusion data. The data from each energy hub station 100 can be stored in a distributed data storage (not shown in the figure, for example but not limited to a distributed database, etc.) after processing.
[0017] The data processing method in the energy internet will be described below in combination with Figure 1 and Figure 2 The data processing method in the energy internet will be described below in combination with
[0018] As Figure 1As shown, the data processing system can include the at least one processor 101 and an interface (not shown in the figure). The interface can be configured to enable the at least one processor 101 to receive data from a plurality of energy hub stations 100 for processing. In some embodiments, various communication interfaces can be provided for each energy hub station 100 so that a distributed data processing system can acquire data from each energy hub station 100 via various networks through the communication interfaces. These networks can employ any of wired public networks, wireless public networks, private dedicated networks, etc. In some embodiments, a communication protocol can be customized or a general communication protocol can be employed for communication between each energy hub station 100 and the data processing system.
[0019] At step 202, the received data can be divided into static data 102 and dynamic data 103 by the at least one processor 101. Dynamic data 103 can represent data that changes over time, and static data 102 can represent data that remains substantially stable over time. By dividing static data 102 and dynamic data 102, the accuracy of data feature extraction can be improved, and the processing efficiency can be improved. For example, as will be described in detail below, the correction based on the sample feature curve has a better correction effect on dynamic data 103 (i.e., time series data) than on static data 102. This correction can not be applied to static data 102, which can be time-consuming (because static data 102 can exhibit a sample feature curve over a long period of time), and can even introduce inappropriate correction of reasonable bias and waste of computing resources of the entire system. In general, static data 102 is more accurate, and a higher accuracy-related parameter 104a can be assigned in advance, and the correction based on the sample feature curve is not performed by default. By dividing dynamic data 102 and implementing the correction based on the sample feature curve specifically for dynamic data 102, the use efficiency of computing resources can be improved and inappropriate correction can be avoided.
[0020] Further, at step 203, the static data 102 can be further divided into various data sub-classes including facility-related data 102a, ownership-related data 102b, and transaction configuration-related data 102c of the energy hub station 100, and the dynamic data 103 can be further divided into various data sub-classes including external environment-related data 103a, internal operation-related data 103b, and transaction dynamic data 103c of the energy hub station 100.
[0021] The static data 102 can represent the stable physical basic information of the energy hub station (e.g., the new energy plant station therein) in a long time period, and can mainly include facility-related data 102a, ownership-related data 102b, and transaction configuration-related data 102c. For example, the facility-related data 102a can include model parameter data of photovoltaic components, wind turbine components, inverters, batteries, etc., and set working condition parameters such as installed capacity, etc. The ownership-related data 102b can represent the ownership relationship of the energy hub station (e.g., the new energy plant station therein) in the geographical and management aspects, including but not limited to data such as station name, location positioning, belonging area, station-line-transformer relationship, belonging power generation group, grid-connected voltage level, consumption mode, unified / non-unified management mode, etc. For example, the transaction configuration-related data 102c can represent the set configuration mode of the energy hub station (e.g., the new energy plant station therein) associated with transactions, including but not limited to data such as whether the power generation plan is medium and long term or day-ahead, transaction mode, transaction type, transaction mode, and transaction evaluation mode, etc.
[0022] The dynamic data 103 can represent the dynamic operation monitoring data generated in the operation and transaction process of the new energy plant station, and mainly involves time series data. In some embodiments, the external environment-related data 103a, the internal operation-related data 103b, and the transaction dynamic data 103c of the energy hub station 100 can constitute the main components of the dynamic data 103. For example, the external environment-related data 103a can include but is not limited to data such as wind speed, light intensity, environmental temperature, humidity, wind power, sunshine hours, etc. For example, the internal operation-related data 103b can include but is not limited to operation state sequence data such as gear speed, component temperature, oil temperature, oil pressure, output power, output voltage, output current, etc., control state data such as action instructions, alarm codes, and on / off, and recording data in emergency and fault states. For example, the transaction dynamic data 103c can include but is not limited to market transaction sequence data such as node electricity price, transaction electricity quantity, transaction price, transaction time, credit evaluation, and data related to new energy such as equipment price, manufacturer stock price, futures, options, etc.
[0023] The above comprehensively considers various business application analyses oriented to the fusion data in the energy internet, and performs various subclass subdivisions on the static data 102 and the dynamic data 103. Thus, most of the various rich data from the energy hub station 100 can be subdivided into various subclasses, and a unified data structure (such as but not limited to an instance table structure, a business table, etc.) is adopted for various data in the same subclass, so as to facilitate subsequent application analysis and correction (if necessary).
[0024] The at least one processor 101 can determine (independently for each subcategory) the accuracy-related parameter 104a, the integrity-related parameter 104b, the consistency-related parameter 104c, and the timeliness-related parameter 104d for the various data, such as but not limited to the facility-related data 102a, the home-related data 102b, the transaction configuration-related data 102c, the external environment-related data 103a, the internal operation-related data 103b, and the transaction dynamics data 103c, at step 204. Note that in Figure 1 the figure reference is only provided for the accuracy-related parameter 104a, the integrity-related parameter 104b, the consistency-related parameter 104c, and the timeliness-related parameter 104d of the facility-related data 102a, while the figure reference for the accuracy-related parameter, the integrity-related parameter, the consistency-related parameter, and the timeliness-related parameter of other subcategories is omitted.
[0025] By performing multi-level evaluation of the various data from the 4 dimensions of accuracy, integrity, consistency, and timeliness, comprehensive quantitative analysis of abnormal data and unreasonable data of the fusion data in the energy internet can be achieved. The accuracy-related parameter 104a, the integrity-related parameter 104b, the consistency-related parameter 104c, and the timeliness-related parameter 104d will be illustrated below. The multi-level evaluation of the 4 dimensions is particularly effective for various causes of abnormality and unreasonableness of the fusion data in the energy internet, such as data missing, data record format not matching business table structure attribute, data being too old or outdated, a few outliers deviating from normal values, etc., and can achieve high detection rate for abnormal data and unreasonable data of the fusion data in the energy internet.
[0026] In some embodiments, the accuracy-related parameter 104a, the integrity-related parameter 104b, the consistency-related parameter 104c, and the timeliness-related parameter 104d can be defined as follows.
[0027] The accuracy-related parameter 104a can be determined based on the ratio of the number of data points deviating from the sample feature curve to the number of data points of the sample curve, so that the higher the ratio, the lower the accuracy-related parameter. The sample feature curve can be determined based on historical data using various clustering algorithms. In some embodiments, the clustering algorithm can employ supervised or unsupervised clustering algorithms. In some embodiments, unsupervised clustering algorithms such as k-means clustering algorithm, graph-based clustering algorithm, etc. can be employed to cope with the problem of lack of ground truth in data of the energy internet.
[0028] In some embodiments, the accuracy analysis can be a statistical analysis of the true level of data records. After calculating the sample data feature curve using the clustering algorithm, the accuracy-related parameter Q1 can be calculated according to formula (1):
[0029]
[0030] where ∑out of characteristic curve can represent the total number of data points deviating from the sample characteristic curve, and ∑number of lineData can represent the total number of data points of the sample curve (i.e. the total number of data points that the accuracy analysis is directed to).
[0031] The integrity-related parameter 104b can be determined based on the ratio of the amount of missing data to the total amount of data, such that the higher the ratio, the lower the integrity-related parameter. In some embodiments, the integrity-related parameter Q2 can be calculated according to formula (2):
[0032]
[0033] where ∑Lines of LossData can represent the amount of missing data (which can also be obtained by summing the records of data loss), and ∑Lines of grossData can represent the total amount of data.
[0034] The consistency-related parameter 104c can represent the degree of matching between the defined attributes (such as but not limited to format) of the load data records and the defined attributes of the load table (such as the business table) structure. In some embodiments, the consistency analysis is a matching analysis of the attribute definitions (such as but not limited to format) of the load data records and the attribute definitions of the load table structure. For example, the consistency-related parameter Q3 can be calculated according to formula (3):
[0035]
[0036] where ∑column of load Data file can represent the attribute definitions of the real load data records, and ∑column of load attributes can represent the attribute definitions of the load table.
[0037] The timeliness-related parameter 104d can be determined based on the ratio of the amount of updated data to the total amount of data. In some embodiments, the timeliness analysis can be a statistical analysis of the updating of the data records. For example, the timeliness-related parameter Q4 can be calculated according to formula (4):
[0038]
[0039] where ∑lines of updateData can represent the amount of updated data, and ∑lines of grossData can represent the total amount of data.
[0040] Then, the quality evaluation parameter of various data can be determined by integrating the accuracy related parameter 104a, the integrity related parameter 104b, the consistency related parameter 104c and the timeliness related parameter 104d at step 205. In this way, good robustness of the quality evaluation parameter of various data can be achieved. In some embodiments, corresponding weights can be applied to the accuracy related parameter, the integrity related parameter, the consistency related parameter and the timeliness related parameter respectively, so as to consider the evaluation indicator parameters of the four aspects differently and comprehensively. In some embodiments, in view of the cause distribution of errors of the data transmitted in the energy internet, the weight of the accuracy related parameter, the weight of the integrity related parameter, the weight of the consistency related parameter and the weight of the timeliness related parameter can be sequentially reduced.
[0041] For example, the quality evaluation parameter Q of various data can be determined by using the weighted sum of Q1, Q2, Q3 and Q4 according to formula (5):
[0042] Q = w1*Q1 + w2*Q2 + w3*Q3 + w4*Q4 Formula (5)
[0043] wherein w1, w2, w3 and w4 are the weights of Q1, Q2, Q3 and Q4 respectively.
[0044] In some embodiments, each weight can be determined respectively by forming a judgment matrix and based on the judgment matrix. Specifically, the relative importance of accuracy, integrity, consistency and timeliness can be judged and represented by a matrix to form a judgment matrix. For example, the relative importance of each level factor can be determined according to the proportion of reports of errors or abnormalities in historical data. The eigenvector corresponding to the maximum eigenvalue of the judgment matrix can be determined, and after normalization, the relative weight values w1, w2, w3 and w4 of each level factor (e.g. the accuracy related parameter 104a, the integrity related parameter 104b, the consistency related parameter 104c and the timeliness related parameter 104d) can be obtained, and it can be verified whether the consistency condition is met and adjusted accordingly to obtain the weights w1, w2, w3 and w4 that meet the consistency condition.
[0045] In some embodiments, in order to calculate the accuracy related parameter 104a, especially the accuracy related parameter 104a of various data subclasses of dynamic data, the sample feature curve of historical data needs to be calculated. As an example, the following steps can be used to calculate via a k value clustering algorithm.
[0046] The sub-flow of calculation can start with reading a sample data set M = {x1, x2, x3, …, xn} from a database, where n is a natural number, and x n = {x i = {x i1x i2 ,…,x it}, the Euclidean distance between each sample data is calculated as the distance metric D(x i ,x j ). Other distance metrics can also be used, and the Euclidean distance is used as an example of a distance metric herein. The sequence of sample data of the i-th class is denoted as x i , and it denotes the number of samples in the i-th class. The number of clusters (i.e., classes) k that is expected to be obtained can be preset.
[0047] k objects can be selected from M as initial cluster centroids, and each object can be classified into the nearest cluster according to the nearest-neighbor classification method.
[0048] The classification can be modified using a batch modification method. The cluster centers and the classification can be modified after all objects are input. Specifically, the steps of the batch modification method can be as follows.
[0049] First, the number of classes k is selected, and then k values are selected from all sample data as initial center points, each center point belongs to an independent class, and the set of classes to which the center points belong can be denoted as S={s1,s2,…,s k}.
[0050] All sample data is classified into the class to which the nearest center point belongs, and the center point of the class is recalculated and replaced with the new center point. For example, if D(x i ,s j ) is the minimum value, then x i ∈s j can be determined, and the new center point can be calculated according to formula (6).
[0051]
[0052] where C k denotes the new center point in the class to which the k-th center point belongs.
[0053] In some embodiments, if the difference between the center point before updating and the center point after updating is small (e.g., less than a certain threshold), the classification can be stopped, otherwise the updating iteration process is continued. In some embodiments, a preset condition can also be set, and the updating iteration process can be ended when the preset condition is met. For example, a squared error criterion function can be used as the preset condition, and k clusters that minimize the squared error criterion function are finally obtained. The historical load data can be processed by the K value clustering to obtain representative load feature curves. For example, but not by way of limitation, the daily load feature curve can be extracted in combination with the effective index criterion.
[0054] Figure 3 An exemplary schematic diagram showing a sub-process of pre-data classification of data from an energy hub station in a data processing method in an energy internet according to an embodiment of the present disclosure is shown, which can be performed in advance of accuracy-related parameters, integrity-related parameters, consistency-related parameters and timeliness-related parameters of various data. As shown, Figure 3 As shown, domain analysis 301 can be performed on the data. The domain analysis 301 is mainly to analyze and identify the source of the data from the energy hub station, to determine whether the data is from internal operation of the station, from external environment, from pre-setting of facilities (such as the station and equipment, etc.), or from pre-setting of transaction, etc. Through the domain analysis 301, various sub-classes of static data 102 and dynamic data 103 can be identified accordingly, for example, data from internal operation of the station can be identified as internal operation-related data 301a, which can be power station production operation data, to provide support for production decision of the superior dispatching management agency, and to accept the authority management of the superior dispatching management agency. For example, data from external environment can be identified as external environment-related data 301b, which can be meteorological environment data closely related to production operation, to provide high-precision weather data for superior dispatching planning arrangement. For example, data from pre-setting of facilities can be identified as facility-related data 301c, such as model data, which can include station model and equipment model, to provide accurate parameters for station-level and regional-level online analysis application. For example, data from pre-setting of transaction can be identified as transaction configuration-related data 301d. Through the domain analysis 301, sub-classes can be efficiently divided in a qualitative analysis manner, facilitating subsequent storage and processing of the data (including but not limited to upper-layer application analysis such as evaluation and correction).
[0055] The data can also be subjected to structure type analysis 302 to divide the data into structured data 302a, unstructured data 302b and semi-structured data 302c. Specifically, different reading modes can be adopted to obtain corresponding data according to the structure type of the data, to form an instance library for data quality analysis; after the domain analysis 301 and the structure type analysis 302, an instance table structure can be formed for different business application analysis, so as to be stored in different business tables after subsequent processing.
[0056] The above division of structure type and the sub-class division based on data source facilitate pre-processing of the data, including data cleaning, data integration, normalization processing and storage management, etc., so as to provide timely and accurate data for subsequent upper-layer application analysis based on the data.
[0057] Specifically, the at least one processor 101 can be used to convert unstructured data 302b and semi-structured data 302c into structured data 302a for uploading. For example, data extraction and filter cleaning operations can be performed on unstructured data 302b involving warehousing requirements, and after forming structured data 302a, a data insertion operation is performed on the corresponding database. For structured data 302a to be uploaded, at least one of the following steps can be performed.
[0058] In some embodiments, in the case of failed uploading or insertion, the number of updated data records (affecting the numerator of formula (4)) can be updated.
[0059] In some embodiments, for periodically uploaded data, the domain range and data type where the data is located can be determined, and in the case of missing data uploaded in high-priority records, the number of missing data records (affecting the numerator of formula (2)) can be updated.
[0060] In some embodiments, for periodically uploaded data, the domain range and data type where the data is located can be determined, and in the case of missing data uploaded in high-priority records, the number of missing data records (affecting the numerator of formula (2)) can be updated.
[0061] In some embodiments, for periodically uploaded data, the domain range and data type where the data is located can be determined, and in the case of missing data uploaded in high-priority records, the number of missing data records (affecting the numerator of formula (2)) can be updated.
[0062] In some embodiments, for periodically uploaded data, the domain range and data type where the data is located can be determined, and in the case of missing data uploaded in high-priority records, the number of missing data records (affecting the numerator of formula (2)) can be updated.
[0063] Further, the association between structured data 302a can be identified and integrated. For example, the association between structured data attributes can be identified, and isolated and one-sided data sets facing business systems can be simplified into a complete and comprehensive set of data, realizing the digital twin corresponding fusion of static data 102 and dynamic data 103, thereby providing timely and accurate data for subsequent data-based upper-layer application analysis.
[0064] Figure 4 A flowchart of a data processing method in an energy internet according to an embodiment of the present disclosure is shown. As shown in FIG. 1, the data processing method includes the following steps. Figure 4As shown, data from the energy hub station can be loaded at step 401. Four levels of analysis 402, i.e. integrity analysis, consistency analysis, timeliness analysis and accuracy analysis, can be performed for each subcategory of the loaded data to obtain accuracy related parameters, integrity related parameters, consistency related parameters and timeliness related parameters. Data quality assessment can be performed at step 403, for example, the accuracy related parameters, integrity related parameters, consistency related parameters and timeliness related parameters can be integrated to determine a quality assessment parameter of the various data.
[0065] Data correction processing can be performed at step 404, which is preferably but not limited to applicable to various data in dynamic data. First, current data can be loaded. A sample curve of the current data can be determined based on the current data by using the at least one processor. The sample curve of the current data can be compared with a sample feature curve of corresponding historical data, and an abnormal data point can be determined based on the comparison result. The abnormal data point can be corrected by using a corresponding segment of the sample feature curve relative to the abnormal data point.
[0066] In some embodiments, determining the abnormal data point based on the comparison result can further include determining a difference between the corresponding data points of the sample curve and the sample feature curve, and in a case where the difference between the corresponding data points exceeds a fluctuation threshold, the data point in the sample curve can be determined as the abnormal data point.
[0067] Specifically, after the feature curve is obtained, its smoothness is used to check abnormal data points in the historical data. Let L d be the sample curve, L t be the sample feature curve. With a certain point (i-point) in the sample curve L d , let the value of this point be L d (i), and let it be compared with the value of the same point of the daily load feature curve L t (i), and then calculate the fluctuation rate δ(i) of the two points, see formula (7) below.
[0068]
[0069] The normal range of the value change rate within a certain range in history can be counted, denoted as [+D, -D]. Then, whether the change rate of the i-th point of the sample data is within the normal range of [+D, -D] is compared, so as to determine whether the point is an abnormal data point. After the abnormal data is found, correction can be performed in time, and the corresponding segment of the feature curve can be shifted to the detected data. In some embodiments, in a case where the two ends of the corresponding segment of the feature curve cannot be exactly shifted to the two ends of the corresponding segment of the detected data (with deviation), the former can also be shifted to a position such that the average distance with the latter is minimized.
[0070] In some embodiments, the correcting the abnormal data points in the sample feature curve using the corresponding segment of the sample feature curve relative to the abnormal data points can further include correcting the abnormal data points according to the following formula:
[0071]
[0072] wherein L r represents the corrected sample curve, L t represents the sample feature curve, L d represents the sample curve before correction, the mth point to the nth point are abnormal data points, and i represents the serial number of the abnormal data points.
[0073] The correction can be mainly performed on the abnormal data points in the data without changing the normal data points, and the corrected data can be stored (step 405) for subsequent application analysis. In this way, the data points with large differences from the feature data can be effectively removed, thereby effectively ensuring the data quality, and the adjusted curve has better similarity and smoothness, which can lay a good foundation for subsequent application analysis.
[0074] In some embodiments, the data quality can be evaluated before and after the correction, which is also called post-evaluation and pre-evaluation. For example, after the abnormal data points are corrected, the at least one processor 101 can determine the accuracy-related parameter 104a, the integrity-related parameter 104b, the consistency-related parameter 104c, and the timeliness-related parameter 104d for each data of the static data 102 and the corrected dynamic data 103, respectively, and determine the quality evaluation parameter of the corresponding data after correction by comprehensively considering the accuracy-related parameter 104a, the integrity-related parameter 104b, the consistency-related parameter 104c, and the timeliness-related parameter 104d. By comparing the pre-evaluation result with the post-evaluation result, the effect of data correction on improving the data quality can be verified. In some embodiments, step 404 can also be iteratively performed and ended under certain conditions, for example, when the difference between the pre-evaluation result of the iterative correction step and the post-evaluation result of the iterative correction step is lower than a certain threshold, it is considered that the iteration has little effect on improving the data quality, and the correction has reached the desired effect, and the iteration can be ended.
[0075] Furthermore, although example embodiments have been described herein, the scope includes any and all embodiments having equivalent elements, modifications, omissions, combinations (e.g., of
[0076] The sequence of various steps in the present disclosure is merely exemplary, rather than limiting. The execution sequence of the steps can be adjusted without affecting the implementation of the present disclosure (without destroying the logical relationship between the required steps), and the various embodiments obtained after the adjustment still fall within the scope of the present disclosure.
[0077] The above description is intended to be illustrative, and not restrictive. For example, the above-described examples (or one or more aspects thereof) can be used in combination with each other. Other embodiments can be used as well, which will be apparent to those of ordinary skill in the art upon reading the above description. Further, in the specific embodiments described above, various features can be grouped together or divided into different features for the purpose of streamlining the disclosure. This should not be interpreted as a requirement that the claimed subject matter must claim features that are grouped together. Rather, the claimed subject matter can claim a less, and more, than all of the features of a particular disclosed embodiment. As such, the following claims are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment. The scope of the application should be determined, not by the embodiment described, but by the claims as follow.
Claims
1. A data processing method in an energy internet, characterized in that, The data processing method comprises: receiving data from a plurality of energy hub stations; wherein the energy hub stations at least involve energy stations, energy storage stations and data stations; dividing, by using at least one processor, the received data into static data and dynamic data, the dynamic data representing data that changes over time, the static data representing data that remains stable over time, further dividing the static data into various data including facility-related data, ownership-related data and transaction configuration-related data of the energy hub stations, and further dividing the dynamic data into various data including external environment-related data, internal operation-related data and transaction dynamic data of the energy hub stations; determining, by using the at least one processor, for the various data, accuracy-related parameters, integrity-related parameters, consistency-related parameters and timeliness-related parameters respectively, and determining quality evaluation parameters of the various data by integrating the accuracy-related parameters, the integrity-related parameters, the consistency-related parameters and the timeliness-related parameters; wherein the timeliness-related parameters are determined based on a ratio of updated data quantity to total data quantity; The data processing method further comprises, for the various data in the dynamic data: loading current data; determining, by using the at least one processor, a sample curve of the current data based on the current data; comparing the sample curve of the current data with a sample feature curve of corresponding historical data; determining abnormal data points based on the comparison result; and correcting the abnormal data points by using corresponding segments of the sample feature curve relative to the abnormal data points; wherein, the correcting the abnormal data points by using corresponding segments of the sample feature curve relative to the abnormal data points further comprises correcting the abnormal data points according to the following formula: Wherein, L r represents the corrected sample curve, L t represents the sample feature curve, L d represents the sample curve before correction, the mth point to the nth point are abnormal data points, and i represents the serial number of the abnormal data points.
2. The data processing method according to claim 1, characterized in that, the sample feature curve is determined based on corresponding historical data by using cluster analysis.
3. The data processing method of claim 1, wherein, Determining abnormal data points based on the comparison result further comprises: determining differences between corresponding data points of the sample curve and the sample feature curve; in the case that the differences between the corresponding data points exceed a fluctuation threshold, determining that the data points in the sample curve are abnormal data points.
4. The data processing method of claim 1, wherein, Further comprising, after correcting the abnormal data points, by using the at least one processor: determining accuracy-related parameters, integrity-related parameters, consistency-related parameters and timeliness-related parameters for the static data and the various data of the corrected dynamic data respectively, and determining quality evaluation parameters of the corresponding various data after correction by integrating the accuracy-related parameters, the integrity-related parameters, the consistency-related parameters and the timeliness-related parameters.
5. The data processing method of claim 1, wherein, The accuracy-related parameters are determined based on a ratio of the number of data points deviating from the sample feature curve to the number of data points of the sample curve, the integrity-related parameters are determined based on a ratio of missing data quantity to total data quantity, the consistency-related parameters represent a matching degree of attribute definitions of load data records and attribute definitions of load tables, and the timeliness-related parameters are determined based on a ratio of updated data quantity to total data quantity.
6. The data processing method according to claim 5, characterized in that, The determining the quality evaluation parameter of various data by synthesizing the accuracy-related parameter, the integrity-related parameter, the consistency-related parameter and the timeliness-related parameter further comprises: forming a judgment matrix; determining the weight of each of the accuracy-related parameter, the integrity-related parameter, the consistency-related parameter and the timeliness-related parameter based on the judgment matrix; and determining the weighted result of the accuracy-related parameter, the integrity-related parameter, the consistency-related parameter and the timeliness-related parameter based on the respective weight as the quality evaluation parameter of various data.
7. The data processing method of claim 1, wherein, Further comprising, before determining the accuracy-related parameter, the integrity-related parameter, the consistency-related parameter and the timeliness-related parameter of various data, further utilizing at least one processor: performing structure type analysis on the received data, the structure type comprising structured data, unstructured data and semi-structured data; utilizing the at least one processor, converting the unstructured data into structured data for uploading, identifying the relevance between structured data and integrating the same.
8. The data processing method according to claim 7, characterized in that, Further comprising, for the structured data to be uploaded, performing at least one of the following steps: updating the number of updated data records in the case of failed uploading or insertion; updating the number of missing data records in the case of missing data uploading in high-priority records for periodically uploaded data; updating the number of data consistency records in the case of duplication with other data records of the same attribute but inconsistent content for periodically uploaded data; updating the number of missing data records in the case of missing data uploading in high-priority records for fault event uploaded data; and updating the number of data consistency records in the case of duplication with other data records of the same attribute but inconsistent content for fault event uploaded data. The at least one processor is in the cloud.
9. The data processing method of claim 1, wherein, The data processing system comprises: 10.A data processing system in an energy internet, characterized by, an interface configured to receive data from a plurality of energy hub stations; wherein the energy hub stations at least involve energy stations, energy storage stations and data stations; at least one processor configured to: divide the received data into static data and dynamic data, the dynamic data representing data that changes over time, and the static data representing data that remains stable over time, further divide the static data into various data including facility-related data of the energy hub stations, ownership-related data and transaction configuration-related data, and further divide the dynamic data into various data including external environment-related data of the energy hub stations, internal operation-related data and transaction dynamic data; for the various data, respectively determine an accuracy-related parameter, an integrity-related parameter, a consistency-related parameter and a timeliness-related parameter, and determine a quality evaluation parameter of various data by synthesizing the accuracy-related parameter, the integrity-related parameter, the consistency-related parameter and the timeliness-related parameter; wherein the timeliness-related parameter is determined based on the ratio of the amount of updated data to the total amount of data; the at least one processor is further configured to, for various data in dynamic data: load current data; determining, with the at least one processor, a sample curve thereof based on the current data; comparing the sample curve of the current data with a corresponding sample feature curve of historical data, and determining an abnormal data point based on a comparison result; and correcting the abnormal data point by using a corresponding segment in the sample feature curve relative to the abnormal data point; wherein, the at least one processor is further configured to correct the abnormal data point according to a formula as follows: wherein L r represents the corrected sample curve, L t represents the sample characteristic curve, L d represents the sample curve before correction, the mth point to the nth point are abnormal data points, and i represents the serial number of the abnormal data points.
Citation Information
Patent Citations
Rapid online assessment method for wide area measurement power big data quality
CN109492683A
Method and device for evaluating operation state of electric meter
CN110488218A