Data quality evaluation method and device of distributed energy storage system and computer equipment
Patent Information
- Application Number
- CN202311130722.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-04
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-09-04
AI Technical Summary
[0004]有鉴于此,本发明提供了分布式储能系统的数据质量评估方法、装置及计算机设备,以解决现有储能技术领域缺少对多维度数据质量评估的问题
Smart Images

Figure CN117033920B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of distributed energy storage technology, and specifically to a data quality assessment method, apparatus, and computer equipment for distributed energy storage systems. Background Technology
[0002] Distributed energy storage features rapid power control and flexible energy throughput, enabling it to participate in grid voltage and frequency regulation and demand response as a flexible resource, promoting the consumption of renewable energy and serving as an important means to address the source-load imbalance problem. On the other hand, the diverse types and dispersed locations of distributed energy storage devices, along with their multi-source heterogeneous data characteristics, pose challenges to the data quality management of distributed energy storage systems.
[0003] Existing data quality assessment methods mainly fall into two categories: one is relatively general and has strong universality, but it makes limited use of domain-specific information; the other is a domain-specific assessment method with unique assessment indicators. Current monitoring methods and quality control in the energy storage technology field focus on the battery level, using a Battery Management System (BMS) to monitor the real-time status of energy storage management units such as individual battery cells and battery clusters. However, existing methods lack multi-dimensional data quality assessment. Summary of the Invention
[0004] In view of this, the present invention provides a data quality assessment method, apparatus and computer equipment for distributed energy storage systems to solve the problem of the lack of multi-dimensional data quality assessment in the existing energy storage technology field.
[0005] In a first aspect, the present invention provides a data quality assessment method for a distributed energy storage system, the method comprising: acquiring data to be assessed from the distributed energy storage system; determining a completeness assessment result based on the data missingness of the data to be assessed; determining a consistency assessment result based on the temporal and numerical relationships of different data sources in the data to be assessed; determining an accuracy assessment result based on the data changes in the data to be assessed; and determining a timeliness assessment result based on the time information of the data to be assessed.
[0006] The data quality assessment method for distributed energy storage systems provided in this invention assesses the data to be evaluated in four dimensions: completeness, consistency, accuracy, and timeliness. This enables multi-dimensional data quality assessment of distributed energy storage systems, which is beneficial for comprehensive prediction and energy coordination management of distributed energy storage systems.
[0007] In one optional implementation, the data to be evaluated includes equipment parameter information and monitoring data of the distributed energy storage system. After obtaining the data to be evaluated from the distributed energy storage system, the method further includes: preprocessing the data to be evaluated; before obtaining the data to be evaluated from the distributed energy storage system, the method further includes: determining evaluation parameters.
[0008] In this embodiment, when acquiring the data to be evaluated, not only monitoring data is acquired, but also the system's equipment parameter information is collected. This solves the problem in related technologies that the supporting equipment of energy storage systems, such as converters, are not included in a unified data quality management framework, which affects the aggregation and fusion of data and is not conducive to the comprehensive prediction and energy coordination management of distributed energy storage systems.
[0009] In one optional implementation, the integrity assessment result is determined based on the data missingness of the data to be assessed, including: calculating the data missing rate based on the actual number of collection points and the theoretical number of collection points obtained from the data to be assessed by the sliding window, and then determining the integrity assessment result.
[0010] In this embodiment, the actual number of data collection points is obtained by using a sliding window method, and the data missing rate is calculated together with the theoretical number of data collection points, thereby realizing the assessment of data integrity.
[0011] In one optional implementation, different data sources are divided according to control instructions, and the consistency evaluation result is determined based on the temporal and numerical relationships of different data sources in the data to be evaluated. The division of different data sources according to control instructions includes: using changes in state variables in the data to be evaluated as signals for control instructions to distinguish different monitoring data tables with the same state variable information in the data to be evaluated as different data sources; determining temporal consistency based on the temporal relationships of the collection points within each data source and the temporal relationships of different data sources; determining numerical consistency based on correlation analysis between different data sources; and determining the consistency evaluation result based on temporal consistency and numerical consistency.
[0012] In this embodiment, consistency assessment was performed from both the time and numerical aspects of the data, which improved the comprehensiveness of the consistency assessment.
[0013] In one optional implementation, determining time consistency based on the time relationships of collection points within each data source and the time relationships of different data sources includes: determining the mean of the collection time interval for each data source based on the time difference between adjacent collection points within each data source; calculating the mean of the collection time interval for all data sources based on the time differences between adjacent collection points within all data sources; calculating the sum of squared errors between groups based on the number of collection points in different data sources, the mean of the collection time interval for each data source, and the mean of the collection time interval for all data sources; calculating the total sum of squared errors based on the time interval for each collection point in each data source and the mean of the collection time interval for all data sources; and determining time consistency based on the ratio of the sum of squared errors between groups to the total sum of squared errors.
[0014] In one optional implementation, the evaluation parameters include a preset interval, a preset distribution model, and a confidence interval. The accuracy evaluation results include point anomaly rate and / or sequence anomaly rate. The point anomaly rate is determined as follows: based on the difference between the monitoring data of adjacent collection points and the relationship between the preset interval, the point anomaly rate is determined as follows: the monitoring data is segmented into multiple segments; the change in the monitoring data of adjacent collection points within each segment and the mean of the change are calculated to obtain the parameters of the sample to be compared; based on the preset distribution model, the relationship between the parameters of the sample to be compared and the confidence interval is used to determine whether each segment is abnormal, thus obtaining the sequence anomaly rate.
[0015] In this embodiment, the accuracy of the data is evaluated from two aspects: point anomaly rate and sequence anomaly rate, which improves the comprehensiveness of the accuracy evaluation.
[0016] In one optional implementation, the evaluation parameters include a preset sampling period, and the timeliness evaluation result is determined based on the time information of the data to be evaluated, including: calculating the time difference between the data recording time and the data generation time of the monitoring data; determining the data delay rate based on the mean of the time difference, the number of monitoring data in each period, and the preset sampling period, and determining the timeliness evaluation result.
[0017] In one optional implementation, the method further includes: obtaining a quality analysis report based on the integrity assessment results, consistency assessment results, accuracy assessment results, and timeliness assessment results; and updating the assessment parameters according to the operating parameters of the distributed energy storage system.
[0018] In this embodiment, by generating a quality analysis report, the evaluation results of the data can be fully displayed; by updating the evaluation parameters, the accuracy of subsequent evaluations is further improved.
[0019] Secondly, the present invention provides a data quality assessment device for a distributed energy storage system. The device includes: a data acquisition module for acquiring data to be assessed from the distributed energy storage system; an integrity assessment module for determining an integrity assessment result based on the data missingness of the data to be assessed; a consistency assessment module for determining a consistency assessment result based on the time and numerical relationships of different data sources in the data to be assessed; an accuracy assessment module for determining an accuracy assessment result based on the data changes in the data to be assessed; and a timeliness assessment module for determining a timeliness assessment result based on the time information of the data to be assessed.
[0020] In one optional implementation, the data to be evaluated includes equipment parameter information and monitoring data of the distributed energy storage system. The device further includes: a preprocessing module for preprocessing the data to be evaluated; and a parameter determination module for determining the evaluation parameters.
[0021] In one optional implementation, the integrity assessment module is specifically used to calculate the data missing rate based on the actual number of collection points and the theoretical number of collection points obtained from the data collection to be assessed by the sliding window, and to determine the integrity assessment result.
[0022] In one optional implementation, the consistency assessment module includes: a data source segmentation module, used to distinguish different monitoring data tables with the same state quantity information in the data to be assessed into different data sources by using changes in state quantities in the data to be assessed as signals for control commands; a time consistency assessment module, used to determine time consistency based on the time relationship between collection points within each data source and the time relationship between different data sources; a numerical consistency assessment module, used to determine numerical consistency based on correlation analysis between different data sources; and an assessment submodule, used to determine the consistency assessment result based on time consistency and numerical consistency.
[0023] In one optional implementation, different data sources are divided according to control instructions, and the time consistency assessment module is specifically used to: determine the mean of the acquisition time interval of each data source based on the time difference between adjacent acquisition points within each data source; calculate the mean of the acquisition time interval of all data sources based on the time difference between adjacent acquisition points within all data sources; calculate the sum of squared errors between groups based on the number of acquisition points of different data sources, the mean of the acquisition time interval of each data source, and the mean of the acquisition time interval of all data sources; calculate the total sum of squared errors based on the time interval of each acquisition point of each data source and the mean of the acquisition time interval of all data sources; and determine time consistency based on the ratio of the sum of squared errors between groups to the total sum of squared errors.
[0024] In one optional implementation, the evaluation parameters include a preset interval, a preset distribution model, and a confidence interval. The accuracy evaluation results include point anomaly rate and / or sequence anomaly rate. The point anomaly rate is determined as follows: based on the difference between the monitoring data of adjacent collection points and the relationship between the preset interval, the point anomaly rate is determined as follows: the monitoring data is segmented into multiple segments; the change in the monitoring data of adjacent collection points within each segment and the mean of the change are calculated to obtain the parameters of the sample to be compared; based on the preset distribution model, the relationship between the parameters of the sample to be compared and the confidence interval is used to determine whether each segment is abnormal, thus obtaining the sequence anomaly rate.
[0025] In one optional implementation, the evaluation parameters include a preset sampling period. The timeliness evaluation module is specifically used to: calculate the time difference between the data recording time and the data generation time of the monitoring data; determine the data delay rate based on the mean of the time difference, the number of monitoring data in each period, and the preset sampling period, and determine the timeliness evaluation result.
[0026] In one optional implementation, the device further includes: a report determination module for obtaining a quality analysis report based on the integrity assessment results, consistency assessment results, accuracy assessment results, and timeliness assessment results; and an update module for updating the assessment parameters according to the operating parameters of the distributed energy storage system.
[0027] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the data quality assessment method for a distributed energy storage system described in the first aspect or any corresponding embodiment thereof.
[0028] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the data quality assessment method for a distributed energy storage system according to the first aspect or any corresponding embodiment described above. Attached Figure Description
[0029] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0030] Figure 1 This is a flowchart illustrating a data quality assessment method for a distributed energy storage system according to an embodiment of the present invention. Figure 2 This is a flowchart illustrating another data quality assessment method for a distributed energy storage system according to an embodiment of the present invention; Figure 3 This is a structural block diagram of a data quality assessment device for a distributed energy storage system according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] According to an embodiment of the present invention, a data quality assessment method for a distributed energy storage system is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0033] This embodiment provides a data quality assessment method for distributed energy storage systems, which can be used in electronic devices such as computers, mobile phones, and tablets. Figure 1 This is a flowchart of a data quality assessment method for a distributed energy storage system according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Obtain the data to be evaluated for the distributed energy storage system.
[0034] Specifically, the operation modes of a distributed energy storage system include the operating modes of the PCS (Power Conversion System, inverter) (off-grid, grid-connected), the charging and discharging modes of the energy storage batteries (constant current, constant voltage, constant power), and the connection methods between the energy storage battery clusters and the PCS. The data to be evaluated specifically includes equipment parameter information and monitoring data of the distributed energy storage system. The equipment in the distributed energy storage system includes various forms of energy storage management units such as battery modules and battery clusters, inverters, distribution transformers, and other auxiliary equipment. Therefore, the equipment parameter information includes equipment type and performance parameters, such as the conventional parameters and input / output parameters of the PCS, and the nominal, structural, and electrical parameters of the energy storage batteries. The monitoring data can be real-time acquired data, specifically including analog quantities such as current, voltage, temperature, state of charge (SOC), and state of health (SOH) of individual battery cells, battery modules, and battery clusters, as well as state quantities such as battery system switching status and communication status, collected by the battery management system; analog quantities such as DC side voltage, current, and power, and AC side three-phase voltage, three-phase current, and input / output power, as well as state quantities such as switching status and charge / discharge status, collected by the energy storage converter; and three-phase connection temperature and three-phase current, collected by the distribution transformer. It should be noted that the acquired data to be evaluated can be stored in chronological order, that is, each timestamp and its corresponding data to be evaluated are stored together, facilitating subsequent evaluation.
[0035] Step S102: Determine the integrity assessment result based on the data missing information of the data to be evaluated.
[0036] Specifically, integrity assessment involves evaluating the completeness of the acquired data by calculating the missing data rate.
[0037] Step S103: Determine the consistency assessment result based on the temporal and numerical relationships of different data sources in the data to be evaluated.
[0038] The control command refers to the instruction information issued from the converter (PCS) to the battery energy management system (BMS), and then from the BMS to the battery clusters, battery modules, or individual battery cells. This command controls the charging and discharging state of the battery by changing the switching state of the battery system. Specifically, based on the issued control commands, different data tables with the same state quantity information in the data to be evaluated are distinguished into multiple data sources. Then, the consistency evaluation of the data to be evaluated is achieved based on the time and numerical relationships between the multiple data sources.
[0039] Step S104: Determine the accuracy assessment result based on the data changes of the data to be evaluated.
[0040] Specifically, the changes in the data include changes between adjacent data to be evaluated, and also changes between multiple segments after the data to be evaluated has been sliced. The accuracy assessment result is determined by the changes.
[0041] Step S105: Determine the timeliness assessment result based on the time information of the data to be evaluated.
[0042] The time information of the data to be evaluated specifically includes the relationship between the data generation and recording times, and the timeliness assessment is based on this relationship.
[0043] The data quality assessment method for distributed energy storage systems provided in this invention assesses the data to be evaluated in four dimensions: completeness, consistency, accuracy, and timeliness. This enables multi-dimensional data quality assessment of distributed energy storage systems, which is beneficial for comprehensive prediction and energy coordination management of distributed energy storage systems.
[0044] In one optional implementation, after acquiring the data to be evaluated of the distributed energy storage system, the method further includes: preprocessing the data to be evaluated; before acquiring the data to be evaluated of the distributed energy storage system, the method further includes: determining evaluation parameters.
[0045] The data preprocessing can be implemented in different ways depending on the evaluation dimensions, including data sampling, cleaning, transformation, and encoding. Specific preprocessing methods will be explained in detail during the subsequent evaluation process. The evaluation parameters are the relevant parameters used in different dimensions of evaluation; these will be discussed in the subsequent evaluation process and will not be elaborated upon here.
[0046] This embodiment provides a data quality assessment method for a distributed energy storage system, the process of which includes the following steps: Step S201: Obtain the data to be evaluated for the distributed energy storage system.
[0047] Step S202: Determine the integrity assessment result based on the data missing information of the data to be evaluated.
[0048] Specifically, step S202 includes: Step S2021: Calculate the data missing rate based on the actual number of data collection points and the theoretical number of data collection points obtained from the data collection to be evaluated through the sliding window, and determine the integrity evaluation result.
[0049] When calculating the missing rate, the sliding window is first fixed. Then randomly select A non-overlapping sliding window; calculate the actual number of data collection points within the deduplicated window. ; Calculate the theoretical number of data collection points within the sliding window based on the equipment's data collection cycle. ; Calculate the data missing rate based on the actual number of data collection points and the theoretical number of data collection points. The formula is as follows:
[0050] Step S203: Determine the consistency assessment result based on the time and numerical relationships of different data sources in the data to be evaluated. Different data sources are divided according to control instructions. Specifically, step S203 includes: Step S2031: Using changes in state variables in the data to be evaluated as signals for control commands, different monitoring data tables with the same state variable information in the data to be evaluated are distinguished into different data sources. Specifically, information collected by physical devices at the same level under the same control command is considered as the same data source. Physical devices at the same level can be multiple battery clusters controlled by the same PCS, multiple battery modules connected in parallel within a battery cluster, or a group of individual batteries connected in series. When the control command changes, the collected monitoring data will change significantly. Therefore, monitoring data tables with the same state information are distinguished into different data sources based on the control command. This control command information can be obtained from the device state variable information in the data to be evaluated. The monitoring data tables store information collected by physical devices. Information collected by a single physical device at the same level under the same control command is stored in one data table. Then, the data tables corresponding to multiple physical devices at the same level under the same control command are divided into one monitoring data table, thus obtaining multiple monitoring data tables.
[0051] Step S2032: Determine time consistency based on the time relationship of the collection points within each data source and the time relationship of different data sources.
[0052] Specifically, step S2032 includes: Step a1: Determine the average collection time interval for each data source based on the time difference between adjacent collection points within each data source. Specifically, determine the time difference between adjacent collection points within each data source based on the timestamps of the acquired data to be evaluated, i.e., subtract the previous time point from the later time point to obtain the collection time interval. Then, calculate all collection time intervals within each data source. The average value of the sampling time intervals is obtained by averaging. .
[0053] Step a2: Calculate the mean of the collection time intervals for all data sources based on the time differences between adjacent collection points within all data sources; specifically, sum the time differences between adjacent collection points within all data sources and divide by the number of time differences to obtain the mean of the collection time intervals for all data sources. .
[0054] Step a3: Calculate the sum of squared errors between groups based on the number of data collection points from different data sources, the mean of the data collection time interval for each data source, and the mean of the data collection time interval for all data sources. Specifically, the sum of squared errors between groups is determined using the following formula:
[0055] In the formula, This represents the number of collection points in the i-th data source, and k represents the total number of data sources.
[0056] Step a4: Calculate the total sum of squared errors based on the time interval between each data source and each collection point, and the mean of the time intervals across all data sources; specifically, the total sum of squared errors is determined using the following formula:
[0057] In the formula, Indicates the first The first data source The time difference between each collection point, specifically the time difference between that collection point and the previous (number) data point within the data source. The time difference between each collection point.
[0058] Step a5: Determine time consistency based on the ratio of the sum of squared errors between groups to the total sum of squared errors. Specifically, time consistency... The following formula is used to determine it:
[0059] Step S2033: Determine numerical consistency based on correlation analysis between different data sources; specifically, before conducting the numerical consistency assessment, preprocess the data by deduplication, filling in missing records, and removing outliers, then perform pairwise correlation analysis on all data sources to obtain a correlation coefficient matrix; based on the number of data sources... The correlation coefficient matrix is used to calculate the consistency rate of numerical values from different data sources. The formula is as follows:
[0060] in, For data source and data source The correlation coefficient between the data. The correlation coefficient can be selected based on the characteristics of the data, such as Pearson correlation coefficient, cosine similarity formula, or Spearman's rank correlation coefficient.
[0061] Step S2034: Determine the consistency assessment result based on temporal consistency and numerical consistency. Specifically, temporal consistency and numerical consistency together constitute the consistency assessment result.
[0062] Step S204: Determine the accuracy assessment result based on the data variation of the data to be evaluated. Before conducting the accuracy assessment, the data undergoes deduplication preprocessing. Assessment parameters include preset intervals, preset distribution models, and confidence intervals. The accuracy assessment result includes point anomaly rate and / or sequence anomaly rate.
[0063] Specifically, the point anomaly rate is determined as follows: Step S2041, the point anomaly rate is determined based on the relationship between the difference in monitoring data of adjacent collection points and a preset interval; specifically, the difference in the data to be evaluated of adjacent collection points can be determined by subtracting the analog quantity of the previous collection point from the analog quantity to be evaluated obtained at each collection point. This difference is compared with a preset interval. When the difference exceeds the preset interval, the analog quantity of the corresponding collection point is considered a point anomaly. Thus, the point anomaly rate is determined using the following formula:
[0064] In the formula, m represents the total number of simulated quantities in the data to be evaluated. Indicates the number of point anomalies.
[0065] The sequence anomaly rate is determined in the following manner: Step S2042 involves segmenting the monitoring data into multiple segments. Specifically, in addition to the point anomaly rate, the sequence anomaly rate can also be calculated during accuracy evaluation. To calculate the sequence anomaly rate, the analog quantities in the monitoring data are first segmented. For example, the charge / discharge state quantities can be used to slice the analog quantities, resulting in multiple charging (or discharging) time series segments. The state quantities are categorized as either 1 or 0 based on the charging and discharging states. Then, a segment of state quantities that is entirely 0 or entirely 1 is selected, and its start and end times are obtained. The analog data segment is then selected as a time series segment based on the start and end times.
[0066] Step S2043: Calculate the change in monitoring data from adjacent collection points within each segment and the mean of these changes to obtain the parameters of the sample to be compared. Specifically, first calculate the change in monitoring data from adjacent collection points within each segment (the difference between the monitoring data from two adjacent collection points), then calculate the mean of all changes within each segment based on this change, and use the mean of the changes as the parameters of the sample to be compared. The preset distribution model can be obtained by observing historical data and combining histograms, kernel density estimation, QQ plots (quantile-quantile plots), or other non-parametric statistical methods to obtain the distribution pattern of the mean. Generally, according to the central limit theorem, the mean distribution approximates a normal distribution. .
[0067] Step S2044: Based on a preset distribution model, determine whether each segment is abnormal according to the relationship between the parameters of the samples to be compared and the confidence interval, and obtain the sequence abnormality rate. Specifically, determine whether the parameters of the samples to be compared for each segment exceed the confidence interval of the preset distribution model. If the value exceeds the limit, the corresponding segment is determined to be an anomalous sequence. The anomalous rate of this sequence is expressed by the following formula:
[0068] In the formula, q represents the total number of segments. Indicates the number of abnormal sequences.
[0069] It should be noted that the above steps do not consider specific distributed energy storage charging and discharging modes, but only provide a general form. In other embodiments, if the distributed energy storage system is in a fixed charging mode, point anomaly rate assessment can be performed on the fixed analog quantity, while a combination of point anomaly rate and sequence anomaly rate assessment can be used for the variable analog quantity. For example, in constant current charging mode, the point anomaly rate of the individual battery current value, the point anomaly rate of the voltage value, and the sequence anomaly rate can be assessed.
[0070] Step S205: Determine the timeliness assessment result based on the time information of the data to be evaluated. Before conducting the timeliness assessment, first determine the sampling period in the assessment parameters to obtain the preset sampling period. Simultaneously, deduplicated data and erroneous data are removed from the data, and missing precision is filled according to the theoretical collection period and precision, thereby achieving data preprocessing.
[0071] Specifically, step S205 includes: Step S2051: Calculate the time difference between the data recording time and the data generation time of the monitored data. Specifically, a corresponding timestamp is generated during the data acquisition process; this timestamp is the data generation time. Simultaneously, the acquired data is transmitted over the network, received and recorded by the database, generating a timestamp, which is called the data recording time. Both times can be recorded in the data table, identified by different value fields. Therefore, the time difference for each data point can be obtained by subtracting the data recording time and data generation time from the data table.
[0072] Step S2055: Determine the data delay rate based on the average time difference, the number of monitoring data points per cycle, and the preset sampling period, and determine the timeliness assessment result. Specifically, the data delay rate is determined using the following formula:
[0073] In the formula, T represents the preset sampling period, and n represents the number of sampling points. This represents the data recording time of the i-th data. This represents the data generation time of the i-th data.
[0074] This embodiment provides a data quality assessment method for a distributed energy storage system, which includes the following steps: Step S301: Obtain the data to be evaluated for the distributed energy storage system; for details, please refer to [link to relevant documentation]. Figure 1 Step S101 of the illustrated embodiment will not be described again here.
[0075] Step S302: Determine the integrity assessment result based on the data missing information in the data to be evaluated; for details, please refer to [link to relevant documentation]. Figure 1 Step S102 of the illustrated embodiment will not be described again here.
[0076] Step S303: Determine the consistency assessment result based on the temporal and numerical relationships between different data sources in the data to be evaluated; for details, please refer to [link to relevant documentation]. Figure 1 Step S103 of the illustrated embodiment will not be described again here.
[0077] Step S304: Determine the accuracy assessment result based on the changes in the data to be evaluated. For details, please refer to [link to relevant documentation]. Figure 1 Step S104 of the illustrated embodiment will not be described again here.
[0078] Step S305: Determine the timeliness assessment result based on the time information of the data to be evaluated. For details, please refer to [link to relevant documentation]. Figure 1 Step S105 of the illustrated embodiment will not be described again here.
[0079] Step S306: Based on the integrity assessment results, consistency assessment results, accuracy assessment results, and timeliness assessment results, a quality analysis report is obtained. Specifically, the assessment results of the distributed energy storage system in four dimensions can be dynamically tracked, and then the assessment results of the distributed energy storage system equipment in each dimension can be obtained according to weekly, monthly, quarterly, and annual cycles. Line charts, box plots, and other forms are used to reflect the data quality fluctuations and trends, and a quality analysis report of the distributed energy storage system is obtained.
[0080] Step S307: Update the evaluation parameters based on the operating parameters of the distributed energy storage system. Specifically, the parameters can be updated based on the operating time of the distributed energy storage system, the cycle number of the energy storage battery, the depth of discharge and aging degree, the inverter operating mode, and the characteristics of the energy storage system operation monitoring data.
[0081] As a specific application embodiment of the present invention, such as Figure 2 As shown, the data quality assessment method for this distributed energy storage system is implemented using the following process: 1. Obtain the equipment parameter information and real-time monitoring data of the distributed energy storage system as the data to be evaluated.
[0082] 2. Define the quality assessment dimensions as including four dimensions: completeness, consistency, accuracy, and timeliness, and conduct quantitative assessment of the data in these four dimensions.
[0083] 3. Dynamically track the data quality assessment results of the distributed energy storage system across four dimensions and output a quality analysis report.
[0084] 4. Update the evaluation parameters.
[0085] This embodiment also provides a data quality assessment device for a distributed energy storage system. This device is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0086] This embodiment provides a data quality assessment device for a distributed energy storage system, such as... Figure 3 As shown, it includes: Data acquisition module 31 is used to acquire the data to be evaluated from the distributed energy storage system; Integrity assessment module 32 is used to determine the integrity assessment result based on the data missing information of the data to be assessed; The consistency assessment module 33 is used to determine the consistency assessment result based on the temporal and numerical relationships of different data sources in the data to be assessed. Accuracy assessment module 34 is used to determine the accuracy assessment result based on the data changes of the data to be assessed; The timeliness assessment module 35 is used to determine the timeliness assessment result based on the time information of the data to be assessed.
[0087] In one optional implementation, the data to be evaluated includes equipment parameter information and monitoring data of the distributed energy storage system. The device further includes: a preprocessing module for preprocessing the data to be evaluated; and a parameter determination module for determining the evaluation parameters.
[0088] In one optional implementation, the integrity assessment module is specifically used to calculate the data missing rate based on the actual number of collection points and the theoretical number of collection points obtained from the data collection to be assessed by the sliding window, and to determine the integrity assessment result.
[0089] In one optional implementation, different data sources are divided according to control commands. The consistency assessment module includes: a data source division module, used to distinguish different monitoring data tables with the same state quantity information in the data to be assessed as different data sources by using the change of state quantity in the data to be assessed as the signal of the control command; a time consistency assessment module, used to determine time consistency based on the time relationship of the collection points within each data source and the time relationship of different data sources; a numerical consistency assessment module, used to determine numerical consistency based on the correlation analysis between different data sources; and an assessment submodule, used to determine the consistency assessment result based on time consistency and numerical consistency.
[0090] In one optional implementation, the time consistency assessment module is specifically used to: determine the mean of the collection time interval for each data source based on the time difference between adjacent collection points within each data source; calculate the mean of the collection time interval for all data sources based on the time difference between adjacent collection points within all data sources; calculate the sum of squared errors between groups based on the number of collection points in different data sources, the mean of the collection time interval for each data source, and the mean of the collection time interval for all data sources; calculate the total sum of squared errors based on the time interval for each collection point in each data source and the mean of the collection time interval for all data sources; and determine time consistency based on the ratio of the sum of squared errors between groups to the total sum of squared errors.
[0091] In one optional implementation, the evaluation parameters include a preset interval, a preset distribution model, and a confidence interval. The accuracy evaluation results include point anomaly rate and / or sequence anomaly rate. The point anomaly rate is determined as follows: based on the difference between the monitoring data of adjacent collection points and the relationship between the preset interval, the point anomaly rate is determined as follows: the monitoring data is segmented into multiple segments; the change in the monitoring data of adjacent collection points within each segment and the mean of the change are calculated to obtain the parameters of the sample to be compared; based on the preset distribution model, the relationship between the parameters of the sample to be compared and the confidence interval is used to determine whether each segment is abnormal, thus obtaining the sequence anomaly rate.
[0092] In one optional implementation, the evaluation parameters include a preset sampling period. The timeliness evaluation module is specifically used to: calculate the time difference between the data recording time and the data generation time of the monitoring data; determine the data delay rate based on the mean of the time difference, the number of monitoring data in each period, and the preset sampling period, and determine the timeliness evaluation result.
[0093] In one optional implementation, the device further includes: a report determination module for obtaining a quality analysis report based on the integrity assessment results, consistency assessment results, accuracy assessment results, and timeliness assessment results; and an update module for updating the assessment parameters according to the operating parameters of the distributed energy storage system.
[0094] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0095] This invention also provides a computer device having the above-described features. Figure 3 The data quality assessment device for the distributed energy storage system shown is shown.
[0096] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 4 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 4 Take a processor 10 as an example.
[0097] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.
[0098] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.
[0099] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device as shown by a landing page for an app. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, which can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0100] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0101] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.
[0102] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.
[0103] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A data quality assessment method for a distributed energy storage system, characterized in that, The method includes: After determining the evaluation parameters, obtain the evaluation data of the distributed energy storage system; The data to be evaluated is preprocessed according to the evaluation dimensions, wherein the integrity index is determined based on the data before preprocessing. Based on the relationship between the actual number of data collection points and the theoretical number of data collection points in the data to be evaluated, the integrity assessment result is determined by calculating the data missing rate index. This process includes: calculating the difference between the actual number of data collection points and the theoretical number of data collection points within multiple sliding windows; calculating the data missing rate index for each sliding window based on the difference; and performing a weighted average of the data missing rates for each sliding window to obtain the integrity assessment result. Based on the temporal and numerical relationships of different data sources in the data to be evaluated, the consistency evaluation result is determined by calculating time consistency indices and numerical consistency indices. This process includes: classifying different data sources based on control command signals; calculating time consistency indices based on the temporal relationships of acquisition points within each data source and the temporal relationships of different data sources; calculating numerical consistency indices based on correlation analysis algorithms between different data sources; and obtaining the consistency evaluation result based on the time consistency indices and numerical consistency indices. Furthermore, classifying different data sources based on control command signals includes using changes in state variables in the data to be evaluated as control commands. The signal divides different monitoring data tables with the same state quantity information into different data sources. The time consistency index is calculated based on the time relationship between collection points within each data source and the time relationship between different data sources. This includes: determining the mean collection time interval of each data source based on the time difference between adjacent collection points within each data source; calculating the mean global collection time interval based on the time difference between adjacent collection points across all data sources; calculating the inter-group error sum of squares based on the number of collection points in different data sources, the mean collection time interval of each data source, and the mean global collection time interval; calculating the total error sum of squares based on the time interval between each collection point in each data source and the mean global collection time interval; and calculating the time consistency index based on the ratio of the inter-group error sum of squares to the total error sum of squares. Based on the changes in data points and data segments in the data to be evaluated, the accuracy assessment result is determined by calculating the point anomaly rate index and the sequence anomaly rate index. Specifically, determining the accuracy assessment result by calculating the point anomaly rate index and the sequence anomaly rate index includes: calculating the point anomaly rate index by comparing the difference between monitoring data from adjacent collection points and the relationship between the difference and a preset interval determined based on the assessment parameters; calculating the sequence anomaly rate index based on the overall change in monitoring data within a segment and the relationship with the assessment parameters; and determining the accuracy assessment result based on the point anomaly rate and the sequence anomaly rate index. Specifically, determining the sequence anomaly rate index based on the overall change in monitoring data within a segment and the relationship with the assessment parameters includes: segmenting the data to be evaluated into multiple segments; calculating the mean of the changes in monitoring data from adjacent collection points within each segment to obtain the parameters of the sample to be compared; and determining whether each segment is anomaly based on a preset distribution model, according to the relationship between the parameters of the sample to be compared and the confidence interval, and calculating the sequence anomaly rate index. The preset distribution model is a normal distribution model. The timeliness assessment result is determined by calculating the data latency rate index based on the time difference between the data recording time and the data generation time in the data to be evaluated. This process includes: calculating the time difference between the data recording time and the data generation time of the monitored data; calculating the data latency rate index based on the average time difference, the number of monitored data points in each sampling period, and the preset sampling period; and determining the timeliness assessment result based on the data latency rate index.
2. The method according to claim 1, characterized in that, The evaluation parameters include a preset interval, a preset distribution model, a confidence interval, and a preset sampling period; Preprocessing the data to be evaluated according to the evaluation dimensions includes: performing different preprocessing on the data to be evaluated according to the evaluation dimensions, wherein the preprocessing includes at least one of data sampling, cleaning, transformation and encoding.
3. The method according to claim 1, characterized in that, The method further includes: A quality analysis report is generated based on the results of the integrity assessment, consistency assessment, accuracy assessment, and timeliness assessment. The evaluation parameters are updated based on the operating parameters of the distributed energy storage system.
4. A data quality assessment device for a distributed energy storage system, characterized in that, The device includes: The data acquisition module is used to acquire the data to be evaluated from the distributed energy storage system after the evaluation parameters are determined. The preprocessing module is used to preprocess the data to be evaluated according to the evaluation dimensions, wherein the integrity index is determined based on the data before preprocessing; The integrity assessment module is used to determine the integrity assessment result by calculating the data missing rate index based on the relationship between the actual number of data collection points and the theoretical number of data collection points in the data to be assessed. Specifically, determining the integrity assessment result by calculating the data missing rate index based on the relationship between the actual number of data collection points and the theoretical number of data collection points in the data to be assessed includes: calculating the difference between the actual number of data collection points and the theoretical number of data collection points within multiple sliding windows; calculating the data missing rate index for each sliding window based on the difference; and performing a weighted average of the data missing rates for each sliding window to obtain the integrity assessment result. The consistency assessment module is used to determine the consistency assessment result by calculating time consistency indices and numerical consistency indices based on the temporal and numerical relationships of different data sources in the data to be assessed. Specifically, determining the consistency assessment result by calculating time consistency indices and numerical consistency indices based on the temporal and numerical relationships of different data sources in the data to be assessed includes: classifying different data sources according to control command signals; calculating time consistency indices based on the temporal relationships of acquisition points within each data source and the temporal relationships of different data sources; calculating numerical consistency indices based on correlation analysis algorithms between different data sources; and obtaining the consistency assessment result based on the time consistency indices and numerical consistency indices. The classification of different data sources according to control command signals includes: using changes in state variables in the data to be assessed as... The control command signals divide different monitoring data tables with the same state quantity information in the data to be evaluated into different data sources. The time consistency index is calculated based on the time relationship between the collection points within each data source and the time relationship between different data sources. This includes: determining the mean collection time interval of each data source based on the time difference between adjacent collection points within each data source; calculating the mean global collection time interval based on the time difference between adjacent collection points across all data sources; calculating the sum of squared errors between groups based on the number of collection points in different data sources, the mean collection time interval of each data source, and the mean global collection time interval; calculating the total sum of squared errors based on the time interval between each collection point in each data source and the mean global collection time interval; and calculating the time consistency index based on the ratio of the sum of squared errors between groups to the total sum of squared errors. An accuracy assessment module is used to determine the accuracy assessment result by calculating point anomaly rate and sequence anomaly rate indicators based on the changes in data points and data segments in the data to be assessed. Specifically, determining the accuracy assessment result by calculating point anomaly rate and sequence anomaly rate indicators based on the changes in data points and data segments in the data to be assessed includes: calculating the point anomaly rate indicator by comparing the difference between monitoring data from adjacent collection points and the relationship between the difference and a preset interval determined based on assessment parameters; calculating the sequence anomaly rate indicator based on the overall change in monitoring data within a segment and the relationship with assessment parameters; and determining the accuracy assessment result based on the point anomaly rate and sequence anomaly rate indicators. Specifically, determining the sequence anomaly rate indicator based on the overall change in monitoring data within a segment and the relationship with assessment parameters includes: segmenting the data to be assessed into multiple segments; calculating the mean of the changes in monitoring data from adjacent collection points within each segment to obtain the parameters of the sample to be compared; and determining whether each segment is anomaly based on a preset distribution model, according to the relationship between the parameters of the sample to be compared and the confidence interval, and calculating the sequence anomaly rate indicator. The preset distribution model is a normal distribution model. The timeliness assessment module is used to determine the timeliness assessment result by calculating the data latency rate index based on the time difference between the data recording time and the data generation time in the data to be assessed. This process includes: calculating the time difference between the data recording time and the data generation time of the monitored data; calculating the data latency rate index based on the average time difference, the number of monitored data points in each sampling period, and the preset sampling period; and determining the timeliness assessment result based on the data latency rate index.
5. The apparatus according to claim 4, characterized in that, The evaluation parameters include a preset interval, a preset distribution model, a confidence interval, and a preset sampling period; Preprocessing the data to be evaluated according to the evaluation dimensions includes: performing different preprocessing on the data to be evaluated according to the evaluation dimensions, wherein the preprocessing includes at least one of data sampling, cleaning, transformation and encoding.
6. The apparatus according to claim 4, characterized in that, The integrity assessment module is specifically used to calculate the difference between the actual number of data collection points and the theoretical number of data collection points within multiple sliding windows; calculate the data missing rate index for each sliding window based on the difference; and perform a weighted average of the data missing rates for each sliding window to obtain the integrity assessment result.
7. The apparatus according to claim 4, characterized in that, The consistency assessment module includes: a partitioning module, used to partition different data sources according to control command signals; a time consistency assessment module, used to calculate time consistency indices based on the time relationships of collection points within each data source and the time relationships of different data sources; a numerical consistency assessment module, used to calculate numerical consistency indices based on correlation analysis algorithms between different data sources; and a first assessment submodule, used to determine the consistency assessment results based on the time consistency indices and numerical consistency indices.
8. The apparatus according to claim 4, characterized in that, The accuracy assessment module includes: a point anomaly rate assessment module, which calculates the point anomaly rate index by comparing the difference between monitoring data from adjacent collection points and the relationship between the preset intervals determined based on assessment parameters; a sequence anomaly rate assessment module, which calculates the sequence anomaly rate index based on the overall change of monitoring data within a segment and the relationship with assessment parameters; and a second assessment submodule, which determines the accuracy assessment result based on the point anomaly rate and the sequence anomaly rate index.
9. The apparatus according to claim 4, characterized in that, The timeliness assessment module is specifically used to: calculate the time difference between the data recording time and the data generation time of the monitoring data; calculate the data delay rate index based on the average time difference, the number of monitoring data in each sampling period, and the preset sampling period; and determine the timeliness assessment result based on the data delay rate index.
10. The apparatus according to claim 4, characterized in that, The device also includes: a report determination module for generating a quality analysis report based on the integrity assessment results, consistency assessment results, accuracy assessment results, and timeliness assessment results; and an update module for updating the assessment parameters according to the operating parameters of the distributed energy storage system.
11. The apparatus according to claim 7, characterized in that, The partitioning module is specifically used to: use changes in state variables in the data to be evaluated as signals for control commands to partition different monitoring data tables with the same state variable information into different data sources.
12. The apparatus according to claim 7, characterized in that, The time consistency assessment module is specifically used for: determining the mean of the collection time interval for each data source based on the time difference between adjacent collection points within each data source; calculating the mean of the global collection time interval based on the time difference between adjacent collection points across all data sources; calculating the sum of squared errors between groups based on the number of collection points in different data sources, the mean of the collection time interval for each data source, and the mean of the global collection time interval; calculating the total sum of squared errors based on the time interval between each collection point in each data source and the mean of the global collection time interval; and calculating the time consistency index based on the ratio of the sum of squared errors between groups to the total sum of squared errors.
13. The apparatus according to claim 8, characterized in that, The sequence anomaly rate assessment module is specifically used for: segmenting the data to be assessed into multiple segments; calculating the mean of the changes in monitoring data of adjacent collection points within each segment to obtain the parameters of the sample to be compared; determining whether each segment is abnormal based on a preset distribution model and the relationship between the parameters of the sample to be compared and the confidence interval, and calculating the sequence anomaly rate index. The preset distribution model is a normal distribution model.
14. A computer device, characterized in that, include: A memory and a processor are interconnected, the memory stores computer instructions, and the processor executes the computer instructions to perform the data quality assessment method for the distributed energy storage system as described in any one of claims 1 to 3.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the data quality assessment method for the distributed energy storage system according to any one of claims 1 to 3.
Citation Information
Patent Citations
System anomaly detection method and device
CN106685750A
Data quality evaluation method and device, terminal equipment and storage medium
CN112506904A