A redundant data cleaning method, device, equipment and medium

CN115481115BActive Publication Date: 2026-09-04STATE GRID BEIJING ELECTRIC POWER CO +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211157916.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-22
Publication Date
2026-09-04
Estimated Expiration
2042-09-22

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提供一种冗余数据清洗方法、装置、设备及介质,以解决现有供电所数据清洗技术缺乏对于数据冗余情况的考虑,导致供电所监测数据的准确性与可靠性低的技术问题

Benefits of technology

[0025] 1. This invention first performs initial filtering on redundant data to filter out some noisy data; based on this, it calculates the fusion weight and information entropy value, and finally obtains the fused data, which further improves the accuracy of the fused data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115481115B_ABST
    Figure CN115481115B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of data cleaning, and particularly relates to a redundant data cleaning method, device, equipment and medium. The method comprises the following steps: obtaining power supply redundant data; performing filtering processing on the power supply redundant data to obtain a filtered data sequence; calculating the information entropy value of each group of filtered data in the filtered data sequence, and calculating the fusion weight corresponding to each group of filtered data according to the information entropy value of each group of filtered data; and superimposing the filtered data according to the fusion weight to obtain fusion data and outputting the fusion data. The present application firstly performs primary filtering processing on the redundant data to filter certain noise data; on this basis, the fusion weight and the information entropy value are calculated, and finally the fusion data is obtained, thereby further improving the accuracy of the fusion data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data cleaning technology, specifically relating to a method, apparatus, equipment, and medium for cleaning redundant data. Background Technology

[0002] Because power supply station systems involve massive amounts of multi-dimensional data from power supply stations, power source side, power consumption side equipment, and climate conditions, data redundancy is prone to occur during data generation, measurement, transmission, and reception. This leads to data distortion when uploaded to the system platform, hindering unified management and business operations of the power supply station system. Therefore, researching redundant data cleaning techniques for power supply stations to obtain more accurate and standardized monitoring data is of great significance for ensuring the safe and efficient operation of power supply stations.

[0003] Domestic and international experts and scholars have conducted some research on power system data processing. Existing technologies include a comprehensive energy data compression and acquisition method based on compressed sensing with a composite data structure, and an anomaly identification method based on improved K-Means clustering, considering data anomalies and missing data. However, this method neglects the consideration of redundant data. Existing technologies also include cloud-based power big data cleaning models, which study data storage, identification, and cleaning of power big data, but lack consideration for data distortion. Addressing the problems of high investment and low efficiency in traditional user anomaly consumption pattern detection models, a full-cycle user anomaly consumption detection model integrating data cleaning, feature selection, and model training is proposed.

[0004] The above research has promoted the application of data cleaning technology in power systems. However, research on data cleaning technology in power supply substation scenarios is still limited and insufficient, lacking consideration for data redundancy, which brings problems and challenges to the safe and efficient operation of power supply substations. Summary of the Invention

[0005] The purpose of this invention is to provide a redundant data cleaning method, apparatus, equipment, and medium to solve the technical problem that existing power supply station data cleaning technologies lack consideration for data redundancy, resulting in low accuracy and reliability of power supply station monitoring data.

[0006] To achieve the above objectives, the present invention employs the following technical solution:

[0007] Firstly, a method for cleaning redundant data includes the following steps:

[0008] Obtain redundant data from the power supply station;

[0009] The redundant data from the power supply station is filtered to obtain a filtered data sequence.

[0010] Calculate the information entropy value of each group of filtered data in the filtered data sequence, and calculate the fusion weight corresponding to each group of filtered data based on the information entropy value of each group of filtered data;

[0011] The filtered data is superimposed according to the fusion weights to obtain fused data, which is then output.

[0012] A further improvement of the present invention is that the redundant data of the power supply station includes the grid voltage amplitude, the active power of the grid node, the reactive power of the grid node, the active load of the grid line, and the reactive load of the grid line.

[0013] A further improvement of the present invention is that the filtering process is a Kalman filter.

[0014] A further improvement of the present invention is that the Kalman filtering process includes prediction processing and correction processing.

[0015] A further improvement of the present invention is that: when calculating the fusion weight corresponding to each group of filtered data based on the information entropy value of each group of filtered data, the output probability is calculated based on the filtered data sequence, the information entropy value is calculated based on the output probability, and finally the fusion weight is calculated based on the information entropy value.

[0016] A further improvement of the present invention is that: when calculating the information entropy value of each group of filtered data in the filtered data sequence, the information entropy theory is used.

[0017] Secondly, a redundant data cleaning device includes:

[0018] Redundant data acquisition module: used to acquire redundant data from the power supply station;

[0019] Filtering module: Used to filter redundant data from the power supply station to obtain a filtered data sequence;

[0020] Fusion weight calculation module: used to calculate the information entropy value of each group of filtered data in the filtered data sequence, and to calculate the fusion weight corresponding to each group of filtered data based on the information entropy value of each group of filtered data;

[0021] Fusion Data Output Module: Used to superimpose filtered data according to fusion weights to obtain fused data and output it.

[0022] Thirdly, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the aforementioned redundant data cleaning method.

[0023] Fourthly, a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned redundant data cleaning method.

[0024] Compared with the prior art, the present invention has at least the following beneficial effects:

[0025] 1. This invention first performs initial filtering on redundant data to filter out some noisy data; based on this, it calculates the fusion weight and information entropy value, and finally obtains the fused data, which further improves the accuracy of the fused data.

[0026] 2. The information entropy theory used in this invention to calculate information entropy not only ensures the effective extraction of information before fusion, but also improves the accuracy of the data after fusion. Attached Figure Description

[0027] The accompanying drawings, which form part of this specification, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0028] Figure 1 This is a flowchart of a redundant data cleaning method according to the present invention;

[0029] Figure 2 This is a relative error diagram of redundant data fusion in Embodiment 1 of the redundant data cleaning method of the present invention;

[0030] Figure 3 This is a system block diagram of a redundant data cleaning device according to the present invention. Detailed Implementation

[0031] The present invention will now be described in detail with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0032] The following detailed description is exemplary and intended to provide further detailed explanation of the invention. Unless otherwise specified, all technical terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this invention is for describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention.

[0033] Example 1

[0034] A method for cleaning redundant data, such as Figure 1 As shown, it includes the following steps:

[0035] S1. Obtain redundant data from the power supply station;

[0036] Redundant data from the power supply station includes grid voltage amplitude U. i Active power P at grid nodes i Reactive power Q at power grid nodes iActive load P of power grid lines ij and reactive load Q of power grid lines ij wait

[0037] S2. Obtain the filtered data sequence by processing the redundant data of the power supply station through Kalman filtering;

[0038] In S2, Kalman filtering can be divided into two parts: prediction and correction, specifically including the following steps:

[0039] S21. For the redundant data of k groups of power supply stations [X1,X2,...,X...] k ], to perform predictive processing;

[0040]

[0041] P i - =AP i-1 A T +Q;

[0042] In the formula: Let i be the prior state estimate at time i; Let A be the posterior state estimate at time i-1; let B be the state transition coefficient from the previous state to the current state; let B be the state transition coefficient from the control input to the current state; u i To control input variables; P i - P represents the prior estimate error covariance. i-1 denoted as posterior estimation error covariance; Q is the process noise covariance.

[0043] S22. Correct the prediction results to obtain the filtered data sequence;

[0044]

[0045]

[0046]

[0047] In the formula: K i Kalman gain; H is the measurement coefficient; R is the measurement noise covariance; z i Let i be the measurement value at time i; I is the unit coefficient.

[0048] Take k sets of redundant data sequences [X1, X2, ..., X...] k The filtered data sequence can be obtained by performing Kalman filtering.

[0049] S3. Calculate the information entropy value of each group of filtered data in the filtered data sequence, and calculate the fusion weight corresponding to each group of filtered data based on the information entropy value of each group of filtered data.

[0050]

[0051]

[0052]

[0053] In the formula: E is the information entropy value; λ represents the value in the i-th row and j-th column of the filtered data sequence; W is the fusion weight; and n is the length of each set of filtered data.

[0054] S4. The filtered data is superimposed according to the fusion weights to obtain the fused data and then output it.

[0055]

[0056] In the formula: This is the data sequence obtained by fusing filtered data.

[0057] After the filtering and fusion of the above four steps, the redundant data of the power supply station is effectively cleaned.

[0058] The following uses actual data from a power supply station as an example to verify the proposed data fusion-based method for cleaning redundant data in power supply stations. In Table 1, Redundant Data 1 and Redundant Data 2 are the current data of a distributed power source in the power supply station system.

[0059] Table 1 Redundancy Data of Power Supply Stations

[0060]

[0061]

[0062]

[0063] Table 1 shows that redundant data 1 and redundant data 2 involve duplicate sampling of current data from a certain distributed power source, and there is a certain error. The results after data cleaning based on the proposed method are as follows: Figure 2 As shown. By Figure 2 It can be seen that the proposed method effectively fuses redundant data, and the fused data has a smaller relative error and is closer to the true value than the redundant data before fusion.

[0064] Example 2

[0065] A redundant data cleaning device, such as Figure 2 As shown, it includes:

[0066] Redundant data acquisition module: used to acquire redundant data from the power supply station;

[0067] Filtering module: Used to process redundant data from the power supply station using Kalman filtering to obtain a filtered data sequence;

[0068] Fusion weight calculation module: used to calculate the information entropy value of each group of filtered data in the filtered data sequence, and to calculate the fusion weight corresponding to each group of filtered data based on the information entropy value of each group of filtered data;

[0069] Fusion Data Output Module: Used to superimpose filtered data according to fusion weights to obtain fused data and output it.

[0070] Example 3

[0071] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned redundant data cleaning method.

[0072] Example 4

[0073] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method for cleaning redundant data.

[0074] As is known from common technical knowledge, this invention can be implemented through other embodiments that do not depart from its spirit or essential characteristics. Therefore, the disclosed embodiments described above are merely illustrative in all respects and are not the only ones. All modifications within the scope of this invention or its equivalents are included in this invention.

[0075] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0076] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0077] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0078] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for cleaning redundant data, characterized in that, Includes the following steps: Obtain redundant data from the power supply station; The redundant data from the power supply station is filtered to obtain a filtered data sequence. Calculate the information entropy value of each group of filtered data in the filtered data sequence, and calculate the fusion weight corresponding to each group of filtered data based on the information entropy value of each group of filtered data: in, For the first The first group of filtered data sequences filter values The corresponding output probability; Calculate the output probability based on the filtered data sequence, and then calculate the information entropy value based on the output probability: in, For the first The information entropy value of the filtered data sequence; n The length of each set of filtered data; Finally, the fusion weights are calculated based on the information entropy values: in, For the first The fusion weights corresponding to the group filtered data sequences; This represents the total number of filtered data sequences that participate in the fusion process during redundant data cleaning. The filtered data is superimposed according to the fusion weights to obtain fused data, which is then output. in, The data sequence obtained by fusing filtered data. For the first i Group filtered data sequence.

2. The redundant data cleaning method according to claim 1, characterized in that, The redundant data of the power supply station includes the grid voltage amplitude, grid node active power, grid node reactive power, grid line active load, and grid line reactive load.

3. The redundant data cleaning method according to claim 1, characterized in that, The filtering process includes prediction processing and correction processing.

4. The redundant data cleaning method according to claim 3, characterized in that, The prediction process includes: in, for i Time-prior state estimator; for i Posterior state estimator at time -1; A is State transition coefficients from the previous state to the current state; B To control the state transition coefficients from the input to the current state; u i To control input variables; The prior estimate is the error covariance; P i-1 The covariance of the posterior estimation error; Q Let be the process noise covariance.

5. The redundant data cleaning method according to claim 4, characterized in that, The correction process includes: ; ; ; in, K i Kalman gain; H For measurement coefficients; R To measure noise covariance; z i for i Time measurement; I The coefficient is a unit coefficient.

6. A redundant data cleaning device, characterized in that, include: Redundant data acquisition module: used to acquire redundant data from the power supply station; Filtering module: Used to filter redundant data from the power supply station to obtain a filtered data sequence; The fusion weight calculation module is used to calculate the information entropy value of each group of filtered data in the filtered data sequence, and to calculate the fusion weight corresponding to each group of filtered data based on the information entropy value of each group of filtered data. in, For the first The first group of filtered data sequences filter values The corresponding output probability; Calculate the output probability based on the filtered data sequence, and then calculate the information entropy value based on the output probability: in, For the first The information entropy value of the filtered data sequence; n The length of each set of filtered data; Finally, the fusion weights are calculated based on the information entropy values: in, For the first The fusion weights corresponding to the group filtered data sequences; This represents the total number of filtered data sequences that participate in the fusion process during redundant data cleaning. The fused data output module is used to superimpose the filtered data according to the fusion weights to obtain fused data and then output it. in, The data sequence obtained by fusing filtered data. For the first i Group filtered data sequence.

7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements a redundant data cleaning method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a redundant data cleaning method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-pilot AUV (Autonomous Underwater Vehicle) collaborative navigation method based on information entropy

    CN113654565A

  • Power distribution network measurement data cleaning method based on space-time combination

    CN114860706A