Big data preprocessing method for main transformer based on big data analysis platform

By implementing the main transformer big data preprocessing method on the big data analysis platform, the problem of data abnormality of the main transformer monitoring is solved, data quality and consistency are improved, and efficient and accurate data support is provided for subsequent data mining and analysis.

CN114840505BActive Publication Date: 2025-05-06GUANGXI POWER GRID CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210263204.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-17
Publication Date
2025-05-06
Estimated Expiration
2042-03-17

AI Technical Summary

Technical Problem

The monitoring data of the main transformer is abnormal due to system errors, network delays and other factors, which increases the difficulty of data analysis, and preprocessing is required to be applied to later data mining and analysis.

Method used

The main transformer big data preprocessing method based on the big data analysis platform is adopted, including data acquisition and storage, duplicate data detection and processing, abnormal data detection and processing, local outlier point detection and processing, data integrity detection and processing. Through these steps, the operation data of the main transformer is processed to ensure the quality and consistency of the data.

Benefits of technology

The quality of the main transformer operating data is improved through big data preprocessing, and standardized, standardized, continuous and accurate large batches of data are obtained, which improves efficiency and accuracy for subsequent main transformer big data mining and analysis, and takes into account the continuity and priori of time when the data is completed missing to ensure that the completed values ​​are more realistic and accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114840505B_ABST
    Figure CN114840505B_ABST
Patent Text Reader

Abstract

The present invention processes the current, voltage, active power and reactive power of the main transformer extracted from the dispatching automation system, including repeated data detection and processing, abnormal data detection and processing, local outlier detection and processing, and data integrity detection and processing, and can handle the noise, abnormality, missing, and repetition problems in the current, voltage, active power, and reactive power of the main transformer operation data. The quality of the data is improved through big data preprocessing, and a large number of standardized, standard, continuous, and accurate data is obtained, which improves the efficiency and accuracy of subsequent main transformer big data mining and analysis. At the same time, when completing the missing data, the continuity and a priori of time are considered, that is, the correlation between the monitoring data of the main transformer before and after the moment, so the weighted summation method is used for calculation, and the trend of the monitoring curve is considered. The completed value is more real and accurate, closer to the actual monitored data, and the accuracy of subsequent data mining is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data preprocessing, and in particular relates to a main transformer big data preprocessing method based on a big data analysis platform. Background Art

[0002] The main transformer, referred to as the main transformer (GSU), is a total step-down transformer mainly used for power transmission and transformation in a unit or substation, and is also the core part of the substation. The transformer is the core equipment of the traction power supply system of the electric locomotive, and is also the key equipment to ensure the safe and stable operation of the traction power supply system. The capacity of the main transformer is generally large, and high reliability is required. Although the failure rate of the main transformer is not high, once a failure occurs, it will cause significant losses. At the least, it may cause equipment failure, and at the worst, it may cause a fire, endangering normal transportation safety. Therefore, it is of great significance to analyze the cause of the transformer failure and take corresponding preventive measures. With the development of society and the advancement of technology, condition-based maintenance, as a maintenance method that can reduce maintenance costs, shorten maintenance outage time, and improve equipment utilization, has become the development direction of maintenance of power equipment such as transformers. And correctly grasping the operating status of the transformer is the key to the success of condition-based maintenance. At present, the correct way to grasp the operating status of the main transformer is to use monitoring equipment to monitor the main transformer and collect and analyze the corresponding monitoring data. However, due to system errors of the monitoring equipment, network delays or other factors, the collected monitoring data are abnormal, which increases the difficulty of subsequent main transformer data analysis. Therefore, it is necessary to preprocess the collected monitoring data to make it suitable for subsequent data mining and analysis. Summary of the invention

[0003] In order to solve the above problems, the present invention provides a main transformer big data preprocessing method based on a big data analysis platform. The specific technical solution is as follows:

[0004] The main transformer big data preprocessing method based on the big data analysis platform includes the following steps:

[0005] Step S1, data collection and storage: extracting the operating data of the main transformer from the dispatching automation system, including the current, voltage, active power and reactive power at each moment of operation, and storing the extracted operating data in the big data analysis platform;

[0006] Step S2, duplicate data detection and processing: the big data analysis platform detects duplicate data in the extracted operation data of the main transformer, retains one of the duplicate data, removes redundant duplicate data, and inputs the processed data into step S3;

[0007] Step S3: abnormal data detection and processing: the big data analysis platform detects whether the extracted operation data of the main transformer is abnormal. If there is abnormal data, the abnormal data is removed and the processed data is input into step S4;

[0008] Step S4: local outlier detection and processing: the big data analysis platform detects whether there are local outliers in the extracted operation data of the main transformer. If there are local outliers, the local outliers are removed;

[0009] Step S5, data integrity detection and processing: The big data analysis platform detects whether the extracted operating data is complete. If there are missing values ​​in the extracted operating data, the missing values ​​are filled in and the filled operating data is output as the final processed data.

[0010] Preferably, the duplicate data detection and processing in step S2 specifically includes the following steps:

[0011] Step S21: Divide the extracted operation data of the main transformer into n data blocks according to type; each data block includes m objects; the recording method of the objects is represents the jth data in the i-th data block in the k-th type of operating data; wherein k=1,2,3,4, respectively representing current, voltage, active power, and reactive power; i=1,2,...n; j=1,2,...m; Represented as a moment object, represented as a moment and value pair;

[0012] Step S22: using an XOR operation to detect whether there is duplicate data between any two objects in each data block in the k-th type of running data; if the operation result is 0, it indicates that duplicate data exists, and one of them needs to be removed; if the operation result is 1, it indicates that there is no duplicate data in the data block;

[0013] Step S23: After removing duplicate data from each data block, an XOR operation is performed on any two data blocks to detect whether there is duplicate data, that is, each object in one data block is XORed with each object in the other data block. If duplicate data exists, only one data is retained.

[0014] Preferably, when performing the XOR operation, the time is XORed to eliminate data of repeated time.

[0015] Preferably, the abnormal data detection and processing in step S3 specifically includes the following steps:

[0016] Step S31: setting the maximum and minimum values ​​of the main transformer current, voltage, active power and reactive power on the big data analysis platform;

[0017] Step S32: Detect whether each type of extracted operating data is between the set corresponding maximum and minimum values. If the corresponding value is not between the minimum and maximum values, it is determined to be abnormal data and the value is eliminated.

[0018] Preferably, the local outlier detection and processing in step S4 specifically includes the following steps:

[0019] Step S41: Divide the extracted operation data of each type of the main transformer into n data blocks; initialize the distance between each object value of the data block and its (m+k) nearest neighbor to the maximum value;

[0020] Step S42: Calculate the distance between each object value of the running data and each object value of the first data block, and update the (m+k) nearest neighbors of each object value in the first data block, and calculate the outlier degree of each object value in real time. When the number of neighbors is less than m+k, the outlier degree is set to infinity, and the outliers less than the initial threshold c are excluded from the data block; the outlier degree of each object value is the sum of the distances between the object value and its m+1 to m+k nearest neighbors;

[0021] Step S43: after processing the first data block, sort the values ​​of the objects that are not excluded in the first data block from large to small according to the outlier degree, take the first n object values ​​and add them to the TOP n outliers, and update the threshold c;

[0022] Step S44: Calculate the distance between each object value in the running data and each object value in the second data block, and update the (m+k) nearest neighbors of each object value in the second data block, and calculate the outlier degree of each object value in real time. When the number of nearest neighbors is less than m+k, the outlier degree is set to infinity, and the outlier degree is less than the threshold value c is excluded from the data block;

[0023] Step S45: after processing the second data block, if the outlier degree of the object value that is not excluded in the second data block is greater than the outlier degree in the TOP n outliers, then update the TOP n outliers and update the threshold c;

[0024] Step S46: for the i-th data block, i=3, 4, 5...n, repeat steps S44-S45 until all data blocks are processed and the TOP n outliers are output;

[0025] In step S43 and step S45, when updating the threshold c, the outlier degree of the nth outlier point in the TOP n outliers is used as the value of the threshold c.

[0026] Preferably, the threshold c in step S42 is set to 0.

[0027] Preferably, the data integrity detection and processing in step S5 specifically includes the following steps:

[0028] Step S51: for each type of operation data, after being processed in steps S2 to S4, data missing includes missing of both time and value, and missing of value only; detecting the integrity of the processed operation data of the main transformer, judging whether the data is complete, i.e., including the corresponding time and value, and if the data is missing, judging the corresponding data missing type;

[0029] Step S52: If both the time and the value are missing, first fill in the corresponding missing time, convert the missing type to only missing value, and then use the method in step S53 to complete the data;

[0030] Step S53: If only a value is missing, extract the N values ​​before and after the corresponding moment of the missing value, calculate the average value Eq of the N values ​​before the corresponding moment and the average value Eh of the N values ​​after the corresponding moment, and use the N values ​​before as the first data and the N values ​​after as the second data, assign a weight λ to the first value adjacent to the missing value in the first data and the second data, and assign a weight a to the second value adjacent to the missing value in the first data and the second data. The weights of N-2 data in the first data and the second data are

[0031] Step S54: multiply the N values ​​in the first data by their corresponding weights and sum them to obtain a first calculated value, multiply the N values ​​in the second data by their corresponding weights and sum them to obtain a second calculated value, average the first calculated value and the second calculated value to obtain an intermediate calculated value, and determine whether the intermediate calculated value is between the average values ​​Eq and Eh. If the intermediate calculated value is between the average values ​​Eq and Eh, use the intermediate calculated value as a supplementary value for the missing value of the data. If the intermediate calculated value is not between the average values ​​Eq and Eh, adjust the values ​​of the weight values ​​λ and a, and repeat the above calculation so that the calculated intermediate calculated value is between the average values ​​Eq and Eh.

[0032] Preferably, if the average values ​​Eq and Eh are equal, the missing value at the corresponding moment is the average value Eq or Eh.

[0033] Preferably, if the missing values ​​are from the initial two moments, the calculation is performed using the next N adjacent values, and it is observed whether the following values ​​increase or decrease with time. If there is an increasing trend, the calculated middle value is set to be smaller than the average value Eh of the next N values ​​at the corresponding moment; if there is a decreasing trend, the calculated middle value is set to be larger than the average value Eh of the next N values ​​at the corresponding moment.

[0034] Preferably, if the missing values ​​are from the last two moments, the calculation is performed using the preceding N values ​​adjacent to them, and it is observed whether the preceding values ​​increase or decrease with time. If there is an increasing trend, the calculated middle value is set to be greater than the average value Eq of the N values ​​following the corresponding moment; if there is a decreasing trend, the calculated middle value is set to be less than the average value Eq of the N values ​​following the corresponding moment.

[0035] The beneficial effects of the present invention are as follows: the present invention processes the current, voltage, active power and reactive power of the main transformer extracted from the dispatching automation system, including duplicate data detection and processing, abnormal data detection and processing, local outlier detection and processing, and data integrity detection and processing, and can handle the noise, abnormality, missing, and repetition problems in the current, voltage, active power and reactive power of the main transformer operation data. The quality of the data is improved through big data preprocessing, and a large number of standardized, standard, continuous and accurate data is obtained, which improves the efficiency and accuracy of subsequent main transformer big data mining and analysis. At the same time, when completing the missing data, the continuity and a priori of time are considered, that is, the correlation between the monitoring data of the main transformer before and after the moment, so the weighted summation method is adopted for calculation, and the trend of the monitoring curve is considered, so that the completed values ​​are more real and accurate, closer to the actual monitored data, and the accuracy of subsequent data mining is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the specific embodiments or the description of the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn according to the actual scale.

[0037] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION

[0038] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0039] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0040] It should also be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an" and "the" are intended to include plural forms unless the context clearly indicates otherwise.

[0041] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0042] like Figure 1 As shown, a specific embodiment of the present invention provides a main transformer big data preprocessing method based on a big data analysis platform, comprising the following steps:

[0043] Step S1, data collection and storage: extract the operating data of the main transformer from the dispatching automation system, including the current, voltage, active power and reactive power at each moment during operation, and store the extracted operating data in the big data analysis platform.

[0044] Step S2, duplicate data detection and processing: The big data analysis platform detects duplicate data in the extracted operating data of the main transformer, retains one of the duplicate data, removes redundant duplicate data, and inputs the processed data into step S3.

[0045] The specific steps include:

[0046] Step S21: Divide the extracted operation data of the main transformer into n data blocks according to type; each data block includes m objects; the recording method of the objects is represents the jth data in the i-th data block in the k-th type of operating data; wherein k=1,2,3,4, respectively representing current, voltage, active power, and reactive power; i=1,2,...n; j=1,2,...m; Represented as a moment object, represented as a moment and value pair;

[0047] Step S22: using an XOR operation to detect whether there is duplicate data between any two objects in each data block in the k-th type of running data; if the operation result is 0, it indicates that duplicate data exists, and one of them needs to be removed; if the operation result is 1, it indicates that there is no duplicate data in the data block;

[0048] Step S23: After removing duplicate data from each data block, an XOR operation is performed on any two data blocks to detect whether there is duplicate data, that is, each object in one data block is XORed with each object in another data block. If duplicate data exists, only one data is retained. When performing the XOR operation, the time is XORed to remove the data of the duplicate time. The time of the monitoring data obtained after this operation is unique, and there is no data of the same time.

[0049] Step S3: Abnormal data detection and processing: The big data analysis platform detects whether the extracted operating data of the main transformer is abnormal. If there is abnormal data, the abnormal data is removed and the processed data is input into step S4. Specifically, the following steps are included:

[0050] Step S31: setting the maximum and minimum values ​​of the main transformer current, voltage, active power and reactive power on the big data analysis platform;

[0051] Step S32: Detect whether each type of extracted operating data is between the set corresponding maximum and minimum values. If the corresponding value is not between the minimum and maximum values, it is determined to be abnormal data and the value is eliminated.

[0052] Step S4: local outlier detection and processing: The big data analysis platform detects whether there are local outliers in the extracted operating data of the main transformer. If there are local outliers, the local outliers are removed.

[0053] The specific steps include:

[0054] Step S41: Divide the extracted operation data of each type of the main transformer into n data blocks; initialize the distance between each object value of the data block and its (m+k) nearest neighbor to the maximum value;

[0055] Step S42: Calculate the distance between each object value of the running data and each object value of the first data block, and update the (m+k) nearest neighbors of each object value in the first data block, and calculate the outlier degree of each object value in real time. When the number of neighbors is less than m+k, the outlier degree is set to infinity, and the outliers less than the initial threshold c are excluded from the data block; the outlier degree of each object value is the sum of the distances between the object value and its m+1 to m+k nearest neighbors; the initial threshold c is set to 0;

[0056] Step S43: after processing the first data block, sort the values ​​of the objects that are not excluded in the first data block from large to small according to the outlier degree, take the first n object values ​​and add them to the TOP n outliers, and update the threshold c;

[0057] Step S44: Calculate the distance between each object value in the running data and each object value in the second data block, and update the (m+k) nearest neighbors of each object value in the second data block, and calculate the outlier degree of each object value in real time. When the number of nearest neighbors is less than m+k, the outlier degree is set to infinity, and the outlier degree is less than the threshold value c is excluded from the data block;

[0058] Step S45: after processing the second data block, if the outlier degree of the object value that is not excluded in the second data block is greater than the outlier degree in the TOP n outliers, then update the TOP n outliers and update the threshold c;

[0059] Step S46: for the i-th data block, i=3, 4, 5...n, repeat steps S44-S45 until all data blocks are processed and the TOP n outliers are output;

[0060] In step S43 and step S45, when updating the threshold c, the outlier degree of the nth outlier point in the TOP n outliers is used as the value of the threshold c.

[0061] Preferably, in step S42

[0062] Step S5, data integrity detection and processing: The big data analysis platform detects whether the extracted operation data is complete. If there are missing values ​​in the extracted operation data, the missing values ​​are filled in and the filled operation data is output as the final processed data. Specifically, the following steps are included:

[0063] Step S51: for each type of operation data, after being processed in steps S2 to S4, data missing includes missing of both time and value, and missing of value only; detecting the integrity of the processed operation data of the main transformer, judging whether the data is complete, i.e., including the corresponding time and value, and if the data is missing, judging the corresponding data missing type;

[0064] Step S52: If both the time and the value are missing, first fill in the corresponding missing time, convert the missing type to only missing value, and then use the method in step S53 to complete the data;

[0065] Step S53: If only a value is missing, extract the N values ​​before and after the corresponding moment of the missing value, calculate the average value Eq of the N values ​​before the corresponding moment and the average value Eh of the N values ​​after the corresponding moment, and use the N values ​​before as the first data and the N values ​​after as the second data, assign a weight λ to the first value adjacent to the missing value in the first data and the second data, and assign a weight a to the second value adjacent to the missing value in the first data and the second data. The weights of N-2 data in the first data and the second data are Among them, λ>a;

[0066] Step S54: multiply the N values ​​in the first data by their corresponding weights and sum them to obtain a first calculated value, multiply the N values ​​in the second data by their corresponding weights and sum them to obtain a second calculated value, average the first calculated value and the second calculated value to obtain an intermediate calculated value, and determine whether the intermediate calculated value is between the average values ​​Eq and Eh. If the intermediate calculated value is between the average values ​​Eq and Eh, use the intermediate calculated value as a supplementary value for the missing value of the data. If the intermediate calculated value is not between the average values ​​Eq and Eh, adjust the values ​​of the weight values ​​λ and a, and repeat the above calculation so that the calculated intermediate calculated value is between the average values ​​Eq and Eh.

[0067] For example, if the missing voltage data is the data at the 12th moment, set N=4, that is, take the voltage values ​​at the 4 moments before the 12th moment as the first data, where the weight of the voltage value at the 11th moment is λ, the weight of the voltage value at the 10th moment is a, and the weights of the voltage values ​​at the 9th and 8th moments are Assume that the voltage values ​​at the 4 moments before the 12th moment are D11, D10, D9, and D8, then Eq = (D11+D10+D9+D8) / 4.

[0068] The voltage values ​​at the 4 moments after the 12th moment are taken as the first data, where the weight of the voltage value at the 13th moment is λ, the weight of the voltage value at the 14th moment is a, and the weights of the voltage values ​​at the 15th and 16th moments are Assume that the voltage values ​​at the 4 moments before the 12th moment are D13, D14, D15, and D16, then Eh = (D13 + D14 + D15 + D16) / 4.

[0069] The calculation of the first intermediate value is:

[0070] The second intermediate value is calculated as:

[0071] If the average values ​​Eq and Eh are equal, the missing value at the corresponding time is the average value Eq or Eh.

[0072] If the missing values ​​are from the initial two moments, the calculation is performed using the next N adjacent values, and it is observed whether the following values ​​increase or decrease with time. If it is an increasing trend, the calculated middle value is set to be smaller than the average value Eh of the next N values ​​at the corresponding moment; if it is a decreasing trend, the calculated middle value is set to be larger than the average value Eh of the next N values ​​at the corresponding moment.

[0073] If the missing values ​​are from the last two moments, the calculation is performed using the preceding N values ​​adjacent to the last two moments, and it is observed whether the preceding values ​​increase or decrease over time. If there is an increasing trend, the calculated middle value is set to be greater than the average value Eq of the N values ​​following the corresponding moment; if there is a decreasing trend, the calculated middle value is set to be less than the average value Eq of the N values ​​following the corresponding moment.

[0074] Those of ordinary skill in the art will appreciate that the units of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition of each example has been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0075] In the embodiments provided in the present application, it should be understood that the division of units is merely a logical function division, and there may be other division methods in actual implementation, for example, multiple units may be combined into one unit, one unit may be split into multiple units, or some features may be ignored.

[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and specification of the present invention.

Claims

1. A main transformer big data preprocessing method based on a big data analysis platform is characterized by: The following steps are involved: Step S1, data collection and storage: extracting the operating data of the main transformer from the dispatching automation system, including the current, voltage, active power and reactive power at each moment of operation, and storing the extracted operating data in the big data analysis platform; Step S2, duplicate data detection and processing: the big data analysis platform detects duplicate data in the extracted operation data of the main transformer, retains one of the duplicate data, removes redundant duplicate data, and inputs the processed data into step S3; Step S3: abnormal data detection and processing: the big data analysis platform detects whether the extracted operation data of the main transformer is abnormal. If there is abnormal data, the abnormal data is removed and the processed data is input into step S4; Step S4: local outlier detection and processing: the big data analysis platform detects whether there are local outliers in the extracted operation data of the main transformer. If there are local outliers, the local outliers are removed; Step S5, data integrity detection and processing: the big data analysis platform detects whether the extracted operation data is complete. If there are missing values ​​in the extracted operation data, the missing values ​​are filled in and the filled operation data is output as the final processed data; The data integrity detection and processing in step S5 specifically includes the following steps: Step S51: for each type of operation data, after being processed in steps S2 to S4, data missing includes missing of both time and value, and missing of value only; detecting the integrity of the processed operation data of the main transformer, judging whether the data is complete, i.e., including the corresponding time and value, and if the data is missing, judging the corresponding data missing type; Step S52: If both the time and the value are missing, first fill in the corresponding missing time, convert the missing type to only missing value, and then use the method in step S53 to complete the data; Step S53: If only a value is missing, extract the N values ​​before and after the corresponding moment of the missing value, calculate the average value Eq of the N values ​​before the corresponding moment and the average value Eh of the N values ​​after the corresponding moment, and use the N values ​​before as the first data and the N values ​​after as the second data, and assign a weight to the first value adjacent to the missing value in the first data and the second data. , assign a weight to the second value in the first data and the second data that is adjacent to the missing value , the weight of N-2 data in the first data and the second data is ; Step S54: multiply the N values ​​in the first data by their corresponding weights and sum them to obtain a first calculated value, multiply the N values ​​in the second data by their corresponding weights and sum them to obtain a second calculated value, average the first calculated value and the second calculated value to obtain an intermediate calculated value, and determine whether the intermediate calculated value is between the average values ​​Eq and Eh. If the intermediate calculated value is between the average values ​​Eq and Eh, use the intermediate calculated value as a supplementary value for the missing value of the data; if the intermediate calculated value is not between the average values ​​Eq and Eh, adjust the weight value. and Repeat the above calculations until the intermediate value obtained is between the average values ​​Eq and Eh. If the average values ​​Eq and Eh are equal, the missing value at the corresponding time is the average value Eq or Eh; If the missing values ​​are from the initial two moments, the calculation is performed using the next N values ​​adjacent to them, and it is observed whether the following values ​​increase or decrease over time. If it is an increasing trend, the calculated middle value is set to be smaller than the average value Eh of the next N values ​​at the corresponding moment; if it is a decreasing trend, the calculated middle value is set to be larger than the average value Eh of the next N values ​​at the corresponding moment; If the missing values ​​are from the last two moments, the calculation is performed using the preceding N values ​​adjacent to the last two moments, and it is observed whether the preceding values ​​increase or decrease over time. If there is an increasing trend, the calculated middle value is set to be greater than the average value Eq of the N values ​​following the corresponding moment; if there is a decreasing trend, the calculated middle value is set to be less than the average value Eq of the N values ​​following the corresponding moment.

2. The main transformer big data preprocessing method based on the big data analysis platform according to claim 1 is characterized in that: The duplicate data detection and processing in step S2 specifically includes the following steps: Step S21: Divide the extracted operation data of the main transformer into n data blocks according to type; each data block includes m objects; the recording method of the objects is , represents the jth data in the ith data block in the kth type of operating data; wherein, k=1,2,3,4, respectively represents current, voltage, active power, and reactive power; i=1,2,···n; j=1,2,···m; Represented as a moment object, represented as a moment and value pair; Step S22: using an XOR operation to detect whether there is duplicate data between any two objects in each data block in the k-th type of running data; if the operation result is 0, it indicates that duplicate data exists, and one of them needs to be removed; if the operation result is 1, it indicates that there is no duplicate data in the data block; Step S23: After removing duplicate data from each data block, an XOR operation is performed on any two data blocks to detect whether there is duplicate data, that is, each object in one data block is XORed with each object in the other data block. If duplicate data exists, only one data is retained.

3. The main transformer big data preprocessing method based on the big data analysis platform according to claim 2 is characterized in that: When performing an XOR operation, the time is XORed and the data of duplicate time is eliminated.

4. The main transformer big data preprocessing method based on the big data analysis platform according to claim 1 is characterized in that: The abnormal data detection and processing in step S3 specifically includes the following steps: Step S31: setting the maximum and minimum values ​​of the main transformer current, voltage, active power and reactive power on the big data analysis platform; Step S32: Detect whether each type of extracted operating data is between the set corresponding maximum and minimum values. If the corresponding value is not between the minimum and maximum values, it is determined to be abnormal data and the value is eliminated.

5. The main transformer big data preprocessing method based on the big data analysis platform according to claim 1 is characterized in that: The local outlier detection and processing in step S4 specifically includes the following steps: Step S41: Divide the extracted operation data of each type of the main transformer into n data blocks; initialize the distance between each object value of the data block and its (m+k) nearest neighbor to the maximum value; Step S42: Calculate the distance between each object value of the running data and each object value of the first data block, and update the (m+k) nearest neighbors of each object value in the first data block, and calculate the outlier degree of each object value in real time. When the number of neighbors is less than m+k, the outlier degree is set to infinity, and the outliers less than the initial threshold c are excluded from the data block; the outlier degree of each object value is the sum of the distances between the object value and its m+1 to m+k nearest neighbors; Step S43: after processing the first data block, sort the values ​​of the objects that are not excluded in the first data block from large to small according to the outlier degree, take the first n object values ​​and add them to the TOP n outliers, and update the threshold c; Step S44: Calculate the distance between each object value in the running data and each object value in the second data block, and update the (m+k) nearest neighbors of each object value in the second data block, and calculate the outlier degree of each object value in real time. When the number of nearest neighbors is less than m+k, the outlier degree is set to infinity, and the outlier degree is less than the threshold value c is excluded from the data block; Step S45: after processing the second data block, if the outlier degree of the object value that is not excluded in the second data block is greater than the outlier degree in the TOP n outliers, then update the TOP n outliers and update the threshold c; Step S46: for the i-th data block, i=3, 4, 5...n, repeat steps S44-S45 until all data blocks are processed and the TOP n outliers are output; In step S43 and step S45, when updating the threshold c, the outlier degree of the nth outlier point in the TOP n outliers is used as the value of the threshold c.

6. The main transformer big data preprocessing method based on the big data analysis platform according to claim 5 is characterized in that: The threshold c in step S42 is set to 0.

Citation Information

Patent Citations

  • A method and apparatus for generating a risk assessment scale

    CN109359850A

  • Method for interpolating and supplementing missing values based on neighbor algorithm in power load prediction

    CN111768034A