Oil chromatography data cleaning method and device based on time sequence, electronic equipment and storage medium
Through the time series-based oil chromatography data cleaning method, the problem that the existing technology cannot effectively capture the data patterns and trend changes in dissolved gas monitoring in oil is solved, and high-precision data cleaning and fault information retention are achieved, providing more reliable data support for transformer fault diagnosis and status evaluation.
Patent Information
- Application Number
- CN202510158033.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-30
AI Technical Summary
Existing dissolved gas data cleaning methods in oil cannot effectively capture the potential laws and trend changes in dissolved gas monitoring data in oil, resulting in the loss of important fault information and affecting fault warning.
The time series-based oil chromatography data cleaning method is used to input the oil chromatography data to be cleaned, pre-processing and cleaning, time series analysis, missing detection and repair are performed, and key features are extracted to achieve high-precision data cleaning.
This method can better capture the time series characteristics of oil chromatography data, improve data cleaning accuracy, and ensure that fault information is not lost, thereby providing more reliable data support for transformer fault diagnosis and status evaluation.
Smart Images

Figure CN120067538A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power equipment status monitoring, and particularly relates to a method, device, electronic device and storage medium for cleaning oil chromatogram data based on time series. Background Art
[0002] Dissolved Gas Analysis (DGA) in oil is one of the important means for power transformer status monitoring and fault diagnosis. By monitoring the concentration of dissolved gases in transformer oil, especially the content changes and trends of gases such as hydrogen (H 2 ), methane (CH 4 ), ethane (C 2 H 6 ), ethylene (C 2 H 4 ), etc., local discharge, overheating, insulation deterioration and other insulation defects and potential fault problems inside the transformer can be detected in time. Therefore, the dissolved gas monitoring data in oil plays a crucial role in transformer status assessment, fault diagnosis and risk warning.
[0003] However, the dissolved gas data in oil often faces various difficulties and challenges in the actual application process, such as low accuracy, insufficient reliability, etc., which are mainly due to limited effective monitoring data and low data quality.
[0004] In order to improve the quality of dissolved gas monitoring data in oil, most traditional methods for cleaning dissolved gas data in oil adopt statistical means, such as removing outliers based on statistical thresholds, linear interpolation, etc. to process data noise and missing values. Although these methods can improve the data quality to a certain extent, due to the lack of consideration of the time series characteristics of monitoring data, they cannot effectively capture the potential laws and trend changes of dissolved gas monitoring data in oil. For example, when dealing with the drastic changes in the concentration of dissolved gases in oil caused by local discharge or overheating, too simple cleaning methods may lead to the loss of important fault information and affect subsequent fault warning.
[0005] Therefore, a more accurate method for cleaning oil chromatogram data is needed. Summary of the Invention
[0006] In view of this, the embodiments of the present application provide a method, device, electronic device and storage medium for cleaning oil chromatogram data based on time series, with high cleaning accuracy, improved data quality, and providing more reliable data support for transformer fault diagnosis and status assessment.
[0007] The first aspect of the embodiments of the present application provides a method for cleaning oil chromatogram data based on time series, including:
[0008] Inputting the oil chromatogram data to be cleaned;
[0009] Preprocess and clean the oil chromatographic data;
[0010] Perform time series analysis on the data after preprocessing and cleaning, identify and nullify outliers, and obtain the change trend of the data after nullifying the outliers;
[0011] Perform missing value detection and missing value repair on the data after nullifying the outliers to obtain repaired data with a trend consistent with the change trend;
[0012] Extract key features from the repaired data to complete the cleaning of the oil chromatographic data.
[0013] In one embodiment, the preprocessing and cleaning of the oil chromatographic data includes:
[0014] Unify the format and data type of each attribute column data element of the oil chromatographic data, and align the oil chromatographic data with the corresponding time stamps;
[0015] Delete duplicate record rows in the oil chromatographic data;
[0016] Adjust the time stamp interval to be consistent with the data acquisition period;
[0017] Identify basic outliers in the oil chromatographic data, mark and nullify the basic outliers.
[0018] In one embodiment, the time series analysis of the data after preprocessing and cleaning, identifying and nullifying outliers, and obtaining the change trend of the data after nullifying the outliers includes:
[0019] Perform time series analysis on the data after preprocessing and cleaning to obtain the predicted trend of the data after preprocessing and cleaning;
[0020] Based on the predicted trend, perform outlier detection on the data after preprocessing and cleaning to identify outliers;
[0021] Mark and nullify the outliers;
[0022] Perform noise smoothing on the data after nullifying the outliers, and extract the long-term trend as the change trend of the data after nullifying the outliers.
[0023] In one embodiment, performing missing value detection and missing value repair on the data after nullifying the outliers to obtain repaired data with a trend consistent with the change trend includes:
[0024] Perform missing value detection on the data after nullifying the outliers to identify missing data;
[0025] Repair the missing data using any one of interpolation method, trend extrapolation method, and long short-term memory network prediction method according to the nature of the missing data;
[0026] Perform trend verification on the repaired data so that the trend of the repaired data is consistent with the change trend.
[0027] The second aspect of the embodiments of the present application provides an oil chromatographic data cleaning device based on time series, including:
[0028] An input module for inputting oil chromatographic data to be cleaned;
[0029] A data preprocessing and cleaning module for preprocessing and cleaning the oil chromatographic data;
[0030] A time series analysis module for performing time series analysis on the data after preprocessing and cleaning, identifying and nullifying outliers, and obtaining the change trend of the data after nullifying outliers;
[0031] A data repair module for performing missing detection and missing repair on the data after nullifying outliers to obtain repaired data whose trend is consistent with the change trend;
[0032] A feature extraction module for extracting key features from the repaired data to complete the cleaning of oil chromatographic data.
[0033] In one embodiment, the time series analysis module includes:
[0034] A trend analysis module for performing time series analysis on the data after preprocessing and cleaning to obtain the predicted trend of the data after preprocessing and cleaning;
[0035] An outlier detection module for detecting outliers in the data after preprocessing and cleaning based on the predicted trend, identifying outliers; and marking and nullifying the outliers;
[0036] A noise smoothing module for performing noise smoothing processing on the data after nullifying outliers and extracting the long-term trend as the change trend of the data after nullifying outliers.
[0037] The third aspect of the embodiments of the present application provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the time series-based oil chromatographic data cleaning method provided in the first aspect of the embodiments of the present application.
[0038] The fourth aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the time-series-based oil chromatographic data cleaning method provided in the first aspect of the embodiments of the present application.
[0039] The time-series-based oil chromatographic data cleaning method provided in the first aspect of the embodiments of the present application includes: inputting the oil chromatographic data to be cleaned; performing preprocessing cleaning on the oil chromatographic data; performing time-series analysis on the data after the preprocessing cleaning to identify and nullify outliers, and obtaining the change trend of the data after nullifying the outliers; performing missing value detection and missing value repair on the data after nullifying the outliers to obtain repaired data whose trend is consistent with the change trend; and extracting key features from the repaired data to complete the cleaning of the oil chromatographic data. It can make full use of the time-series characteristics of the oil chromatographic data, analyze and predict historical data using a time-series model, so as to achieve data cleaning. Compared with traditional methods, they can better capture time dependence and complex patterns, and improve the cleaning accuracy. Oil chromatographic data usually has the characteristics of sudden anomalies and trend changes. Introducing a time-series model for cleaning is not only an extension of conventional analysis, but also can more accurately locate and repair data deviations caused by external interference or sensor failures. Predicting the reasonable range of data (such as future trends) based on the time-series model and dynamically adjusting the repair method of outliers, this prediction-based cleaning strategy cannot be achieved by traditional static methods. The high-quality data cleaned by this method provides a solid data foundation for subsequent transformer fault diagnosis and condition assessment, helps to more accurately evaluate the health status of the transformer, detect potential faults in advance, extend the service life of the equipment and reduce maintenance costs.
[0040] It can be understood that the beneficial effects of the above second to fourth aspects can be referred to the relevant descriptions in the above first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 FIG. is a schematic flowchart of a time-series-based oil chromatographic data cleaning method provided by an embodiment of the present application;
[0043] Figure 2It is a schematic flowchart of a method for cleaning oil chromatographic data based on time series provided by another embodiment of the present application;
[0044] Figure 3 It is a schematic diagram of the LSTM model architecture;
[0045] Figure 4 It is a schematic flowchart of outlier detection provided by an embodiment of the present application;
[0046] Figure 5 It is a schematic structural diagram of an apparatus for cleaning oil chromatographic data based on time series provided by an embodiment of the present application;
[0047] Figure 6 It is a schematic structural diagram of an input module provided by an embodiment of the present application. Detailed implementation manners
[0048] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are presented in order to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of the present application.
[0049] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0050] References to "one embodiment" or "some embodiments" etc. described in the specification of the present application mean that specific features, structures, or characteristics described in conjunction with the embodiment are included in one or more embodiments of the present application. Thus, the phrases "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0051] The method for cleaning oil chromatographic data based on time series provided by the embodiments of the present application can be executed by the processor of an electronic device when running a computer program with corresponding functions. By inputting the oil chromatographic data to be cleaned, preprocessing and cleaning the oil chromatographic data, performing time series analysis on the preprocessed and cleaned data to identify and nullify outliers, and obtaining the change trend of the data after nullifying the outliers, performing missing value detection and missing value repair on the data after nullifying the outliers to obtain repaired data with a trend consistent with the change trend, and extracting key features from the repaired data to complete the cleaning of the oil chromatographic data. It can make full use of the time series characteristics of the oil chromatographic data, analyze and predict historical data using a time series model, thereby realizing data cleaning. Compared with traditional methods, they can better capture time dependence and complex patterns, and improve the cleaning accuracy. Oil chromatographic data usually has the characteristics of sudden anomalies and trend changes. Introducing a time series model for cleaning is not only an extension of conventional analysis but also can more accurately locate and repair data deviations caused by external interference or sensor failures. Predicting the reasonable range of data (such as future trends) based on a time series model and dynamically adjusting the repair method of outliers, this cleaning strategy based on prediction cannot be achieved by traditional static methods. The high-quality data cleaned by this method provides a solid data foundation for subsequent transformer fault diagnosis and condition assessment, helps to more accurately evaluate the health status of the transformer, discover potential faults in advance, extend the service life of the equipment and reduce maintenance costs.
[0052] As Figure 1 shown, the method for cleaning oil chromatographic data based on time series provided by the embodiments of the present application includes the following steps S101 to S105:
[0053] Step S101: Input the oil chromatographic data to be cleaned.
[0054] Step S102: Perform preprocessing and cleaning on the oil chromatographic data.
[0055] Step S103: Perform time series analysis on the preprocessed and cleaned data, identify and nullify outliers, and obtain the change trend of the data after nullifying the outliers.
[0056] Step S104: Perform missing value detection and missing value repair on the data after nullifying the outliers to obtain repaired data with a trend consistent with the change trend.
[0057] Step S105: Extract key features from the repaired data to complete the cleaning of the oil chromatographic data.
[0058] By combining historical patterns with trend analysis, the embodiments of this application can effectively identify and clean the noise and outliers in oil chromatographic data, ensuring that key fault information is not lost during the data processing process, achieving more accurate noise processing and missing data repair. During data cleaning, key fault signals and characteristic information are retained, providing a more reliable basis for the fault diagnosis of transformers, thereby improving the safety and stability of equipment operation and avoiding false alarms or missed alarms. It supports dynamic processing of real-time collected oil chromatographic data, has the ability of online data cleaning and repair, and can synchronously filter and repair noise during the data collection process, ensuring the efficient operation and rapid response of the monitoring system. This method can not only handle various outliers and noise in oil chromatographic data, but also cope with complex scenarios such as missing data and trend drift.
[0059] In one embodiment, step S102 includes the following steps S1021 to S1024:
[0060] Step S1021: Unify the format and data type of the data elements in each attribute column of the oil chromatographic data, and align the oil chromatographic data with the corresponding timestamps.
[0061] Step S1022: Delete the duplicate record rows in the oil chromatographic data.
[0062] In applications, for the rows of all columns containing timestamps, if there are duplicate row records, retain one record.
[0063] Step S1023: Adjust the timestamp interval to be consistent with the data collection period.
[0064] In applications, fix the data timestamp interval according to the data collection period to ensure that the timestamp intervals of the time series are equal.
[0065] Step S1024: Identify the basic outliers in the oil chromatographic data, mark the basic outliers and set them to null.
[0066] In one embodiment, the basic outliers include the data where all monitored quantities are the minimum values of the sensor range and the data where all monitored quantities remain unchanged within a set time period.
[0067] In applications, the basic outliers include two forms. One is that all monitored quantities are -99999 that frequently appear due to abnormal device circuits or uncalculated results in the internal digital circuit. The other form is that all monitored quantities remain completely unchanged for a long time. The reason for this situation is often still that the device fails to measure the actual data, and due to the sensor's own settings, the previous moment's value is displayed and falls into an infinite loop. In step 214, for the first type of basic outliers, set them to null, and for the second type, retain the first data in the time period when all monitored quantities remain unchanged for a long time and set the data for the rest of this period to null.
[0068] By processing the above two forms of basic outliers in the embodiments of the present application, the data quality can be effectively improved, interference caused by noise data to subsequent analysis can be avoided, and the authenticity and reliability of the data can be ensured. At the same time, it helps to improve the accuracy and processing efficiency of subsequent time series analysis and anomaly detection, and reduce misleading conclusions.
[0069] In one embodiment, as Figure 2 shown, step S103 includes the following steps S1031 to S1034:
[0070] Step S1031: Perform time series analysis on the preprocessed and cleaned data to obtain the predicted trend of the preprocessed and cleaned data.
[0071] In application, step S1031 specifically includes the following steps:
[0072] Step S311: Based on the preprocessed and cleaned data, decompose the time series of each gas, decompose the data into a trend part, a periodic part, and a random fluctuation part, and apply a moving average smoothing algorithm to remove short-term random fluctuations and extract the long-term trend;
[0073] Step S312: Extract the trend of the gas concentration. If the change of the gas concentration over time is linear, perform regression analysis on the gas concentration to construct a linear trend model. If the change of the gas concentration shows a non-linear trend, use a non-linear model for fitting, and use the fast Fourier transform (FFT) to extract the periodic components and identify whether there are periodic fluctuations in the gas concentration;
[0074] Step S313: Convert the time series data into a sliding window form. For example, input the data of the past n steps to predict the next step or the future m steps. Use the LSTM time series prediction model, introduce prior knowledge such as industry rules and expert thresholds into the traditional LSTM for model training, adjust the hyperparameters, evaluate the model, select the optimal parameter model, predict the future change trend of the gas concentration, identify potential fault signs, set a threshold in combination with historical data, and trigger a fault warning when the gas concentration exceeds the normal range;
[0075] In application, step S313, predicting the future change trend of the gas concentration is achieved by constructing an LSTM time series prediction model based on oil chromatogram data, which includes:
[0076] Constructing an LSTM Time Series Prediction Model: LSTM has the ability to remember long and short term information and can be used to process sequential data. In time series prediction, LSTM can be used both as a multi-variable prediction mechanism and as a single unit prediction mechanism. In the process of predicting the future change trend of gas concentration, LSTM can be constructed as a single unit prediction model to predict the future value of gas concentration. Specifically, the historical data of gas concentration is used as the input of LSTM, and its future value is used as the output of LSTM. During the training process, the error backpropagation algorithm is used to update the parameters of LSTM to optimize the prediction performance of the model.
[0077] The basic structure of the LSTM time series prediction model is as Figure 3 shown, where the black lines represent the transfer of vectors between different nodes, the "+" represents vector arithmetic operations, each rectangle represents a different neural network layer, the converging lines represent the connection of vectors, and the diverging lines indicate that the vector content is copied and sent to different locations. Here, x represents the input historical data, and h represents the output data stored in the module. During the LSTM training process, the output parameter h is stored through deep learning, the updated state is saved, and three gates are added, which include:
[0078] (1) Forget Gate: The forget gate determines the information that the model discards. The forget gate reads the output h t-1 of the previous sequence model and the current model input X t to control whether each number in the cell state is retained, as shown in the figure. Its transfer function is as follows:
[0079] f t = α(W f [h t-1 , x t + b f )
[0080] In the formula, f t represents the output of the forget gate, α represents the activation function, W f represents the weight of the forget gate, x t represents the current model input, h t-1 represents the output of the previous sequence model, and b f represents the bias of the forget gate.
[0081] (2) Input Gate: The input gate is as shown in the figure. The input gate can be divided into two parts. One part is to find the data states that need to be updated. The other part is to update the information that needs to be updated into the module. Its structure can be expressed by the following formula:
[0082] i t = α(W i · [h t-1 , x t+b i )
[0083] C t =tanh(W C ·[h t-1 ,x t +b C )
[0084] In the formula, i t represents the data state to be updated, α represents the activation function, x t represents the input of the current model, h t-1 represents the output of the previous sequence model, W t represents the weight for calculating i t , b t represents the bias for calculating i t , C t represents the data state created using the tanh function, W C represents the weight for calculating C t , b C represents the bias for calculating C t After the forget gate finds the information f t to be forgotten, it multiplies it with the old state to discard the information that is determined to be discarded. (If the weight at the corresponding position needs to be discarded, it is set to 0), then, the result is added with i t *C t to enable the module state to obtain new information. In this way, the update of the data within the module is completed
[0085] (3) Output gate: In the output gate, through an activation function layer (using the Sigmoid activation function) to determine which part of the information will be output, then the cell state is processed through tanh (obtaining a value between -1 and 1), and it is multiplied with the output of the Sigmoid gate to obtain the final output part. The output gate structure is represented by the following formula:
[0086] O t =α(W O ·[ht- 1 ,x t +b O )
[0087] h t =O t ×tanh(B t )
[0088] In the formula, O t represents the output information, α represents the activation function, W O represents the weight for calculating O t , b O represents the bias for calculating O tThe bias, B t Indicates the state of the module after update, h t Indicates the output result of the current sequence model.
[0089] After constructing the LSTM time series prediction model, calculate the weights and biases manually, input the oil chromatographic data and train until the future change trend of the gas concentration that meets the expectations is output, and complete the training of the LSTM time series prediction model based on the oil chromatographic data.
[0090] During trend analysis, plot the time series trend diagrams of each gas, including the actual value, the smoothed trend value, and the predicted value of the gas concentration, compare the trends of different gases, and find abnormal increases or decreases in the gas concentration.
[0091] Step S1032: Based on the predicted trend, perform outlier detection on the preprocessed and cleaned data to identify outliers.
[0092] In one embodiment, based on the predicted trend, perform at least one of correlation anomaly detection, volatility anomaly detection, level shift anomaly detection, outlier anomaly detection, and periodic anomaly detection on the preprocessed and cleaned data. When the data is determined to be abnormal in any one or more detections, identify the data as an outlier.
[0093] In the application, as Figure 4 shown, outlier detection includes five anomaly patterns: level shift anomaly, volatility anomaly, outlier anomaly, periodic anomaly, and correlation anomaly. The steps of its detection are as follows:
[0094] Step S321: Correlation anomaly monitoring. According to IEC60599, for the gas concentration in the oil chromatographic data of the same normally operating transformer, the carbon dioxide is greater than carbon monoxide and much greater than other gases, and the content of acetylene is the lowest. If the relative relationship of their values violates the above common sense, regard it as an outlier;
[0095] Step S322: Volatility anomaly data detection. For an online DGA sensor, an index to characterize its stability is the coefficient of variation statistic of the sequence in the tracking time window, which is the ratio of the standard deviation to the mean of the sequence and is used to characterize the standard deviation rate of different sequences. For an oil chromatographic time series with a non-zero mean, the greater the coefficient of variation, the more unstable the sensor. The time window is a time range selected from the time series data;
[0096]
[0097] In the formula, σ is the standard deviation of the sequence; μ is the mean of the sequence; c v is the coefficient of variation.
[0098] Step S323, Detection of horizontal shift abnormal data. In some cases, whether a time point is normal depends on whether the time value is consistent with its nearest time. Slide two time windows side by side and continue to track the difference between their means or medians. As long as the statistical data in the left and right windows are significantly different, it indicates that a sudden change has occurred near this time point. The length of the time window controls the time scale of the change to be detected: for horizontal shift transformation, both windows should be long enough to capture the steady state. In the case of on-line monitoring data of transformer oil chromatography, this situation may be the inherent deviation of the sensor or a change in the state of the transformer. It is necessary to perform correlation analysis on the oil chromatography data to determine whether it is an accuracy deviation;
[0099] For horizontal shift abnormal data, it is necessary to determine whether this abnormal point is due to the weakening of data accuracy caused by the failure of the oil chromatography sensor or is unrelated to data quality and belongs to the abnormality of the transformer itself. At this time, the time series of a single gas concentration needs to be analyzed for similarity at the abnormal point with the time series of other gas concentrations and other physical quantities for transformer condition monitoring (such as oil temperature, etc.) to obtain the true state of this point.
[0100] Step S324, Detection of outlier abnormal data. An outlier is a data point whose value is significantly different from other values. Without considering the time relationship between data points, the outlier in the time series exceeds the normal range of this series. Similar to the detection of horizontal shift abnormal data, it is also necessary to determine whether this abnormal point is due to the weakening of data accuracy caused by the failure of the oil chromatography sensor or is unrelated to data quality and belongs to the abnormality of the transformer itself;
[0101] Step S325, Periodic abnormality monitoring. When time series data is affected by periodic factors, there are often periodic patterns. To extract and analyze these periodic patterns, the classical periodic decomposition method STL is used. Through STL decomposition, the original time series is decomposed into a trend, a periodic component, and a residual part. After extracting the periodic component, the residual series can be used to detect abnormalities. By analyzing the deviation of the residual part, the time periods when the time series deviates from its periodic pattern can be identified, and potential abnormal points can be found;
[0102] Step S1033, Mark and nullify the abnormal values.
[0103] In the application, for the detected abnormal value data, first mark the abnormal value, and then nullify it.
[0104] Step S1034, Perform noise smoothing on the data after nullifying the abnormal values, and extract the long-term trend as the change trend of the data after nullifying the abnormal values.
[0105] In the application, the data sequence processed in this step is the data sequence after nullifying the outlier values. By selecting an appropriate time window and combining with the moving average method, the short-term random fluctuations are removed, the long-term trend is extracted, and the long-term trend of each gas is displayed.
[0106] In one embodiment, step S104 includes the following steps S1041 to S1043:
[0107] Step S1041: Perform missing detection on the data after nullifying the outlier values to identify the missing data.
[0108] In the application, check the data after nullifying the outlier values to identify the missing values. It includes record missing (all gas concentrations at a certain timestamp are missing) and data element missing (a certain gas data at a certain timestamp is missing).
[0109] Step S1042: Repair the missing data by using any one of the interpolation method, trend extrapolation method, and long short-term memory network prediction method according to the nature of the missing data.
[0110] In the application, select an appropriate completion method for data completion. According to the nature of the missing data and the characteristics of the historical data, select an appropriate completion method. If the data change is relatively stable, use the interpolation method to complete the missing values between two adjacent valid data. If the data change has an obvious trend or periodicity, use the trend extrapolation method to complete the missing values. For complex time series data, use the LSTM network to predict the missing values based on the historical data.
[0111] Step S1043: Perform trend verification on the repaired data so that the trend of the repaired data is consistent with the change trend.
[0112] In the application, for trend verification, after completing the data, compare the completion result with the trend of the historical data to ensure that the completed data maintains the same change trend as the overall historical data. After the data repair is completed, the system provides feedback to verify the effectiveness of the repair and makes adjustments if necessary.
[0113] In one embodiment, step S105 extracts features from the data sequence returned after data repair. Based on the oil chromatographic data, the features that can be extracted include:
[0114] Average concentration. The average concentration refers to the arithmetic average of the gas concentration values within a specific time period. It can reflect the typical concentration level of the gas during this time period and serve as a representative value of the overall gas behavior.
[0115] Maximum concentration. The highest value of the gas concentration within a certain time period. This value usually represents the extreme state of the gas concentration and may indicate the peak of a specific component in the sample.
[0116] Minimum concentration. The lowest value of the gas concentration within a certain period. This value reveals the presence of low-concentration components in the gas sample or reflects the gas concentration during a stable period;
[0117] Concentration change range. The range of gas concentration changes over a period of time, which refers to the difference between the maximum concentration and the minimum concentration. This indicator reflects the degree of fluctuation of the gas concentration;
[0118] Skewness and kurtosis. Skewness is a measure describing the symmetry of the gas concentration distribution. Positive skewness indicates that the data is skewed to the right, and negative skewness indicates that the data is skewed to the left. Kurtosis describes the sharpness of the gas concentration distribution. Higher kurtosis indicates the presence of sudden events or outliers in the distribution;
[0119] Gas concentration change rate. The gas concentration change rate refers to the rate of change of the gas concentration over time, which can help judge the speed of gas generation, release, or consumption;
[0120] Gas concentration standard deviation. The gas standard deviation is the degree of dispersion of the concentration data, which reveals the volatility and stability of the gas concentration changes. A high standard deviation indicates large fluctuations in the gas concentration and is usually used to evaluate the stability of equipment operation;
[0121] Coefficient of variation of gas concentration. The coefficient of variation of gas concentration is the ratio of the standard deviation to the average gas concentration and is used to compare the degree of variation of gas concentrations in different samples;
[0122] Gas concentration trend. The concentration trend refers to the change pattern of the gas over time, and long-term change rules are identified through sequence trend analysis;
[0123] Concentration mutation in time. The concentration mutation that occurs in a short period may indicate a sudden event, and this mutation may indicate a rapid fault within the transformer.
[0124] In one embodiment, step S105 also outputs the sequence data returned after data repair. This sequence data is high-quality sequence data after cleaning and can be applied to subsequent feature extraction or other scenarios.
[0125] In one embodiment, step S105 also extracts key features from the sequence data returned after data repair and then outputs them. The output features can be applied to model training for various scenarios such as subsequent transformer fault diagnosis and fault warning.
[0126] In one embodiment, the sequence data returned after data repair is also visually output. The visual content is as follows:
[0127] Comparison between the original data and the cleaned data. Display the oil chromatographic data before and after cleaning. By comparing the original data and the cleaned data, the data cleaning effect can be intuitively understood, such as the removal of noise and the filling of missing values;
[0128] Gas concentration trend analysis. Display the concentration change trends of different gases (such as H 2 , CH 4 , C 2 H 4 , CO, etc.) over time, which can help identify potential trends and anomalies;
[0129] Abnormal detection results. Highlight the abnormal points or abnormal patterns in the data, and identify the detected gas concentration mutations, abnormal gas generation rates, etc.;
[0130] Data distribution and statistical characteristics. Display the statistical characteristics of the gas concentration data, including the distribution, mean, standard deviation, etc. of the concentration;
[0131] Periodic analysis and residual analysis. Visualize the periodic patterns and the results of residual analysis, and display the deviation of the gas concentration from the periodic pattern in different time periods to help users discover abnormal points;
[0132] In addition to the above visual content, the gas concentration change amplitude and volatility, gas generation rate and concentration change rate, etc. can also be displayed, which can help comprehensively understand the effect of oil chromatographic data cleaning, timely discover potential faults, and provide intuitive support for the assessment of the transformer operation status.
[0133] In one embodiment, it also includes in-depth analysis and interpretation of the cleaned oil chromatographic data to assist in transformer fault diagnosis and status assessment. The main contents included are as follows:
[0134] Statistical analysis of the cleaned data. Conduct statistical analysis on the cleaned oil chromatographic data, display the statistical information of the data, and based on time series analysis, identify the long-term trends and short-term fluctuations of the gas concentrations in the oil;
[0135] Multi-gas ratio analysis. Analyze the ratios of different gas concentrations (such as C 2 H 2 / CH 4 , C 2 H 4 / H 2 , etc.), and combine classical oil chromatographic analysis methods such as the ratio method and the triangle method to diagnose potential faults of the transformer.
[0136] It should be understood that the sequence numbers of the steps in the above embodiments do not indicate the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0137] The embodiments of the present application further provide an oil chromatographic data cleaning device based on time series, which is used to execute the steps in the embodiments of the above-mentioned oil chromatographic data cleaning method based on time series. The oil chromatographic data cleaning device based on time series can be a virtual appliance in an electronic device, which is run by the processor of the electronic device, or the electronic device itself.
[0138] As Figure 5 shown, the oil chromatographic data cleaning device based on time series provided by the embodiments of the present application includes:
[0139] An input module 1, which is used to input the oil chromatographic data to be cleaned;
[0140] A data preprocessing and cleaning module 2, which is used to preprocess and clean the oil chromatographic data;
[0141] A time series analysis module 3, which is used to perform time series analysis on the data after preprocessing and cleaning, identify and set null the outliers, and obtain the change trend of the data after setting null the outliers;
[0142] A data repair module 4, which is used to perform missing detection and missing repair on the data after setting null the outliers, and obtain repaired data whose trend is consistent with the change trend;
[0143] A feature extraction module 5, which is used to extract key features from the repaired data to complete the cleaning of the oil chromatographic data.
[0144] The embodiments of the present application adopt a modular design, which is convenient for integration with existing monitoring systems, and at the same time has flexible scalability, and can be customized according to actual needs. For the input data to be cleaned, first perform data preprocessing and cleaning, secondly perform time series analysis on the data to identify the change trend in the data, based on trend analysis, further identify the outliers and the values deviating from the normal trend, and then for the missing values in the data and the missing values caused by processing outliers and noise, use a time series prediction model to complete, ensure the continuity of the data, the repaired data should be consistent with the historical trend, then extract key features from the repaired data, such as the change rate, peak value, mean value, etc. of the gas concentration, this step provides important parameters for the subsequent transformer fault diagnosis, and finally visually display, store and output the repaired data and the extracted feature data.
[0145] In the data cleaning process of the present invention, time series analysis and prediction techniques are innovatively combined, which can not only effectively handle noise, outliers and missing values, but also ensure that the repaired data is consistent with the historical trend. By using time series prediction models such as ARIMA model and LSTM network, the missing data is filled in, so as to avoid the accuracy of transformer fault diagnosis being affected by discontinuous or abnormal fluctuations of monitoring data.
[0146] Through in-depth feature extraction of the repaired oil chromatogram data, such as the change rate, peak value, mean value, variance, etc. of gas concentration, the present invention provides multi-dimensional data analysis means and provides key parameter support for subsequent fault diagnosis. Different from traditional simple statistical analysis, the present invention can identify long-term trends, abnormal fluctuations and potential fault signals in the data, providing a more comprehensive and accurate basis for transformer condition assessment.
[0147] In one embodiment, the data input module 1 includes:
[0148] A data display module 11 for displaying part of the data and the dimension size of the data;
[0149] A basic attribute module 12 for displaying the column names and data types of each attribute column of the data set;
[0150] A missing situation module 13, and the missing situation of the data will display the overall missing situation of the data set and the missing situation of each attribute column;
[0151] An acquisition information module 14 for displaying information such as the data acquisition period;
[0152] A data distribution module 15 for displaying the time series diagram, data distribution diagram and statistical information of each attribute column.
[0153] In one embodiment, the data preprocessing and cleaning module 2 is further used for:
[0154] Unify the format and data type of each attribute column data element of the oil chromatogram data, and align the oil chromatogram data with the corresponding time stamp;
[0155] Delete the duplicate record rows in the oil chromatogram data;
[0156] Adjust the time stamp interval to be consistent with the data acquisition period;
[0157] Identify the basic outliers in the oil chromatogram data, mark the basic outliers and set them to null.
[0158] In one embodiment, the basic outliers include the data where all monitored quantities are the minimum values of the sensor range and the data where all monitored quantities remain unchanged within a set time period.
[0159] In one embodiment, the timing analysis module 3 includes:
[0160] A trend analysis module 31 for performing time series analysis on the preprocessed and cleaned data to obtain the predicted trend of the preprocessed and cleaned data;
[0161] An outlier detection module 32 for performing outlier detection on the preprocessed and cleaned data based on the predicted trend, identifying outliers; and performing marking and nulling processing on the outliers;
[0162] A noise smoothing module 33 for performing noise smoothing processing on the data after nulling the outliers, and extracting the long-term trend as the change trend of the data after nulling the outliers.
[0163] In one embodiment, the outlier detection module is further configured to:
[0164] Perform at least one of correlation anomaly detection, volatility anomaly detection, level shift anomaly detection, outlier anomaly detection, and periodic anomaly detection on the preprocessed and cleaned data based on the predicted trend. When the data is determined to be abnormal in any one or more detections, identify the data as an outlier.
[0165] In one embodiment, the data repair module is further configured to:
[0166] Perform missing detection on the data after nulling the outliers, and identify the missing data;
[0167] Repair the missing data using any one of interpolation method, trend speculation method, and long short-term memory network prediction method according to the nature of the missing data;
[0168] Perform trend verification on the repaired data to make the trend of the repaired data consistent with the change trend.
[0169] In one embodiment, the feature extraction module 5 includes:
[0170] A feature extraction module 51 for extracting features from the data sequence returned after data repair. Based on the oil chromatographic data, the features that can be extracted include:
[0171] Average concentration. The average concentration refers to the arithmetic average of the gas concentration values within a specific time period. It can reflect the typical concentration level of the gas during this time period and serve as a representative value of the overall gas behavior;
[0172] Maximum concentration. The highest value of the gas concentration within a certain time period. This value usually represents the extreme state of the gas concentration and may indicate the peak value of a specific component in the sample;
[0173] Minimum concentration. The lowest value of the gas concentration within a certain period of time. This value reveals the presence of low-concentration components in the gas sample or reflects the gas concentration during a stable period;
[0174] Concentration change range. The range of gas concentration changes over a period of time, which refers to the difference between the maximum concentration and the minimum concentration. This indicator reflects the degree of fluctuation of the gas concentration;
[0175] Skewness and kurtosis. Skewness is a measure describing the symmetry of the gas concentration distribution. Positive skewness indicates that the data is skewed to the right, and negative skewness indicates that the data is skewed to the left. Kurtosis describes the sharpness of the gas concentration distribution. A higher kurtosis indicates the presence of sudden events or outliers in the distribution;
[0176] Gas concentration change rate. The gas concentration change rate refers to the rate of change of the gas concentration over time, which can help judge the speed of gas generation, release, or consumption;
[0177] Gas concentration standard deviation. The gas standard deviation is the degree of dispersion of the concentration data, which reveals the volatility and stability of the gas concentration changes. A high standard deviation indicates large fluctuations in the gas concentration and is usually used to evaluate the stability of equipment operation;
[0178] Gas concentration coefficient of variation. The gas concentration coefficient of variation is the ratio of the standard deviation to the average gas concentration and is used to compare the degree of variation of gas concentrations in different samples;
[0179] Gas concentration trend. The concentration trend refers to the change pattern of the gas over time, and the long-term change law is identified through sequence trend analysis;
[0180] Concentration mutation in time. The concentration mutation that occurs in a short period of time may indicate a sudden event, and this mutation may indicate a rapid fault in the transformer.
[0181] Data output module 52 is used to output the sequence data returned after data repair. This sequence data is high-quality sequence data after cleaning and can be applied to later feature extraction or other scenarios.
[0182] Eigenvalue output module 53 is used to extract key features from the sequence data returned after data repair and then output them. The output features can be applied to model training in various scenarios such as later transformer fault diagnosis and fault warning.
[0183] In one embodiment, it further includes a data visualization module 6, which includes a visualization module 61 and a result analysis module 62.
[0184] Visualization module 61 is used to also perform visual output on the sequence data returned after data repair. The visualization content is as follows:
[0185] Comparison between the original data and the cleaned data. The oil chromatographic data before and after cleaning are shown. By comparing the original data and the cleaned data, the data cleaning effect can be intuitively understood, such as the removal of noise and the filling of missing values;
[0186] Gas concentration trend analysis. The concentration change trends of different gases (such as H 2 , CH 4 , C 2 H 4 , CO, etc.) over time are shown, which can help identify potential trends and anomalies;
[0187] Abnormality detection results. The abnormal points or abnormal patterns in the data are highlighted, and the detected gas concentration mutations, abnormal gas generation rates, etc. are identified;
[0188] Data distribution and statistical characteristics. The statistical characteristics of the gas concentration data are shown, including the distribution, mean, standard deviation, etc. of the concentration;
[0189] Periodicity analysis and residual analysis. The periodic patterns and the results of residual analysis are visualized, and the deviation of the gas concentration from the periodic pattern in different time periods is shown to help users discover abnormal points;
[0190] In addition to the above visual content, the gas concentration change amplitude and volatility, gas generation rate and concentration change rate, etc. can also be shown, which can help comprehensively understand the effect of oil chromatographic data cleaning, timely discover potential faults, and provide intuitive support for the evaluation of the transformer operation status.
[0191] The result analysis module 62 is used to deeply analyze and interpret the cleaned oil chromatographic data to assist in transformer fault diagnosis and status evaluation. The main contents included are as follows:
[0192] Statistical analysis of the cleaned data. The cleaned oil chromatographic data are statistically analyzed, the statistical information of the data is shown, and based on time series analysis, the long-term trends and short-term fluctuations of the gas concentrations in the oil are identified;
[0193] Multi-gas ratio analysis. The ratios of different gas concentrations (such as C 2 H 2 / CH 4 , C 2 H 4 / H 2 , etc.) are analyzed, and combined with classical oil chromatographic analysis methods such as the ratio method and the triangle method, potential faults of the transformer are diagnosed.
[0194] In applications, each module in the oil chromatographic data cleaning device based on time series can be a software program module, can also be implemented by different logic circuits integrated in the processor, or can also be implemented by multiple distributed processors.
[0195] As shown in the figure, an embodiment of the present application further provides an electronic device, including: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, the steps in the above method embodiments are implemented.
[0196] In applications, the electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 6 merely examples of the electronic device, which do not constitute a limitation to the electronic device, may include more or fewer components than shown in the figure, or combine certain components, or different components.
[0197] In applications, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0198] In applications, the memory may be an internal storage unit of the electronic device in some embodiments, such as the hard disk or memory of the electronic device. The memory may also be an external storage device of the electronic device in other embodiments. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory may also include both the internal storage unit and the external storage device of the electronic device. The memory is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of a computer program. The memory may also be used to temporarily store data that has been output or will be output.
[0199] It should be noted that the content such as information interaction and execution process between the above devices / units, due to being based on the same concept as the method embodiments of the present application, for its specific functions and the technical effects brought, please refer to the method embodiment part for details, and will not be elaborated here.
[0200] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above-mentioned system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0201] The embodiment of this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the foregoing method embodiments can be implemented.
[0202] The embodiment of this application provides a computer program product, including a computer program. When the computer program product runs on an electronic device, the electronic device is enabled to execute the steps in the foregoing method embodiments.
[0203] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned method embodiments of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the foregoing method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the device / electronic device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0204] In the above embodiments, the descriptions of the various embodiments have their own focuses. For parts not described in detail or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0205] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0206] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0207] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0208] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of this application, and should all be included in the protection scope of this application.
Claims
1. A time series-based oil chromatography data cleaning method, characterized in that: include: Input the oil chromatogram data to be cleaned; Preprocessing and cleaning the oil chromatogram data; Performing time series analysis on the pre-processed and cleaned data, identifying and removing outliers, and obtaining a change trend of the data after removing outliers; Perform missing detection and missing repair on the data after the outliers are cleared, and obtain repair data with a trend consistent with the change trend; Key features are extracted from the repair data to complete oil chromatography data cleaning.
2. The time series-based oil chromatography data cleaning method according to claim 1, characterized in that: The pre-processing and cleaning of the oil chromatogram data comprises: Unifying the format and data type of each attribute column data element of the oil chromatogram data, and aligning the oil chromatogram data with a corresponding timestamp; Deleting duplicate record rows in the oil chromatogram data; Adjust the timestamp interval to be consistent with the data collection period; Identify basic outliers in the oil chromatogram data, mark the basic outliers and set them to zero.
3. The time series-based oil chromatography data cleaning method according to claim 2, characterized in that: The basic abnormal values include data in which all monitored quantities are the minimum values of the sensor range and data in which all monitored quantities remain unchanged within a set time period.
4. The time series-based oil chromatography data cleaning method according to claim 1, characterized in that: The performing of time series analysis on the pre-processed and cleaned data, identifying and removing outliers, and obtaining a change trend of the data after removing outliers, includes: Performing time series analysis on the preprocessed and cleaned data to obtain a predicted trend of the preprocessed and cleaned data; Based on the predicted trend, performing outlier detection on the pre-processed and cleaned data to identify outliers; Marking and blanking the abnormal values; The data after the outliers are removed are subjected to noise smoothing, and the long-term trend is extracted as the changing trend of the data after the outliers are removed.
5. The time series-based oil chromatography data cleaning method according to claim 4, characterized in that: Based on the predicted trend, performing outlier detection on the pre-processed and cleaned data to identify outliers includes: Based on the predicted trend, at least one of correlation anomaly detection, volatility anomaly detection, level shift anomaly detection, outlier anomaly detection and periodic anomaly detection is performed on the pre-processed and cleaned data. When the data is determined to be abnormal in any one or more of the detections, the data is identified as an outlier.
6. The time series-based oil chromatography data cleaning method according to claim 1, characterized in that: Perform missing detection and missing repair on the data after the outliers are removed to obtain repaired data whose trend is consistent with the change trend, including: Perform missing data detection on the data after removing outliers to identify missing data; According to the nature of the missing data, the missing data is repaired by using any one of an interpolation method, a trend inference method, and a long short-term memory network prediction method; The repaired data is trend checked to ensure that the trend of the repaired data is consistent with the change trend.
7. An oil chromatography data cleaning system based on time series, characterized in that: include: An input module, used for inputting the chromatographic data of the oil to be cleaned; A data preprocessing and cleaning module, used for preprocessing and cleaning the oil chromatogram data; A time series analysis module is used to perform time series analysis on the pre-processed and cleaned data, identify and clear outliers, and obtain a change trend of the data after clearing outliers; A data repair module, used to perform missing detection and missing repair on the data after the outliers are emptied, to obtain repaired data whose trend is consistent with the change trend; The feature extraction module is used to extract key features from the repair data to complete the oil chromatography data cleaning.
8. The time series-based oil chromatography data cleaning system according to claim 7, characterized in that: The data preprocessing and cleaning module is also used for: Unifying the format and data type of each attribute column data element of the oil chromatogram data, and aligning the oil chromatogram data with a corresponding timestamp; Deleting duplicate record rows in the oil chromatogram data; Adjust the timestamp interval to be consistent with the data collection period; Identify basic outliers in the oil chromatogram data, mark the basic outliers and set them to zero.
9. The time series-based oil chromatography data cleaning system according to claim 8, characterized in that: The basic abnormal values include data in which all monitored quantities are the minimum values of the sensor range and data in which all monitored quantities remain unchanged within a set time period.
10. The time series-based oil chromatography data cleaning system according to claim 7, characterized in that: The timing analysis module comprises: A trend analysis module, used for performing time series analysis on the pre-processed and cleaned data to obtain a predicted trend of the pre-processed and cleaned data; An outlier detection module is used to perform outlier detection on the pre-processed and cleaned data based on the predicted trend, identify outliers, and mark and clear the outliers; The noise smoothing module is used to perform noise smoothing on the data after the outliers are removed, and extract the long-term trend as the change trend of the data after the outliers are removed.
11. The oil chromatography data cleaning system based on time series according to claim 10, characterized in that: The outlier detection module is also used for: Based on the predicted trend, at least one of correlation anomaly detection, volatility anomaly detection, level shift anomaly detection, outlier anomaly detection and periodic anomaly detection is performed on the pre-processed and cleaned data. When the data is determined to be abnormal in any one or more of the detections, the data is identified as an outlier.
12. The time series-based oil chromatography data cleaning system according to claim 7, characterized in that: The data repair module is also used for: Perform missing data detection on the data after removing outliers to identify missing data; According to the nature of the missing data, the missing data is repaired by using any one of an interpolation method, a trend inference method, and a long short-term memory network prediction method; The repaired data is trend checked to ensure that the trend of the repaired data is consistent with the change trend.
13. An electronic device, characterized in that: The electronic device comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device implements the method as claimed in any one of claims 1 to 6.
14. A computer-readable storage medium storing a computer program, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.