Data cleaning and steady-state detection method and system based on robust time series modeling
Through the method of robust time series modeling, a robust model is constructed and combined with data cleaning and machine learning models, the accuracy and robustness of traditional steady-state detection methods under outliers and noise interference are solved, and high accuracy detection of data in complex industrial process is achieved.
Patent Information
- Application Number
- CN202510837171.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-06-23
AI Technical Summary
When traditional steady-state detection methods have low accuracy when there are outliers or noise interference, they cannot effectively deal with complex industrial process data, and are poorly robust.
Through a method based on robust time series modeling, a linear trend model is constructed and robust optimization is performed, and the steady state of time series data is judged by combining nonlinear optimization solutions and data cleaning and machine learning models.
It improves the accuracy and robustness of steady-state detection, can adapt to various complex industrial process data, and ensures the reliability of the final detection results.
Smart Images

Figure CN120354242B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of time series analysis and steady-state detection, and in particular to a data cleaning and steady-state detection method and system based on robust time series modeling. Background Art
[0002] Steady-state detection is a crucial task in data analysis and industrial control. For example, applications such as chemical process monitoring, financial market forecasting, and equipment status diagnosis require determining whether a system is in steady-state to facilitate effective decision-making. However, traditional steady-state detection methods (such as mean-variance analysis and maximum likelihood estimation) are prone to failure in the presence of outliers or noise, leading to misjudgments and low steady-state detection accuracy.
[0003] Currently, linear regression methods and time series modeling methods are widely used for trend analysis and steady-state detection. Among them, the linear trend model based on the least squares method (OLS) can effectively fit time series data and determine whether the data is in a steady state. However, the least squares method is extremely sensitive to outliers and is easily affected by extreme data points, resulting in misjudgment of steady-state detection and low accuracy of steady-state detection. Therefore, robust statistical methods have gradually attracted attention, among which the method based on ρ The robust regression method of the function performs well in resisting outliers. In addition, the traditional steady-state detection method cannot effectively cope with various complex industrial process data when dealing with complex low-quality data, and the steady-state detection results are inaccurate and unreliable.
[0004] Based on this, how to provide a steady-state detection method with higher accuracy, stronger robustness and stronger adaptability to meet the needs of modern engineering applications has become a technical problem that needs to be solved urgently in this field. Summary of the Invention
[0005] The purpose of this application is to provide a data cleaning and steady-state detection method and system based on robust time series modeling, which can adapt to various complex industrial process data and effectively improve the accuracy and robustness of steady-state detection.
[0006] To achieve the above objectives, this application provides the following solutions.
[0007] In a first aspect, the present application provides a data cleaning and steady-state detection method based on robust time series modeling, which includes the following steps.
[0008] Get the target time series data.
[0009] A linear trend model is constructed based on the target time series data; the linear trend model refers to a linear model including an intercept term, a slope term, and a random error term.
[0010] Robust optimization is performed on the linear trend model to obtain a robust model.
[0011] The robust model is subjected to nonlinear optimization and solution to determine robust parameters; the robust parameters include intercept term parameters, slope term parameters and fitting y values.
[0012] Based on the robust parameters, it is determined whether the target time series data is in a steady state, and a preliminary steady state detection result is obtained.
[0013] Based on the preliminary steady-state detection results, a method combining data cleaning and machine learning models is used to determine the final steady-state detection results.
[0014] In the second aspect, the present application provides a data cleaning and steady-state detection system based on robust time series modeling, comprising: a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein the processor executes the computer program to implement the data cleaning and steady-state detection method based on robust time series modeling.
[0015] According to the specific embodiments provided in this application, this application has the following technical effects.
[0016] The present application provides a data cleaning and steady-state detection method and system based on robust time series modeling, which obtains target time series data, constructs a linear trend model and performs robust optimization on it, thereby obtaining a robust model. By performing nonlinear optimization on the robust model, the robust parameters of the robust model can be determined, so that the target time series data can be judged to be in a steady state according to the robust parameters, and a preliminary steady-state detection result is obtained. Then, on the basis of the preliminary steady-state detection result, the final steady-state detection result is determined by combining data cleaning with a machine learning model. It can be seen that the present application combines linear trend modeling, robust optimization and nonlinear optimization solution based on target time series data, and on the basis of target time series data and its linear trend model, a robust model is constructed by robust optimization, thereby achieving the purpose of robust time series modeling, and further nonlinear optimization solves the robust parameters, thereby improving the robustness of the steady-state detection of the target time series data and improving the accuracy of the steady-state detection result. Based on the preliminary steady-state detection results, data cleaning technology and machine learning technology were introduced. Combining data cleaning technology and machine learning technology, data cleaning and steady-state detection based on machine learning models are performed simultaneously, thus ensuring the accuracy and reliability of the final steady-state detection results, further improving the accuracy of steady-state detection. In addition, this method is widely applicable to various target time series data, including steady-state detection of a variety of typical process variables such as temperature, flow, and liquid level. It can adapt to various complex industrial process data and can solve the problems of traditional steady-state detection methods that are unable to effectively cope with various complex industrial process data, and have poor robustness and low accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 This is an application environment diagram of a data cleaning and steady-state detection method based on robust time series modeling provided in one embodiment of the present application.
[0019] Figure 2 A flowchart of a data cleaning and steady-state detection method based on robust time series modeling is provided in one embodiment of the present application.
[0020] Figure 3 A schematic diagram of a data cleaning and steady-state detection method based on robust time series modeling provided in one embodiment of the present application.
[0021] Figure 4 A schematic diagram of the original temperature data of the delayed coking process logistics flow data of the factory provided in one embodiment of the present application.
[0022] Figure 5 This is a data diagram of the delayed coking process logistics flow data of the factory provided in one embodiment of the present application after adding abnormal values.
[0023] Figure 6 A schematic diagram of steady-state detection results based on raw temperature data of a robust model provided in one embodiment of the present application.
[0024] Figure 7 A schematic diagram of steady-state detection results based on raw temperature data using a filtering method provided in one embodiment of the present application.
[0025] Figure 8 A schematic diagram of the steady-state detection results of the KMeans clustering algorithm based on raw temperature data provided in one embodiment of the present application.
[0026] Figure 9 A schematic diagram of steady-state detection results of the robust model provided in one embodiment of the present application based on data with outliers added.
[0027] Figure 10 A schematic diagram of steady-state detection results of data after adding outliers based on the filtering method provided in one embodiment of the present application.
[0028] Figure 11 A schematic diagram of the steady-state detection results of the KMeans clustering algorithm provided in one embodiment of the present application based on data with outliers added.
[0029] Figure 12 A schematic diagram of the data cleaning results of the robust model provided in one embodiment of the present application.
[0030] Figure 13 A schematic diagram of the steady-state detection results of the KMeans clustering algorithm based on the fitted y value provided in one embodiment of the present application.
[0031] Figure 14 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0033] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0034] The data cleaning and steady-state detection method based on robust time series modeling provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the target time series data to the server 104. After the server 104 receives the target time series data, for the target time series data, the server 104 constructs a linear trend model based on the target time series data; performs robust optimization on the linear trend model to obtain a robust model; performs nonlinear optimization on the robust model to determine the robust parameters; based on the robust parameters, it is determined whether the target time series data is in a steady state to obtain a preliminary steady-state detection result; based on the preliminary steady-state detection result, a method combining data cleaning and machine learning model is used to determine the final steady-state detection result. The server 104 can feed back the final steady-state detection result to the terminal 102. In addition, in some embodiments, the data cleaning and steady-state detection method based on robust time series modeling can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly perform data cleaning and steady-state detection on the target time series data, or the server 104 can obtain the target time series data from the data storage system and perform data cleaning and steady-state detection on the target time series data.
[0035] The terminal 102 may be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, and IoT devices. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers, or a cloud server.
[0036] In an exemplary embodiment, Figure 2 As shown, a data cleaning and steady-state detection method based on robust time series modeling is provided. The method is executed by a computer device, specifically a computer device such as a terminal or a server, or a terminal and a server. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in FIG. 1 is used as an example to illustrate the method, which includes the following steps S1 to S6.
[0037] Step S1: Obtain target time series data.
[0038] In this embodiment, the target time series data may be time series data corresponding to various typical process variables such as temperature, flow rate, and liquid level.
[0039] Step S2: constructing a linear trend model based on the target time series data, wherein the linear trend model refers to a linear model including an intercept term, a slope term, and a random error term.
[0040] Step S3: Perform robust optimization on the linear trend model to obtain a robust model.
[0041] Step S4: Perform nonlinear optimization on the robust model to determine robust parameters, wherein the robust parameters include intercept term parameters, slope term parameters, and fitting y values.
[0042] Step S5: judging whether the target time series data is in a steady state according to the robust parameters, and obtaining a preliminary steady state detection result.
[0043] Step S6: Based on the preliminary steady-state detection results, a method combining data cleaning and machine learning model is used to determine the final steady-state detection results.
[0044] By implementing steps S1 to S6 above, the target time series data is acquired, a linear trend model is constructed, and robust optimization is performed on it, thereby obtaining a robust model. By performing nonlinear optimization on the robust model, the robust parameters of the robust model can be determined. Based on the robust parameters, whether the target time series data is in a steady state can be determined, thereby obtaining a preliminary steady-state detection result. Based on the preliminary steady-state detection result, the final steady-state detection result is determined by combining data cleaning with a machine learning model.
[0045] This embodiment combines linear trend modeling, robust optimization, and nonlinear optimization solutions based on target time series data. Based on the target time series data and its linear trend model, a robust model is constructed through robust optimization to achieve the purpose of robust time series modeling. Furthermore, robust parameters are solved through nonlinear optimization, thereby improving the robustness of steady-state detection of the target time series data and improving the accuracy of the steady-state detection results. Based on the preliminary steady-state detection results, data cleaning technology and machine learning technology are introduced. The data cleaning technology and machine learning technology are combined to simultaneously perform data cleaning and steady-state detection based on the machine learning model, thereby ensuring the accuracy and reliability of the final steady-state detection results and further improving the accuracy of steady-state detection. In addition, the method can be widely applied to various target time series data, including steady-state detection of a variety of typical process variables such as temperature, flow, and liquid level. It can adapt to various complex industrial process data and can solve the problems that traditional steady-state detection methods cannot effectively cope with various complex industrial process data and have poor robustness and low accuracy in steady-state detection.
[0046] In this embodiment, step S2 constructs a linear trend model based on the target time series data, which specifically includes the following steps.
[0047] Step S21: Use a sliding window method to segment the target time series data to obtain a number of time windows.
[0048] Step S22: construct a linear trend model based on each of the time windows.
[0049] In this embodiment, step S3 performs robust optimization on the linear trend model to obtain a robust model, which specifically includes the following steps.
[0050] Step S31: Based on ρ Function, establish a robust objective function. Among them, ρ The function is the logistic function, Cauchy function, or Welsch function.
[0051] Step S32: Perform robust optimization on the linear trend model according to the robust objective function to obtain a robust model.
[0052] In this embodiment, step S4 performs nonlinear optimization on the robust model to determine robust parameters, which specifically includes the following steps.
[0053] The robust model is solved by nonlinear optimization using a trust region method or an interior point optimization method (IPOPT) to obtain robust parameters.
[0054] In this embodiment, step S5 determines whether the target time series data is in a steady state according to the robust parameter to obtain a preliminary steady state detection result, which specifically includes the following steps.
[0055] Step S51: Calculate the t-test statistic value of the slope term parameter according to the robust parameter.
[0056] Step S52: judging whether the target time series data is in a steady state based on the t-test statistic value and the critical threshold of the slope term parameter, and obtaining a preliminary steady-state detection result.
[0057] In this embodiment, step S52 actually determines whether the target time series data is in a steady state based on the relationship between the absolute value of the t-test statistic of the slope term parameter and the critical threshold, which specifically includes the following two situations.
[0058] (1) When < , it is determined that the target time series data is in a steady state.
[0059] (2) When ≥ , it is determined that the target time series data is in a non-stationary state.
[0060] in, is the t-test statistic value of the slope term parameter, Indicates taking the absolute value, is the critical threshold.
[0061] In this embodiment, step S6 determines the final steady-state detection result based on the preliminary steady-state detection result by combining data cleaning with a machine learning model, which specifically includes the following steps.
[0062] Step S61: Based on the preliminary steady-state detection results, a sliding window method is used to traverse all data points in the target time series data. After each window slide, the preliminary steady-state detection results corresponding to a preset number of data points in the current window are retained, and the final output cleaning data is updated step by step to obtain the cleaned data.
[0063] Step S62: Determine the cleaned fitting y value based on the cleaned data.
[0064] Step S63: Using the cleaned fitting y value as the input of the KMeans clustering algorithm, performing steady-state detection on the cleaned data to obtain the final steady-state detection result.
[0065] Among them, the final steady-state detection result is relative to the preliminary steady-state detection result. The preliminary steady-state detection result is the steady-state detection result determined after constructing a robust model and solving the robust parameters through nonlinear optimization. The final steady-state detection result is obtained by introducing a method combining data cleaning and machine learning model on the basis of the preliminary steady-state detection result. While cleaning the data, the machine learning model (i.e., KMeans clustering algorithm) is simultaneously used to perform steady-state detection on the cleaned data, so that the final steady-state detection result is more accurate and reliable.
[0066] In order to make the technical solution of this embodiment clearer, the specific implementation process of the technical solution of this embodiment is described in detail below in the form of examples.
[0067] like Figure 3 As shown, the present embodiment proposes a data cleaning and steady-state detection method based on robust time series modeling. The basic principles of data cleaning and steady-state detection are as follows: first, K groups of time series data are input; then, based on the input K groups of time series data, the standard deviation of the residual is calculated; then, the sliding window T and the step size S are determined, thereby determining the window data (0+N×S, T+N×S), where N is the window size; then, nonlinear optimization is performed to solve the objective function and calculate the linear trend model parameters. and ; Then perform t-test to determine the steady state of the data; determine whether T+N×S is less than KT / 2. If so, output the steady state test result. Otherwise, the window size is N=N+1, and return to the step of "determining window data (0+N×S, T+N×S)". and At the same time, the calculated fitting y value is used as the input of the KMeans clustering algorithm to perform steady-state detection and output the final steady-state detection result. The method specifically includes the following implementation steps.
[0068] Step (1): Data acquisition.
[0069] This embodiment acquires target time series data, which may be various complex industrial process data, such as target temperature time series data, target flow rate time series data, target liquid level time series data, etc.
[0070] Step (2): Sliding window division.
[0071] This embodiment uses a sliding window method to segment the target time series data in step (1), with a window size of N and a step size of 1.
[0072] In this example, a sliding window approach is used to extract local features from the target time series data. The window size N is selected to capture local pattern features, balancing local pattern capture with computational efficiency. An overlapping window design with a step size of 1 ensures continuous data partitioning and maximizes information utilization, effectively avoiding the potential pattern omissions caused by non-overlapping windows. The data within the window must meet the minimum sample size requirement for trend analysis; it is generally recommended to cover 2-3 characteristic periods to capture the complete fluctuation pattern, thereby ensuring the accuracy of steady-state detection.
[0073] Step (3): Establish a linear trend model.
[0074] In this embodiment, a linear trend model is constructed in each time window, which includes the intercept term , slope term and random error terms .
[0075] In step (3), a linear trend model is established in each time window. The linear trend model consists of a deterministic part and a random part, where the deterministic part includes the intercept term and the slope term , the intercept term Used to represent the mean level, slope term Used to describe long-term trend. The random part includes random error terms , random error term The white noise process reflects unpredictable short-term fluctuations. This semi-parametric modeling approach not only retains the main characteristics of the linear trend, but also absorbs high-frequency interference through the white noise term, thereby enhancing the adaptability of the linear trend model.
[0076] In this embodiment, the linear trend model is expressed as follows.
[0077] (1).
[0078] in, is the data fitting value of the linear trend model, represents a zero-mean white noise process with constant variance . is the signal intercept or baseline level, i.e., the intercept term; represents the linear trend captured, represents the slope, i.e. the slope term; Relative time index within the window, ranging from 0 to N-1.
[0079] Step (4): Perform robust optimization on the linear trend model in step (3) to obtain a robust model.
[0080] This embodiment introduces ρ The robust objective function is constructed by the function, and the Logistic function is used to replace the traditional least squares loss to reduce the impact of outliers on model parameter estimation and enhance the robustness of the linear trend model. ρ Functions include Logistic function, Cauchy function, or Welsch function.
[0081] In step (4), in order to enhance the robustness of the linear trend model to outliers, the Logistic function is introduced to replace the traditional least squares loss, which is in the following form.
[0082] (2).
[0083] in, is the value of the Logistic function, is the robustness tuning parameter, is the residual, is the residual standard deviation. The Logistic function imposes a logarithmic penalty on large residuals, effectively suppressing the influence of outliers. It can balance robustness and estimation efficiency. It is usually taken as 0.602, which corresponds to the equivalent effect of Gaussian kernel.
[0084] At the same time, the linear trend model based on the target time series data is combined with ρ Function combination, ρ The function is a commonly used tool in robust statistics and optimization problems to construct robust objective functions or loss functions. ρ The function is a monotonically increasing function with the following core characteristics: it is sensitive to small deviations (residuals) and has the ability to suppress large deviations (such as outliers), thereby improving the robustness of the linear trend model to noise or abnormal data. ρ The function is as follows.
[0085] 1) Logistic function, its expression is as follows.
[0086] (3).
[0087] 2) Cauchy function, whose expression is as follows.
[0088] (4).
[0089] 3) Welsch function, whose expression is as follows.
[0090] (5).
[0091] in, 、 、 Both ρ The specific values of the function parameters need to be set and adjusted according to the actual situation.
[0092] In this embodiment, for non-dimensional ρ Function calculation results and other data are z-score standardized. This method is suitable for scenarios where errors or residuals are continuous variables and have an approximately normal distribution. The standardized data are dimensionless values.
[0093] Step (5): Perform nonlinear optimization solution based on robust optimization, that is, perform nonlinear optimization solution on the robust model and determine the robust parameters.
[0094] This embodiment uses nonlinear optimization algorithms such as interior point optimization method to minimize the objective function and solve (i.e., the intercept parameter value), (slope term parameter value) and the fitted y value. The nonlinear optimization solution algorithm can use the trust region method or interior point optimization method for robust parameter estimation to obtain robust parameters.
[0095] Step (5) performs nonlinear optimization on the robust model to determine the robust parameters. The specific steps are as follows.
[0096] Step (51): Parameter estimation is performed using a nonlinear optimization algorithm. The core of this algorithm is to find the optimal parameter combination by minimizing the sum of squares of the prediction errors. For the linear trend model in formula (1), this method can be solved to obtain and The estimated value of has the advantages of high computational efficiency and good unbiasedness. This method performs well in trend-dominated data and can effectively separate trend components from random disturbances.
[0097] Step (52): Optimize the objective function and update the trend value. Use the nonlinear optimization algorithm to minimize the Logistic function and iteratively solve the optimal trend value, that is, the optimal data fitting value of the linear trend model. After each iteration, the model parameters are updated and the residuals are recalculated. , until the parameter and Converge or reach the preset accuracy. This process achieves accurate extraction of trend components while ensuring strong anti-interference ability against outliers.
[0098] Step (53): Using a nonlinear optimization method, the linear trend model in formula (1) and the logistic function in formula (3) are combined to obtain a robust model. The objective function and constraints of the robust model are as follows.
[0099] In this embodiment, the objective function adopts the form of a Logistic function and is defined as follows.
[0100] (6).
[0101] in, For robustness adjustment parameters, it determines the strength of suppression of outliers; is the residual, which represents the deviation between the actual value and the trend value.
[0102] In this embodiment, the following constraints are added to the model optimization problem.
[0103] 1) Linear trend constraint, the expression is as follows.
[0104] (7).
[0105] This constraint represents the trend portion of the time series It's time A linear function of is the signal intercept or baseline level. is the signal intercept, is the slope.
[0106] 2) The residual definition constraint is expressed as follows.
[0107] (8).
[0108] Among them, the residual is the actual value and trend value The difference between them is used to measure the degree of deviation of the data.
[0109] Step (6): Calculate the statistical significance of the slope. Calculate the slope parameter The t-test statistic , determine whether it is in steady state based on the significance level α.
[0110] In step (6), when calculating the statistical significance of the slope, the slope parameter is evaluated by t-test The statistical significance of is, and its test statistic is ,in is the standard deviation of the slope estimate, is the slope.
[0111] Step (7): Determine whether the target time series data is in a steady state. Less than critical threshold When , the target time series data is determined to be in a steady state. Greater than or equal to the critical threshold When , the target time series data is determined to be in a non-stationary state.
[0112] In step (7), when judging whether the target time series data is in a steady state, the comprehensive slope significance test results are used. < When , it indicates that the trend component is not significant, and the window data (i.e. the target time series data in the current window) is in a steady state. ≥ When , it is considered that there is a significant linear trend, that is, the window data is judged to be in a non-stationary state. This t-test controls the probability of type I error, where , is the critical threshold, which represents the preset threshold for judging whether the target time series data is in a steady state, and is determined by the degrees of freedom Obtained by looking up the table, is the t-test statistic value of the slope term parameter, represents the absolute value, α represents the significance level, and in this embodiment, the significance level α is 0.05 or 0.01.
[0113] Step (8): Combine data cleaning with machine learning models to determine the final steady-state detection results.
[0114] In this embodiment, the data cleaning method adopts the sliding window method. After each window sliding, only the preliminary steady-state detection results corresponding to the first 10 data points in the current window are retained, and the cleaned data are finally outputted by step-by-step iterative updating. Then, the fitted y value in the robust parameter of the robust model in step (5) is used as the data cleaning result (i.e., the fitted y value in the cleaned data) as the input of the KMeans clustering algorithm, and the steady-state detection is performed on the cleaned data, and the final steady-state detection result is outputted.
[0115] In step (8), the KMeans clustering algorithm is an unsupervised learning algorithm that divides data points into K different categories by minimizing the squared distance from the data point to the center of the cluster to which it belongs. In time series steady-state detection, the KMeans clustering algorithm identifies patterns and changes in data by iteratively updating the cluster centers. This method does not require a mechanistic model and can adapt to complex process data. It is widely used in chemical process monitoring, equipment health management, and quality optimization. However, the KMeans clustering algorithm requires reasonable parameter settings, such as the number of clusters K and the steady-state judgment threshold, and is often combined with statistical analysis, physical modeling, and nonlinear optimization methods to improve the accuracy and robustness of steady-state detection.
[0116] The fitted y-values in the cleaned data of the robust model are used as the input of the KMeans clustering algorithm to obtain the final steady-state detection results. It is found that the final steady-state detection results are basically consistent with the original data detection results (i.e., the original data without cleaning). This effectively proves that the cleaning results of the robust model (i.e., the cleaned data) effectively retain the main distribution characteristics of the original data. It also shows that the data after robust processing has good usability and can well support subsequent modeling and steady-state detection processing tasks, providing a reliable data foundation for practical applications.
[0117] Step (9): Data analysis and storage. Analyze and discuss the data cleaning results and steady-state detection results, and store the relevant data.
[0118] This embodiment uses a high-precision temperature sensor to monitor the temperature changes of the reaction system in real time for the delayed coking process, and combines key process variables such as pressure and flow to perform multi-parameter collaborative analysis. Due to factors such as the periodic switching of the coke tower, the load adjustment of the heating furnace, and the fluctuation of the properties of the raw materials, the system will inevitably experience non-steady-state operating conditions (such as instantaneous fluctuations in temperature and pressure or flow redistribution). In order to fully capture such dynamic characteristics, this embodiment uses a high-frequency data acquisition system to continuously record temperature, pressure and logistics flow data at a sampling frequency of once per minute, and continuously accumulates 20 days of operating data to form a high-resolution time series database. This embodiment will conduct case analysis for nearly 1 day, and the data is as follows: Figure 4 shown.
[0119] In order to achieve robust modeling and steady-state detection of dynamic process data, this embodiment adopts the sliding window method to perform segmented modeling and optimization analysis on the target time series data. Specifically, a data window with a length of 20 is used to sequentially intercept data subsets from the target time series data. The target time series data in the window is used as the input of the robust model for parameter estimation and state identification. The objective function uses the Logistic function with robust characteristics. The robust model is under the constraint condition Fitting is performed to solve the parameters and , to construct a fitted value sequence , and evaluate the deviation from the original data.
[0120] Then, calculate the slope estimate Standard deviation , used to assess the significance of linear trends: ,in, is the average value of the time index within the window. Based on the previous results, construct the t-test statistic To test the statistical significance of the slope: , under the null hypothesis, the signal is stationary (i.e. ), the t-test statistic follows a t-distribution with n-2 degrees of freedom. The critical threshold is denoted as , by the significance level Decision. If ≥ , then the null hypothesis is rejected and the window data is classified as non-stationary. Otherwise, the window data is considered to be stationary.
[0121] The entire target time series data is traversed through the sliding window method, and each time the window moves forward one step (step size is 1), the above optimization process is repeated, and the segmented modeling of the entire data set of the target time series data is gradually completed. In terms of result output, the detection results within the window are strategically retained: for the first calculation window, the steady-state detection results of the first 10 time points are retained; for the last window, the steady-state detection results of the last 10 data points are retained; and for the middle window, only the steady-state detection results and fitting results of the 10th data point in the middle position of each window are retained, and the steady-state detection results of the robust model and the fitting y value of the data cleaning are saved, such as Figure 6 、 Figure 7 and Figure 8 as well as Figure 12 This strategy can effectively avoid redundancy caused by data overlap while taking into account the continuity of boundary data, ensuring that the final steady-state detection results have good temporal consistency and representativeness.
[0122] In order to verify the effect of the robust model, this embodiment uses two models, the filtering method and the KMeans clustering algorithm, to compare the steady-state detection results with the robust model.
[0123] Filtering is a traditional steady-state detection method. It calculates the smoothed value and rate of change of data within a sliding window, identifying steady-state and transient states based on statistical characteristics. Its computational simplicity and strong real-time performance make it suitable for scenarios such as online data cleaning, process monitoring, and anomaly detection. For example, low-pass filtering can be used to detect step changes and long-term drift in process data, facilitating process modeling and identifying abnormal operations.
[0124] The KMeans clustering algorithm is an unsupervised learning algorithm that classifies data points into K distinct categories by minimizing the squared distance from each data point to the center of its cluster. In time series steady-state detection, KMeans clustering identifies patterns and changes in the data by iteratively updating cluster centers. This method, which does not require a mechanistic model and can adapt to complex process data, is widely used in fields such as chemical process monitoring, equipment health management, and quality optimization. However, KMeans clustering requires the proper setting of parameters, such as the number of clusters K and the steady-state determination threshold, and is often combined with statistical analysis, physical modeling, and nonlinear optimization methods to improve the accuracy and robustness of steady-state detection.
[0125] In this embodiment, the delayed coking process flow data is introduced into the filtering method and KMeans clustering algorithm models to obtain steady-state detection results, such as Figure 6 、 Figure 7 and Figure 8 It can be clearly observed that the steady-state detection results of the three models are basically consistent, which shows that the steady-state detection of the robust model under normal data is accurate.
[0126] In order to verify the robustness of the robust model under complex data with multiple outliers, this embodiment adds outliers to the data. The principle of adding outliers is to randomly select a certain proportion of data points in the original data and add outliers to these data points. The added outliers are expressed as: .in, is a uniformly distributed random number in the interval [0, 1]; To control the outlier amplitude parameter. In this experiment, the number of outliers is set to 10% of the original data, and the outlier amplitude parameter is set to 5, and the data is as follows Figure 5 shown.
[0127] In this embodiment, the processed delayed coking process flow data is substituted into the filtering method, KMeans clustering algorithm and robust model, and the above steps are repeated to obtain the steady-state detection results of the three models, as shown in FIG. Figure 9 、 Figure 10 and Figure 11 As shown in the results, it can be clearly observed that when faced with data with outliers, the filtering method detects most of the non-steady-state time periods as steady-state time periods, while the steady-state detection result of the KMeans clustering algorithm detects most of the data as non-steady-state time periods, which obviously has lost the function of steady-state detection. However, the robust model can distinguish the steady-state and non-steady-state time periods very well, and the steady-state detection results are consistent with the Figure 6 The steady-state detection results of the original data are basically consistent. The accuracy of the steady-state detection results of the three models with and without outliers are compared, and the accuracy comparison results are shown in Table 1.
[0128] Table 1 Accuracy comparison results
[0129]
[0130] Judging from the overall accuracy in Table 1, the robust model performed the best, reaching 98.53%, significantly outperforming the filtering method's 73.25% and the KMeans clustering algorithm's 44.64%. The robust model's overall accuracy improved by 25.28% over the filtering method and 53.89% over the KMeans clustering algorithm, demonstrating its overwhelming superiority in global recognition capabilities. Regarding the accuracy during the steady-state period, while the filtering method achieved a slightly higher accuracy of 98.67%, the robust model also achieved 97.56%, a difference of only 1.11%, essentially matching the robust model. The KMeans clustering algorithm, on the other hand, performed extremely poorly, achieving only 3.93%, 93.63% lower than the robust model. This demonstrates that both the robust model and the filtering method possess strong capabilities in recognizing steady-state data segments, while the KMeans clustering algorithm completely fails. The robust model maintained a high accuracy of 98.77% for non-stationary periods, a 79.18% improvement over the filtering method's 19.59%, but slightly lower than the KMeans clustering algorithm's 99.89% by 1.12%. This indicates that the KMeans clustering algorithm tends to identify data as non-stationary, resulting in a higher accuracy rate for this type of data. However, this also further demonstrates its extremely weak ability to detect steady-state periods, severely impacting overall accuracy.
[0131] In summary, the robust model significantly outperforms other methods in terms of overall accuracy and non-steady-state identification accuracy, while also matching the best method in steady-state identification, demonstrating strong overall performance. Compared to filtering and the KMeans clustering algorithm, the robust model offers greater stability and accuracy when handling different state segments, making it a more reliable tool for data cleaning and steady-state detection.
[0132] observe Figure 12 In order to verify the data cleaning effect of the robust model, the present embodiment performs statistical analysis on the fitted y values. The statistical analysis results are shown in Table 2.
[0133] Table 2 Statistical analysis results of fitted y values
[0134]
[0135] Table 2 clearly demonstrates that, by comparing the statistical changes between the original and cleaned data, the robust model's data cleaning method performs well in improving data quality. First, the mean of the cleaned data decreases slightly from 28.1345 to 28.1307, a negligible change of only 0.0038. This indicates that the robust model handles outliers without disrupting the overall trend of the data. The standard deviation increases slightly, from 5.6037 to 5.6405, an improvement of 0.0368. This suggests that a certain degree of volatility may have been retained during the data cleaning process to avoid oversmoothing, which helps preserve true variation. The skewness decreases from 0.8793 to 0.8771, slightly reducing the right skewness of the original data. The kurtosis decreases from -0.0977 to -0.0996, showing little change, indicating that the kurtosis of the data distribution remains largely unchanged. The cleaned data showed little change in mean, increased standard deviation, regression of skewness, and stabilization of kurtosis. This demonstrates that the method in this example not only effectively preserves the original structure and trend of the data, but also significantly improves its stability and modelability. This result validates the effectiveness of robust models and data cleaning methods in reducing abnormal interference and optimizing data quality.
[0136] Finally, the data cleaning results of the robust model are fitted with y values and used as the input of the KMeans clustering algorithm to obtain the steady-state detection results as follows: Figure 13 As shown in the figure, it is found that the steady-state detection results based on the cleaned data are basically consistent with the steady-state detection results of the original data. This effectively proves that the data cleaning results of the robust model effectively retain the main distribution characteristics of the original data. It also shows that the data after robust optimization processing has good usability and can well support subsequent modeling and steady-state detection and other processing tasks, providing a reliable data foundation for practical applications.
[0137] This embodiment proposes a steady-state detection and data cleaning method based on robust time series modeling, which combines sliding window technology, ρThis paper proposes a robust function and nonlinear optimization strategy to improve the robustness of steady-state detection in time series data, particularly for industrial process data with noise and outliers. First, a sliding window method is used to extract local features from the target time series data and construct a linear trend model consisting of an intercept term, a slope term, and a random noise term. Robust regression is then performed using the logistic function to reduce the impact of outliers on trend estimation. A nonlinear optimization solution is then used to determine the robustness parameter, and the statistical significance of the slope parameter is calculated. Finally, a t-test statistic is used to determine whether the data is in steady-state. This method can effectively improve the accuracy and robustness of steady-state detection and can be used for data cleaning to provide high-quality data. Furthermore, the data values (i.e., fitted y values) fitted by the robust model can be used as input to the KMeans clustering algorithm to implement machine learning hybrid modeling. Experimental results demonstrate that this method has strong anti-interference capabilities for complex dynamic process data. It is suitable for the chemical industry and embodies the deep integration of artificial intelligence and process industry.
[0138] This embodiment proposes a steady-state detection and data cleaning method based on robust time series modeling. This method combines the sliding window technology, the monotonically increasing function ( ρ Functions) and nonlinear optimization strategies are used to improve the robustness of steady-state detection. The data cleaning results can be used as input to machine learning models to determine the steady-state detection results of the target time series data. The combination of the two represents a significant step forward in the chemical industry's move toward intelligent, data-driven development. This eliminates the reliance on single indicators or manually set rules for steady-state detection, instead leveraging the inherent structure of historical data to achieve more intelligent and generalizable operating condition identification. This not only improves the accuracy and robustness of steady-state detection, but also provides a critical data foundation and decision-making support for real-time monitoring, intelligent operation and maintenance, and digital twin systems of chemical processes, demonstrating the application prospects of the deep integration of artificial intelligence and process industries.
[0139] This embodiment introduces a robust model, a loss function, and a sliding window mechanism, which can accurately identify data trends and perform real-time steady-state judgments in the presence of a large number of outliers. Compared with the traditional "clean first, then detect" separation strategy, the method of this embodiment realizes the integrated processing of data cleaning and steady-state detection under the same modeling framework, significantly reduces error transmission and steady-state identification delay, and improves the accuracy and real-time performance of steady-state detection. In addition, the method has good interpretability and adaptability, and can be widely applied to a variety of typical process variables including temperature, flow, liquid level, etc. Experimental results show that the method can still stably output reliable results under high noise background, providing high-quality data support for data-driven modeling, optimization control, and fault warning, thereby enhancing the operational stability and economic benefits in complex industrial processes, and has good engineering application prospects.
[0140] In an exemplary embodiment, a data cleaning and steady-state detection system based on robust time series modeling is provided. The system can be a computer device, which can be a server or a terminal. The internal structure diagram can be as follows: Figure 14 As shown. The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, memory and input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store target time series data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a data cleaning and steady-state detection method based on robust time series modeling is implemented.
[0141] Those skilled in the art will understand that Figure 14 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0142] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0143] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0144] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A data cleaning and steady-state detection method based on robust time series modeling, characterized in that: The data cleaning and steady-state detection method based on robust time series modeling is applicable to steady-state detection of various target time series data, including temperature, flow rate and liquid level; The data cleaning and steady-state detection method based on robust time series modeling includes: Get target time series data; Constructing a linear trend model based on the target time series data; the linear trend model refers to a linear model including an intercept term, a slope term, and a random error term; Performing robust optimization on the linear trend model to obtain a robust model; Performing nonlinear optimization on the robust model to determine robust parameters; the robust parameters include intercept term parameters, slope term parameters and fitting y values; Judging whether the target time series data is in a steady state according to the robust parameters, and obtaining a preliminary steady-state detection result; Based on the preliminary steady-state detection results, a method combining data cleaning and machine learning model is used to determine the final steady-state detection results; According to the target time series data, a linear trend model is constructed, specifically including: Using a sliding window method, the target time series data is segmented to obtain a number of time windows; Based on each of the time windows, respectively construct a linear trend model; The expression of the linear trend model is: ; in, is the data fitting value of the linear trend model, represents a zero-mean white noise process, represents the signal intercept or baseline level, represents the linear trend captured, is the slope, Represents the relative time index within the window; Based on the robust parameters, it is determined whether the target time series data is in a steady state, and a preliminary steady state detection result is obtained, which specifically includes: Calculating a t-test statistic value of the slope term parameter based on the robust parameter; Determine whether the target time series data is in a steady state based on the t-test statistic value of the slope term parameter and a critical threshold value, and obtain a preliminary steady-state detection result; the critical threshold value is a preset threshold value for determining whether the target time series data is in a steady state; Based on the t-test statistic value and critical threshold of the slope term parameter, it is determined whether the target time series data is in a steady state, and a preliminary steady-state detection result is obtained, which specifically includes: when < When , it is determined that the target time series data is in a steady state; when ≥ When , it is determined that the target time series data is in a non-stationary state; in, is the t-test statistic value of the slope term parameter, Indicates taking the absolute value, is the critical threshold.
2. The data cleaning and steady-state detection method based on robust time series modeling according to claim 1 is characterized in that: Performing robust optimization on the linear trend model to obtain a robust model specifically includes: based on ρ Function, establish robust objective function; According to the robust objective function, the linear trend model is robustly optimized to obtain a robust model.
3. The data cleaning and steady-state detection method based on robust time series modeling according to claim 2 is characterized in that: described ρ The function is the logistic function, Cauchy function, or Welsch function.
4. The data cleaning and steady-state detection method based on robust time series modeling according to claim 1, characterized in that: Performing nonlinear optimization on the robust model to determine robust parameters, specifically including: The robust model is solved by nonlinear optimization using a trust region method or an interior point optimization method to obtain robust parameters.
5. The data cleaning and steady-state detection method based on robust time series modeling according to claim 1 is characterized in that: Based on the preliminary steady-state detection results, a method combining data cleaning and machine learning models is used to determine the final steady-state detection results, specifically including: Based on the preliminary steady-state detection results, a sliding window method is used to traverse all data points in the target time series data, retaining the preliminary steady-state detection results corresponding to a preset number of data points in the current window after each window sliding, and gradually iteratively updating the final output cleaned data to obtain cleaned data; Determining the cleaned fitting y value according to the cleaned data; The cleaned fitting y value is used as the input of the KMeans clustering algorithm, and a steady-state test is performed on the cleaned data to obtain the final steady-state test result.
6. A data cleaning and steady-state detection system based on robust time series modeling, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the data cleaning and steady-state detection method based on robust time series modeling according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method for steady-state detection and related equipment
CN108491357A
Predictive monitoring of the glucose-insulin endocrine metabolic regulatory system
US20220039758A1