Data generation model-based data governance method for continuous casting production process

By combining the conditional diffusion model and time series decomposition method, the problem of data missing during continuous casting of steel is solved, and high-accurate data interpolation is achieved, errors are reduced, and the fitting ability of the data model is improved.

CN120296312APending Publication Date: 2025-07-11AUTOMATION RES & DESIGN INST OF METALLURGICAL IND
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510337714.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

During the continuous casting of steel, the monitoring data has problems such as data interruption, data jump or zeroing, resulting in data loss and affecting the fitting ability of the data model. The existing technology interpolation method has a large error or a offset when complex waveforms are encountered.

Method used

Combining the conditional diffusion model and the time series decomposition method, the data is cleaned through the shutdown data identification algorithm, and the decomposed data is into trend components and residual components. The missing data is processed separately using the three-time Hermit interpolation and conditional diffusion models to ensure the accuracy of interpolation.

Benefits of technology

It effectively reduces the interpolation error under high variance data, improves the interpolation accuracy of missing data, and ensures the completeness and accuracy of the data model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296312A_ABST
    Figure CN120296312A_ABST
Patent Text Reader

Abstract

The invention discloses a data generation model-based data governance method for a continuous casting production process, which belongs to the field of data governance in a metallurgical production process, and specifically comprises the following steps of: firstly, aligning and combining a plurality of acquired single-variable steel continuous casting data into a multivariable data sequence by taking time as a standard; and a shutdown data identification algorithm is used for cleaning. Then, decomposing the cleaned data set into a trend component and a residual component by using a sliding window averaging method, standardizing the residual component, and inputting a conditional diffusion model for training to obtain a residual component interpolation model; secondly, decomposing missing data needing to be interpolated into a trend component Mt and a residual component Mr; performing prediction by adopting cubic Hermite interpolation to obtain a missing value of the trend component Mt; and meanwhile, predicting a missing value of the residual component Mr by using the residual component interpolation model. And finally, summing the missing values of the two parts to obtain complete prediction data of the missing part. According to the invention, the accuracy of missing data interpolation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data governance in the metallurgical production process, and specifically relates to a data governance method based on a data generation model for the continuous casting production process. Background Art

[0002] The metallurgical industry is facing challenges in technological innovation and digital transformation; in order to fully tap the potential value of data and meet intelligent requirements such as intelligent diagnosis of equipment failures, optimization of process parameters, and optimization of production processes, the metallurgical industry needs high-quality data sets to support data analysis and intelligent decision-making.

[0003] In the steel continuous casting process, affected by instrument failures, network failures, or host computer failures, there are problems such as data interruption, data jump, or normal / abnormal zeroing in the monitoring data, and there will also be cases of missing data in the collected process data. These missing data will lead to sample bias and affect the fitting ability of the data model to the overall data.

[0004] Traditional data imputation methods, such as mean or forward imputation, obtain information from a single variable dimension for imputation, and have good results in solving the missing values of smooth waveform data, but have large imputation errors when facing steel data variables with complex waveforms.

[0005] Using a conditional diffusion model for missing value imputation has advantages in predicting complex waveform changes, but will produce offset errors when facing high-variance data. Summary of the Invention

[0006] In order to solve the problem of missing data in steel continuous casting, the present invention provides a data governance method based on a data generation model for the continuous casting production process, which combines a conditional diffusion model with a time series decomposition method, reduces the offset error generated by predicting high-variance data using the conditional diffusion model, and improves the accuracy of missing data imputation.

[0007] The data governance method based on a data generation model for the continuous casting production process includes the following steps:

[0008] Step 1: Obtain the time series data of steel continuous casting, align multiple single-variable data according to time, and merge them into a multi-variable data sequence.

[0009] Step 2: Use a downtime data identification algorithm to delete the data corresponding to the downtime on the multi-variable data sequence for cleaning.

[0010] The specific cleaning process is as follows:

[0011] First, use the sliding window average algorithm to calculate the average value of the data within the window to smooth the data, thereby extracting the trend component of the "middle package temperature" variable, and then calculate the average value and standard deviation of the trend component;

[0012] Then, traverse the trend component in chronological order. If a data point with a deviation from the average value by 6 times the standard deviation appears, mark all the time axis positions with the same monotonic change trend connected to this point.

[0013] The specific formula is:

[0014] x is the value of the trend component being traversed, is the average value, and σ is the standard deviation.

[0015] Finally, after the traversal, delete the variables at the marked time axis positions in the multivariate sequence of the original data.

[0016] Step 3: Use the sliding window average method to decompose the cleaned data set into a trend component and a residual component.

[0017] Decomposition process:

[0018] D t = W(D)

[0019] D - D t = D r

[0020] Let D represent the cleaned data set, W represent the sliding window average function. After passing through the sliding window average algorithm, D becomes the trend component D t , and the residual component is D r refers to the difference between the original data set and the trend component;

[0021] Step 4: Standardize the residual component, input it into the conditional diffusion model for training, and optimize the model parameters through backpropagation to obtain the residual component interpolation model.

[0022] Residual component standardization is to subtract the average value of the original data from the residual component and then divide by the standard deviation to obtain data with a mean of 0 and a standard deviation of 1; the specific formula is:

[0023]

[0024] D s is the residual component, D rn is the standardized residual component, is the average value of the residual component, is the standard deviation of the residual component;

[0025] Step 5: Decompose the missing data M to be interpolated collected in the continuous casting process of steel into a new trend component M t and a new residual component M r ;

[0026] The decomposition formula is:

[0027] M t = W(M)

[0028] M - M t = M r

[0029] W represents the moving window average function. The missing data M passes through the moving window average algorithm to obtain the trend component M t ;

[0030] Step 6: Perform cubic Hermite interpolation on the new trend component M t to obtain the missing values of the trend component; meanwhile, use the trained residual component interpolation model to predict the missing values of the new residual component M r .

[0031] The cubic Hermite interpolation process is as follows: Input the value and corresponding derivative at the starting point of the missing part, the value and corresponding derivative at the ending point of the missing part, and interpolate the cubic curve function of the missing segment according to the above four input information to obtain the predicted values of the intermediate missing points;

[0032] The function formula is

[0033]

[0034] where f(a) represents the value at the starting point a of the missing part, f′(a) represents the derivative value at the starting point a of the missing part; f(b) represents the value at the ending point b of the missing part, and f′(b) represents the derivative value at the ending point b of the missing part;

[0035] The input of the residual component interpolation model includes the missing residual component M r and the conditional information. The conditional information refers to a matrix composed of 0 and 1 that marks the position of the missing part, where 0 represents missing and 1 represents not missing; from the output, select and retain the missing values of the predicted residual component M r .

[0036] Step 7: Sum the predicted missing values of the trend component and the missing values of the residual component to obtain the complete predicted data for the missing part.

[0037] The advantages of the present invention are:

[0038] 1. A data governance method based on a data generation model for the continuous casting production process. Compared with the existing technology, in the steel time series data, the data in the downtime part has a significant difference in value from the data in the normal part, resulting in an increased error in the interpolation result of the normal data. The present invention proposes a downtime data identification algorithm, which can retain the fluctuating data in other normal ranges when completely removing the downtime data, ensuring the accuracy of the data.

[0039] 2. A data governance method based on a data generation model for the continuous casting production process. Compared with the existing technology, the steel data has a large change range and a large variance, resulting in the problem of interpolation value deviation in the conditional diffusion model after data standardization and scaling. The present invention proposes an interpolation method combining the time series decomposition algorithm and the conditional diffusion model, decomposes the original data into a trend component and a residual component, calculates the missing trend component using the cubic Hermite interpolation algorithm, predicts the missing residual component using the conditional diffusion model, and finally adds the two components to obtain a complete prediction component, reducing the offset error when using the diffusion model alone. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a flowchart of a data governance method based on a data generation model for the continuous casting production process according to the present invention;

[0041] Figure 2 It is a flowchart of the downtime data identification method described in the present invention;

[0042] Figure 3 It is a structural diagram of the conditional diffusion model described in an embodiment of the present invention;

[0043] Figure 4 It is a network structure diagram of the residual module of the conditional diffusion model described in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0044] The following further details the specific implementation method of the present invention with reference to the accompanying drawings.

[0045] In the existing technology, the form of using the conditional diffusion model for missing interpolation is to directly input the data with missing values and predict the missing part. The present invention proposes a data governance method based on a data generation model for the continuous casting production process, decomposes the missing part into two components (trend + residual), then uses the conditional diffusion model to predict the residual among them, calculates the trend component using the traditional cubic interpolation, and finally sums the two components. In this way, since the variance of the residual component is smaller than the original data, the influence of high variance on the model can be avoided; and the trend component itself is a smooth curve, which is also suitable for calculation using the traditional cubic interpolation algorithm.

[0046] Such as Figure 1As shown, the data governance method for the continuous casting production process based on data generation models includes the following steps:

[0047] Step 1: Obtain the time-series data of steel continuous casting, align multiple single-variable data based on time, and merge them into a multi-variable data sequence.

[0048] Step 2: Use the downtime data identification algorithm to clean the multi-variable data sequence by deleting the data corresponding to the downtime.

[0049] As Figure 2 shown, the specific cleaning process is as follows:

[0050] First, use the moving window average algorithm to calculate the average value of the data within the window to smooth the data, thereby extracting the trend component of the "tundish temperature" variable, and then calculate the average value and standard deviation of the trend component;

[0051] Tundish temperature: The tundish temperature refers to the temperature of the molten steel in the tundish during the continuous casting production process. It is generally continuously measured by a thermocouple or a radiation temperature measurement system to obtain a time-series data sequence.

[0052] Then, traverse the trend component in chronological order. If a data point that deviates from the average value by 6 times the standard deviation appears, mark all the time-axis positions with the same monotonic change trend connected to this point.

[0053] The specific formula is:

[0054] x is the value of the trend component being traversed, is the average value, and σ is the standard deviation.

[0055] Finally, after the traversal, delete the variables at the marked time-axis positions in the multi-variable sequence of the original data.

[0056] Step 3: Use the moving window average method to decompose the cleaned data set into a trend component and a residual component.

[0057] Decomposition process:

[0058]

[0059] Let D represent the cleaned data set, W represent the moving window average function. After passing through the moving window average algorithm, D becomes the trend component D t , and the residual component is D r which refers to the difference between the original data set and the trend component;

[0060] Step 4: Standardize the residual component, input it into the conditional diffusion model for training, and optimize the model parameters through backpropagation to obtain the residual component interpolation model.

[0061] Residual component standardization is to subtract the mean of the original data from the residual component and then divide by the standard deviation to obtain data with a mean of 0 and a standard deviation of 1. The specific formula is:

[0062]

[0063] D s is the residual component, D rn is the standardized residual component. is the mean of the residual component, is the standard deviation of the residual component;

[0064] The training process of the model is specifically as follows: When the complete data is input into the conditional diffusion model, a random missing matrix is generated to simulate the missing state, and the conditional diffusion model uses the known part of the data to calculate the complete data. By calculating the loss function between the output missing prediction value and the corresponding true value, the parameters of the conditional diffusion model are optimized through backpropagation until the loss function converges. The finally obtained conditional diffusion model with optimal parameters is called the residual component imputation model.

[0065] Step 5: Decompose the missing data M to be imputed collected in the steel continuous casting process into a new trend component M t and a new residual component M r ;

[0066] The decomposition formula is:

[0067] M t = W(M)

[0068] M - M t = M r

[0069] W represents the moving window average function. The data M with missing values becomes the trend component M after passing through the moving window average algorithm. t ;

[0070] Step 6: Perform cubic Hermite interpolation on the new trend component M t to obtain the missing values of the trend component. At the same time, use the trained residual component imputation model to predict the missing values of the new residual component M r .

[0071] Cubic Hermite interpolation is an algorithm that uses the information of the starting point and ending point of the missing segment for interpolation. The process is as follows: Input the value and corresponding derivative at the starting point of the missing segment, the value and corresponding derivative at the ending point of the missing segment, and interpolate the cubic curve function of the missing segment based on the above four inputs to obtain the predicted values of the intermediate missing points;

[0072] The function formula is

[0073]

[0074] Among them, f(a) represents the value at the missing starting point a, and f′(a) represents the derivative value at the missing starting point a; f(b) represents the value at the missing ending point b, and f′(b) represents the derivative value at the missing ending point b.

[0075] The calculation methods of the derivatives at the starting point and the ending point are the same, and are specifically as follows:

[0076] The values at the starting point and the ending point are both known. Since it is discrete data and there is no exact derivative, it is necessary to perform linear regression on n points near the starting point and the ending point to calculate the approximate derivative.

[0077] Taking the calculation of the derivative at the starting point a as an example, let the adjacent points be x1, x2,..., x i , and the calculation formula is

[0078]

[0079] y0 drift is the derivative at the x0 point, and y i is the value of the corresponding point. n is the number of adjacent points participating in the calculation, and n takes 4 in this method.

[0080] The residual component interpolation model is based on the conditional diffusion model, so when making predictions, it is necessary to input the original data and conditional information for prediction. The original data refers to the residual component M with missing data r , and the conditional information refers to the matrix composed of 0 and 1 that marks the position of the missing part, where 0 represents missing and 1 represents non-missing; the output is the complete residual component after replacing the missing part of the residual component with the model prediction value.

[0081] Assume that X is the original data and Q is the matrix representing the missing positions.

[0082]

[0083]

[0084] The model output will be the complete data after replacing the missing part of the original data with the model's predicted value;

[0085] Since the data shape output by the conditional diffusion neural network model is consistent with the input, the output actually predicts all the data. However, as a missing imputation task, only the predicted values of the missing part are important. Therefore, only this part participates in the calculation when calculating the loss function. In the last step before the residual component interpolation model outputs, the values corresponding to the missing part of the model output are retained, and the known parts in the input data are used to replace the areas outside the missing part.

[0086] Step 7: Sum the missing values of the predicted trend component (obtained by cubic interpolation) and the missing values of the residual component (obtained by the diffusion model) to obtain the complete predicted data for the missing part.

[0087] Example:

[0088] Step 1: (Data collection) Align the collected univariate sequences with time as the standard and merge them into a multivariate sequence.

[0089] Since the collected variables are univariate sequences with time information, and the variables in the metallurgical production process have strong coupling characteristics, univariate data governance often cannot reflect the correlation between data, resulting in low data governance accuracy. Therefore, the univariate sequences are merged into a multivariate sequence.

[0090] Step 2: (Data cleaning) Use the moving window average algorithm to calculate the trend component of the "tundish temperature" variable, and then calculate the average value and standard deviation of the trend component. After that, traverse the trend component in chronological order. If there is a data point that deviates from the average value by 6 standard deviations, mark all the time axis positions with the same monotonic change trend connected to this point. After the traversal, delete the variables at the marked time axis positions in the original multivariate sequence of data.

[0091] The tundish temperature is selected to identify the downtime because after the downtime, the tundish temperature will continuously and monotonically decrease and deviate significantly from the working temperature until the temperature rises again and returns to the normal temperature when starting up again. The downtime can be distinguished based on the characteristic of deviating from the working temperature, and the entire rising and falling processes can be identified through the characteristic of monotonic change. The 6-fold standard deviation is selected because fluctuations with a small deviation amplitude may be caused by other changes during normal operation, and using a larger deviation amount can distinguish the downtime from these normal changes.

[0092] Specifically for the moving window average algorithm, the average value of each point and the nearest 4 points is used as the trend component at this point. To ensure that the sequence length remains unchanged before and after the calculation, padding is performed at the beginning and end of the sequence before the calculation. The value of the first position is copied forward by two positions, and the value of the last position is copied backward by two positions. The waveform of the calculated trend component will be smoother and will not be damaged by individual abnormal points in the monotonic change trend of the data.

[0093] All data with the same monotonic change trend are marked because the downtime starts when the data is decreasing. If this part of the data is not excluded, a section of data from the normal value to the deviation of 6 standard deviations during the decrease process will be retained, which will affect the training effect of the diffusion model.

[0094] Step 3: (Data decomposition) Use the moving window average algorithm to calculate the trend component of the entire dataset, and subtract the trend component from the original dataset to obtain the residual component of the dataset.

[0095] Step 4: (Model Training) Standardize the residual components, divide the standardized results into a training set and a test set in a ratio of 8:1, and input them into the conditional diffusion model for training to obtain an imputation model for the residual components.

[0096] Standardization means subtracting the original data from the mean first and then dividing by the standard deviation to obtain data with a mean of 0 and a standard deviation of 1. The input of standardized data is a necessary condition for the training of the diffusion neural network.

[0097] Step 5: (Missing Value Prediction) Take the data that needs to be imputed, calculate the existing partial trend components, and take the difference to obtain the residual components. Through the existing partial trend components, perform cubic Hermite interpolation on the missing part to obtain the smooth trend components of the missing part; input the existing partial residual components into the conditional diffusion model to predict the residual components of the missing part. Sum them with the corresponding residual components to obtain the complete predicted data of the missing part.

[0098] The structure of the conditional diffusion model used in this example is as Figure 3 shown, where the network structure of the residual module of this model is as Figure 4 shown.

[0099] Since the conditional diffusion model will have the problem of imputation value deviation after data standardization and scaling, and it is positively correlated with the variance of the data. And the time series data of continuous casting of steel with trend components has a large distribution range and a large variance, and there will be a serious deviation problem after standardization and scaling. Therefore, the original data is decomposed into trend components and residual components for separate imputation. The trend components represent the fluctuations of the data on a large scale, with a smooth waveform and a large variance, while the residual components represent the small-scale fluctuations above and below the trend components of the data, with a small variance. Use smooth cubic Hermite interpolation to predict the trend components. Since the waveform of the trend components is smooth, the prediction accuracy can be improved; use the conditional diffusion model to predict the residual components. Since the variance is small, the prediction accuracy can be improved.

Claims

1. A data governance method based on data generation model for continuous casting production process, characterized in that It includes the following steps: Step 1: Obtain the time-series data of continuous casting of steel, align multiple single-variable data with time as the standard, and merge them into a multi-variable data sequence; Step 2: Use the downtime data identification algorithm to clean the multi-variable data sequence by deleting the data corresponding to the downtime; The specific cleaning process is as follows: First, use the moving window average algorithm to calculate the average value of the data within the window to smooth the data, so as to extract the trend component of the "tundish temperature" variable, and then calculate the average value and standard deviation of the trend component; Then, traverse the trend component in chronological order. If a data point with a deviation from the average value of 6 times the standard deviation appears, mark all the time-axis positions with the same monotonic change trend connected to this point; Finally, after the traversal, delete the variables at the marked time-axis positions in the multi-variable sequence of the original data; Step 3: Use the moving window average method to decompose the cleaned data set into a trend component and a residual component; Step 4: Standardize the residual component, input it into the conditional diffusion model for training, and optimize the model parameters by backpropagation to obtain the residual component interpolation model; Step 5: Decompose the missing data M to be interpolated collected in the steel continuous casting process into a new trend component M t and a new residual component M r ; Step 6: For the new trend component M t Use cubic Hermite interpolation for prediction to obtain the missing values of the trend component; meanwhile, use the trained residual component interpolation model to predict the missing values of the new residual component M r ; The cubic Hermite interpolation process is: input the value and corresponding derivative at the missing starting point, the value and corresponding derivative at the missing ending point, and interpolate the cubic curve function of the missing segment according to the above four input information to obtain the predicted values of the intermediate missing points; The function formula is where f(a) represents the value at the missing starting point a, and f′(a) represents the derivative value at the missing starting point a; f(b) represents the value at the missing ending point b, and f′(b) represents the derivative value at the missing ending point b; The input of the residual component interpolation model includes the residual component M with missing values r and conditional information, where the conditional information refers to a matrix of 0s and 1s that marks the positions of the missing parts, with 0 representing missing and 1 representing non-missing; from the output, select the predicted residual component M r and retain the missing values; Step 7: Sum the predicted missing values of the trend component and the missing values of the residual component to obtain the complete predicted data for the missing part.

2. The data governance method according to claim 1, wherein In step 2, the formula for calculating the marked data points is as follows: x is the value of the trend component being traversed, is the average value, and σ is the standard deviation.

3. The data governance method according to claim 1, wherein In the said Step 3, the decomposition process: D t = W(D) D-D t = D r Let \(D\) denote the cleaned data set and \(W\) denote the moving window average function. After passing \(D\) through the moving window average algorithm, the trend component is \(D\). t The residual component is \(D\). r It refers to the difference between the original data set and the trend component.

4. The data governance method according to claim 1, wherein In the said Step 4, the standardization of the residual component is to subtract the average value of the original data from the residual component and then divide by the standard deviation to obtain data with a mean of 0 and a standard deviation of 1; the specific formula is: D s is the residual component, D rn is the standardized residual component, is the average value of the residual component, is the standard deviation of the residual component.

Citation Information

Cited By

  • Multi-periodic ecological environment index anomaly identification and deletion interpolation method

    CN120724360A

  • Image generation method, image synthesis method, computing device and electronic device

    CN121329804A