Converter steelmaking end-point carbon and temperature prediction method and device based on adaptive data augmentation
Through adaptive data augmentation technology, the interpolated samples are generated using isolated forests and adaptive SMOTE algorithms, which solves the data imbalance and noise problems in the carbon temperature prediction of the end point of the converter steelmaking, improves the generalization ability and prediction accuracy of the model, and reduces the risk of overfitting.
Patent Information
- Application Number
- CN202411695219.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-11-25
AI Technical Summary
The existing carbon temperature prediction method for converter steelmaking endpoint carbon temperature in the uneven data and high noise environment, the model generalization ability is insufficient, making it difficult to accurately cover all working conditions, resulting in low prediction accuracy and risk of overfitting.
Adaptive data augmentation technology is adopted to select features through isolated forest algorithm detection outliers and recursive feature elimination method, and combined with adaptive SMOTE data augmentation algorithm to generate interpolated samples in sparse areas, dynamically adjust the number of interpolations, and use a random forest model to predict.
The generalization ability of the model under sparse and uneven data conditions is improved, the risk of overfitting is reduced, and the accuracy and stability of endpoint carbon temperature prediction is improved.
Smart Images

Figure CN119514379B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of converter steelmaking, and particularly to a method and device for predicting the end-point carbon temperature of converter steelmaking based on adaptive data augmentation. Background Art
[0002] With the rapid development of China's economy, the steel industry has become increasingly important in the national economy, providing key raw material guarantees for national construction and development. The emergence of the oxygen top-blown converter steelmaking technology has greatly promoted the transformation and upgrading of the steelmaking industry and brought about a comprehensive, sustainable and considerable steel industry structure, thus greatly increasing the overall supply capacity of steel.
[0003] As an important link in steel production, the stability of the converter steelmaking process and the accuracy of end-point prediction play a crucial role in production efficiency and molten steel quality. The carbon content in molten steel is one of the key indicators of steel quality, affecting physical properties such as the hardness, strength, plasticity and endurance of steel. Excessive or too low carbon content will affect the properties of steel; the molten steel temperature affects the overall metallurgical reaction rate in the smelting process, the dissolution of alloying elements and the fluidity of molten steel, and ultimately affects the partial organizational structure and mechanical properties of raw materials. Accurately predicting the end-point molten steel temperature and carbon content is not only beneficial for operators to make timely adjustments during the production process, avoid the occurrence of extreme situations, ensure the smooth progress of the smelting process, improve molten steel quality and production efficiency, but also can improve molten steel quality and production efficiency, reduce production costs, and promote the automation and environmental protection development of steel production.
[0004] Early studies usually adopted traditional mechanism models, mainly divided into static models and dynamic models. The static model is based on the physical and chemical reactions in the steelmaking process, calculates the addition amounts of various raw materials and auxiliary materials required to achieve the end-point carbon content and temperature, and the blowing amount of gas, so as to establish a model; the dynamic model corrects the blowing conditions based on the time-series process data information on the basis of the static model to achieve the control of the end-point. However, the traditional models cannot fully adapt to the instability of the physical and chemical reactions in the converter steelmaking process, and there are great limitations in the control and prediction of the end-point. With the development of automation technology, on-line flue gas analysis and sublance detection technology have been applied in the control and prediction of the converter steelmaking end-point. The sublance detection technology makes the probe enter the molten steel to directly detect the state of the molten steel. However, the sublance probe is expensive, consumes a large amount, and can only be applied to large-scale equipment; on-line flue gas analysis predicts the molten steel composition and temperature by detecting and analyzing the flue gas information during the steelmaking process. However, since the detection accuracy is easily affected by the cleanliness of the converter flue gas, the prediction accuracy of the model is extremely unstable. Therefore, traditional models and analysis methods based on automation equipment are difficult to cope with the complexity of the working conditions and the limitations of detection means in the steelmaking process, resulting in inaccurate prediction of the end-point carbon content and molten steel temperature.
[0005] In recent years, intelligent steelmaking technologies have gradually developed, and the accumulation of a large amount of data has enabled the application of data-driven modeling methods. By applying machine learning and deep learning algorithms to the prediction of the end-point of converter steelmaking, the accuracy of prediction can be effectively improved. However, the high temperature and high pressure in the steelmaking environment cause data acquisition devices to operate under extreme conditions. The collected data usually contains noise and has missing features. This situation seriously affects the training of machine learning models and easily leads to problems such as overfitting and low hit rates in end-point prediction. Therefore, how to use data augmentation methods to address issues such as uneven data distribution, significant noise effects, and data sparsity in the converter steelmaking process, improve the prediction accuracy of the model, and thus achieve the control of the steelmaking process is of great significance. Data augmentation techniques originated in the field of computer vision. As early as the 1980s, Brutlag et al. proposed methods for transforming images such as rotation and scaling to augment the dataset and improve the generalization ability of the model, laying the foundation for subsequent research. In 2002, N. V. Chawla first proposed the SMOTE (Synthetic Minority Oversampling Technique) method, which was applied to the classification problem of tabular data. For different tasks, the SMOTE data augmentation algorithm has also been applied to regression problems where the target variable is a continuous variable. In addition, researchers have explored the application of SMOTE and its variant algorithms in combination with different domain backgrounds to improve the performance of the model in practical applications. Currently, no SMOTE data augmentation model has been applied to the prediction of the end-point carbon temperature in converter steelmaking.
[0006] With the wide application of machine learning in the field of data mining, the prediction method for the end-point carbon temperature in converter steelmaking has shifted from traditional mechanism model prediction to a prediction method based on a data-driven model. Currently, the prediction method for the end-point of converter steelmaking based on machine learning mainly directly establishes a prediction model according to the process data of converter steelmaking. First, methods such as data preprocessing and correlation analysis are used to process and analyze the data, and then prediction models are established for different tasks to achieve the prediction of the end-point molten steel temperature and composition.
[0007] First of all, most of the existing methods use Pearson correlation analysis and partial correlation analysis methods to select input features to further improve the generalization ability and accuracy of the model. In the case of high data quality and complete collection, such methods often perform well. However, due to the complex physical and chemical reactions involved in the converter steelmaking process, the characteristics between various features are non-linear, strongly coupled, and high-dimensional. Therefore, such methods cannot accurately mine the relationships between features.
[0008] Secondly, the data required for the BOF steelmaking prediction task usually comes from multi-source databases. However, during the process of data-driven modeling, manual collation and merging of this data are needed, which is not only time-consuming and laborious, but also there are varying degrees of missing data from different sources for different heats. Therefore, the amount of high-quality data available for model training is small, and the model cannot learn accurate information, resulting in a low prediction accuracy of the model.
[0009] Finally, in the process control and endpoint prediction tasks of BOF steelmaking, different models need to be established for different types of tasks, making full use of the distribution and characteristics of the data itself for processing so that it can adapt to different scenarios, thereby controlling and predicting the steelmaking process and endpoint. Summary of the Invention
[0010] To solve the technical problem of how to use adaptive data augmentation and noise control means to improve the accuracy of predicting the end-point carbon content and molten steel temperature in BOF steelmaking, the embodiments of the present invention provide a method and device for predicting the end-point carbon and temperature in BOF steelmaking based on adaptive data augmentation. The present invention aims to solve the limitations of data sparsity and noise impact by improving data augmentation and feature selection methods. The technical solutions are as follows:
[0011] On the one hand, a method for predicting the end-point carbon and temperature in BOF steelmaking based on adaptive data augmentation is provided. This method is implemented by an end-point carbon and temperature prediction device for BOF steelmaking, and the method includes:
[0012] S1. Obtain the historical production data set of the BOF steelmaking process, and preprocess the data in the historical production data set to obtain a preprocessed data set.
[0013] S2. Process the preprocessed data set through the adaptive SMOTE data augmentation technique to obtain a processed data set.
[0014] S3. Build a prediction model for the end-point carbon and temperature in steelmaking based on random forest.
[0015] S4. Train the prediction model for the end-point carbon and temperature in steelmaking according to the processed data set to obtain a trained prediction model for the end-point carbon and temperature in steelmaking with adaptive data augmentation.
[0016] S5. Obtain the production data of the BOF steelmaking process to be predicted, and input the production data into the trained prediction model for the end-point carbon and temperature in steelmaking with adaptive data augmentation to obtain the prediction result of the end-point carbon and temperature in BOF steelmaking.
[0017] Optionally, the preprocessing of the data in the historical production data set in S1 to obtain a preprocessed data set includes:
[0018] S11. Introduce a noise control and outlier detection module, use the Isolation Forest algorithm to identify outliers, and remove the outliers to obtain the screened data.
[0019] S12. Use the Recursive Feature Elimination (RFE) method to perform feature selection on the screened data to obtain the preprocessed dataset.
[0020] Optionally, in S2, the preprocessed dataset is processed by the Adaptive Synthetic Minority Over-sampling Technique (SMOTE) to obtain the processed dataset, including:
[0021] S21. Calculate the k-nearest neighbor average distance for each data in the preprocessed dataset, use the k-nearest neighbor average distance as a measure of the sample data distribution density, divide the data in the preprocessed dataset into a dense region and a sparse region according to the measure of the sample data distribution density, and calculate the interpolation weight for each data in the dense region and the sparse region. .
[0022] S22. Dynamically determine the optimal number of sample interpolations according to the historical production dataset, and use the Adaptive SMOTE data augmentation technique based on the optimal number of sample interpolations to interpolate new data according to the weights for the sparse region and the dense region to obtain the processed dataset.
[0023] Optionally, the calculation method of the measure of the sample data distribution density in S21 is as shown in the following formula (1):
[0024] (1)
[0025] In the formula, represents the measure of the sample data distribution density, represents the maximum number of neighboring samples, represents the data and the data the distance between.
[0026] Optionally, the calculation method of the interpolation weight in S21 is as shown in the following formula (2):
[0027] (2)
[0028] In the formula, represents the maximum measure of the sample data distribution density in the preprocessed dataset, represents the average value of the measure of the sample data distribution density.
[0029] Optionally, the new data in S22 is as shown in the following formula (3):
[0030] (3)
[0031] In the formula, represents the new data synthesized between and ; represents the -th attribute value of the -th data in the minority class; represents the -th -th nearest neighbor data of the data; , is the maximum number of nearest neighbor samples.
[0032] Among them, has the expression of formula (4), is inversely proportional to ; represents , the distance between. The greater the distance between the sample point and its nearest neighbor sample point, the smaller the interpolation range, which is used to ensure that the new sample is within the effective space and to avoid the generated samples concentrating in a certain direction;
[0033] (4).
[0034] On the other hand, a converter steelmaking end-point carbon temperature prediction device based on adaptive data augmentation is provided. This device is applied to the converter steelmaking end-point carbon temperature prediction method based on adaptive data augmentation. The device includes:
[0035] A data preprocessing module, which is used to obtain the historical production data set of the converter steelmaking process, preprocess the data in the historical production data set, and obtain the preprocessed data set.
[0036] An adaptive data augmentation module, which is used to process the preprocessed data set through the adaptive SMOTE data augmentation technology to obtain the processed data set.
[0037] A model building module, which is used to build a steelmaking end-point carbon temperature prediction model based on random forest.
[0038] A model training module, which is used to train the steelmaking end-point carbon temperature prediction model according to the processed data set to obtain a trained converter steelmaking end-point carbon temperature prediction model with adaptive data augmentation.
[0039] A model inference module, which is used to obtain the production data of the converter steelmaking process to be predicted, input the production data into the trained converter steelmaking end-point carbon temperature prediction model with adaptive data augmentation, and obtain the converter steelmaking end-point carbon temperature prediction result.
[0040] Optionally, the data preprocessing module is further configured to:
[0041] S11. Introduce a noise control and outlier detection module, identify outliers using the isolation forest algorithm, and remove the outliers to obtain the screened data.
[0042] S12. Use the recursive feature elimination method (RFE) to perform feature selection on the screened data to obtain the preprocessed data set.
[0043] Optionally, the adaptive data augmentation module is further configured to:
[0044] S21. Calculate the k-nearest neighbor average distance for each data in the preprocessed data set, use the k-nearest neighbor average distance as a measure of the sample data distribution density, divide the data in the preprocessed data set into a dense region and a sparse region according to the measure of the sample data distribution density, and calculate the interpolation weight for each data in the dense region and the sparse region .
[0045] S22. Dynamically determine the optimal number of sample interpolations based on the historical production data set, and use the adaptive SMOTE data augmentation technique based on the optimal number of sample interpolations to interpolate and generate new data for the sparse region and the dense region according to the weights to obtain the processed data set.
[0046] Optionally, the calculation method of the measure of the sample data distribution density is as shown in the following formula (1):
[0047] (1)
[0048] In the formula, represents the measure of the sample data distribution density, represents the maximum number of neighboring samples, represents the data and the data the distance between them.
[0049] Optionally, the calculation method of the interpolation weight is as shown in the following formula (2):
[0050] (2)
[0051] In the formula, represents the maximum measure of the sample data distribution density in the preprocessed data set, represents the average value of the measures of the sample data distribution density.
[0052] Optionally, the new data is as shown in the following formula (3):
[0053] (3)
[0054] In the formula, represents the synthesized new data between and ; represents the -th attribute value of the -th data in the minority class; represents the -th nearest neighbor data of data ; , is the maximum number of nearest neighbor samples.
[0055] Among them, is expressed by formula (4), is inversely proportional to ; represents , the distance between. The greater the distance between a sample point and its nearest neighbor sample points, the smaller the interpolation range, which is used to ensure that the new sample is within the effective space and avoid the generated samples concentrating in a certain direction;
[0056] (4).
[0057] On the other hand, a converter steelmaking end-point carbon temperature prediction device is provided. The converter steelmaking end-point carbon temperature prediction device includes: a processor; a memory, and computer-readable instructions are stored on the memory. When the computer-readable instructions are executed by the processor, any one of the methods in the above converter steelmaking end-point carbon temperature prediction method based on adaptive data augmentation is implemented.
[0058] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the methods in the above converter steelmaking end-point carbon temperature prediction method based on adaptive data augmentation.
[0059] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0060] In the present invention, in the problem of predicting the end point of converter steelmaking, how to improve the generalization ability of the model and avoid overfitting in an environment of data imbalance and high noise is the core issue. The current mainstream approach is mainly to directly perform conventional cleaning on the original data set and establish a single prediction model based on machine learning methods to predict the carbon content at the end point of the converter and the molten steel temperature. Such methods can achieve good results when the data is sufficient and evenly distributed, but in the steelmaking process, the data samples under extreme conditions are scarce and unevenly distributed, making it difficult for the model to accurately cover all working conditions. The problems of data sparsity and imbalance not only affect the generalization ability of the model but also increase the risk of the model in practical applications. The present invention believes that the SMOTE data augmentation method can solve the problems such as the difficulty in obtaining high-quality data sets and high labor costs in practical problems. By using the feature information in the original data to generate appropriate interpolation samples in the sparse data area. This method can effectively increase the sample size and diversity of the data when the data is scarce, which helps the model to learn various working conditions more effectively, reducing the bias and overfitting risks.
[0061] Due to the characteristics of non-linearity, strong coupling, and high dimensionality of the converter steelmaking process data, most traditional data preprocessing methods rely on simple correlation analysis and cannot accurately identify abnormal data and capture the relationships between complex high-dimensional features. The present invention designs a data preprocessing scheme that uses the Isolation Forest algorithm to detect sample points deviating from the data distribution and uses the RFE feature selection method to automatically evaluate the importance of each feature. This data preprocessing scheme can improve the quality of the data set and avoid feature redundancy for further analysis by the subsequent model.
[0062] Applying the SMOTE data augmentation algorithm in combination with the characteristics of the data itself can give full play to the advantages of this algorithm to a greater extent. The present invention designs a data distribution density weight adjustment mechanism, which enables the SMOTE algorithm to perform dynamic interpolation adjustment according to feature sparsity and data density, avoiding redundant interpolation and introducing noise in the data-dense area. This not only effectively enhances the prediction ability of the model but also can more accurately reflect the data characteristics under different working conditions. This mechanism is applicable to various data augmentation algorithms.
[0063] The samples generated by SMOTE data augmentation provide a more balanced data distribution for the model, reducing the impact of the original data imbalance. The present invention also proposes an adaptive mechanism, enabling the interpolated samples to adapt to the dynamic working condition changes in converter steelmaking and ensuring the rationality and effectiveness of the augmented data. Finally, while maintaining the prediction accuracy, the method of the present invention reduces the overfitting risk of the model, improves the generalization ability and stability of the model, and provides an efficient and practical technical path for the carbon temperature prediction problem. Description of the Drawings
[0064] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0065] Figure 1 It is a flowchart of a converter steelmaking end-point carbon temperature prediction method based on adaptive data augmentation provided by an embodiment of the present invention;
[0066] Figure 2 It is a flowchart of RFE feature selection provided by an embodiment of the present invention;
[0067] Figure 3 It is a schematic diagram of data distribution density weight adjustment provided by an embodiment of the present invention;
[0068] Figure 4 It is a schematic diagram of adaptive SMOTE data augmentation provided by an embodiment of the present invention;
[0069] Figure 5 It is a general flowchart of a converter steelmaking end-point carbon temperature prediction method based on adaptive data augmentation provided by an embodiment of the present invention;
[0070] Figure 6 It is a general structure diagram of a converter steelmaking end-point carbon temperature prediction method based on adaptive data augmentation provided by an embodiment of the present invention;
[0071] Figure 7 It is a block diagram of a converter steelmaking end-point carbon temperature prediction device based on adaptive data augmentation provided by an embodiment of the present invention;
[0072] Figure 8 It is a schematic diagram of the structure of a converter steelmaking end-point carbon temperature prediction device provided by an embodiment of the present invention. Detailed implementation manners
[0073] The following describes the technical solutions in the present invention in conjunction with the accompanying drawings.
[0074] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0075] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.
[0076] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.
[0077] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0078] The embodiment of the present invention provides a method for predicting the end point carbon temperature of converter steelmaking based on adaptive data enhancement. The method can be implemented by a device for predicting the end point carbon temperature of converter steelmaking. The device for predicting the end point carbon temperature of converter steelmaking can be a terminal or a server. Figure 1 The flowchart of the method for predicting the end point carbon temperature of converter steelmaking based on adaptive data enhancement is shown in FIG. The processing flow of the method may include the following steps:
[0079] S1. Obtain a historical production data set of a converter steelmaking process, and preprocess the data in the historical production data set to obtain a preprocessed data set.
[0080] In a feasible implementation manner, the data preprocessing stage of the present invention can use data cleaning, feature selection and other technologies to perform outlier processing and feature selection on the data in the converter steelmaking process production data set to obtain a preprocessed data set.
[0081] Specifically, the preprocessing of the data in the historical production data set in the above step S1 to obtain the preprocessed data set may include the following steps S11-S12:
[0082] S11. Introduce the noise control and outlier detection module, use the isolation forest algorithm to identify outliers, and remove outliers to obtain filtered data.
[0083] In a feasible implementation method, data cleaning is performed by introducing a noise control and anomaly detection module, and an Isolation Forest algorithm is used to screen and remove abnormal data, remove data that does not conform to normal operating conditions, and improve the quality of input data.
[0084] First, select from the training dataset of the converter steelmaking process dataset Take a sample as a sample subset, extract some samples from it and select some features to construct an isolation tree iTree. Use the binary search tree structure of the isolation tree iTree to isolate the samples. Repeat the above operations until all data are isolated to obtain T isolation trees, forming an isolation forest.
[0085] Secondly, substitute each sample point into each isolation tree in the forest to calculate the PathLength. The expression is as shown in Equation (1):
[0086] (1)
[0087] In the formula, is the number of edges experienced by the sample from the root node to the leaf node of the tree. is the number of samples in the same leaf node as the sample . represents the average path length of constructing a binary tree with samples. The expression is Equation (2):
[0088] (2)
[0089] Then calculate the outlier score of each sample point. The expression is Equation (3):
[0090] (3)
[0091] In the formula, is the mean value of PathLength of the sample in T isolation trees.
[0092] S12. Adopt the recursive feature elimination method RFE to perform feature selection on the screened data to obtain the preprocessed data set.
[0093] In a feasible implementation manner, for feature selection, RFE (Recursive Feature Elimination) screens out the features that have a significant impact on the prediction of the end point carbon content and molten steel temperature. The RFE feature selection method is as Figure 2 shown. Specifically, first, based on the process and mechanism principle of converter steelmaking, analyze the factors affecting the end point carbon content and molten steel temperature, including the addition amount of raw materials, auxiliary materials, and gas injection amount, etc.; secondly, to eliminate the influence of low-correlation features on the model, by training a random forest model, calculate the influence of each feature on the model performance, and gradually eliminate the features with small contribution to the prediction until the optimal feature subset is retained.
[0094] The present invention designs a data preprocessing method for outlier rejection and feature selection. A method for identifying and rejecting outliers based on the Isolation Forest algorithm is adopted to reject data that does not conform to normal working conditions; a RFE feature selection module based on prior knowledge is designed to automatically evaluate the importance of each feature and eliminate feature redundancy.
[0095] S2. Process the preprocessed data set through the adaptive SMOTE data augmentation technology to obtain a processed data set.
[0096] In a feasible implementation, in the adaptive SMOTE data augmentation stage, an adaptive SMOTE interpolation method based on data sparsity and feature space density is used to dynamically adjust the amount of interpolated data for different tasks. Interpolated data is preferentially generated in the sparse regions of the data set to ensure the uniformity of data distribution, avoid generating too much data in the sample-dense regions, and prevent the introduction of noise. The obtained processed data set is used in step S3.
[0097] Specifically, the above step S2 may include the following steps S21 - S22:
[0098] S21. Calculate the k-nearest neighbor average distance for each data in the preprocessed data set, use the k-nearest neighbor average distance as a measure of the sample data distribution density, divide the data in the preprocessed data set into dense regions and sparse regions according to the measure of the sample data distribution density, and calculate the interpolation weight for each data in the dense region and the sparse region .
[0099] Specifically, the distribution density of the sample is determined by adjusting the density weight, and the specific method is as Figure 3 shown. Calculate the k-nearest neighbor average distance for each sample point as a measure of the sample data distribution density, as shown in the following formula (4):
[0100] (4)
[0101] In the formula, represents the measure of the sample data distribution density, represents the maximum number of neighboring samples, represents the data and the distance between the data . In a certain region, the closer the data samples are, the greater the density value, indicating that the region is a dense region, and vice versa, it is a sparse region.
[0102] To further determine the interpolation region and the distribution of new sample data, calculate the interpolation weight for each sample, so that the weight of the sparse region is larger and the weight of the dense region is smaller, and use the average value in the distribution density values as the threshold , if it is greater than this value, it is a dense area and a smaller weight is assigned; otherwise, it is a sparse area and a larger weight is assigned. Calculate the interpolation weight It is expressed as Equation (5):
[0103] (5)
[0104] In the formula, represents the maximum density value in the dataset.
[0105] S22. According to the historical production dataset, dynamically determine the optimal number of sample interpolations. Based on the optimal number of sample interpolations, adopt the adaptive SMOTE data augmentation technique to interpolate the sparse area and the dense area according to the weights to generate new data, and obtain the processed dataset.
[0106] In a feasible implementation, the adaptive SMOTE data augmentation interpolation technique is adopted to preferentially interpolate in the sparse sample area, avoiding the generation of redundant data and the accumulation of noise due to excessive interpolation in the dense area. In addition, according to the inherent characteristics of the original dataset, dynamically determine the optimal number of sample interpolations to synthesize new sample data. The new sample data synthesis formula is Equation (6):
[0107] (6)
[0108] In the formula, , is the -th attribute value of the -th sample in the minority class, , The expression of is Equation (7); is the -th nearest neighbor sample of sample is the new sample synthesized between sample and , , is the maximum number of nearest neighbor samples.
[0109] (7)
[0110] is inversely proportional to the distance between and . The greater the distance between the sample point and its nearest neighbor sample point, the smaller the interpolation range, ensuring that the new sample is within the effective space and at the same time avoiding the generation of samples concentrated in a certain direction.
[0111] The specific process of the SMOTE algorithm is as follows: first, select minority class samples, and calculate the distance between each sample in the minority class and its neighboring samples; second, select neighboring samples, randomly select a minority class sample, and find its nearest k neighboring samples, usually k=5, for the selected minority class sample, randomly select a sample from its k nearest neighbors; then, on the line segment between the sample and the selected sample, calculate the distance between the sample and the neighboring samples; second, select neighboring samples, randomly select a minority class sample, and find its nearest k neighboring samples, usually k=5, for the selected minority class sample, randomly select a sample from its k nearest neighbors; then, calculate the distance between the sample and the selected sample according to A new feature variable is generated by interpolating the random values of , this vector is located at and neighbor samples Between, then according to The random value interpolation generates a corresponding target variable , and finally generate and The new samples are merged to obtain the new samples, and the eigenvalues of the new samples are linear combinations of the eigenvalues of the original samples, with the aim of not changing the distribution of the original samples as much as possible. Finally, the above steps are repeated until the predetermined maximum number of samples is reached.
[0112] The present invention adopts the SMOTE data enhancement algorithm in combination with the converter steelmaking process, uses the characteristic information in the original data to generate samples with a smooth distribution, fills the data missing and sparse areas, which can not only effectively enhance the prediction ability of the model, but also more accurately reflect the data characteristics under different working conditions.
[0113] A data distribution density weight adjustment method is designed. This method can perform dynamic interpolation adjustment based on feature sparsity and data density. This not only effectively enhances the prediction ability of the model, but also more accurately reflects the data characteristics under different working conditions.
[0114] An adaptive interpolation mechanism is proposed, which enables the interpolated samples to adapt to the dynamic changes in converter steelmaking conditions, ensuring the rationality and effectiveness of the enhanced data. While maintaining the prediction accuracy, it reduces the risk of overfitting of the model and improves the generalization ability and stability of the model.
[0115] S3. Build a steelmaking endpoint carbon temperature prediction model based on random forest.
[0116] In a feasible implementation manner, the model building in the present invention adopts a random forest model. The random forest consists of multiple decision trees, each of which is trained based on subsamples randomly extracted from the training data and randomly selected features. During the prediction process, each tree will predict the input sample separately, and then obtain the final prediction result by voting or taking the average. It is suitable for converter steelmaking process data with high dimensionality, strong nonlinearity, and high noise. The random forest algorithm is used to construct a prediction model for the end point carbon content and molten steel temperature.
[0117] S4. Train the prediction model for the carbon content and temperature at the end of steelmaking based on the processed data set to obtain a trained prediction model for the carbon content and temperature at the end of steelmaking with adaptive data augmentation.
[0118] In a feasible implementation manner, the model training in the above step S4 may include the following steps S41 - S42:
[0119] S41. Pre - train the adaptive SMOTE data augmentation model. The specific method is as Figure 4 shown. First, divide the original data set into a training set and a test set. Secondly, set the SOMTE algorithm parameters, including the random seed, interpolation step size step, maximum interpolation number n max , the number of nearest neighbor samples k for synthesizing new samples. The interpolation numbers in the dense area and the sparse area are determined by Equation (8); then set the key parameters of the random forest model, including the number of decision trees (n_estimators), the depth of the tree (max_depth), the minimum number of samples in the leaf node (min_samples_leaf), etc. Use the training set data to train the model, traverse all the interpolation numbers, obtain the best interpolation number and the newly synthesized data based on the original training set, and save the model.
[0120] (8)
[0121] Among them, is the interpolation number of the current interpolation at the sample point, and the value range of .
[0122] S42. Use the converter steelmaking data to train the random forest prediction model. Integrate the training set of the original data and the newly generated samples, use the model parameters saved in step S41 as the initial values, and select the random search parameter tuning method to train the model. Finally, retain the best model during the validation process for inference in step S5.
[0123] S5. Obtain the production data of the converter steelmaking process to be predicted, input the production data into the trained prediction model for the carbon content and temperature at the end of steelmaking with adaptive data augmentation, and obtain the prediction result of the carbon content and temperature at the end of converter steelmaking.
[0124] In a feasible implementation manner, the model inference in the above step S5 may include the following steps S51 - S53:
[0125] S51. Perform data pre - processing on the input data, the same as step S1;
[0126] S52. Load the trained model saved in step S42 from the specified path. Ensure that the model parameters are correct, and prepare to input the processed test data into the model for prediction. Input the preprocessed test data into the loaded model, and the model will perform inference based on the parameters it has learned to calculate the predicted values of the end-point molten steel temperature and carbon content for each sample.
[0127] S53. Obtain the true values of the test set for comparison with the model's prediction results; use the true values to compare with the model's prediction results, calculate the prediction hit rates of the model for the end-point molten steel temperature and carbon content, and analyze the performance of the model based on the prediction hit rates.
[0128] Figure 5 It is a schematic flowchart of the method according to the embodiment of the present invention. The whole process can be divided into five steps: data preprocessing, adaptive data augmentation, model construction, model training, and model inference.
[0129] For example, the present invention uses the actual production data from the converter steelmaking process of a certain steel plant as the test data set, and the input features mainly include the addition amounts of raw materials, auxiliary materials, gases, etc. and equipment information in the steelmaking process. To verify the effectiveness of the model, the end-point molten steel temperature and carbon content of converter steelmaking are used as target variables for training respectively, as Figure 6 shown.
[0130] Furthermore, details of data preprocessing: For step S11, initially use box plots to detect outliers and delete them; then use the Isolation Forest algorithm to find and remove outliers in the data. The parameters of the Isolation Forest are set as: n_estimator = 100, max_samples = 'auto', contamination = float(0.1).
[0131] For step S12, use a random forest model for training, select features according to the model metrics, and use the random search method to search for the model parameters and retain the optimal results.
[0132] Furthermore, details of adaptive SMOTE data augmentation: For step S22, select the k nearest neighbor samples of the samples in the sparse region, set k to 5, the interpolation step size to 5, and the maximum interpolation number to 1000.
[0133] The ratio of the training set to the test set is 8:2.
[0134] Furthermore, details of the model parameters: The model parameters in step S12 are used as the initial parameters of the prediction model, and the parameter tuning is performed using the random search method. The parameters of the end-point molten steel carbon content prediction model, namely n_estimators, max_depth, max_features, and random_state, are 81, 10, 17, and 7 respectively; the parameters of the end-point molten steel temperature prediction model, namely n_estimators, max_depth, max_features, and random_state, are 71, 15,'sqrt', and 80 respectively.
[0135] Furthermore, experimental settings: The present invention uses the TensorFlow 2.11.0 learning framework. All experiments are carried out on a machine equipped with 4 NVIDIA Titan XP GPUs. The data processing is performed during the inference process, and the process is as Figure 5 shown.
[0136] The technical solution closest to the present invention:
[0137] The paper (Torgo L, Ribeiro R P, Pfahringer B, et al. Smote for regression[C] / / Portuguese conference on artificial intelligence. Berlin, Heidelberg: Springer Berlin Heidelberg, 2013: 378-389.) is the basis of the present invention. The present invention uses this paper to construct the SMOTE data augmentation algorithm in regression problems, combines the paper (Chen B, Xia S, Chen Z, et al. RSMOTE: A self-adaptive robust SMOTE for imbalanced problems with label noise[J]. Information Sciences, 2021, 553: 397-428.) to construct a data distribution density weight adjustment method, and adds an adaptive mechanism to the data augmentation strategy, enabling it to dynamically determine the interpolation strategy and apply it to the data augmentation of the converter steelmaking process.
[0138] The main differences between the present invention and the patent (CN118869313A, a high-precision intrusion detection method and system for network security communication based on a multi-scale convolutional neural network) are as follows: 1) The isolated forest algorithm is used for data outlier processing, and an RFE feature extraction method based on the prior knowledge of the converter steelmaking process is proposed; 2) An adaptive SMOTE data augmentation method is designed for regression problems with continuous variables. The distribution density of samples is determined by adjusting the density weight, and the optimal interpolation strategy is dynamically determined for different working conditions to synthesize new sample data; 3) The application fields are different. The present invention is used for the prediction of the end-point carbon temperature in converter steelmaking.
[0139] In the embodiment of the present invention, in the problem of converter steelmaking end-point prediction, how to improve the generalization ability of the model and avoid overfitting in an environment of data imbalance and high noise is the core issue. The current mainstream approach is mainly to directly perform conventional cleaning on the original data set and establish a single prediction model based on machine learning methods to predict the end-point carbon content and molten steel temperature of the converter. Such methods can achieve good results when the data is sufficient and evenly distributed. However, in the steelmaking process, the data samples under extreme conditions are scarce and unevenly distributed, resulting in the difficulty for the model to accurately cover all working conditions. The problems of data sparsity and imbalance will not only affect the generalization ability of the model but also increase the risk of the model in practical applications. The present invention believes that the SMOTE data augmentation method can solve the problems such as the difficulty in obtaining high-quality data sets and high labor costs in practical problems. By using the feature information in the original data to generate appropriate interpolation samples in the sparse data area. This method can effectively increase the sample size and diversity of the data when the data is scarce, which helps the model to learn various working conditions more effectively, reducing the bias and overfitting risks.
[0140] Due to the characteristics of non-linearity, strong coupling, and high dimensionality of the converter steelmaking process data, most traditional data preprocessing methods rely on simple correlation analysis and cannot accurately identify abnormal data and capture the relationships between complex high-dimensional features. The present invention designs a data preprocessing scheme that uses the isolated forest algorithm to detect sample points deviating from the data distribution and uses the RFE feature selection method to automatically evaluate the importance of each feature. This data preprocessing scheme can improve the quality of the data set and avoid feature redundancy for further analysis by the subsequent model.
[0141] Applying the SMOTE data augmentation algorithm in combination with the distribution characteristics of the data itself can give full play to the advantages of the algorithm to a greater extent. The present invention designs a data distribution density weight adjustment mechanism, which enables the SMOTE algorithm to perform dynamic interpolation adjustment according to feature sparsity and data density, avoiding redundant interpolation in dense data distribution areas and introducing noise. This not only effectively enhances the prediction ability of the model but also can more accurately reflect the data characteristics under different working conditions. This mechanism is applicable to various data augmentation algorithms.
[0142] The samples generated by SMOTE data augmentation provide a more balanced data distribution for the model, reducing the impact of the imbalance of the original data. The present invention also proposes an adaptive mechanism, enabling the interpolated samples to adapt to the dynamic working condition changes in converter steelmaking, ensuring the rationality and effectiveness of the augmented data. Finally, while maintaining the prediction accuracy, the method of the present invention reduces the overfitting risk of the model, improves the generalization ability and stability of the model, and provides an efficient and practical technical path for the carbon temperature prediction problem.
[0143] Figure 7 It is a block diagram of a converter steelmaking end-point carbon temperature prediction device shown according to an exemplary embodiment, and this device is used for the converter steelmaking end-point carbon temperature prediction method based on adaptive data augmentation. Referring to Figure 7 , this device includes a data preprocessing module 310, an adaptive data augmentation module 320, a model building module 330, a model training module 340, and a model inference module 350. Among them:
[0144] The data preprocessing module 310 is used to obtain the historical production data set of the converter steelmaking process, preprocess the data in the historical production data set, and obtain the preprocessed data set.
[0145] The adaptive data augmentation module 320 is used to process the preprocessed data set through the adaptive SMOTE data augmentation technology to obtain the processed data set.
[0146] The model building module 330 is used to build a steelmaking end-point carbon temperature prediction model based on random forest.
[0147] The model training module 340 is used to train the steelmaking end-point carbon temperature prediction model according to the processed data set to obtain a trained converter steelmaking end-point carbon temperature prediction model with adaptive data augmentation.
[0148] The model inference module 350 is used to obtain the production data of the converter steelmaking process to be predicted, input the production data into the trained converter steelmaking end-point carbon temperature prediction model with adaptive data augmentation, and obtain the converter steelmaking end-point carbon temperature prediction result.
[0149] In the embodiments of the present invention, in the problem of predicting the end point of converter steelmaking, how to improve the generalization ability of the model and avoid overfitting in an environment of data imbalance and high noise is the core issue. The current mainstream approach is mainly to directly perform conventional cleaning on the original data set and establish a single prediction model based on machine learning methods to predict the carbon content at the end point of the converter and the molten steel temperature. Such methods can achieve good results when the data is sufficient and evenly distributed, but in the steelmaking process, the data samples under extreme conditions are scarce and unevenly distributed, resulting in the model being difficult to accurately cover all working conditions. The problems of data sparsity and imbalance will not only affect the generalization ability of the model but also increase the risk of the model in practical applications. The present invention believes that the SMOTE data augmentation method can solve the problems such as the difficulty in obtaining high-quality data sets and high labor costs in practical problems. By using the feature information in the original data to generate appropriate interpolation samples in the sparse data area. This method can effectively increase the sample size and diversity of the data when the data is scarce, which helps the model to learn various working conditions more effectively, reduce bias and overfitting risks.
[0150] Due to the characteristics of nonlinearity, strong coupling, and high dimensionality of the converter steelmaking process data, most traditional data preprocessing methods rely on simple correlation analysis and cannot accurately identify abnormal data and capture the relationships between complex high-dimensional features. The present invention designs a data preprocessing scheme that uses the Isolation Forest algorithm to detect sample points deviating from the data distribution and uses the RFE feature selection method to automatically evaluate the importance of each feature. This data preprocessing scheme can improve the quality of the data set and avoid feature redundancy for further analysis by subsequent models.
[0151] Applying the SMOTE data augmentation algorithm in combination with the characteristics of the data itself can give full play to the advantages of the algorithm to a greater extent. The present invention designs a data distribution density weight adjustment mechanism, which enables the SMOTE algorithm to perform dynamic interpolation adjustment according to feature sparsity and data density, avoiding redundant interpolation and introducing noise in the dense data distribution area. This not only effectively enhances the prediction ability of the model but also can more accurately reflect the data characteristics under different working conditions. This mechanism is applicable to various data augmentation algorithms.
[0152] The samples generated by SMOTE data augmentation provide a more balanced data distribution for the model, reducing the impact of the original data imbalance. The present invention also proposes an adaptive mechanism that enables the interpolated samples to adapt to the dynamic working condition changes in converter steelmaking, ensuring the rationality and effectiveness of the augmented data. Finally, while maintaining the prediction accuracy, the method of the present invention reduces the overfitting risk of the model, improves the generalization ability and stability of the model, and provides an efficient and practical technical path for the carbon temperature prediction problem.
[0153] Figure 8 This is a schematic structural diagram of a converter steelmaking end-point carbon temperature prediction device provided by an embodiment of the present invention. As shown in Figure 8 the figure, the converter steelmaking end-point carbon temperature prediction device may include the above-mentioned Figure 7 converter steelmaking end-point carbon temperature prediction device based on adaptive data augmentation shown in the figure. Optionally, the converter steelmaking end-point carbon temperature prediction device 410 may include a first processor 2001.
[0154] Optionally, the converter steelmaking end-point carbon temperature prediction device 410 may further include a memory 2002 and a transceiver 2003.
[0155] Among them, the first processor 2001 is connected to the memory 2002 and the transceiver 2003, and can be connected through a communication bus, for example.
[0156] Next, the various components of the converter steelmaking end-point carbon temperature prediction device 410 will be specifically introduced in conjunction with Figure 8 the figure:
[0157] Among them, the first processor 2001 is the control center of the converter steelmaking end-point carbon temperature prediction device 410, and can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can also be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0158] Optionally, the first processor 2001 can execute various functions of the converter steelmaking end-point carbon temperature prediction device 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0159] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, for example Figure 8 the CPU0 and CPU1 shown in
[0160] In a specific implementation, as an embodiment, the converter steelmaking end-point carbon temperature prediction device 410 may also include multiple processors, for example Figure 8The first processor 2001 and the second processor 2004 shown in []. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0161] Among them, the memory 2002 is used to store the software program for implementing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation method can refer to the above method embodiment and will not be elaborated here.
[0162] Optionally, the memory 2002 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 can be integrated with the first processor 2001 or exist independently and is coupled to the first processor 2001 through the interface circuit of the converter steelmaking end-point carbon temperature prediction device 410 ( Figure 8 not shown in []). The embodiments of the present invention do not make specific limitations on this.
[0163] The transceiver 2003 is used to communicate with network devices or with terminal devices.
[0164] Optionally, the transceiver 2003 can include a receiver and a transmitter ( Figure 8 not shown separately in []). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0165] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently and is coupled to the first processor 2001 through the interface circuit of the converter steelmaking end-point carbon temperature prediction device 410 ( Figure 8 not shown in []). The embodiments of the present invention do not make specific limitations on this.
[0166] It should be noted that Figure 8 the structure of the converter steelmaking end-point carbon temperature prediction device 410 shown in does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0167] In addition, the technical effects of the converter steelmaking end-point carbon temperature prediction device 410 can refer to the technical effects of the converter steelmaking end-point carbon temperature prediction method based on adaptive data augmentation described in the above method embodiments, and will not be elaborated here.
[0168] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0169] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0170] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0171] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.
[0172] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0173] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0174] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0175] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices, apparatuses, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0176] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0177] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0178] In addition, the functional units in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0179] If the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0180] As described above, the above are only specific implementation manners of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for predicting the end-point carbon temperature of converter steelmaking based on adaptive data augmentation, characterized in that, The method includes: S1. Obtain the historical production dataset of the converter steelmaking process, preprocess the data in the historical production dataset to obtain a preprocessed dataset; S2. Process the preprocessed dataset through the adaptive SMOTE data augmentation technique to obtain a processed dataset; S3. Build a steelmaking end-point carbon temperature prediction model based on random forest; S4. Train the steelmaking end-point carbon temperature prediction model according to the processed dataset to obtain a trained steelmaking end-point carbon temperature prediction model with adaptive data augmentation; S5. Obtain the production data of the converter steelmaking process to be predicted, input the production data into the trained steelmaking end-point carbon temperature prediction model with adaptive data augmentation to obtain the prediction result of the converter steelmaking end-point carbon temperature; The process of processing the preprocessed dataset through the adaptive SMOTE data augmentation technique in S2 to obtain a processed dataset includes: S21. Calculate the k-nearest neighbor average distance for each data in the preprocessed dataset, use the k-nearest neighbor average distance as a measure of the sample data distribution density, divide the data in the preprocessed dataset into a dense region and a sparse region according to the measure of the sample data distribution density, and calculate the interpolation weight for each data in the dense region and the sparse region S22. Dynamically determine the optimal sample interpolation number according to the historical production dataset, and use the adaptive SMOTE data augmentation technique based on the optimal sample interpolation number to interpolate and generate new data according to weights in the sparse area and the dense area to obtain a processed dataset; The calculation method of the sample data distribution density measurement index in S21 is shown in the following formula (1): wherein, Density(x i ) represents the measure index of the sample data distribution density, k represents the maximum number of neighboring samples, and d(x i , x j ) represents the distance between the data x i and the data x j ; The interpolation weights in S21 are calculated as shown in the following formula (2): In the formula, Density max represents the measure index of the maximum sample data distribution density in the preprocessed dataset, and σ represents the average value of the measure index of the sample data distribution density; The new data in S22 is shown in the following formula (3): x new,feat = x i,feat + (x ij,feat - x i,feat ) × γ (3) where x new,feat represents the synthesized new data between data x i,feat and x ij,feat ; x i,feat represents the feat-th attribute value of the i-th data in the minority class; x ij,feat represents the j-th nearest neighbor data of data x i , where j = 1, 2,..., k, and k is the maximum number of nearest neighbor samples of x i,feat ; Among them, the expression of γ is Equation (4), and γ is inversely proportional to d(x i,feat , x ij,feat ). d(x i,feat , x ij,feat ) represents the distance between x i,feat and x ij,feat . The greater the distance between a sample point and its nearest neighbor sample point, the smaller the interpolation range, which is used to ensure that the new sample is within the effective space and to avoid the generated samples concentrating in a certain direction; 2. The method for predicting the end-point carbon temperature of converter steelmaking based on adaptive data augmentation according to claim 1, wherein The process of preprocessing the data in the historical production dataset in S1 to obtain a preprocessed dataset includes: S11. Introduce a noise control and outlier detection module, use the isolation forest algorithm to identify outliers, and remove the outliers to obtain filtered data; S12. Use the recursive feature elimination method RFE to perform feature selection on the filtered data to obtain a preprocessed dataset.
3. A converter steelmaking end-point carbon temperature prediction device based on adaptive data augmentation, wherein the converter steelmaking end-point carbon temperature prediction device based on adaptive data augmentation is used to implement the converter steelmaking end-point carbon temperature prediction method according to any one of claims 1-2, and is characterized in that, The device includes: A data preprocessing module, configured to obtain the historical production dataset of the converter steelmaking process, and preprocess the data in the historical production dataset to obtain a preprocessed dataset; An adaptive data augmentation module, configured to process the preprocessed dataset through the adaptive SMOTE data augmentation technique to obtain a processed dataset; A model building module, configured to build a steelmaking end-point carbon temperature prediction model based on random forest; A model training module, configured to train the steelmaking end-point carbon temperature prediction model according to the processed dataset to obtain a trained steelmaking end-point carbon temperature prediction model with adaptive data augmentation; A model inference module, configured to obtain the production data of the converter steelmaking process to be predicted, and input the production data into the trained steelmaking end-point carbon temperature prediction model with adaptive data augmentation to obtain the prediction result of the converter steelmaking end-point carbon temperature.
4. A converter steelmaking end-point carbon and temperature prediction device, characterized in that The converter steelmaking end-point carbon temperature prediction device includes: A processor; A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 2 is implemented.
5. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code, and the program code can be called by a processor to execute the method according to any one of claims 1 to 2.
Citation Information
Patent Citations
Network security communication high-precision intrusion detection method and system based on multi-scale convolutional neural network
CN118869313A
Risk identification method and device for target object
CN117312857A
Industrial mother machine processing workpiece quality prediction method based on adaptive period discovery
CN117495211A