Multi-production-line data set acquisition method for predicting mechanical properties of hot-rolled strip steel
By establishing a list of element components and process parameters, performing augmentation and similarity calculations, optimizing process parameters, and building a multi-production line data set, the dimension inconsistency and imbalance of the multi-production line hot-rolled strip data set is solved, the accuracy and efficiency of model training are improved, and the accurate prediction of the mechanical properties of hot-rolled strip is achieved.
Patent Information
- Application Number
- CN202510345598.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-08-15
AI Technical Summary
When processing the data set of hot-rolled strips in the production line, the prior art has problems such as insufficient data volume, unbalanced data and inconsistent dimensions, resulting in poor model training effect and low accuracy, and the inability to accurately predict the mechanical properties of hot-rolled strips.
By establishing the element component list and process parameter list, performing augmentation processing and similarity calculation, ensuring the dimensional consistency of each production line data set, and training the neural network model through the data dimension elimination method, optimizing the process parameter set, and finally building a multi-production line data set.
It improves the accuracy and efficiency of model training, accurately predicts the mechanical properties of hot-rolled strips, improves product quality and production efficiency, and solves the problems of inconsistent and imbalance in data dimensions.
Smart Images

Figure CN120492792A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of hot-rolled strip steel, and in particular to a method for acquiring a multi-production line data set for predicting the mechanical properties of hot-rolled strip steel. Background Art
[0002] With the rapid development of the steel industry, hot-rolled strip, a key steel product, has been widely used in numerous fields. Its mechanical properties, such as yield strength, tensile strength, and elongation, are crucial for ensuring product quality and subsequent processing performance. Therefore, accurately predicting the mechanical properties of hot-rolled strip not only significantly improves production efficiency and reduces scrap, but also provides a precise control reference for the production process, thereby promoting overall process improvement. To achieve this goal, advanced data analysis techniques must be used to accurately evaluate and optimize these performance indicators.
[0003] Currently, data preprocessing technology for hot-rolled strip steel mainly focuses on solving problems in single-line datasets, such as data redundancy and outlier removal. Common methods include hierarchical cluster analysis and outlier removal. However, when it comes to multi-line datasets, existing practices are mostly limited to simple data splicing or weighted splicing. Although this method can integrate data resources from different production lines to a certain extent, it does not fully consider the dimensional inconsistencies that may exist between production lines, which often leads to poor model training results and low accuracy. In addition, the problems of insufficient data volume and data imbalance often faced by single-line datasets also limit the effectiveness and reliability of the model.
[0004] The shortcomings of the aforementioned existing technologies demonstrate the need for a more comprehensive and flexible data processing approach to overcome these challenges. An ideal approach should be able to effectively integrate data from multiple production lines while addressing issues such as insufficient data volume, data imbalance, and inconsistent dimensionality. By developing new dataset acquisition and preprocessing techniques, the accuracy and efficiency of model training can be significantly improved, enabling more precise predictions of the mechanical properties of hot-rolled strip. Such improvements will not only promote intelligent and refined management of steel production, but also further enhance product quality and meet the growing market demand for high-performance steel. Summary of the Invention
[0005] In order to effectively integrate data from multiple production lines and solve the problems of insufficient data volume, data imbalance, and inconsistent dimensions, the present invention proposes a multi-production line dataset acquisition method for hot-rolled strip mechanical property prediction, comprising:
[0006] Obtain the parameter data set corresponding to each hot-rolled strip production line; each parameter data set includes n iParameter data; i = 1, 2, ..., m, where m represents the total number of hot-rolled strip production lines; each parameter data includes: element composition set, process parameter set and mechanical parameter set;
[0007] Establish an element component list containing all element components in each data set, and determine the standard dimension of the augmented matrix based on the element component list. Based on this standard dimension, augment the element component set of each parameter data in each parameter data set, and obtain the augmented element component set as the target element set;
[0008] A process parameter list containing all process parameters in each parameter data set is constructed, and the standard dimension of the augmented matrix is determined based on the list; based on the standard dimension, the process parameter set in each parameter data set is augmented, and the missing process parameters are set to zero; then, a similarity calculation method is used to estimate and complete the process parameters initialized to zero to obtain a set to be optimized; the set to be optimized includes the completed process parameter set and the augmented process parameter set that does not need to be estimated and completed;
[0009] All the candidate sets are processed by data dimension elimination, and a neural network model is trained on the processed candidate sets to obtain multiple target models. The target process parameters are set based on the prediction accuracy of each target model, and the target process parameters in all the candidate sets are eliminated to obtain the target process parameter set.
[0010] A multi-line dataset containing multiple samples is constructed, where each sample consists of a target element set, and its corresponding target process parameter set and mechanical parameter set.
[0011] Furthermore, obtaining the target element set specifically includes:
[0012] Create an empty list;
[0013] Traverse each parameter data in each parameter data set, add all the traversed element components to the created empty list, and form an element component list after removing duplicates;
[0014] Determine the standard dimension of the augmented matrix, i.e. the order of elements and the total number of elements, according to the element component list;
[0015] For each element component set in each parameter data set, augmentation processing is performed according to the determined element order and total element number, wherein the missing element components are supplemented by zero values.
[0016] Furthermore, obtaining the set to be optimized includes:
[0017] Create an empty collection;
[0018] Traverse each parameter data in each parameter data set, add all the traversed process parameters to the created empty set, and form a process parameter list after removing duplicates;
[0019] Determine the standard dimensions of the augmented matrix according to the process parameter list, i.e., the process parameter sequence and the total number of process parameters;
[0020] Augmenting the process parameter set in each parameter data set based on the determined process parameter sequence and the total number of process parameters, and setting the missing process parameters to zero values;
[0021] The similarity calculation method is used to estimate and complete the process parameters initialized to zero values.
[0022] Furthermore, the similarity calculation method is used to estimate and complete the process parameters initialized to zero values, specifically including:
[0023] For each process parameter initialized to zero, set the process parameter as a parameter to be estimated, and define the process parameter set to which it belongs as a set to be estimated;
[0024] Based on the parameters to be estimated, C target process parameter sets are screened from all parameter data sets, where each target process parameter set meets the condition of containing the parameters to be estimated with non-zero values; C represents the number of process parameter sets that meet this condition;
[0025] Each target process parameter set and the set to be estimated constitute a process parameter set pair;
[0026] Calculate the similarity between each pair of process parameter sets, excluding the influence of the parameters to be estimated during the calculation process;
[0027] According to the similarity sorting, a preset number of target process parameter sets closest to the target process parameter set are selected;
[0028] Averaging the values of the parameters to be estimated in the selected target process parameter set to obtain a mean value;
[0029] Use the mean to complete the parameters to be estimated in the set to be estimated.
[0030] Furthermore, the method of eliminating data dimensions is used to process all the sets to be optimized, and the neural network model is trained by the processed sets to be optimized to obtain multiple target models, specifically:
[0031] Based on all the sets to be optimized, p comprehensive data sets are obtained by using the data dimension elimination method, where p represents the total number of process parameters in the set to be optimized; for each comprehensive data set, the data set is divided into a training set and a test set, and the neural network model is trained with the training set to obtain the corresponding target model, and the prediction accuracy corresponding to the target model is calculated with the test set.
[0032] Furthermore, the method of eliminating data dimensions based on all the sets to be optimized is used to obtain p comprehensive data sets, specifically:
[0033] S1: Set loop variable j=1 to identify the index of the process parameter currently being rejected;
[0034] S2: Select the jth process parameter in the process parameter list as the parameter to be eliminated;
[0035] S3: removing the parameters to be eliminated from all the sets to be optimized to form a new subset of process parameters;
[0036] S4: construct a comprehensive dataset through the process parameter subset and mark it as the jth comprehensive dataset;
[0037] S5: Increase the loop variable j, i.e. j=j+1;
[0038] S6: Determine whether j is less than or equal to p. If so, return to step S2; if not, end the loop.
[0039] Furthermore, the method of eliminating data dimensions is used to process all the sets to be optimized, and a neural network model is trained by the processed sets to be optimized to obtain multiple target models. The target process parameters are set based on the prediction accuracy of each target model, and the target process parameters in all the sets to be optimized are eliminated to obtain a target process parameter set, which is specifically:
[0040] A1: Initialize the prediction accuracy of the previous model to 0;
[0041] A2: Set loop variable j = 1 to identify the index of the process parameter currently being rejected;
[0042] A3: Select the jth process parameter in the process parameter list as the parameter to be eliminated;
[0043] A4: Remove the parameters to be eliminated from all the sets to be optimized to form a new subset of process parameters;
[0044] A5: Build a comprehensive dataset using process parameter subsets and divide it into training and test sets;
[0045] A6: Train the neural network model using the training set to obtain the corresponding target model, and calculate the prediction accuracy of the target model using the test set.
[0046] A7: Determine whether accuracy is greater than previous_accuracy. If so, update the target process parameters to the current parameters to be eliminated, and update previous_accuracy = accuracy. If not, keep the target process parameters and previous_accuracy unchanged.
[0047] A8: Determine whether |accuracy - previous_accuracy| < ∈ holds. If so, eliminate all target process parameters in the set to be optimized, obtain the target process parameter set, and end the loop. If not, proceed to the next step; ∈ represents the set threshold.
[0048] A9: Increase the loop variable j, i.e. j = j + 1;
[0049] A10: Determine whether j is less than or equal to p. If so, return to step A3; if not, end the loop. p represents the total number of process parameters in the set to be optimized.
[0050] Furthermore, the comprehensive data set is constructed by using the process parameter subsets, specifically:
[0051] The corresponding process parameter subset is marked by the mechanical parameter set, and the marked process parameter subset is set as a sample; a comprehensive data set is formed through each sample.
[0052] Furthermore, the set threshold ∈=0.005.
[0053] Furthermore, the neural network model is a random forest model.
[0054] Compared with the prior art, the present invention has at least the following beneficial effects:
[0055] (1) The present invention first standardizes the dimensions of each production line data set by establishing an element composition list and a process parameter list, and augments the element composition set and process parameter set in each parameter data to ensure the consistency of all data sets; for missing process parameters, a similarity calculation method is used to estimate and complete them to form a set to be optimized; then, these sets to be optimized are processed using a data dimension elimination method, multiple target models are obtained by training a neural network model, and the process parameter set is optimized based on the prediction accuracy of each target model; finally, a multi-production line data set containing multiple samples is constructed, each sample consisting of a target element set, an optimized target process parameter set, and a mechanical parameter set; this method not only solves the problem of inconsistent data dimensions, but also improves the accuracy and efficiency of model training, thereby more accurately predicting the mechanical properties of hot-rolled strip, providing strong support for improving product quality and production efficiency.
[0056] (2) The method proposed in this application standardizes the dimensions of each data set by constructing a list of elemental components and a list of process parameters, ensuring that data from different production lines can be effectively integrated. This process not only solves the problem of inconsistent data dimensions caused by direct splicing or weighted splicing in traditional methods, but also overcomes the problem of insufficient and unbalanced data from a single production line. By unifying all parameters to the same dimensional standard through augmentation processing, the resulting multi-production line data set is more effective and more accurate for model training.
[0057] (3) For process parameters initialized to zero, the present invention uses a similarity calculation method for estimation and completion. This method selects a target process parameter set from all parameter datasets based on the parameters to be estimated and uses the similarity between these sets to complete the missing values. This strategy significantly improves the quality and completeness of the dataset and reduces model prediction errors caused by missing data.
[0058] (4) The present invention processes all the candidate sets through the method of data dimension elimination, and sets the target process parameters based on the prediction accuracy of each target model, and finally obtains the optimal target process parameter set. This method can not only effectively identify and eliminate those process parameters that have little contribution to the model prediction accuracy or even have a negative effect, but also dynamically adjust to find the best parameter combination. This makes the final multi-line data set more efficient for model training, improves the model prediction accuracy and efficiency, and is of great significance for improving product quality, reducing scrap rate, and achieving precise control in the production process. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 This is a flow chart of a method for acquiring a multi-production line data set for hot-rolled strip mechanical property prediction according to an embodiment of the present invention;
[0060] Figure 2 Flowchart for obtaining p comprehensive data sets in an embodiment of the present invention;
[0061] Figure 3 This is a flow chart for obtaining a target process parameter set in an embodiment of the present invention. DETAILED DESCRIPTION
[0062] The following are specific embodiments of the present invention and the accompanying drawings to further describe the technical solutions of the present invention, but the present invention is not limited to these embodiments.
[0063] Example 1
[0064] In order to effectively integrate data from multiple production lines and solve the problems of insufficient data volume, data imbalance and inconsistent dimensions, such as Figure 1As shown, the present invention proposes a multi-production line data set acquisition method for hot-rolled strip mechanical property prediction, comprising:
[0065] Obtain the parameter data set corresponding to each hot-rolled strip production line; each parameter data set includes n i Parameter data; i = 1, 2, ..., m, where m represents the total number of hot-rolled strip production lines; each parameter data includes: an element composition set, a process parameter set, and a mechanical parameter set;
[0066] In this embodiment, the parameter data sets corresponding to m hot-rolled strip production lines are represented as: D1, D2, ..., D m , the i-th parameter data set includes n i Parameter data, specifically recorded as;
[0067]
[0068] Where, d i,j =(x i,j ,y i,j ), j=1,2…n i ,in:
[0069] x i,j =[x i,j,1 ,x i,j,2 ,…,x i,j,qi ,z i,j,1 ,z i,j,2 ,…,z i,j,pi ] T ;
[0070] because:
[0071]
[0072]
[0073] So, we have:
[0074]
[0075] Where, Indicates that the set type (superscript 1) is an element component set of element components, Indicates that the set type (superscript 2) is a process parameter set of process parameters, q i Represents the number of element components in the element component set, p i Indicates the number of process parameters in the process parameter set.
[0076] y i,j =[y i,j,1 ,y i,j,2 ,y i,j,3 ]T ;
[0077] Among them, y i,j,1 represents the yield strength, y i,j,2 Indicates tensile strength, y i,j,3 Indicates elongation.
[0078] The above formula can be used to i,j Written as:
[0079]
[0080] Since the data collected from different production lines may have differences in elemental composition dimensions and order, a series of preprocessing steps must be performed on the elemental composition data to ensure that all elemental composition sets have consistent elemental composition dimensions and order, thereby ensuring the consistency and comparability of data from different production lines. Specifically:
[0081] Establish an element component list containing all element components in each data set, and determine the standard dimension of the augmented matrix based on the element component list. Based on this standard dimension, augment the element component set of each parameter data in each parameter data set, and obtain the augmented element component set as the target element set;
[0082] The acquisition of the target element set specifically includes:
[0083] Create an empty list ε final ; ε final =φ;
[0084] Traverse each parameter data in each parameter data set, add all the traversed element components to the created empty list, and form an element component list after removing duplicates;
[0085] Specifically, start traversing from the first data in all parameter data sets, assuming that the first data is d i,j , the elemental composition it contains is in:
[0086] x i,j,k (k=1,2,...,q i ) is the element component that appears in the first data. By traversing all parameter data sets, the element components in each parameter data set are added to ε final In this process, we can express it through set union operation as:
[0087]
[0088] Determine the standard dimension of the augmented matrix, i.e. the order of elements and the total number of elements, according to the element component list;
[0089] For each element component set in each parameter data set, augmentation processing is performed according to the determined element order and total element number, wherein the missing element components are supplemented by zero values.
[0090] After traversing all the data, the final generated element component list ε final contains all the element components that have appeared, and its size is:
[0091] M final =|ε final |;
[0092] That is, the final augmented data dimension (standard dimension of the augmented matrix) M final Equal to the element composition list ε final The number of element components in . Take the element component list ε final The order of elements in is the order of elements in the augmented matrix. For example, the final order of elements is:
[0093]
[0094] Determine the element composition list ε containing all element compositions final After that, each element component set can be expanded to a unified dimension M final , for the jth data d in the i-th production line i,j , assuming that the element composition it contains is in ε final The chemical elements are already arranged in order, but some dimensions are missing. The augmented data can be represented as a matrix with a length of M. final A vector of . The missing mass percentages of the elements are filled with 0.
[0095] Specifically, the augmented element component set is an M final A vector of dimension, according to ε final The order of the elements in .
[0096] When integrating parameter data sets from multiple production lines, each parameter data set may contain different process parameters because different production lines may use different sensors or process conditions. To solve this problem, it is necessary to define a unified initial dimension that covers the process parameters appearing in all parameter data sets. Specifically:
[0097] A process parameter list containing all process parameters in each parameter data set is constructed, and the standard dimension of the augmented matrix is determined based on the list; based on the standard dimension, the process parameter set in each parameter data set is augmented, and the missing process parameters are set to zero; then, a similarity calculation method is used to estimate and complete the process parameters initialized to zero to obtain a set to be optimized; the set to be optimized includes the completed process parameter set and the augmented process parameter set that does not need to be estimated and completed;
[0098] It should be noted that the data in the set to be optimized have been normalized.
[0099] The acquisition of the set to be optimized includes:
[0100] Create an empty set P init =φ;
[0101] Traverse each parameter data in each parameter data set, add all the traversed process parameters to the created empty set, and form a process parameter list after removing duplicates;
[0102] Specifically, for example, for the kth parameter data set D k Traverse, for data d i,j , which contains the process parameter set: P i,j ={z i,j,1 ,z i,j,2 ,...,z i,j,pi}, P i,j Add to set P init middle,
[0103] All the data in the data set are traversed, and after the traversal is completed, a set of process parameters in all the data is finally obtained, that is, the process parameter list P init .
[0104] Determine the standard dimensions of the augmented matrix according to the process parameter list, i.e., the process parameter sequence and the total number of process parameters;
[0105] Augmenting the process parameter set in each parameter data set based on the determined process parameter sequence and the total number of process parameters, and setting the missing process parameters to zero values;
[0106] The similarity calculation method is used to estimate and complete the process parameters initialized to zero values.
[0107] By constructing a process parameter list, an initial process parameter dimension is determined that encompasses all process parameters present in the parameter dataset. To unify all process parameter sets onto this standard dimension, missing process parameter sets need to be completed. However, simply setting these missing process parameters to zero is inappropriate, as in real production environments, these parameter values are not truly zero; they are simply not collected. Therefore, this method uses a similarity calculation method to estimate and complete these process parameters initialized to zero.
[0108] The similarity calculation method is used to estimate and complete the process parameters initialized to zero, specifically including:
[0109] For each process parameter initialized to zero, set the process parameter as a parameter to be estimated, and define the process parameter set to which it belongs as a set to be estimated;
[0110] Based on the parameters to be estimated, C target process parameter sets are screened from all parameter data sets, where each target process parameter set meets the condition of containing the parameters to be estimated with non-zero values; C represents the number of process parameter sets that meet this condition;
[0111] Specifically, from all parameter data sets, all process parameter sets whose parameters to be estimated are not zero values are screened out as target process parameter sets.
[0112] Each target process parameter set and the set to be estimated constitute a process parameter set pair;
[0113] Calculate the similarity between each pair of process parameter sets, excluding the influence of the parameters to be estimated during the calculation process;
[0114] Specifically, for example: data d i,j The qth parameter is missing from the process parameter set, and the data d m,n The parameter is not missing in the process parameter set (the set that does not lack the parameter is the target process parameter set), and the similarity (Euclidean distance) between the two is calculated:
[0115]
[0116] Where z i,j,l With z m,n,l The data z i,j With data z m,n At the value of the lth process parameter, p represents the standard dimension of the augmented matrix determined according to the process parameter list, that is, the number of process parameters in the process parameter list. Since the influence of the qth parameter needs to be excluded, only the distances of other parameters are calculated, so l≠q.
[0117] Sort by similarity (sort all data in ascending order based on the calculated Euclidean distance) and select the nearest preset number k of target process parameter sets;
[0118] The values of the parameters to be estimated in the selected target process parameter set are averaged to obtain the mean value; specifically:
[0119] For the missing qth process parameter z i,j,q , is completed by the mean of the same parameter of the k nearest neighbors. The specific process is:
[0120]
[0121] Where p r,q is the value of the rth nearest neighbor data at the qth process parameter.
[0122] Use this mean Complete the parameters to be estimated in the set to be estimated.
[0123] Right i,j Repeat the above steps to complete all missing parameters in d i,j All data in have completed the completion process.
[0124] All the candidate sets are processed by data dimension elimination, and a neural network model is trained on the processed candidate sets to obtain multiple target models. The target process parameters are set based on the prediction accuracy of each target model, and the target process parameters in all the candidate sets are eliminated to obtain the target process parameter set.
[0125] The method of eliminating data dimensions is used to process all the sets to be optimized, and the neural network model is trained by the processed sets to be optimized to obtain multiple target models, specifically:
[0126] Based on all the sets to be optimized, we use the data dimension elimination method to obtain p comprehensive data sets, where p represents the total number of process parameters in the set to be optimized; for each comprehensive data set, we divide the data set into training set X train With the test set X test , through the training set X train Train the neural network model to obtain the corresponding target model, and pass the test set X test Calculate the prediction accuracy corresponding to the target model.
[0127] The neural network model is a random forest model.
[0128] like Figure 2 As shown, the method of eliminating data dimensions based on all the sets to be optimized is used to obtain p comprehensive data sets, specifically:
[0129] S1: Set loop variable j=1 to identify the index of the process parameter currently being rejected;
[0130] S2: Select the jth process parameter in the process parameter list as the parameter to be eliminated;
[0131] S3: removing the parameters to be eliminated from all the sets to be optimized to form a new subset of process parameters;
[0132] S4: construct a comprehensive dataset through the process parameter subset and mark it as the jth comprehensive dataset;
[0133] S5: Increase the loop variable j, i.e. j=j+1;
[0134] S6: Determine whether j is less than or equal to p. If so, return to step S2; if not, end the loop.
[0135] For example, assuming there are 5 initial process parameters in the set to be optimized, the set to be optimized can be expressed as X current ={p1,p2,p3,p4,p5}; then:
[0136] In progress Figure 2 For the first cycle shown, all subsets of process parameters can be expressed as:
[0137]
[0138] In progress Figure 2 For the second cycle shown, all subsets of process parameters can be expressed as:
[0139]
[0140] And so on, in the Figure 2 For the fifth cycle shown, all subsets of process parameters can be expressed as:
[0141]
[0142] The comprehensive data set is constructed by using a subset of process parameters, specifically:
[0143] The corresponding process parameter subset is marked by the mechanical parameter set, and the marked process parameter subset is set as a sample; a comprehensive data set is formed through each sample.
[0144] For each comprehensive data set, the data set is divided into training set X train With the test set X test , through the training set X train Train the neural network model to obtain the corresponding target model, and pass the test set X test Calculate the prediction accuracy corresponding to the target model, specifically:
[0145] Divide the training set X into 80% and 20% train With the test set X test , the training set is used for model training, and the test set is used for performance evaluation;
[0146] Through the training set X train Train the neural network model to obtain the corresponding target model M train ;
[0147] Through the test set X test Calculate the prediction accuracy corresponding to the target model. The calculation formula for the prediction accuracy is:
[0148]
[0149] In the formula is the indicator function, when the true value y i and predicted value If they are the same, take 1; if they are different, take 0. test is the number of samples in the test set.
[0150] The target process parameters are set based on the prediction accuracy of each target model, and all target process parameters in the set to be optimized are eliminated. Specifically, the parameters to be eliminated corresponding to the target model with the highest prediction accuracy are selected as the target process parameters, and all target process parameters in the set to be optimized are eliminated.
[0151] A multi-line dataset containing multiple samples is constructed, where each sample consists of a target element set, and its corresponding target process parameter set and mechanical parameter set.
[0152] The present invention first standardizes the dimensions of each production line data set by establishing an element composition list and a process parameter list, and augments the element composition set and process parameter set in each parameter data to ensure the consistency of all data sets; for missing process parameters, a similarity calculation method is used to estimate and complete them to form a set to be optimized; then, these sets to be optimized are processed using a data dimension elimination method, multiple target models are obtained by training a neural network model, and the process parameter set is optimized based on the prediction accuracy of each target model; finally, a multi-production line data set containing multiple samples is constructed, each sample consisting of a target element set, an optimized target process parameter set, and a mechanical parameter set; this method not only solves the problem of inconsistent data dimensions, but also improves the accuracy and efficiency of model training, thereby more accurately predicting the mechanical properties of hot-rolled strip, providing strong support for improving product quality and production efficiency.
[0153] Example 2
[0154] In order to effectively identify and eliminate process parameters that have little contribution to or even negative effects on model prediction accuracy, and to dynamically adjust to find the optimal parameter combination, the present invention also proposes another implementation method, specifically:
[0155] A method for acquiring a multi-production line data set for hot-rolled strip mechanical property prediction includes:
[0156] Obtain the parameter data set corresponding to each hot-rolled strip production line; each parameter data set includes n i Parameter data; i = 1, 2, ..., m, where m represents the total number of hot-rolled strip production lines; each parameter data includes: an element composition set, a process parameter set, and a mechanical parameter set;
[0157] Establish an element component list containing all element components in each data set, and determine the standard dimension of the augmented matrix based on the element component list. Based on this standard dimension, augment the element component set of each parameter data in each parameter data set, and obtain the augmented element component set as the target element set;
[0158] The acquisition of the target element set specifically includes:
[0159] Create an empty list;
[0160] Traverse each parameter data in each parameter data set, add all the traversed element components to the created empty list, and form an element component list after removing duplicates;
[0161] Determine the standard dimension of the augmented matrix, i.e. the order of elements and the total number of elements, according to the element component list;
[0162] For each element component set in each parameter data set, augmentation processing is performed according to the determined element order and total element number, wherein the missing element components are supplemented by zero values.
[0163] A process parameter list containing all process parameters in each parameter data set is constructed, and the standard dimension of the augmented matrix is determined based on the list; based on the standard dimension, the process parameter set in each parameter data set is augmented, and the missing process parameters are set to zero; then, a similarity calculation method is used to estimate and complete the process parameters initialized to zero to obtain a set to be optimized; the set to be optimized includes the completed process parameter set and the augmented process parameter set that does not need to be estimated and completed;
[0164] The acquisition of the set to be optimized includes:
[0165] Create an empty collection;
[0166] Traverse each parameter data in each parameter data set, add all the traversed process parameters to the created empty set, and form a process parameter list after removing duplicates;
[0167] Determine the standard dimensions of the augmented matrix according to the process parameter list, i.e., the process parameter sequence and the total number of process parameters;
[0168] Augmenting the process parameter set in each parameter data set based on the determined process parameter sequence and the total number of process parameters, and setting the missing process parameters to zero values;
[0169] The similarity calculation method is used to estimate and complete the process parameters initialized to zero values.
[0170] The method proposed in this application standardizes the dimensionality of each dataset by constructing lists of elemental components and process parameters, ensuring that data from different production lines can be effectively integrated. This process not only solves the problem of inconsistent data dimensions caused by direct or weighted splicing in traditional methods, but also overcomes the problems of insufficient and unbalanced data from a single production line. By unifying all parameters to the same dimensionality standard through augmentation, the resulting multi-production line dataset is more effective and more accurate for model training.
[0171] The similarity calculation method is used to estimate and complete the process parameters initialized to zero, specifically including:
[0172] For each process parameter initialized to zero, set the process parameter as a parameter to be estimated, and define the process parameter set to which it belongs as a set to be estimated;
[0173] Based on the parameters to be estimated, C target process parameter sets are screened from all parameter data sets, where each target process parameter set meets the condition of containing the parameters to be estimated with non-zero values; C represents the number of process parameter sets that meet this condition;
[0174] Each target process parameter set and the set to be estimated constitute a process parameter set pair;
[0175] Calculate the similarity between each pair of process parameter sets, excluding the influence of the parameters to be estimated during the calculation process;
[0176] According to the similarity sorting, a preset number of target process parameter sets closest to the target process parameter set are selected;
[0177] Averaging the values of the parameters to be estimated in the selected target process parameter set to obtain a mean value;
[0178] Use the mean to complete the parameters to be estimated in the set to be estimated.
[0179] For process parameters initialized to zero, this paper uses a similarity calculation method to estimate and complete missing values. This method selects a target set of process parameters from all parameter datasets based on the parameters to be estimated and uses the similarity between these sets to complete missing values. This strategy significantly improves the quality and completeness of the dataset and reduces model prediction errors caused by missing data.
[0180] All the sets to be optimized are processed by the method of data dimension elimination. The neural network model is trained with the processed sets to be optimized to obtain multiple target models. The target process parameters are set based on the prediction accuracy of each target model. The target process parameters in all the sets to be optimized are eliminated to obtain the target process parameter set. Figure 3 As shown, specifically:
[0181] A1: Initialize the prediction accuracy of the previous model to 0;
[0182] A2: Set loop variable j = 1 to identify the index of the process parameter currently being rejected;
[0183] A3: Select the jth process parameter in the process parameter list as the parameter to be eliminated;
[0184] A4: Remove the parameters to be eliminated from all the sets to be optimized to form a new subset of process parameters;
[0185] A5: Build a comprehensive dataset using process parameter subsets and divide it into training and test sets;
[0186] A6: Train the neural network model using the training set to obtain the corresponding target model, and calculate the prediction accuracy of the target model using the test set.
[0187] A7: Determine whether accuracy is greater than previous_accuracy. If so, update the target process parameters to the current parameters to be eliminated, and update previous_accuracy = accuracy. If not, keep the target process parameters and previous_accuracy unchanged.
[0188] There is another implementation method of step A7, specifically: updating the target process parameters according to the prediction accuracy. The update formula of the target process parameters can be expressed as:
[0189]
[0190] Where, Indicates the updated target process parameters; arg max: This is a mathematical operator used to find the independent variable that makes a function reach its maximum value. In this embodiment, it is used to find the prediction accuracy Acc i The largest parameter to be eliminated;
[0191] i∈{1, 2, ..., P current} represents the target model number. A1-A10 cycles generate a target model. The numbers start from 1 and increase in sequence. current Indicates the number of the target model corresponding to the current cycle. In other words, the purpose of the target process parameter update formula is to Update to: the parameters to be removed that maximize the model prediction accuracy.
[0192] A8: Determine whether |accuracy - previous_accuracy| < ∈ holds. If so, eliminate all target process parameters in the set to be optimized, obtain the target process parameter set, and end the loop. If not, proceed to the next step; ∈ represents the set threshold.
[0193] The set threshold ∈=0.005.
[0194] A9: Increase the loop variable j, i.e. j = j + 1;
[0195] A10: Determine whether j is less than or equal to p. If so, return to step A3; if not, end the loop. p represents the total number of process parameters in the set to be optimized.
[0196] The process from A1 to A10 not only dynamically adjusts to find the optimal parameter combination, but also effectively identifies parameters that do not significantly contribute to the model's prediction accuracy or even have a negative impact, thereby eliminating them (A8). After traversing all process parameters or when the judgment condition of A8 is met, the iteration stops, ensuring the efficiency and rationality of the algorithm. This strategy not only improves the model's prediction performance and reduces unnecessary computational overhead, but also provides a more precise control basis for the production process, helping to improve product quality, reduce scrap rates, and promote intelligent management of steel production. In other words, steps A1 to A10, by using both the loop variable j and the prediction accuracy for dual judgment, not only ensure that all process parameters are systematically evaluated and optimized, but also effectively identify and eliminate parameters that do not contribute to the improvement of model performance, thereby dynamically finding the optimal parameter combination and significantly improving the model's prediction accuracy and efficiency.
[0197] The neural network model is a random forest model.
[0198] The comprehensive data set is constructed by using a subset of process parameters, specifically:
[0199] The corresponding process parameter subset is marked by the mechanical parameter set, and the marked process parameter subset is set as a sample; a comprehensive data set is formed through each sample.
[0200] A multi-line dataset containing multiple samples is constructed, where each sample consists of a target element set, and its corresponding target process parameter set and mechanical parameter set.
[0201] This method processes all candidate sets through a data dimension elimination method and sets target process parameters based on the prediction accuracy of each target model, ultimately obtaining the optimal set of target process parameters. This approach not only effectively identifies and eliminates process parameters that contribute little to or even negatively impact model prediction accuracy, but also dynamically adjusts to find the optimal parameter combination. This makes the resulting multi-line dataset more efficient for model training, improving model prediction accuracy and efficiency, which is of great significance for improving product quality, reducing scrap rates, and achieving precise control during the production process.
[0202] It should be noted that all directional indications in the embodiments of the present invention (such as up, down, left, right, front, back, etc.) are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.
[0203] In addition, in the present invention, descriptions such as "first," "second," and "one" are for descriptive purposes only and should not be understood to indicate or imply their relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0204] In the present invention, unless otherwise specified or limited, the terms "connection" and "fixation" should be understood in a broad sense. For example, "fixation" can mean fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; internal communication between two elements or interaction between two elements, unless otherwise specified. Those skilled in the art will be able to understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0205] In addition, the technical solutions between the various embodiments of the present invention can be combined with each other, but it must be based on the fact that ordinary technicians in this field can implement it. When the combination of technical solutions is mutually contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
Claims
1. A method for acquiring a multi-line data set for hot-rolled strip mechanical property prediction, characterized in that: include: Obtain the parameter data set corresponding to each hot-rolled strip production line; each parameter data set includes n i Parameter data; i=1,2,...,m, where m represents the total number of hot-rolled strip production lines; each parameter data includes: element composition set, process parameter set and mechanical parameter set; Establish an element component list containing all element components in each data set, and determine the standard dimension of the augmented matrix based on the element component list. Based on this standard dimension, augment the element component set of each parameter data in each parameter data set, and obtain the augmented element component set as the target element set; A process parameter list containing all process parameters in each parameter data set is constructed, and the standard dimension of the augmented matrix is determined based on the list; based on the standard dimension, the process parameter set in each parameter data set is augmented, and the missing process parameters are set to zero; then, a similarity calculation method is used to estimate and complete the process parameters initialized to zero to obtain a set to be optimized; the set to be optimized includes the completed process parameter set and the augmented process parameter set that does not need to be estimated and completed; All the candidate sets are processed by data dimension elimination, and a neural network model is trained on the processed candidate sets to obtain multiple target models. The target process parameters are set based on the prediction accuracy of each target model, and the target process parameters in all the candidate sets are eliminated to obtain the target process parameter set. A multi-line dataset containing multiple samples is constructed, where each sample consists of a target element set, and its corresponding target process parameter set and mechanical parameter set.
2. The method for acquiring a multi-production line data set for hot-rolled strip mechanical property prediction according to claim 1, characterized in that: The acquisition of the target element set specifically includes: Create an empty list; Traverse each parameter data in each parameter data set, add all the traversed element components to the created empty list, and form an element component list after removing duplicates; Determine the standard dimension of the augmented matrix, i.e. the order of elements and the total number of elements, according to the element component list; For each element component set in each parameter data set, augmentation processing is performed according to the determined element order and total element number, wherein the missing element components are supplemented by zero values.
3. The method for acquiring a multi-production line data set for hot-rolled strip mechanical property prediction according to claim 1, characterized in that: The acquisition of the set to be optimized includes: Create an empty collection; Traverse each parameter data in each parameter data set, add all the traversed process parameters to the created empty set, and form a process parameter list after removing duplicates; Determine the standard dimensions of the augmented matrix according to the process parameter list, i.e., the process parameter sequence and the total number of process parameters; Augmenting the process parameter set in each parameter data set based on the determined process parameter sequence and the total number of process parameters, and setting the missing process parameters to zero values; The similarity calculation method is used to estimate and complete the process parameters initialized to zero values.
4. The method for acquiring a multi-production line data set for hot-rolled strip mechanical property prediction according to claim 3, characterized in that: The similarity calculation method is used to estimate and complete the process parameters initialized to zero, specifically including: For each process parameter initialized to zero, set the process parameter as a parameter to be estimated, and define the process parameter set to which it belongs as a set to be estimated; Based on the parameters to be estimated, C target process parameter sets are screened from all parameter data sets, where each target process parameter set meets the condition of containing the parameters to be estimated with non-zero values; C represents the number of process parameter sets that meet this condition; Each target process parameter set and the set to be estimated constitute a process parameter set pair; Calculate the similarity between each pair of process parameter sets, excluding the influence of the parameters to be estimated during the calculation process; According to the similarity sorting, a preset number of target process parameter sets closest to the target process parameter set are selected; Averaging the values of the parameters to be estimated in the selected target process parameter set to obtain a mean value; Use the mean to complete the parameters to be estimated in the set to be estimated.
5. The method for acquiring a multi-production line data set for hot-rolled strip mechanical property prediction according to claim 4, characterized in that: The method of eliminating data dimensions is used to process all the sets to be optimized, and the neural network model is trained by the processed sets to be optimized to obtain multiple target models, specifically: Based on all the sets to be optimized, p comprehensive data sets are obtained by using the data dimension elimination method, where p represents the total number of process parameters in the set to be optimized; for each comprehensive data set, the data set is divided into a training set and a test set, and the neural network model is trained with the training set to obtain the corresponding target model, and the prediction accuracy corresponding to the target model is calculated with the test set.
6. The method for acquiring a multi-production line data set for hot-rolled strip mechanical property prediction according to claim 5, characterized in that: The method of eliminating data dimensions based on all the candidate sets is used to obtain p comprehensive data sets, specifically: S1: Set loop variable j=1 to identify the index of the process parameter currently being rejected; S2: Select the jth process parameter in the process parameter list as the parameter to be eliminated; S3: removing the parameters to be eliminated from all the sets to be optimized to form a new subset of process parameters; S4: construct a comprehensive dataset through the process parameter subset and mark it as the jth comprehensive dataset; S5: Increase the loop variable j, i.e. j=j+1; S6: Determine whether j is less than or equal to p. If so, return to step S2; if not, end the loop.
7. The method for acquiring a multi-production line data set for hot-rolled strip mechanical property prediction according to claim 4, characterized in that: The method of eliminating data dimensions is used to process all the sets to be optimized, and a neural network model is trained on the processed sets to be optimized to obtain multiple target models. The target process parameters are set based on the prediction accuracy of each target model, and the target process parameters in all the sets to be optimized are eliminated to obtain a target process parameter set, which is specifically: A1: Initialize the prediction accuracy of the previous model to 0; A2: Set loop variable j = 1 to identify the index of the process parameter currently being rejected; A3: Select the jth process parameter in the process parameter list as the parameter to be eliminated; A4: Remove the parameters to be eliminated from all the sets to be optimized to form a new subset of process parameters; A5: Build a comprehensive dataset using process parameter subsets and divide it into training and test sets; A6: Train the neural network model using the training set to obtain the corresponding target model, and calculate the prediction accuracy of the target model using the test set. A7: Determine whether accuracy is greater than previous_accuracy. If so, update the target process parameters to the current parameters to be eliminated, and update previous_accuracy = accuracy. If not, keep the target process parameters and previous_accuracy unchanged. A8: Determine whether |accuracy - previous_accuracy| < ∈ holds. If so, eliminate all target process parameters in the set to be optimized, obtain the target process parameter set, and end the loop. If not, proceed to the next step; ∈ represents the set threshold. A9: Increase the loop variable j, i.e. j = j + 1; A10: Determine whether j is less than or equal to p. If so, return to step A3; if not, end the loop. p represents the total number of process parameters in the set to be optimized.
8. The method for acquiring a multi-production line dataset for hot-rolled strip mechanical property prediction according to claim 6 or 7, wherein the comprehensive dataset is constructed using a subset of process parameters, specifically: The corresponding process parameter subset is marked by the mechanical parameter set, and the marked process parameter subset is set as a sample; a comprehensive data set is formed through each sample.
9. The method for acquiring a multi-production line data set for hot-rolled strip mechanical property prediction according to claim 7, characterized in that: The set threshold ∈=0.
005.
10. The method for acquiring a multi-production line data set for hot-rolled strip mechanical property prediction according to any one of claims 1 to 9, characterized in that: The neural network model is a random forest model.